SimpleRNN • LSTM • GRU • DistilBERT • HuggingFace • TensorFlow • Gradio
A Deep Learning and NLP project that automatically classifies customer complaint narratives into the correct complaint category.
The project compares traditional recurrent neural networks with a fine-tuned Transformer model to determine the best-performing architecture for complaint classification.
- Complete NLP preprocessing pipeline
- Text cleaning and normalization
- Class balancing
- Tokenization & Sequence Padding
- Word Embedding Layer
- Built and trained from scratch:
- SimpleRNN
- LSTM
- GRU
- Fine-tuned DistilBERT Transformer
- Performance comparison between all models
- Interactive Gradio web application
- Real-time complaint prediction with confidence scores
| Model | Accuracy |
|---|---|
| SimpleRNN | 78.9% |
| GRU | 81.8% |
| LSTM | 82.5% |
| DistilBERT (Fine-tuned) | 85.1% |
Consumer Complaints Dataset for NLP
Contains real customer complaints from financial services.
Classes include:
- Credit Reporting
- Credit Card
- Debt Collection
- Mortgages and Loans
- Retail Banking
Dataset
↓
Text Preprocessing
↓
Class Balancing
↓
Tokenization
↓
Sequence Padding
↓
Embedding Layer
↓
Train Deep Learning Models
↓
Evaluate Performance
↓
Fine-tune DistilBERT
↓
Model Comparison
↓
Gradio Deployment
- Accuracy
- Precision
- Recall
- F1-score
- Classification Report
- Confusion Matrix
The project includes an interactive Gradio interface that allows users to:
- Enter a complaint narrative
- Predict the complaint category
- Display prediction confidence
- Compare probabilities across all categories
---
- Python
- TensorFlow / Keras
- HuggingFace Transformers
- PyTorch
- Scikit-learn
- NLTK
- NumPy
- Pandas
- Matplotlib
- Seaborn
- Gradio
Consumer-Complaint-Classification
│
├── consumer-complaint-classification.ipynb
├── gradio_app.py
├── requirements.txt
├── tokenizer.pkl
├── label_encoder.pkl
├── SimpleRNN_model.keras
├── LSTM_model.keras
├── GRU_model.keras
│
└── transformer_complaint_model_final
├── config.json
├── tokenizer.json
├── tokenizer_config.json
└── model.safetensors
The Fine-tuned DistilBERT model achieved the highest performance and was selected for deployment in the Gradio application.
- 💻 GitHub Repository: https://github.com/basmalakhaled20/Consumer-Complaint-Classification
- 📒 Kaggle Notebook: https://www.kaggle.com/code/basmalakhaled20/consumer-complaint-classification
- 📊 Dataset: https://www.kaggle.com/datasets/shashwatwork/consume-complaints-dataset-fo-nlp
Basmala Khaled
AI Engineer

