- Project Scope
- Audience-Friendly Description
- Business Impact
- Project Description
- Folder Structure
- Flow Diagram
- Demo Video
- Output
- How to Run & Dependencies
- Contribution Guidelines
- License & Credits
- Next Steps
This project implements an end-to-end NLP pipeline for abstractive text summarization using the PEGASUS transformer model.
It covers the complete machine learning lifecycle:
- Remote data ingestion
- Dataset validation
- Tokenization and transformation
- Model training using Hugging Face Trainer
- Model evaluation using ROUGE metrics
The pipeline is modular, configurable, and suitable for production-scale workflows.
Think of long chat conversations that need to be summarized into short, meaningful insights.
This project:
- Takes raw conversational data
- Prepares it for machine learning
- Trains a state-of-the-art NLP model
- Generates concise summaries
- Evaluates summary quality automatically
All steps are automated and reusable.
- Automates summarization of conversations and transcripts
- Reduces manual summarization effort
- Improves operational efficiency
- Scales easily for large datasets
- Leverages pretrained models to reduce cost and time
Applicable Use Cases:
- Customer support chat summarization
- Call center analytics
- Conversational AI systems
- Documentation automation
The system consists of five core stages:
- Downloads dataset from a remote URL
- Avoids re-downloading existing files
- Extracts compressed ZIP files
- Checks presence of required dataset files
- Writes validation status to a report file
- Prevents invalid data from entering the pipeline
- Tokenizes dialogue and summary pairs
- Converts text into model-compatible tensors
- Stores transformed datasets on disk
- Fine-tunes the PEGASUS model
- Uses Hugging Face Trainer
- Supports GPU acceleration
- Saves trained model and tokenizer
- Generates summaries on test data
- Computes ROUGE-1, ROUGE-2, ROUGE-L, ROUGE-Lsum
- Stores evaluation metrics as CSV
├── artifacts/ │ ├── data_ingestion/ │ │ └── samsum_dataset/ │ ├── data_transformation/ │ │ └── samsum_dataset/ │ ├── model_trainer/ │ │ └── pegasus-samsum-model/ │ └── model_evaluation/ │ └── rouge_scores.csv │ ├── config/ │ ├── data_ingestion.yaml │ ├── data_validation.yaml │ ├── data_transformation.yaml │ ├── model_trainer.yaml │ └── model_evaluation.yaml │ ├── src/ │ ├── data_ingestion.py │ ├── data_validation.py │ ├── data_transformation.py │ ├── model_trainer.py │ └── model_evaluation.py │ ├── README.md ├── requirements.txt └── main.py
- Trained PEGASUS model
- Tokenizer files
Saved as CSV:
artifacts/model_evaluation/rouge_scores.csv
ROUGE Metrics:
- ROUGE-1
- ROUGE-2
- ROUGE-L
- ROUGE-Lsum
- Python 3.8+
- CUDA-enabled GPU (optional)
pip install -r requirements.txt
python main.py
- transformers
- datasets
- torch
- evaluate
- pandas
- tqdm
- Fork the repository
- Create a feature branch
- Follow PEP8 standards
- Add tests where applicable
- Submit a pull request
License: MIT
Credits:
- Hugging Face Transformers
- SAMSum Dataset
- PEGASUS Research
Planned improvements:
- Add inference API using FastAPI
- Experiment with T5 and BART models
- Add MLflow experiment tracking
- Hyperparameter tuning
- Dockerize the pipeline
- Cloud deployment (Azure / AWS)

