This project performs topic modeling (LDA and prodLDA) and sentiment analysis (TweetNLP) on a dataset of 70,000 tweets collected from US politicians' accounts regarding the 2024 elections.
The project is organized into six main phases:
Collection of tweets from US politicians' Twitter accounts using custom scraping tools.
Text preprocessing pipeline applied to the collected tweets:
- Language detection and filtering (English only)
- URL, hashtag, emoji, and special character removal
- Text normalization (lowercase, contractions expansion)
- Punctuation and number removal
- Stop word removal (326 common words + domain-specific terms)
- Lemmatization and stemming
- Single character and short tweet removal (≤5 characters)
Final dataset: 70,000 cleaned tweets saved in doc/cleaned.csv
Implementation of Latent Dirichlet Allocation using TF-IDF representation:
- Tested configurations: 5-11 topics with varying epochs (1-350)
- Evaluation metrics: coherence scores (c_v, c_umass, c_uci, c_npmi)
- Model selection based on coherence metrics and topic interpretability
Implementation of Product of Experts LDA:
- Tested configurations: 5-10 topics, 50 epochs
- Hyperparameters: learning rate 1e-3, batch size 32
- Evaluation using same coherence metrics as LDA
Selection of the best model (LDA with 8 topics, 150 epochs) and topic assignment:
Identified Topics:
- American/economics/health
- War
- News/radio/livestream
- Republicans vs Democrats
- Border/community/family
- Election/debate
- Abortion/rights/guns
- Infrastructure/job/energy
Each tweet was assigned primary and secondary topics with associated probabilities. Results saved in doc/results.csv.
Multi-dimensional sentiment analysis using TweetNLP:
- Sentiment detection: positive, neutral, negative
- Emotion recognition: anger, optimism, anticipation, joy, disgust, fear, sadness, surprise, love, pessimism, trust
- Hate speech detection: hate, not-hate
- Offensive language identification: offensive, non-offensive
Analysis performed on:
- All tweets
- Specific political candidates
- Timeline analysis (10-day intervals from March 6, 2023 to October 22, 2023)
- Successfully identified 8 distinct topics in US political discourse
- Comprehensive sentiment and emotion profiles for each topic
- Candidate-specific sentiment analysis with temporal trends
- Interactive visualizations available in
5_processing_results/results html
Topic Distribution (Best Model: 8 topics, 150 epochs)

Sentiment Distribution by Topic

Hate Speech Detection by Topic

- Python 3.10
- Natural Language Processing: spaCy, NLTK, Gensim
- Topic Modeling: LDA (Gensim), prodLDA (Pyro)
- Sentiment Analysis: TweetNLP
- Visualization: pyLDAvis, Matplotlib
- Data Processing: Pandas, scikit-learn
doc/cleaned.csv: Preprocessed tweetsdoc/results.csv: Topic assignments for all tweets6_sentiment_analysis/result_candidates_and_timeline: Candidate and timeline-specific analyses




