-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathREADME.html
More file actions
193 lines (108 loc) · 10.9 KB
/
Copy pathREADME.html
File metadata and controls
193 lines (108 loc) · 10.9 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
<h1 id="newsclusters">NewsClusters</h1>
<h2 id="background">Background</h2>
<p>Recently my girlfriend was hired as a media strategist in the northeast region for an environmental organization. She'd never worked in the NY area before, so she faced the tough task of learning about the different reporters specific to the area. Individuals who work in communications and public relations, often need to find reporters to write about issues that are important to their organization. But how could they find these new authors without actually reading their stories? I had an idea - what about a machine learning application that would intake text, assign topics based on content, and allow her to search for unknown reporters based on topics they write about. She could then also see which other reporters write about similar topics, allowing her to find out about new reporters without actually searching through the news and reading all of their stories.</p>
<h2 id="approach">Approach</h2>
<ul>
<li>Data Collection</li>
<li>Data Cleaning</li>
<li>Topic Modeling</li>
<li>Clustering</li>
<li>Web Interface</li>
</ul>
<p><img src="plots/technologies.png" alt="alt text" title="Technology Stack" /></p>
<h2 id="datacollection">Data Collection</h2>
<p>I decided to use the last five years of New York Times articles, but only the headline and lead of each article. The data is offered for free by the NYTimes in json format using their Archive API. I set up an EC2 instance using AWS, and created a python script that connected to the API, downloaded the archives one month at a time, and then pre-cleaned and combined the months into year-long dataframes.</p>
<p>In all, I collected more than 230,000 articles written by over 13k different authors going back through 2014.</p>
<h2 id="datacleaning">Data Cleaning</h2>
<p>After gathering the data, several steps were taken to clean the data for topic analysis.</p>
<ul>
<li>Unneeded columns were removed.</li>
<li>Articles with more than one author were removed.</li>
<li>Author IDs were created using first, middle, and last name.</li>
<li>Article IDs were created using the tail of the URL.</li>
<li>Text was cleaned for punctuation and case.</li>
<li>Headline and lead were combined into one column for NLP.</li>
<li>Duplicate articles were removed.</li>
</ul>
<p>After cleaning, I ended up with 197,603 articles written by 12,701 different authors.</p>
<p>I then created a PostgreSQL database with author and article tables and imported all of the data from each article to the articles table.</p>
<h2 id="topicmodeling">Topic Modeling</h2>
<p>I decided to use LDA to perform topic modeling. I started by using Gensim's LDA package, and I performed the following preprocessing tasks:</p>
<ul>
<li>Stopwords were removed.</li>
<li>Text was stemmed and lemmatized using gensim's default settings.</li>
<li>Bigrams and trigrams were created.</li>
<li>Created word corpus.</li>
</ul>
<p>I randomly chose 8, 10, 16, and 20 topics, and I didn't get very good results. My best result was a coherence score of .23 using 10 topics. You can see from the plot below (created using pyLDAvis) that the topics are overlapping one another, and there is very little clarity in the topics - the words don't allow for any kind of logical inference.</p>
<iframe src = "plots/lda_gensim.html" width = "1250" height = "875">
Sorry your browser does not support inline frames.
<a href="plots/lda_gensim.html">Try this link.</a>
</iframe>
<p>I then created a for loop to try different numbers of topics, and recorded the coherence score for each number of topics, and I discovered that there was a significant increase in coherence score until about 34 topics, at which point the score leveled off.</p>
<p><img src="plots/coherence_10-44.png" alt="alt text" title="Coherence Scores" /></p>
<p>So I then again tried the Gensim LDA package, this time with 34 topics, but I still had poor results.</p>
<iframe src = "plots/lda_gensim2.html" width = "1250" height = "875">
Sorry your browser does not support inline frames.
<a href="plots/lda_gensim2.html">Try this link.</a>
</iframe>
<p>So I decided to try using the Mallet LDA package, which is created by UMASS and Gensim has a wrapper for it so that you can easily apply it on top of the Gensim pipeline. I also made the following tweaks to preprocesing and recreated the corpus.</p>
<ul>
<li>Text was stemmed and lemmatized, but only nouns and verbs were included. Adjectives and adverbs were ignored.</li>
<li>Extreme words were removed. Words that occurred in more than 50% of documents were ignored, and words that were in less than 15 documents total were ignored.</li>
</ul>
<p>I then ran the model using the preprocessing settings, Mallet LDA, and 34 topics, and this produced extremely clear results. The coherence score more than doubled to .54, and you can see the clarity in the topics below.</p>
<iframe src = "plots/lda_mallet2.html" width = "1250" height = "875">
Sorry your browser does not support inline frames.
<a href="plots/lda_mallet2.html">Try this link.</a>
</iframe>
<p>It was extremely easy to infer human-useable terms to represent each topic. Below are the terms I used to describe each of the 34 topics, and a sample of the words that had the most relevance to that topic:</p>
<ul>
<li><strong>Presidential Politics</strong> - trump, president, policy, obama, donald_trump, administration...<br></li>
<li><strong>Society</strong> - women, change, america, man, power, focus, girl, society, history, nation...<br></li>
<li><strong>Hope and Resilience</strong> - give, lead, hope, lose, remain, end, start, struggle...<br></li>
<li><strong>Performing Arts</strong> - review, play, theater, festival, dance, work, stage, broadway...<br></li>
<li><strong>Economy</strong> - pay, bank, money, market, price, percent, cut, raise, cost, tax...<br></li>
<li><strong>Technology</strong> - service, company, video, technology, datum, facebook, apple, internet...<br></li>
<li><strong>World News</strong> - report, news, week, test, american, accord, happen, talk...<br></li>
<li><strong>National News</strong> - rule, plan, bill, car, law, power, limit, ban, effort, measure...<br></li>
<li><strong>Court and Law</strong> - case, court, judge, charge, accuse, lawyer, claim, trial...<br></li>
<li><strong>Disaster</strong> - home, people, move, fire, town, community, california, area, build...<br></li>
<li><strong>Elections</strong> - state, party, election, vote, republican, debate, campaign, race...<br></li>
<li><strong>Conflict</strong> - fight, war, group, force, country, battle, europe, fear, aid...<br></li>
<li><strong>Business</strong> - company, deal, business, sell, buy, offer, investment, firm...<br></li>
<li><strong>Education</strong> - school, student, program, college, learn, university, give...<br></li>
<li><strong>Fashion</strong> - show, fashion, designer, line, style, brand, collection...<br></li>
<li><strong>Food</strong> - food, restaurant, open, serve, bar, chef, cook, eat, drink, recipe...<br></li>
<li><strong>Family</strong> - life, family, love, friend, mother, live, story, son, father, boy...<br></li>
<li><strong>Art</strong> - work, art, artist, museum, history, exhibition, culture, master, gallery...<br></li>
<li><strong>Transportation</strong> - day, land, air, train, minute, flight, strike, storm, crash...<br></li>
<li><strong>Announcements</strong> - year, return, leave, end, meet, join, club, announce, replace...<br></li>
<li><strong>Housing</strong> - york, city, brooklyn, street, manhattan, park, building, project, queens...<br></li>
<li><strong>Crime</strong> - man, kill, police, death, attack, people, protest, shoot, arrest, officer...<br></li>
<li><strong>Movements</strong> - day, event, draw, celebrate, year, moment, summer, washington, king, march...<br></li>
<li><strong>Government and Regulation</strong> - face, challenge, call, issue, security, official, agency...<br></li>
<li><strong>Music</strong> - music, album, award, song, record, voice, singer, band, rock, concert...<br></li>
<li><strong>Innovation</strong> - world, make, thing, watch, idea, rise, surprise, mystery, decade...<br></li>
<li><strong>Film and TV</strong> - show, film, star, series, review, tv, role, actor, character, director...<br></li>
<li><strong>Science and Research</strong> - find, child, study, age, scientist, researcher, animal, scince...<br></li>
<li><strong>Foreign Affairs</strong> - china, leader, country, trade, government, russia, britain, israel...<br></li>
<li><strong>Local Sports</strong> - run, season, yankee, met, start, game, win, series, giant, baseball...<br></li>
<li><strong>Travel</strong> - plan, hotel, offer, travel, tour, trip, water, island, road, visit, mexico...<br></li>
<li><strong>Football (and Futbol)</strong> - team, game, player, win, coach, soccer, match, nfl, world_cup...<br></li>
<li><strong>Health</strong> - health, people, risk, drug, doctor, heart, care, expert, union, hospital...<br></li>
<li><strong>Books and Writing</strong> - book, write, story, writer, talk, life, author, discuss, cover...<br></li>
</ul>
<h2 id="webinterface">Web Interface</h2>
<p>In order to search the database and see which topics are being assigned to each article and author, I created a flask application that integrates with the PostgreSQL database. <br />
I created a search function that allows you to search by keywords in the headline, or the name of the author, or by topic.</p>
<p>Here are some of the results when I searched for "dogs" in the headline. If you read the headline, you can see that the topics are accurately describing the content of each article.
<img src="plots/proof_dogs.png" alt="alt text" title="Proof Dogs" /></p>
<h2 id="clustering">Clustering</h2>
<p>I then created an authors table in the PostgreSQL database, and added a column for each topic, and entered the percentage of each author's articles for the column in which that topic was the dominant topic. I was able to then use hierarchical LDA clustering to group authors based on their distribution of topics.</p>
<p>To more easily visualize this, I clustered the 100 authors with the most articles. You can see that it is correctly grouping together authors that write about similar topics.</p>
<p><img src="plots/dendrogram_top100.png" alt="alt text" title="Dendrogram" /></p>
<p><img src="plots/dendro_zoomed.png" alt="alt text" title="Proof" /></p>
<h2 id="nextsteps">Next Steps</h2>
<p>I'd like to make this much more user-friendly, and in order to do that I need to make the flask search much smoother and more useful by including more information about each author, and by showing author-to-author search results to find authors that write about topics similar to one another.
Additionally, the code needs to be written in a way that is more modular, so that it can be more easily used for other applications, and I'd like to include news from many more major news outlets in order to make this more worthwhile.</p>