Thread Intel is a content research and generation engine that turns live community discussions into Twitter/X-ready content ideas for niche creators. It collects signal from Reddit and optional secondary sources, extracts pain points and strong opinions with AI, generates posts and thread ideas, scores them, and packages the best ones into a daily posting plan.
The project includes two interfaces:
- A CLI for one-off runs with sample, file-based, interactive, or live-scraped input
- A FastAPI web app for onboarding a creator profile, running the engine, and reviewing saved results
Given a niche, an offer, and a stream of community conversations, Thread Intel:
- collects source material from Reddit and optional non-Reddit communities
- cleans and normalizes post text and comments
- extracts pain points, recurring patterns, strong opinions, and content angles with AI
- generates tweet candidates and optional visual/code assets
- scores each candidate for hook strength, clarity, depth, relatability, shareability, and monetization potential
- expands top posts into short threads
- builds a simple daily posting schedule
- saves each run as JSON and Markdown for later review
The main pipeline lives in reddit_intel/main.py.
-
Input It loads posts from a JSON file, an interactive CLI prompt, or live Reddit scraping. In web mode it can also merge in Hacker News, Dev.to, and Stack Exchange results based on the selected niche.
-
Analysis The analysis module batches posts, interleaves communities for better cross-community comparisons, and sends structured prompts to the configured AI provider. The output is an
AnalysisReportcontaining pain points, opinions, patterns, opportunity signals, and tweet-ready content angles. -
Generation The generator turns those content angles into a required mix of Growth, Authority, and Monetization posts. It can also create separate assets like code snippets or diagrams that support hybrid content.
-
Scoring Each generated post is scored on six weighted dimensions. The engine then buckets results into
POST_NOW,SCHEDULE,HIGH_RISK, andARCHIVE. -
Thread expansion The best immediate posts are expanded into 4-5 tweet threads with a fixed structure: hook, problem, insight, example, and soft CTA.
-
Scheduling and output The scheduler assigns top posts to UTC time slots, and the output writer saves timestamped Markdown and JSON files under
reddit_intel/outputs/.
The AI layer is abstracted in reddit_intel/ai_client.py. You can use either:
- Claude via
ANTHROPIC_API_KEY - OpenAI via
OPENAI_API_KEY
Provider selection is controlled by AI_PROVIDER or --provider openai.
The sourcing side works like this:
- Reddit scraping uses PRAW in reddit_intel/scraper.py
- already-seen Reddit post IDs are tracked to reduce duplicate processing
- niche profiles in niches/catalog.py define audience, persona, subreddit set, hooks, and optional secondary-source queries
- web runs can enrich Reddit input with Hacker News, Dev.to, and Stack Exchange via sources/init.py
thread-intel/
├── run.py # CLI entrypoint
├── requirements.txt # Python dependencies
├── source.json # fallback subreddit list for Reddit scraping
├── .env.example # sample environment variables
├── reddit_intel/ # core engine pipeline
│ ├── ai_client.py # Claude/OpenAI wrapper
│ ├── analysis.py # signal extraction
│ ├── generator.py # tweet and asset generation
│ ├── scoring.py # scoring and bucketing
│ ├── threads.py # thread expansion
│ ├── scheduler.py # daily posting plan
│ ├── scraper.py # Reddit ingestion
│ ├── output_writer.py # JSON/Markdown output
│ ├── renderer.py # asset rendering
│ ├── data/ # sample input and seen-post state
│ └── outputs/ # generated run artifacts
├── niches/ # niche catalog and audience profiles
├── sources/ # Hacker News / Dev.to / Stack Exchange connectors
└── web/ # FastAPI app, templates, CSS, SQLite helpers
python3 -m venv .venv
source .venv/bin/activatepip install -r requirements.txtCopy .env.example to .env and fill in the keys you need.
Minimum options:
- for AI:
ANTHROPIC_API_KEYorOPENAI_API_KEY - for live Reddit scraping:
REDDIT_CLIENT_IDandREDDIT_CLIENT_SECRET
The real .env file is intentionally ignored by git.
Use the sample input:
python run.pyRun against live Reddit data:
python run.py --scrapeUse a custom JSON file:
python run.py --input reddit_intel/data/sample_reddit.jsonOpen the interactive prompt:
python run.py --interactiveForce OpenAI and change tweet count:
python run.py --provider openai --tweets 20Start the server:
uvicorn web.app:app --reload --port 8000Then open http://127.0.0.1:8000.
Web flow:
- choose a niche during onboarding
- describe the offer and any extra creator context
- trigger a run from the dashboard
- inspect saved results and historical runs
The file-based engine expects JSON shaped like this:
{
"posts": [
{
"subreddit": "r/ExperiencedDevs",
"title": "Post title",
"body": "Post body text",
"comments": ["Top comment 1", "Top comment 2"]
}
],
"niche": "backend engineering, system design",
"offer": "system design course for senior engineers",
"extra_context": "Target audience details or positioning notes"
}reddit_intel/outputs/stores Markdown and JSON run exportsreddit_intel/outputs/images/stores rendered visual assetsweb/run_cache/stores cached web run payloadsthread_intel.dbstores web profiles and run history
- The web app uses SQLite and creates
thread_intel.dbautomatically. - The engine defaults to Claude unless you override
AI_PROVIDERor pass--provider openai. - Secondary sources are optional. Failures there do not stop the run.
- Generated output folders, caches, and the local database are ignored in git by default.