A sophisticated AI-powered research agent that takes user prompts and generates comprehensive, structured research reports. Built with LangChain, LangGraph, and Hugging Face APIs.
- Intelligent Prompt Analysis: Parses user prompts to identify main topics and subtopics using Chain-of-Thought reasoning
- Dual-Mode Research:
- Detailed mode: Respects user-specified subtopics
- Inference mode: Automatically identifies relevant subtopics when only a topic is provided
- Web Research Integration: Searches the web using DuckDuckGo for current, relevant data
- Advanced Prompting Techniques:
- β Chain-of-Thought prompting for logical reasoning
- β Few-Shot prompting for high-quality synthesis
- β Least-to-Most prompting for structured report generation
- β Role-based prompting for specialized perspectives
- LCEL Pipeline: Fully implemented LangChain Expression Language for composable chains
- Structured Report Generation: Produces professional, well-formatted reports
- Report Validation: Ensures output meets quality standards (2+ A5 pages, ~2000+ words)
Each generated report includes:
- Executive Summary - High-level overview and key takeaways
- Introduction - Background and research objectives
- Detailed Analysis - In-depth coverage of each subtopic
- Key Findings - Most important discoveries and insights
- Implications & Recommendations - Practical applications and future directions
- Conclusion - Summary and overall significance
ResearchAgent (Orchestrator)
β
LangGraph Workflow (5 Nodes)
ββ analyze_prompt_node: Parse and structure user input
ββ research_node: Conduct web searches
ββ synthesize_research_node: Transform raw data into insights
ββ generate_report_node: Create structured report
ββ validate_report_node: Quality assurance
β
Advanced Prompting Templates
ββ Prompt Analysis Template (Chain-of-Thought + Role)
ββ Research Synthesis Template (Few-Shot)
ββ Report Generation Template (Least-to-Most)
β
Hugging Face LLM (Mistral-7B)
β
WebSearcher (DuckDuckGo)
ResearchState = {
"user_prompt": str, # Original user input
"topic": str, # Identified main topic
"subtopics": list[str], # All subtopics to research
"prompt_analysis": dict, # Structured analysis output
"research_data": dict, # Web search results
"structured_report": str, # Final output report
"validation_result": dict # Quality metrics
}- Python 3.8+
- Hugging Face API token (free tier available at https://huggingface.co/settings/tokens)
- Internet connection (for web searches)
# Clone or navigate to the project directory
cd "C:\Users\HP\Desktop\Om_programming\AgenticAI\Research Agent"
# Install dependencies
pip install -r requirements.txtCreate a .env file in the project root:
# Copy the example
copy .env.example .env
# Edit .env and add your Hugging Face API token
# HUGGINGFACEHUB_API_TOKEN=your_token_hereGet your token at: https://huggingface.co/settings/tokens
from agent import ResearchAgent
# Initialize agent
agent = ResearchAgent()
# Research with explicit subtopics
prompt = """
Research on Machine Learning with focus on:
- Deep Learning architectures
- Training methodologies
- Real-world applications
"""
result = agent.research(prompt)
print(result["structured_report"])python -c "from agent import ResearchAgent; agent = ResearchAgent(); result = agent.research('Artificial Intelligence'); print(result['structured_report'])"agent = ResearchAgent()
result = agent.research("""
Research Quantum Computing with focus on:
- Quantum bits and superposition
- Quantum gates and circuits
- Current quantum computers
- Applications and challenges
""")agent = ResearchAgent()
result = agent.research("What are the impacts of blockchain technology?")
# Agent automatically identifies relevant subtopicsresult = agent.research("Your research prompt")
# Access different components
print(result["topic"]) # Main topic
print(result["subtopics"]) # List of subtopics
print(result["prompt_analysis"]) # Analysis details
print(result["research_data"]) # Web search results
print(result["structured_report"]) # Final report
print(result["validation_result"]) # Quality metricsUsed in prompt analysis to break down reasoning:
- Step 1: Extract main topic
- Step 2: Identify explicit subtopics
- Step 3: Infer missing subtopics
- Step 4: Validate relevance
Applied in research synthesis with examples of good synthesis:
- Shows how to transform raw data into coherent insights
- Demonstrates professional writing style
- Provides templates for structured output
Implemented in report generation:
- STEP 1: Organize content into sections
- STEP 2: Enhance with insights
- STEP 3: Create conclusion
- STEP 4: Apply professional formatting
System prompts assign specific expertise:
- "You are an expert research analyst"
- "You are an expert research synthesizer"
- "You are a professional research report writer"
Reports are generated in text format with:
- Length: 2,000-2,500+ words (exceeds 2 A5 pages)
- Sections: 6 comprehensive sections
- Formatting: Professional headers, proper spacing
- Quality: Validated for content coverage and length
Edit the code to customize:
# In agent.py - WebSearcher class
max_results=5 # Number of search results per subtopic
# In HuggingFaceModelManager class
"temperature": 0.7 # Creativity level (0=deterministic, 1=random)
"top_p": 0.9 # Diversity of output
"max_new_tokens": 1000 # Maximum generation length
# Model selection
repo_id="mistralai/Mistral-7B-Instruct-v0.2" # Can change to other HF modelsThe agent is tested with:
- Primary: mistralai/Mistral-7B-Instruct-v0.2
- Alternative: meta-llama/Llama-2-7b-hf
- Alternative: tiiuae/falcon-7b
To use a different model, modify repo_id in HuggingFaceModelManager.__init__().
User Input
β
[Node 1: analyze_prompt] β Parse topic and subtopics
β
[Node 2: research] β Search web for each subtopic
β
[Node 3: synthesize] β Transform raw data to insights
β
[Node 4: generate_report] β Create structured report
β
[Node 5: validate] β Check quality metrics
β
Output Report + Metadata
Solution:
# Add to .env file
HUGGINGFACEHUB_API_TOKEN=hf_xxxxxxxxxxSolution:
- Increase
max_resultsin WebSearcher - Try different search terms
- Check internet connection
Solution:
- Increase subtopics (more comprehensive research)
- Adjust
max_new_tokensin model configuration - Modify report template to require more detailed sections
Solution:
- Reduce
max_resultsin WebSearcher - Use a smaller/faster model
- Reduce
max_new_tokens
| Package | Purpose |
|---|---|
| langchain | Core LLM framework |
| langchain-community | Community integrations |
| langchain-huggingface | HF model integration |
| langgraph | Graph-based workflows |
| duckduckgo-search | Web searching |
| huggingface-hub | HF API access |
| transformers | LLM tokenization |
| pydantic | Data validation |
| python-dotenv | Environment management |
| requests | HTTP requests |
The agent can be easily integrated as a library in any Python project:
# In your Python script
from agent import ResearchAgent
agent = ResearchAgent()
result = agent.research("Your research topic")
# Use the results
print(result["structured_report"])
print(result["topic"])
print(result["subtopics"])Open source for educational purposes.
Research Agent/
βββ agent.py # Main research agent
βββ config.py # Configuration management
βββ utils.py # Utility functions
βββ requirements.txt # Python dependencies
βββ .env.example # Environment template
βββ README.md # This file
This project demonstrates:
- LangChain's LCEL (Expression Language)
- LangGraph for state management and workflows
- Advanced prompting techniques for better LLM outputs
- Integration with Hugging Face models
- Web data aggregation and synthesis
- Structured report generation
Possible improvements:
- Multi-document synthesis
- Source citation and attribution
- PDF report generation
- Interactive Q&A on generated reports
- Fact-checking and verification
- Multiple language support
- Real-time streaming output
- Custom report templates
- Be Specific: More detailed prompts β better subtopics
- Use Industry Terms: Technical prompts get better results
- Specify Focus Areas: Listing subtopics guides research
- Monitor API Usage: HF free tier has rate limits
- Review Output: Validate reports match your needs
Created: December 2025
Framework: LangChain + LangGraph
Model: Hugging Face (Mistral-7B)
Search: DuckDuckGo API