Skip to content

Latest commit

Β 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ€– Intelligent Research Agent

A sophisticated AI-powered research agent that takes user prompts and generates comprehensive, structured research reports. Built with LangChain, LangGraph, and Hugging Face APIs.

🎯 Features

Core Capabilities

  • Intelligent Prompt Analysis: Parses user prompts to identify main topics and subtopics using Chain-of-Thought reasoning
  • Dual-Mode Research:
    • Detailed mode: Respects user-specified subtopics
    • Inference mode: Automatically identifies relevant subtopics when only a topic is provided
  • Web Research Integration: Searches the web using DuckDuckGo for current, relevant data
  • Advanced Prompting Techniques:
    • βœ… Chain-of-Thought prompting for logical reasoning
    • βœ… Few-Shot prompting for high-quality synthesis
    • βœ… Least-to-Most prompting for structured report generation
    • βœ… Role-based prompting for specialized perspectives
  • LCEL Pipeline: Fully implemented LangChain Expression Language for composable chains
  • Structured Report Generation: Produces professional, well-formatted reports
  • Report Validation: Ensures output meets quality standards (2+ A5 pages, ~2000+ words)

Report Structure

Each generated report includes:

  1. Executive Summary - High-level overview and key takeaways
  2. Introduction - Background and research objectives
  3. Detailed Analysis - In-depth coverage of each subtopic
  4. Key Findings - Most important discoveries and insights
  5. Implications & Recommendations - Practical applications and future directions
  6. Conclusion - Summary and overall significance

πŸ› οΈ Architecture

System Components

ResearchAgent (Orchestrator)
    ↓
LangGraph Workflow (5 Nodes)
    β”œβ”€ analyze_prompt_node: Parse and structure user input
    β”œβ”€ research_node: Conduct web searches
    β”œβ”€ synthesize_research_node: Transform raw data into insights
    β”œβ”€ generate_report_node: Create structured report
    └─ validate_report_node: Quality assurance
    ↓
Advanced Prompting Templates
    β”œβ”€ Prompt Analysis Template (Chain-of-Thought + Role)
    β”œβ”€ Research Synthesis Template (Few-Shot)
    └─ Report Generation Template (Least-to-Most)
    ↓
Hugging Face LLM (Mistral-7B)
    ↓
WebSearcher (DuckDuckGo)

State Management

ResearchState = {
    "user_prompt": str,        # Original user input
    "topic": str,              # Identified main topic
    "subtopics": list[str],    # All subtopics to research
    "prompt_analysis": dict,   # Structured analysis output
    "research_data": dict,     # Web search results
    "structured_report": str,  # Final output report
    "validation_result": dict  # Quality metrics
}

πŸ“‹ Prerequisites

πŸš€ Quick Start

1. Installation

# Clone or navigate to the project directory
cd "C:\Users\HP\Desktop\Om_programming\AgenticAI\Research Agent"

# Install dependencies
pip install -r requirements.txt

2. Configuration

Create a .env file in the project root:

# Copy the example
copy .env.example .env

# Edit .env and add your Hugging Face API token
# HUGGINGFACEHUB_API_TOKEN=your_token_here

Get your token at: https://huggingface.co/settings/tokens

3. Basic Usage

from agent import ResearchAgent

# Initialize agent
agent = ResearchAgent()

# Research with explicit subtopics
prompt = """
Research on Machine Learning with focus on:
- Deep Learning architectures
- Training methodologies
- Real-world applications
"""

result = agent.research(prompt)
print(result["structured_report"])

4. Run Your First Research

python -c "from agent import ResearchAgent; agent = ResearchAgent(); result = agent.research('Artificial Intelligence'); print(result['structured_report'])"

πŸ“ Usage Examples

Example 1: Detailed Topic with Subtopics

agent = ResearchAgent()
result = agent.research("""
Research Quantum Computing with focus on:
- Quantum bits and superposition
- Quantum gates and circuits
- Current quantum computers
- Applications and challenges
""")

Example 2: General Topic (Auto-Inference)

agent = ResearchAgent()
result = agent.research("What are the impacts of blockchain technology?")
# Agent automatically identifies relevant subtopics

Example 3: Access Detailed Results

result = agent.research("Your research prompt")

# Access different components
print(result["topic"])                    # Main topic
print(result["subtopics"])                # List of subtopics
print(result["prompt_analysis"])          # Analysis details
print(result["research_data"])            # Web search results
print(result["structured_report"])        # Final report
print(result["validation_result"])        # Quality metrics

🧠 Advanced Prompting Techniques Explained

1. Chain-of-Thought (CoT)

Used in prompt analysis to break down reasoning:

  • Step 1: Extract main topic
  • Step 2: Identify explicit subtopics
  • Step 3: Infer missing subtopics
  • Step 4: Validate relevance

2. Few-Shot Prompting

Applied in research synthesis with examples of good synthesis:

  • Shows how to transform raw data into coherent insights
  • Demonstrates professional writing style
  • Provides templates for structured output

3. Least-to-Most Prompting

Implemented in report generation:

  • STEP 1: Organize content into sections
  • STEP 2: Enhance with insights
  • STEP 3: Create conclusion
  • STEP 4: Apply professional formatting

4. Role-Based Prompting

System prompts assign specific expertise:

  • "You are an expert research analyst"
  • "You are an expert research synthesizer"
  • "You are a professional research report writer"

πŸ“Š Report Output

Reports are generated in text format with:

  • Length: 2,000-2,500+ words (exceeds 2 A5 pages)
  • Sections: 6 comprehensive sections
  • Formatting: Professional headers, proper spacing
  • Quality: Validated for content coverage and length

πŸ”§ Configuration Options

Edit the code to customize:

# In agent.py - WebSearcher class
max_results=5  # Number of search results per subtopic

# In HuggingFaceModelManager class
"temperature": 0.7      # Creativity level (0=deterministic, 1=random)
"top_p": 0.9           # Diversity of output
"max_new_tokens": 1000 # Maximum generation length

# Model selection
repo_id="mistralai/Mistral-7B-Instruct-v0.2"  # Can change to other HF models

🌐 Supported Models

The agent is tested with:

  • Primary: mistralai/Mistral-7B-Instruct-v0.2
  • Alternative: meta-llama/Llama-2-7b-hf
  • Alternative: tiiuae/falcon-7b

To use a different model, modify repo_id in HuggingFaceModelManager.__init__().

πŸ“ˆ Workflow Execution Flow

User Input
    ↓
[Node 1: analyze_prompt] β†’ Parse topic and subtopics
    ↓
[Node 2: research] β†’ Search web for each subtopic
    ↓
[Node 3: synthesize] β†’ Transform raw data to insights
    ↓
[Node 4: generate_report] β†’ Create structured report
    ↓
[Node 5: validate] β†’ Check quality metrics
    ↓
Output Report + Metadata

🚨 Troubleshooting

Issue: "Hugging Face API token not found"

Solution:

# Add to .env file
HUGGINGFACEHUB_API_TOKEN=hf_xxxxxxxxxx

Issue: "No search results found"

Solution:

  • Increase max_results in WebSearcher
  • Try different search terms
  • Check internet connection

Issue: Report too short

Solution:

  • Increase subtopics (more comprehensive research)
  • Adjust max_new_tokens in model configuration
  • Modify report template to require more detailed sections

Issue: Slow execution

Solution:

  • Reduce max_results in WebSearcher
  • Use a smaller/faster model
  • Reduce max_new_tokens

πŸ“¦ Dependencies

Package Purpose
langchain Core LLM framework
langchain-community Community integrations
langchain-huggingface HF model integration
langgraph Graph-based workflows
duckduckgo-search Web searching
huggingface-hub HF API access
transformers LLM tokenization
pydantic Data validation
python-dotenv Environment management
requests HTTP requests

🀝 Integration with Other Systems

The agent can be easily integrated as a library in any Python project:

# In your Python script
from agent import ResearchAgent

agent = ResearchAgent()
result = agent.research("Your research topic")

# Use the results
print(result["structured_report"])
print(result["topic"])
print(result["subtopics"])

πŸ“„ License

Open source for educational purposes.

πŸ“ File Structure

Research Agent/
β”œβ”€β”€ agent.py              # Main research agent
β”œβ”€β”€ config.py             # Configuration management  
β”œβ”€β”€ utils.py              # Utility functions
β”œβ”€β”€ requirements.txt      # Python dependencies
β”œβ”€β”€ .env.example          # Environment template
└── README.md             # This file

πŸŽ“ Educational Value

This project demonstrates:

  • LangChain's LCEL (Expression Language)
  • LangGraph for state management and workflows
  • Advanced prompting techniques for better LLM outputs
  • Integration with Hugging Face models
  • Web data aggregation and synthesis
  • Structured report generation

πŸ”¬ Future Enhancements

Possible improvements:

  • Multi-document synthesis
  • Source citation and attribution
  • PDF report generation
  • Interactive Q&A on generated reports
  • Fact-checking and verification
  • Multiple language support
  • Real-time streaming output
  • Custom report templates

πŸ’‘ Tips for Best Results

  1. Be Specific: More detailed prompts β†’ better subtopics
  2. Use Industry Terms: Technical prompts get better results
  3. Specify Focus Areas: Listing subtopics guides research
  4. Monitor API Usage: HF free tier has rate limits
  5. Review Output: Validate reports match your needs

Created: December 2025
Framework: LangChain + LangGraph
Model: Hugging Face (Mistral-7B)
Search: DuckDuckGo API

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages