Skip to content

About

🎣 ML-driven phishing URL & email scoring engine β€” scikit-learn + Flask + Real-time analysis dashboard

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

PhishGuard

Email threat detection engine combining calibrated ensemble learning, header verification, and rule-based heuristics.

graph TD
    A[Incoming Email: EML / Text / IMAP] --> B[Preprocessor]
    B --> C1[1. Machine Learning Ensemble\nLogistic Regression, Random Forest, Gradient Boosting]
    B --> C2[2. Structural Heuristics\nUrgency Phrases, Typosquatting, Subdomain Spoofing]
    B --> C3[3. Header Authentication\nSPF, DKIM, DMARC, Display Name Mismatches]
    B --> C4[4. URL Risk Engine\nLookalike Domains, IP Links, Target Brands]
    
    C1 --> D[Scoring Engine]
    C2 --> D
    C3 --> D
    C4 --> D
    
    D --> E{Confidence Gate}
    E -->|Score >= 60 & Conf >= 60%| F[Phishing: Quarantine]
    E -->|Score >= 35 or Low Conf| G[Suspicious: Warning Banner]
    E -->|Score < 35| H[Safe: Delivered]
    
    F --> I[(Async SQLite)]
    G --> I
    H --> I
    I --> J[Web Dashboard & REST API]
Loading

Scoring Architecture

PhishGuard calculates a composite threat score from 0 to 100 using four independent evaluation layers:

Layer Max Points Evaluation Scope
Machine Learning Ensemble 35 Calibrated soft-voting classifier (LogisticRegression, RandomForest, GradientBoosting) trained on TF-IDF word vectors and metadata features.
Heuristics Engine 30 Urgency dictionaries, social engineering patterns, excessive capitalization, and subdomain brand spoofing (e.g. paypal.com.evil-site.xyz).
Header Authentication 15 Authentication-Results, Received-SPF, DKIM-Signature, DMARC, and sender display name vs. Reply-To mismatches.
URL Risk 20 Typosquatting via Levenshtein distance, direct IP addresses, and brand impersonation.
Whitelist Offset -25 Trusted domains receive a score reduction to minimize false positives.

Confidence Gating

To prevent false quarantines, emails scoring $\ge 60$ require at least $60%$ model confidence. Messages failing this confidence floor downgrade to Suspicious.


Repository Structure

phishing-analyzer/
β”œβ”€β”€ main.py                  FastAPI application and endpoints
β”œβ”€β”€ DESIGN.md                Warm-canvas editorial design system specification
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ analysis.py          Multi-signal scoring and confidence evaluation
β”‚   β”œβ”€β”€ heuristics.py        Rule-based heuristics, brand spoofing, and urgency detection
β”‚   β”œβ”€β”€ model.py             Calibrated machine learning ensemble
β”‚   β”œβ”€β”€ preprocessing.py     MIME email parsing, headers extraction, and URL extraction
β”‚   β”œβ”€β”€ database.py          Async SQLite database (aiosqlite) with WAL mode
β”‚   β”œβ”€β”€ mail_scanner.py      IMAP mailbox polling and quarantine automation
β”‚   └── config.py            Thresholds, weights, and whitelist definitions
β”œβ”€β”€ static/
β”‚   β”œβ”€β”€ index.html           Editorial dashboard layout
β”‚   β”œβ”€β”€ app.js               Dashboard logic, Chart.js, and API bindings
β”‚   └── style.css            Editorial CSS conforming to DESIGN.md
β”œβ”€β”€ tests/
β”‚   β”œβ”€β”€ __init__.py          Test package marker
β”‚   └── test_analyzer.py     Pytest automated test suite (14 test cases)
β”œβ”€β”€ test_scores.py           Benchmark suite for real-world email scenarios
β”œβ”€β”€ trained_model.joblib     Persisted scikit-learn model pipeline
└── requirements.txt         Python package requirements

Setup and Installation

Requirements

  • Python 3.10 or higher
  • Git

Installation Steps

git clone https://github.com/hackergovind/phishing-analyzer.git
cd phishing-analyzer

python -m venv .venv

# Windows:
.venv\Scripts\activate

# Linux / macOS:
# source .venv/bin/activate

pip install -r requirements.txt

Starting the Server

python main.py

The application starts on http://127.0.0.1:8000:

  • Web Dashboard: http://127.0.0.1:8000/
  • OpenAPI Documentation: http://127.0.0.1:8000/docs
  • ReDoc Specification: http://127.0.0.1:8000/redoc

REST API Reference

1. Analyze Plain Text

curl -X POST http://127.0.0.1:8000/analyze/text \
  -F "text=URGENT: Verify your account immediately: http://paypa1.com.verify-access.xyz/login"

Response:

{
  "status": "Phishing",
  "threat_score": 60.8,
  "confidence": 0.88,
  "action": "Quarantine email and block sender",
  "breakdown": {
    "ml_score": 28.5,
    "heuristic_score": 18.0,
    "header_auth": 0.0,
    "url_risk": 14.3,
    "whitelist_offset": 0.0
  },
  "evidence": [
    { "source": "heuristics", "detail": "Urgency phrase: 'verify your account immediately'", "contribution": 10.0 },
    { "source": "url_risk", "detail": "Target brand 'paypal' detected in non-brand domain", "contribution": 14.3 }
  ]
}

2. Analyze Raw .EML File

curl -X POST http://127.0.0.1:8000/analyze \
  -F "file=@sample.eml"

3. Server Health

curl http://127.0.0.1:8000/health

4. Aggregate Metrics

curl http://127.0.0.1:8000/api/stats

5. Audit Results Log

curl "http://127.0.0.1:8000/api/results?limit=25"

IMAP Mailbox Scanner

PhishGuard can run in the background against any standard IMAP server (e.g. Gmail, Outlook, Fastmail).

Configuration Parameters

  • Host: imap.gmail.com
  • Port: 993
  • Folder: INBOX
  • Poll Interval: 30 seconds
  • Quarantine Folder: Quarantine

Gmail Setup

  1. Enable 2-Step Verification in Google Account Settings.
  2. Generate an App Password under Security β†’ 2-Step Verification β†’ App Passwords.
  3. Use the 16-character token in the dashboard or via POST /scanner/start.

Automated Verification

Run unit and integration tests:

python -m pytest tests/ -v

Run benchmark scoring tests across known phishing samples:

python test_scores.py

About

🎣 ML-driven phishing URL & email scoring engine β€” scikit-learn + Flask + Real-time analysis dashboard

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages