Skip to content

Latest commit

 

History

144 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ContainSAGES: Container Forensics and Threat Analysis Workspace

ContainSAGES is a practical digital forensics workspace for analyzing container snapshots, detecting suspicious artifacts, and exploring forensic reports through Streamlit-based interfaces.

This README is written as an operator guide so a new contributor can understand what exists, what to run, and what output to expect.

1) Project Scope

The repository combines 3 major workflows:

  1. Forensic Collection and Triage
  2. Open Source Security Analysis (ClamAV, YARA, and companion tooling)
  3. Forensic Report Search and Visualization (Streamlit + AWS/OpenAI)

2) Current Repository Snapshot (Measured)

Measured from this workspace state:

  1. Full workspace files: 21,436
  2. Curated maintainable files (excluding heavy caches/dumps like .git, .venv, DownloadsTemp, test-container-snapshot, and yara-rules): 65
  3. Curated Python source files: 12
  4. Curated Dockerfiles: 2
  5. Root-level Python entry files: 4
  6. ContainSAGES/Analysis_Module Python files: 5
  7. ContainSAGES/Web_Application Python files: 3

Dependency profile (non-comment lines in requirements files):

  1. requirements.txt: 23
  2. requirements-simple.txt: 19
  3. ContainSAGES/Web_Application/requirements.txt: 21
  4. ContainSAGES/Web_Application/requirements_enhanced.txt: 21
  5. ContainSAGES/Web_Application/requirements_streamlit.txt: 16

3) High-Level Architecture

flowchart TD
		A[Container Snapshot / Tar Artifact] --> B[forensic_collector.py]
		B --> C[open_source_analyzer.py]
		C --> D[ClamAV / YARA / File Intelligence]
		D --> E[Findings and Reports CSV/PDF]
		E --> F[S3 Reports Bucket optional]
		F --> G[ContainSAGES Web App forensic_search_app.py]
		E --> H[Root Streamlit App streamlit_app.py]
Loading

4) Repository Layout and What Each Part Does

Path Role Typical Use
forensic_collector.py Main root forensic pipeline Analyze snapshots and generate findings
open_source_analyzer.py Open-source tooling integration layer ClamAV/YARA/tool-based enrichment
streamlit_app.py Root Streamlit search UI with API endpoint support UI-driven report lookup and stats
Dockerfile Containerized runtime for root collector Build reproducible analysis environment
ContainSAGES/Analysis_Module Moduleized analysis toolchain Extended collector/report/timeline logic
ContainSAGES/Web_Application Streamlit forensic dashboard with S3 + embeddings Search, analytics, and interactive report review
DownloadsTemp Local evidence and test captures Not for source control
SnapshotExtractionModule Snapshot extraction and supporting automation Local extraction workflows

5) Main Entry Points

Use these as the primary start points:

  1. Root collector: python forensic_collector.py
  2. Root Streamlit app: streamlit run streamlit_app.py
  3. Web app module: streamlit run ContainSAGES/Web_Application/forensic_search_app.py
  4. Dockerized collector:
    • docker build -t containsages-forensics .
    • docker run --rm containsages-forensics

6) Prerequisites

Mandatory:

  1. Python 3.10+
  2. pip

Optional but common:

  1. Docker
  2. AWS credentials (for S3-backed workflows)
  3. OpenAI API key (for embedding/search and AI-assisted analysis)

7) Environment Variables Reference

The project uses overlapping but not identical variables across modules.

Variable Required Default Used In
AWS_REGION Recommended ap-south-1 Collector and web modules
AWS_ACCESS_KEY_ID For S3 access none Web app S3 client
AWS_SECRET_ACCESS_KEY For S3 access none Web app S3 client
S3_BUCKET Yes for S3 workflows module-specific Collector and web modules
S3_PREFIX Optional reports/ Web app report discovery
S3_OUTPUT_PREFIX Optional empty Collector report upload namespace
EXECUTION_ID Required for collector naming derived/unknown-container fallback behavior Collector output naming
TAR_S3_KEY Optional but common for tar-based mode EXECUTION_ID Collector snapshot input selection
OPENAI_API_KEY Needed for web embedding/search none ContainSAGES/Web_Application
OPEN_AI_KEY Needed for root AI analysis path none Root collector
OPENAI_MODEL Optional text-embedding-ada-002 Web app embeddings
OPENAI_CHAT_MODEL Optional gpt-3.5-turbo Web app chat analysis
OPENAI_TEMPERATURE Optional 0 Web app response behavior
OPENAI_MAX_TOKENS Optional 1000 Web app response length
VIRUS_TOTAL_KEY Optional none Open source analyzer enrichment
API_ENDPOINT Optional deployed AWS API URL in root app streamlit_app.py

8) Setup Guide

8.1 Local Python Environment

Windows PowerShell:

python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt

Linux/macOS:

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

For web app module dependencies:

pip install -r ContainSAGES/Web_Application/requirements.txt

8.2 Docker Setup (Collector Path)

docker build -t containsages-forensics .
docker run --rm containsages-forensics

The startup sequence invokes startup.sh, performs runtime security tool initialization, and then executes forensic_collector.py.

9) Execution Runbooks

Runbook A: Root Collector Pipeline

  1. Export required environment variables (S3_BUCKET, EXECUTION_ID, optional TAR_S3_KEY).
  2. Run collector.
  3. Review generated findings and uploaded artifacts (if S3 configured).

Command:

python forensic_collector.py

Runbook B: Root Streamlit Search UI

  1. Set API_ENDPOINT if using a non-default backend.
  2. Launch app.
  3. Use health/status panel and search interface.

Command:

streamlit run streamlit_app.py

Runbook C: ContainSAGES Web Application

  1. Configure AWS credentials and OPENAI_API_KEY.
  2. Ensure S3 bucket/prefix contain forensic report CSV data.
  3. Launch app and index/search reports.

Command:

streamlit run ContainSAGES/Web_Application/forensic_search_app.py

Runbook D: Containerized Collector

  1. Build image.
  2. Provide environment variables at runtime.
  3. Execute forensic workflow in isolated runtime.

Commands:

docker build -t containsages-forensics .
docker run --rm \
	-e EXECUTION_ID=<id> \
	-e S3_BUCKET=<bucket> \
	-e AWS_REGION=ap-south-1 \
	containsages-forensics

10) Expected Output Artifacts

Common output categories:

  1. Findings CSV files (for suspicious file indicators, risk categories, and signatures)
  2. Intermediate snapshot-derived text output
  3. Optional PDF-style report outputs
  4. Optional S3-hosted report files for downstream search

Common naming/pattern examples in this workspace:

  1. forensic_findings_*.csv
  2. test-container-snapshot*.csv
  3. test-container-snapshot*.txt
  4. snapshot-*.txt

11) Git Ignore and Data Hygiene Policy

This repository intentionally ignores local-heavy and generated artifacts so Git history stays code-focused.

Ignored classes include:

  1. Packet captures and raw forensic dumps (.pcap, .pcapng, .dump, .img)
  2. Generated analysis outputs and temporary runtime folders
  3. Office/PDF local research documents
  4. Virtual environments, caches, logs, and editor temporary files
  5. Local secret files (.env*, Streamlit secrets)

Recommendation:

  1. Before committing, run git status --short.
  2. If a recurring generated artifact appears untracked, add a focused ignore rule.

12) Troubleshooting

Issue: Streamlit app starts but cannot find reports

Checks:

  1. Verify S3_BUCKET and S3_PREFIX values.
  2. Confirm AWS credentials have s3:ListBucket and s3:GetObject.
  3. Validate report files are valid CSV under the expected prefix.

Issue: OpenAI-related failures

Checks:

  1. Set the correct key variable expected by that module (OPENAI_API_KEY or OPEN_AI_KEY).
  2. Confirm outbound network/API access.
  3. Re-check model variable values if overridden.

Issue: Collector exits early on startup

Checks:

  1. Ensure EXECUTION_ID and S3_BUCKET are present for the selected mode.
  2. If tar-based mode is used, verify TAR_S3_KEY exists in bucket.
  3. Review startup logs for tool availability warnings (ClamAV/YARA paths).

13) Security Notes

  1. Never commit API keys or cloud credentials.
  2. Use environment variables or local .env files only.
  3. Keep generated forensic evidence out of source control unless explicitly required by policy.
  4. Rotate credentials immediately if accidental exposure occurs.

14) Quick Start for New Readers

If you are completely new to this project, do these 5 steps:

  1. Read sections 4, 5, and 7 in this README.
  2. Create a virtual environment and install dependencies.
  3. Run streamlit_app.py locally to verify app boot.
  4. Configure S3/OpenAI variables and run forensic_search_app.py.
  5. Run the collector pipeline once with controlled test input and inspect output artifacts.

This sequence gives a full end-to-end understanding of both analysis and search layers.

About

ISFCR Internship in Digital forensics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages