ContainSAGES is a practical digital forensics workspace for analyzing container snapshots, detecting suspicious artifacts, and exploring forensic reports through Streamlit-based interfaces.
This README is written as an operator guide so a new contributor can understand what exists, what to run, and what output to expect.
The repository combines 3 major workflows:
- Forensic Collection and Triage
- Open Source Security Analysis (ClamAV, YARA, and companion tooling)
- Forensic Report Search and Visualization (Streamlit + AWS/OpenAI)
Measured from this workspace state:
- Full workspace files: 21,436
- Curated maintainable files (excluding heavy caches/dumps like
.git,.venv,DownloadsTemp,test-container-snapshot, andyara-rules): 65 - Curated Python source files: 12
- Curated Dockerfiles: 2
- Root-level Python entry files: 4
ContainSAGES/Analysis_ModulePython files: 5ContainSAGES/Web_ApplicationPython files: 3
Dependency profile (non-comment lines in requirements files):
requirements.txt: 23requirements-simple.txt: 19ContainSAGES/Web_Application/requirements.txt: 21ContainSAGES/Web_Application/requirements_enhanced.txt: 21ContainSAGES/Web_Application/requirements_streamlit.txt: 16
flowchart TD
A[Container Snapshot / Tar Artifact] --> B[forensic_collector.py]
B --> C[open_source_analyzer.py]
C --> D[ClamAV / YARA / File Intelligence]
D --> E[Findings and Reports CSV/PDF]
E --> F[S3 Reports Bucket optional]
F --> G[ContainSAGES Web App forensic_search_app.py]
E --> H[Root Streamlit App streamlit_app.py]
| Path | Role | Typical Use |
|---|---|---|
forensic_collector.py |
Main root forensic pipeline | Analyze snapshots and generate findings |
open_source_analyzer.py |
Open-source tooling integration layer | ClamAV/YARA/tool-based enrichment |
streamlit_app.py |
Root Streamlit search UI with API endpoint support | UI-driven report lookup and stats |
Dockerfile |
Containerized runtime for root collector | Build reproducible analysis environment |
ContainSAGES/Analysis_Module |
Moduleized analysis toolchain | Extended collector/report/timeline logic |
ContainSAGES/Web_Application |
Streamlit forensic dashboard with S3 + embeddings | Search, analytics, and interactive report review |
DownloadsTemp |
Local evidence and test captures | Not for source control |
SnapshotExtractionModule |
Snapshot extraction and supporting automation | Local extraction workflows |
Use these as the primary start points:
- Root collector:
python forensic_collector.py - Root Streamlit app:
streamlit run streamlit_app.py - Web app module:
streamlit run ContainSAGES/Web_Application/forensic_search_app.py - Dockerized collector:
docker build -t containsages-forensics .docker run --rm containsages-forensics
Mandatory:
- Python 3.10+
- pip
Optional but common:
- Docker
- AWS credentials (for S3-backed workflows)
- OpenAI API key (for embedding/search and AI-assisted analysis)
The project uses overlapping but not identical variables across modules.
| Variable | Required | Default | Used In |
|---|---|---|---|
AWS_REGION |
Recommended | ap-south-1 |
Collector and web modules |
AWS_ACCESS_KEY_ID |
For S3 access | none | Web app S3 client |
AWS_SECRET_ACCESS_KEY |
For S3 access | none | Web app S3 client |
S3_BUCKET |
Yes for S3 workflows | module-specific | Collector and web modules |
S3_PREFIX |
Optional | reports/ |
Web app report discovery |
S3_OUTPUT_PREFIX |
Optional | empty | Collector report upload namespace |
EXECUTION_ID |
Required for collector naming | derived/unknown-container fallback behavior |
Collector output naming |
TAR_S3_KEY |
Optional but common for tar-based mode | EXECUTION_ID |
Collector snapshot input selection |
OPENAI_API_KEY |
Needed for web embedding/search | none | ContainSAGES/Web_Application |
OPEN_AI_KEY |
Needed for root AI analysis path | none | Root collector |
OPENAI_MODEL |
Optional | text-embedding-ada-002 |
Web app embeddings |
OPENAI_CHAT_MODEL |
Optional | gpt-3.5-turbo |
Web app chat analysis |
OPENAI_TEMPERATURE |
Optional | 0 |
Web app response behavior |
OPENAI_MAX_TOKENS |
Optional | 1000 |
Web app response length |
VIRUS_TOTAL_KEY |
Optional | none | Open source analyzer enrichment |
API_ENDPOINT |
Optional | deployed AWS API URL in root app | streamlit_app.py |
Windows PowerShell:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txtLinux/macOS:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtFor web app module dependencies:
pip install -r ContainSAGES/Web_Application/requirements.txtdocker build -t containsages-forensics .
docker run --rm containsages-forensicsThe startup sequence invokes startup.sh, performs runtime security tool initialization, and then executes forensic_collector.py.
- Export required environment variables (
S3_BUCKET,EXECUTION_ID, optionalTAR_S3_KEY). - Run collector.
- Review generated findings and uploaded artifacts (if S3 configured).
Command:
python forensic_collector.py- Set
API_ENDPOINTif using a non-default backend. - Launch app.
- Use health/status panel and search interface.
Command:
streamlit run streamlit_app.py- Configure AWS credentials and
OPENAI_API_KEY. - Ensure S3 bucket/prefix contain forensic report CSV data.
- Launch app and index/search reports.
Command:
streamlit run ContainSAGES/Web_Application/forensic_search_app.py- Build image.
- Provide environment variables at runtime.
- Execute forensic workflow in isolated runtime.
Commands:
docker build -t containsages-forensics .
docker run --rm \
-e EXECUTION_ID=<id> \
-e S3_BUCKET=<bucket> \
-e AWS_REGION=ap-south-1 \
containsages-forensicsCommon output categories:
- Findings CSV files (for suspicious file indicators, risk categories, and signatures)
- Intermediate snapshot-derived text output
- Optional PDF-style report outputs
- Optional S3-hosted report files for downstream search
Common naming/pattern examples in this workspace:
forensic_findings_*.csvtest-container-snapshot*.csvtest-container-snapshot*.txtsnapshot-*.txt
This repository intentionally ignores local-heavy and generated artifacts so Git history stays code-focused.
Ignored classes include:
- Packet captures and raw forensic dumps (
.pcap,.pcapng,.dump,.img) - Generated analysis outputs and temporary runtime folders
- Office/PDF local research documents
- Virtual environments, caches, logs, and editor temporary files
- Local secret files (
.env*, Streamlit secrets)
Recommendation:
- Before committing, run
git status --short. - If a recurring generated artifact appears untracked, add a focused ignore rule.
Checks:
- Verify
S3_BUCKETandS3_PREFIXvalues. - Confirm AWS credentials have
s3:ListBucketands3:GetObject. - Validate report files are valid CSV under the expected prefix.
Checks:
- Set the correct key variable expected by that module (
OPENAI_API_KEYorOPEN_AI_KEY). - Confirm outbound network/API access.
- Re-check model variable values if overridden.
Checks:
- Ensure
EXECUTION_IDandS3_BUCKETare present for the selected mode. - If tar-based mode is used, verify
TAR_S3_KEYexists in bucket. - Review startup logs for tool availability warnings (ClamAV/YARA paths).
- Never commit API keys or cloud credentials.
- Use environment variables or local
.envfiles only. - Keep generated forensic evidence out of source control unless explicitly required by policy.
- Rotate credentials immediately if accidental exposure occurs.
If you are completely new to this project, do these 5 steps:
- Read sections 4, 5, and 7 in this README.
- Create a virtual environment and install dependencies.
- Run
streamlit_app.pylocally to verify app boot. - Configure S3/OpenAI variables and run
forensic_search_app.py. - Run the collector pipeline once with controlled test input and inspect output artifacts.
This sequence gives a full end-to-end understanding of both analysis and search layers.