Modern dark-themed Streamlit analytics platform with automated insights, 30+ interactive Plotly charts, and full Kubernetes/Minikube deployment.
| Category | Details |
|---|---|
| Data Ingestion | CSV, Excel (.xlsx/.xls), JSON, SQLite .db, Website URL |
| Auto Insights | Intelligent Finds cards, Executive Summary, skewness/kurtosis analysis |
| Data Cleaning | Missing value imputation, duplicate removal, IQR outlier detection & treatment |
| Univariate | Histograms, KDE density, box plots, violin plots, statistical summary |
| Bivariate/Multi | Scatter+regression, line plots, heatmaps, SPLOM, 3D scatter, joint plots |
| Advanced Charts | Sunburst, treemaps, choropleth maps, animated scatter (Play button), bar charts, pie/donut |
| Tech Stack | Python 3.11 Β· Streamlit Β· Pandas Β· Plotly Express Β· SciPy Β· Scikit-learn Β· Requests/BeautifulSoup (for Website URL ingestion) |
| Deployment | Docker Β· Kubernetes Β· Minikube ready |
| Health Check | /healthz JSON endpoint on port 8502 |
Install:
- Docker (Docker Desktop)
- Minikube
kubectl
Then verify they work:
docker --version
minikube version
kubectl version --client --short# Clone / copy project files to a directory
cd InsightForge/ # adjust to your local clone path
# Make script executable
chmod +x minikube-apply-all.sh
# Run! (starts Minikube, builds Docker image, applies all YAML, prints URL)
./minikube-apply-all.sh
# Windows note:
# If you're running from PowerShell without bash support, use WSL or Git Bash:
# bash ./minikube-apply-all.shThe script will:
- β Check prerequisites
- π Start Minikube cluster (2 CPUs, 4 GB RAM)
- π³ Build
insightforge:latestDocker image inside Minikube - π¦ Apply ConfigMap β PVC β Deployment β Service
- β³ Wait for pod readiness
- π Print the dashboard URL and test Streamlitβs built-in
/_stcore/health
After deployment, the app health endpoint is also available at:
http://<minikube-ip>:30502/healthz
minikube start --cpus=2 --memory=4096 --driver=docker# Point your local Docker CLI to Minikube's daemon (bash/zsh/Git Bash)
eval $(minikube docker-env)
# Build the image
docker build -t insightforge:latest .# Order matters: ConfigMap β PVC β Deployment β Service
kubectl apply -f configmap.yaml
kubectl apply -f pvc.yaml
kubectl apply -f deployment.yaml
kubectl apply -f service.yaml# Check pod status
kubectl get pods -l app=insightforge
# Wait for pod to be Ready
kubectl rollout status deployment/insightforge
# Check all resources
kubectl get all -l app=insightforge# Get the URL (Minikube handles NodePort forwarding)
minikube service insightforge --url
# Output example:
# http://192.168.49.2:30501 β Dashboard (Streamlit)
# http://192.168.49.2:30502/healthz β App healthOpen the dashboard URL in your browser.
curl http://$(minikube ip):30502/healthz
# Expected: {"status": "healthy", "app": "InsightForge"}# Build
docker build -t insightforge:latest .
# Run
docker run -p 8501:8501 -p 8502:8502 insightforge:latest
# Access
# Open in your browser: http://localhost:8501
curl http://localhost:8502/healthzRequires Python 3.11+.
# Create venv + install deps
py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements.txt
# Run
streamlit run app.py --server.port=8501 --server.address=0.0.0.0 --server.headless=true --browser.gatherUsageStats=falseHealth check:
curl http://localhost:8502/healthzInsightForge/
βββ app.py # Main Streamlit application (all analysis code)
βββ requirements.txt # Python dependencies
βββ Dockerfile # Container image definition
βββ deployment.yaml # K8s Deployment (1 replica, probes, resources)
βββ service.yaml # K8s NodePort Service (port 8501, 8502)
βββ configmap.yaml # K8s ConfigMap (app configuration)
βββ pvc.yaml # K8s PersistentVolumeClaim (mounted at /data/uploads; current version reads uploads in-memory)
βββ minikube-apply-all.sh # One-click deploy script
βββ sample_data.csv # Sample retail dataset for testing
βββ README.md # This file
- Intelligent Finds cards (auto-detected issues & insights)
- Executive Summary paragraph
- Key dataset metrics (rows, columns, missing %)
- Quick distribution overview of numeric columns
- Dataset shape, dtypes, null counts, unique values
- Interactive missing value bar chart
- Per-column imputation (mean/median/mode/ffill/drop)
- Duplicate detection and removal
- IQR outlier detection with box plot + cap/remove treatment
- Statistical summary table (mean, median, std, skew, kurtosis)
- Histogram + box marginals
- Kernel Density Estimation (KDE) with mean/median lines
- Box plot + Violin plot (side by side)
- Categorical: frequency bar chart + pie/donut chart
- Auto-generated distribution interpretation paragraph
- Correlation heatmap (Pearson, interactive hover)
- Scatter + OLS regression line (colour by category)
- Time-series line plot (auto-detected datetime columns)
- Box & Violin by category (group comparison)
- Pair plot / SPLOM (scatter plot matrix, up to 6 variables)
- Joint plot (scatter + marginal histograms)
- 3D scatter (3 numeric axes + colour)
- Multi-column histogram overlay (KDE overlay)
- Aggregated bar chart (mean/sum/count/median by category)
- Subplot grid (4-panel distribution overview)
- Sunburst + Treemap (2-level hierarchical)
- Choropleth map (if geographic columns detected)
- Animated scatter (time-based with Play button)
- Interactive hover scatter (extra column hover data)
Every chart is accompanied by a concise auto-generated insight paragraph.
Kubernetes loads environment variables from configmap.yaml (see deployment.yaml β envFrom).
| Key | Default | Description |
|---|---|---|
STREAMLIT_SERVER_PORT |
8501 |
Streamlit port inside the container |
STREAMLIT_SERVER_ADDRESS |
0.0.0.0 |
Bind address |
STREAMLIT_SERVER_HEADLESS |
true |
Run without opening a browser |
STREAMLIT_SERVER_MAX_UPLOAD_SIZE |
200 |
Max upload size (Streamlit units) |
STREAMLIT_BROWSER_GATHER_USAGE_STATS |
false |
Disable analytics |
STREAMLIT_SERVER_ENABLE_CORS |
false |
Disable CORS |
| Key | Default | Description |
|---|---|---|
APP_NAME |
InsightForge |
Display name |
APP_VERSION |
1.0.0 |
App version (informational) |
UPLOAD_DIR |
/data/uploads |
Upload directory (currently not used by the in-memory upload flow) |
HEALTH_PORT |
8502 |
Port for the custom /healthz endpoint |
MAX_ROWS_DISPLAY |
50000 |
Max rows for chart rendering |
DEFAULT_SAMPLE_SIZE |
50000 |
Sample size for large datasets |
OUTLIER_IQR_MULTIPLIER |
1.5 |
IQR fence multiplier |
CORRELATION_THRESHOLD |
0.75 |
Strong correlation threshold |
MISSING_WARN_THRESHOLD |
0.20 |
Missing % warning threshold |
kubectl scale deployment insightforge --replicas=3In pvc.yaml, replace storageClassName: standard with:
- AWS EKS:
gp3 - GKE:
premium-rwo - AKS:
managed-premium
Add an Ingress resource with your domain and TLS cert for production access:
minikube addons enable ingress
# Then create an ingress.yaml with your domainAdjust in deployment.yaml based on dataset sizes:
- Small datasets (<100k rows): 512Mi/250m
- Medium (<1M rows): 1Gi/500m
- Large (>1M rows): 4Gi/2000m
Endpoints:
GET /healthzon port8502:{"status": "healthy", "app": "InsightForge"}- Streamlit health:
GET /_stcore/healthon port8501(used by Kubernetes probes)
Kubernetes probes (in deployment.yaml) are configured against Streamlit:
- Startup probe: ~10s interval, up to ~2 min for cold start
- Liveness probe: every 30s
- Readiness probe: every 10s
To test both from outside the cluster (Minikube):
curl http://$(minikube ip):30502/healthz
curl http://$(minikube ip):30501/_stcore/health# Remove all InsightForge resources
kubectl delete -f deployment.yaml -f service.yaml -f pvc.yaml -f configmap.yaml
# Or delete everything
kubectl delete all,pvc,configmap -l app=insightforge
# Stop Minikube
minikube stop
# Delete cluster entirely
minikube deleteThe app generates synthetic sample data at runtime when you click βLoad Sample Datasetβ:
- 500 rows (seed = 42)
- Columns:
date,category,region,revenue,units_sold,discount_pct,customer_rating,profit_margin,return_rate
The repo also includes sample_data.csv; you can upload it from the sidebar like any other dataset.
MIT License β free to use, modify, and deploy.