I build evidence-first Data Science systems — from large-scale analytics and statistical experiments to ML reliability and grounded AI.
B.Tech CSE (AI & ML) · Uttaranchal University · 2023–2027 · CGPA 8.7/10
Dehradun, India
Portfolio · LinkedIn · LeetCode · Kaggle · Docker Hub · Email
| 15.95M | +0.7692 pp | 0.947 | 483 |
|---|---|---|---|
| complaint records processed | A/B absolute conversion uplift | strict hybrid RAGAS composite | Python3 LeetCode problems solved |
Evidence before claims · Baselines before complexity · Limitations documented, not hidden.
I like problems where the hard part is not only training a model — it is building a trustworthy path from raw evidence to a defensible decision.
QUESTION → DATA → METHOD → EVALUATE → DECIDE
My work sits across four connected areas:
- Large-scale Data Science — converting multi-million-row data into reusable analytics, risk signals and forecasts.
- Statistical Experimentation — estimating treatment effects with uncertainty, effect size and practical significance.
- ML Reliability — checking data quality, leakage, imbalance, baselines and explainability before trusting model results.
- Applied AI — combining deterministic analytics with retrieval, reranking, orchestration and grounded generation.
Large-Scale Data Science · NLP · Forecasting
Built an end-to-end intelligence platform on the CFPB Consumer Complaint dataset, turning an 8–9 GB raw CSV into reusable analytical layers, NLP routing, forecasting and decision views.
15.95M complaints · 3.57% forecast MAPE · 75.28% product classifier accuracy · 14/14 core tests
Architecture
8–9 GB CSV
↓
Chunk + validate
↓
Parquet analytical layers
↓
Analytics · NLP · Forecasting
↓
Decision views
Key decision — pre-aggregate expensive analysis and reuse validated Parquet outputs instead of loading the multi-GB raw CSV inside the application.
Boundary — complaint volume is not normalized by company customer base; forecasts and risk labels support investigation, not causal claims.
Stack — Python · Pandas · NumPy · PyArrow · Parquet · scikit-learn · TF-IDF · Logistic Regression · Prophet · Streamlit · Docker
Repository · Walkthrough · Evaluation
Statistics · Experimentation · Business Decision-Making
Analyzed a controlled advertising experiment to determine whether the advertisement treatment produced a meaningful conversion improvement over a PSA control.
588,101 users · 2.5547% vs 1.7854% conversion · +0.7692 pp uplift · +43.09% relative uplift
95% uplift interval — +0.5951 to +0.9434 percentage points
Method
Experiment data
↓
Validation + group checks
↓
Conversion estimates
↓
Two-proportion test
↓
Confidence intervals + effect sizes
↓
Simulation + logistic consistency
↓
Power planning
↓
Decision
Key principle — never present only a p-value. Report absolute uplift, relative uplift, uncertainty, effect size and practical significance together.
Boundary — ROI is not claimed because campaign cost, incremental revenue and customer lifetime value are unavailable.
Stack — Python · Pandas · SciPy · Statsmodels · Matplotlib · Logistic Regression · Statistical Inference
ML Reliability · Human-in-the-Loop · MLOps
Built a deterministic-first pre-training audit workflow that checks whether tabular data is sufficiently reliable for responsible baseline modeling.
Dataset
↓
Profiling
↓
Data Quality · Leakage · Imbalance
↓
Risk Aggregation
↓
Human Review Gate
↓
Baselines
↓
MLflow · SHAP
↓
Grounded Report + Q&A
Key design — Python owns profiling, checks, calculations, routing and model evaluation. The LLM is restricted to explanation, report generation and follow-up Q&A.
Engineering proof — human review gate · baseline comparison · MLflow tracking · SHAP evidence · FastAPI · pytest · Docker · LangGraph
Boundary — this is not AutoML, a governance certification platform or a replacement for domain review.
Stack — Python · scikit-learn · LangGraph · MLflow · SHAP · FastAPI · Streamlit · pytest · Ruff · Docker
Repository · Live App · Walkthrough
Hybrid Retrieval · Exact Analytics · Grounded AI
Built a document-intelligence system that separates exact structured analytics from semantic retrieval, instead of forcing every question through one RAG path.
|
Structured path |
Semantic path |
Strict hybrid RAGAS — 0.947 composite · 0.966 faithfulness · 1.000 context precision · 1.000 context recall
A separate 1,642-case production benchmark passed its defined factual, source, safety, conversation and HTTP checks. Final benchmark acceptance remains intentionally open because the latency gate has not yet passed.
Key decision — deterministic Pandas analytics answer structured questions; hybrid retrieval + reranking handles semantic questions.
Stack — Python · FastAPI · React · Vite · BGE · BM25 · RRF · CrossEncoder · ChromaDB · Pandas · PostgreSQL · Docker · RAGAS
Repository · Live Demo · Walkthrough · Evaluation
| Area | Tools / Concepts |
|---|---|
| Data | Python · SQL · Pandas · NumPy · PyArrow · Parquet |
| Statistics | A/B Testing · Confidence Intervals · Hypothesis Testing · SciPy · Statsmodels · Forecasting |
| Machine Learning | scikit-learn · Classification · Regression · Cross-validation · Feature Engineering · Model Evaluation · SHAP |
| NLP / Applied AI | TF-IDF · Topic Modeling · Embeddings · BM25 · Hybrid Retrieval · CrossEncoder · RAG · LangGraph |
| Engineering | FastAPI · REST APIs · React · Vite · Streamlit · Docker · MLflow · pytest · Ruff · GitHub Actions · MS SQL Server |
01 · Baseline before complexity
Start with the simplest defensible method. Add complexity only when measurement justifies it.
02 · Evaluation before claims
Use holdouts, uncertainty, baselines, failure analysis and explicit acceptance criteria.
03 · Deterministic code before LLM judgment
Calculations, validation and business rules belong in code. LLMs are useful when language understanding or explanation genuinely adds value.
04 · Boundaries are part of the result
A strong project states what was measured, what remains uncertain and what the system cannot prove.
The flagship projects above are portfolio systems. The work below is ongoing practice — kept separate so exercises are not inflated into project claims.
| Track | Current work |
|---|---|
| SQL → Pandas | joins · CTEs · windows · aggregation · groupby · merge · filtering |
| DSA | arrays · strings · hashing · stack/queue · linked lists · trees · heaps · binary search · DFS/BFS |
| Python | OOP · dunder methods · iterators · generators · decorators · debugging · clean code |
| Engineering Labs | validation · duplicate handling · idempotency · chunk processing · retries · failure isolation · storage choices |
| Ship | FastAPI · REST · SQL databases · Git · Docker · YAML · CI/CD · environment variables · MLflow |
LeetCode public language counts
483 Python3 · 75 MS SQL Server · 36 Pandas
100 Days Badge 2026
I am pursuing B.Tech Computer Science Engineering (AI & ML) at Uttaranchal University, Dehradun.
My strongest work sits where Data Science meets engineering: large-scale analytics, statistical experimentation, ML reliability, NLP/retrieval and grounded AI systems.
I care about the part after “the model works” — validation, failure modes, reproducibility, interfaces, testing and clearly explaining what the result does not prove.



