███╗ ███╗██╗ ██╗███████╗██╗ ██╗ ██████╗ ██████╗ ██████╗
████╗ ████║██║ ██║██╔════╝██║ ██║██╔═══██╗██╔═══██╗██╔══██╗
██╔████╔██║██║ ██║███████╗███████║██║ ██║██║ ██║██║ ██║
██║╚██╔╝██║██║ ██║╚════██║██╔══██║██║ ██║██║ ██║██║ ██║
██║ ╚═╝ ██║╚██████╔╝███████║██║ ██║╚██████╔╝╚██████╔╝██████╔╝
╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚═╝ ╚═╝ ╚═════╝ ╚═════╝ ╚═════╝
Founder, Haga — the independent verification layer for physical AI
Sim-first robot policies are trained inside world models nobody independently audits. Haga stress-tests both and reports what breaks — reproducible numbers, shown failures, never a self-graded pass.
Nearly 7 years in AI engineering across two companies — Confiz and Afiniti — moving from bilingual NLP classifiers, to production RAG and fraud-detection systems, to multi-agent LLM orchestration and on-prem fine-tuning infrastructure. I'm most interested in the layer most people skip once a demo works: inference architecture, concurrency, cost per token, and failure recovery under real traffic.
As of mid-2026 I'm full-time on Haga, sole-founder and building the independent physics-verification layer for world models and robot policies — the audit layer the physical-AI stack doesn't have yet.
- 🔬 Building Haga — adversarial physics stress-testing for sim-first robotics, spanning both robot-policy behavior and generative world-model output under one methodology
- 🧠 ~7 years in AI/ML engineering: RAG, fine-tuning (LoRA/QLoRA), agentic systems (LangGraph, MCP), fraud/risk ML, MLOps
- 🎓 B.S. Software Engineering (Hons.), GIFT University, Gujranwala, Pakistan
- 📍 Based in Pakistan, working async with a global network of design partners and researchers
World models are being trained into the robots that will act in the physical world, but nobody independently checks whether their physics holds up, and policies that look robust in a flawed simulator often fail on deployment. Haga runs one adversarial methodology against both artifacts that matter — the generated world and the policy acting inside it — and ships reproducible numeric reports with defined thresholds and shown failure cases, not a self-graded pass.
Where the market is right now:
| Signal | Figure | Source |
|---|---|---|
| Capital into physical-AI/robotics startups (2025) | $27.6B across 1,009 deals — more than 2× 2024 | PitchBook / Mean CEO, May 2026 |
| AI world-models market, 2025 → 2034 | $5.8B → $28.6B (58.2% CAGR) | MarketIntelo |
| Sim benchmark success vs. real household task success | 89.4% → 12% | Stanford AI Index, 2026 |
Where Haga's own methodology already stands (shipped, not roadmap):
- Policy-stress benchmark across four Robosuite tasks (Lift, Stack, PickPlaceCan, Door), tiered mass/friction perturbations, 50 trials × 4 tasks with Wilson confidence intervals
- Physics-consistency checker v0 — 1.000 recall on physics-violation detection (teleportation, anti-gravity, causeless impulses, interpenetration) with zero false positives under tracking noise
- Real-video vs. generative cohort evaluation (CoTracker3 + Physics-IQ protocol): real footage scored 0% failure vs. a CogVideoX I2V cohort at 100% failure via a documented static-hover freeze mode (n=6–9, protocol v1)
Next phase extends the generative-model side toward NVIDIA Cosmos and larger scenario sweeps, then continuous API-based scoring for release pipelines.
| Impact | How |
|---|---|
| Cut document-processing time from 3 days to <4 hours (~92%) | 6-agent LangGraph orchestration for intake, extraction, and compliance review |
| Raised extraction accuracy 81% → 96% F1 | QLoRA-fine-tuned LLaMA-3 on a single on-prem A100 to meet a data-residency requirement |
| Cut inference latency 2.1s → <180ms (~92%), compute cost −64% | Migrated on-prem batch inference to a GPU-backed streaming service on Azure |
| Cut false-positive fraud alerts 14% → 3.5%, raised catch-rate +22% | Real-time fraud-scoring pipeline across 1.2M+ daily transactions |
| Cut agent hallucination rate 11% → <2% | Citation-grounding and self-verification step in a production agent pipeline |
| Cut integration onboarding ~2 weeks → 2 days across 9 internal APIs | Company-wide MCP-based tool-routing layer |
| Grew a pod from 2 → 5 engineers in 10 months | Owned team roadmap and led hiring |
Full detail in my résumé and on LinkedIn.
AI / ML & LLM Engineering
Backend & Infra
Cloud
Search & Data
Haga's own sim stack — MuJoCo, Robosuite, MuJoCo Menagerie, JAX (CPU-only on Apple Silicon)
Open to design partners for Haga, researchers with artifacts worth stress-testing, and conversations with anyone deep in physical AI / world models.



