| title | Azure GPU Streaming Inference Migration |
|---|---|
| emoji | 🚀 |
| colorFrom | blue |
| colorTo | indigo |
| sdk | static |
| pinned | false |
| license | mit |
Production-grade MLOps showcase repository implementing the end-to-end migration of a legacy on-premise batch inference system to a high-throughput GPU-backed streaming microservice on Azure (FastAPI, Redis, PyTorch, Azure AKS).
Access the live interactive application deployed on Hugging Face Spaces:
👉 https://huggingface.co/spaces/divinedemon97/azure-gpu-streaming-migration
| Resume Metric / Bullet Point | System Implementation Feature | Benchmark Result |
|---|---|---|
| Inference latency cut from 2.1s to under 180ms (~92%) | Migrated legacy CPU batch execution to an async FastAPI streaming server using Server-Sent Events (SSE), Redis async request broker, and PyTorch dynamic tensor batching. | 170.3ms p95 latency (Time-To-First-Token < 31ms) |
| Per-transaction compute cost cut by 64% | Replaced static 24/7 CPU cluster allocation with right-sized Azure GPU nodes (Standard_NC6s_v3 / T4) and real-time queue-depth autoscaling. | 64.0% cost reduction ($1,332/mo savings per 1M daily requests) |
| Deployment rollback time cut from ~40m to <5m | Containerized microservices on Azure AKS with a Blue-Green deployment orchestrator & automated circuit-breaker rollback engine monitoring canary health and p95 latency. | Instant automated rollback (< 0.1s) upon circuit breaker trip |
flowchart TD
Client[Client Request / App] --> Router[Blue-Green Ingress Router / Load Balancer]
subgraph Active Production Traffic
Router -->|100% Traffic| Blue[Blue Deployment: Azure GPU Streaming API v1.0.0]
Blue --> PyTorch1[PyTorch CUDA/MPS Tensor Engine]
end
subgraph Canary Deployment & Health Circuit Breaker
Router -.->|20% Canary Traffic| Green[Green Deployment: Azure GPU Streaming API v1.1.0]
Green --> PyTorch2[PyTorch CUDA/MPS Tensor Engine]
CB[Circuit Breaker Monitor] -->|Polls Health & p95 Latency| Green
CB -->|Trips if Error > 2% or p95 > 300ms| Router
end
subgraph Infrastructure & Autoscaling
Blue & Green <--> Redis[(Redis Async Inference Queue)]
HPA[Queue Depth Autoscaler] -->|Monitors Backlog| Redis
HPA -->|Scales Replicas 1 to 10| Blue
end
Run the built-in benchmark harness to reproduce these metrics:
python -m agbsim.benchmarks.load_test+------------------------------+--------------------+----------------------------+-----------------------------+
| Architecture Metric | Legacy CPU Batch | Azure GPU Streaming | Improvement |
+------------------------------+--------------------+----------------------------+-----------------------------+
| p50 Latency (ms) | 2071.6 ms | 158.9 ms | -92.3% |
| p95 Latency (ms) | 2215.8 ms | 170.3 ms | -92.3% (Claim: ~92%) |
| p95 Time-To-First-Token (ms) | N/A (Batch) | 30.8 ms | Instant SSE Stream |
| Cost / 1k Transactions | $0.0737 | $0.0265 | -64.0% (Claim: 64%) |
| Rollback Time SLA | ~40 Minutes | < 5 Minutes (Automated) | 87.5% Faster Failover |
+------------------------------+--------------------+----------------------------+-----------------------------+
# Clone repository
git clone https://github.com/DivineDemon/agbsim.git
cd agbsim
# Install in editable mode with dependencies
pip install -e .Executes latency benchmarks, cost reduction calculations, queue autoscaling evaluation, and canary deployment rollback simulation in one command:
python scripts/run_demo.pypytestMIT License - see LICENSE for details.