Skip to content
DivineDemonPublic

About

Azure GPU-Backed Streaming Inference Migration

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

title Azure GPU Streaming Inference Migration
emoji 🚀
colorFrom blue
colorTo indigo
sdk static
pinned false
license mit

Azure GPU-Backed Streaming Inference Migration (agbsim)

HuggingFace Space Python 3.10+ FastAPI PyTorch License: MIT

Production-grade MLOps showcase repository implementing the end-to-end migration of a legacy on-premise batch inference system to a high-throughput GPU-backed streaming microservice on Azure (FastAPI, Redis, PyTorch, Azure AKS).


🌐 Live Interactive Demo

Access the live interactive application deployed on Hugging Face Spaces:
👉 https://huggingface.co/spaces/divinedemon97/azure-gpu-streaming-migration


🎯 Resume Impact & Technical Highlights

Resume Metric / Bullet Point System Implementation Feature Benchmark Result
Inference latency cut from 2.1s to under 180ms (~92%) Migrated legacy CPU batch execution to an async FastAPI streaming server using Server-Sent Events (SSE), Redis async request broker, and PyTorch dynamic tensor batching. 170.3ms p95 latency (Time-To-First-Token < 31ms)
Per-transaction compute cost cut by 64% Replaced static 24/7 CPU cluster allocation with right-sized Azure GPU nodes (Standard_NC6s_v3 / T4) and real-time queue-depth autoscaling. 64.0% cost reduction ($1,332/mo savings per 1M daily requests)
Deployment rollback time cut from ~40m to <5m Containerized microservices on Azure AKS with a Blue-Green deployment orchestrator & automated circuit-breaker rollback engine monitoring canary health and p95 latency. Instant automated rollback (< 0.1s) upon circuit breaker trip

🏗 System Architecture

flowchart TD
    Client[Client Request / App] --> Router[Blue-Green Ingress Router / Load Balancer]
    
    subgraph Active Production Traffic
        Router -->|100% Traffic| Blue[Blue Deployment: Azure GPU Streaming API v1.0.0]
        Blue --> PyTorch1[PyTorch CUDA/MPS Tensor Engine]
    end

    subgraph Canary Deployment & Health Circuit Breaker
        Router -.->|20% Canary Traffic| Green[Green Deployment: Azure GPU Streaming API v1.1.0]
        Green --> PyTorch2[PyTorch CUDA/MPS Tensor Engine]
        CB[Circuit Breaker Monitor] -->|Polls Health & p95 Latency| Green
        CB -->|Trips if Error > 2% or p95 > 300ms| Router
    end

    subgraph Infrastructure & Autoscaling
        Blue & Green <--> Redis[(Redis Async Inference Queue)]
        HPA[Queue Depth Autoscaler] -->|Monitors Backlog| Redis
        HPA -->|Scales Replicas 1 to 10| Blue
    end
Loading

📊 Performance & Cost Benchmark Comparison

Run the built-in benchmark harness to reproduce these metrics:

python -m agbsim.benchmarks.load_test
+------------------------------+--------------------+----------------------------+-----------------------------+
| Architecture Metric          | Legacy CPU Batch   | Azure GPU Streaming        | Improvement                 |
+------------------------------+--------------------+----------------------------+-----------------------------+
| p50 Latency (ms)             | 2071.6 ms          | 158.9 ms                   | -92.3%                      |
| p95 Latency (ms)             | 2215.8 ms          | 170.3 ms                   | -92.3% (Claim: ~92%)        |
| p95 Time-To-First-Token (ms) | N/A (Batch)        | 30.8 ms                    | Instant SSE Stream          |
| Cost / 1k Transactions       | $0.0737            | $0.0265                    | -64.0% (Claim: 64%)         |
| Rollback Time SLA            | ~40 Minutes        | < 5 Minutes (Automated)    | 87.5% Faster Failover       |
+------------------------------+--------------------+----------------------------+-----------------------------+

🚀 Quick Start & Installation

1. Local Python Setup

# Clone repository
git clone https://github.com/DivineDemon/agbsim.git
cd agbsim

# Install in editable mode with dependencies
pip install -e .

2. Run Interactive Demo

Executes latency benchmarks, cost reduction calculations, queue autoscaling evaluation, and canary deployment rollback simulation in one command:

python scripts/run_demo.py

3. Run Automated Tests

pytest

📜 License

MIT License - see LICENSE for details.

About

Azure GPU-Backed Streaming Inference Migration

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages