I build production infrastructure for AI systems.
My work focuses on three areas:
-
LLM Serving & Inference Infrastructure
- vLLM, KServe, GPU inference, model serving
- Performance optimization, KV Cache, Prefix Cache, TTFT/TPOT tuning
-
Cloud Native Infrastructure
- Kubernetes, Istio, Observability, Platform Engineering
- Building and operating large-scale production platforms
-
AI Agent Infrastructure
- AI Gateway, MCP Runtime, Agent Observability
- Reliable agent systems and autonomous RCA workflows
I work primarily with:
- Go, Python
- Kubernetes, Docker, Linux
- vLLM, KServe, CUDA
- MCP, LangGraph, OpenTelemetry




