Senior Software Engineer focused on AI evaluation, RL/agent environments, and production Python systems.
I build reproducible benchmarks, verifiers, and AI-ready data platforms used to evaluate and operate LLM/agent systems.
- EvalForge — Production LLM evaluation platform (FastAPI, workers, quality gates, metrics)
- ledger-repair — RL environment for coding agents with oracle grading and soundness checks
- Basalt — Enterprise knowledge pipelines + agent ops (RAG, PII/quality, LLMOps)
Core: Python, SQL, FastAPI, Docker, PostgreSQL, Redis
AI / Eval: LLM evaluation, RL/agent environments, oracle grading, RAG, pgvector
Data: Spark, Airflow, Kafka, AWS (S3, Lambda, EC2)
- RL/agent evaluation environments and verifier soundness
- LLM evaluation pipelines and benchmark engineering
- AI-ready data platforms for retrieval and agent workflows