I am an Agentic LLM Engineer / AI Product Engineer at Fortune Tavern Ltd (UK) focused on production multi-agent systems, RAG, LLM reliability, backend architecture, and AI-native development.
I work across the full lifecycle: business discovery, architecture, tool and data contracts, implementation, evaluation, CI/CD, observability, staged rollout, customer feedback, and production support.
- Designed and operated a 7-agent production workflow across research, generation, QA, approval, execution, and feedback
- Reduced an end-to-end process from ~5 days to under 1 day and average LLM cost per completed run by ~35%
- Built RAG evaluation over 54 labeled questions, improving paraphrase Recall@1 from 0.50 to 0.65 and MRR from 0.62 to 0.72
- Delivered AI automations for external businesses across marketing, sales, analytics, and operations
- Operated a distributed production platform for 3,000+ users, 7 nodes, and 5 client channels
- Built a controlled AgentOps workflow with isolated branches/worktrees, tests, security checks, staging evidence, and human production approval
Detailed production-oriented portfolio covering:
- multi-step agent orchestration and state transitions;
- tool/function calling and structured Pydantic contracts;
- context and memory boundaries;
- retries, fallbacks, stop conditions, and degraded states;
- evaluation, tracing, regression scenarios, and human review;
- hybrid retrieval, reranking boundaries, Recall@k, and MRR;
- the serving boundary: capability routing, fallback chains, and capacity trade-offs.
A Model Context Protocol server with a security model, and the tests that prove it. An MCP server is a remote-code-execution surface driven by a language model on the user's behalf, so the interesting engineering is not how to expose a tool but how to expose one that a confused or manipulated model cannot misuse.
- three permission tiers:
SAFEauto-allowed,GUARDEDrequires an explicit policy grant,PRIVILEGEDrequires a human approval callback and is refused when none is wired; - a path jail that resolves and then verifies containment component-wise rather than by
string prefix, opened with
O_NOFOLLOWandO_NONBLOCKso neither a symlink nor a FIFO can defeat or wedge it, with an adversarial traversal battery in the tests; - tool results wrapped as untrusted data with a per-call nonce, and an injection detector that reports rather than filters, with the reasoning written down;
- a token-bucket rate limiter with an injected clock, a concurrency cap and a per-session byte budget, so the limits are deterministic and testable;
- an append-only, hash-chained audit log that stores argument digests and never raw arguments;
- a 47-check protocol conformance suite driven over the wire: handshake ordering, version negotiation, framing, id echo, JSON-RPC error codes.
974 tests, no network and no API keys.
THREAT_MODEL.md
states the attacker, the trust boundaries, the residual risk per threat and what is
explicitly out of scope.
A durable, typed runtime for multi-step LLM agents. The premise: the hard part of an agent is not the prompt, it is state, failure and cost.
- a typed state machine over an append-only event log, so a run can be resumed after a crash and replayed deterministically against recorded model and tool results;
- discriminated-union transitions, so a step cannot forget to say what happens next;
- four independent stop conditions (step limit, budget, terminal transition, no-progress
detection), with
DEGRADEDas a first-class outcome carrying a machine-readable reason; - budget checked before every model call rather than after, across tokens, cost and calls;
- tools validated against a JSON Schema before execution, with permission levels and a human approval gate for destructive actions;
- structured output with a single repair round that feeds the validation error back;
- guardrails that deliberately do not use the model's self-reported confidence as a signal;
- a regression suite of 16 scripted failure scenarios asserting 92 named invariants.
243 tests, no network and no API keys.
DESIGN.md
states each decision as problem, chosen solution, rejected alternative and the cost of the
choice.
One request surface in front of several LLM providers. A consumer presents a gateway key and asks for a logical model with the capabilities it needs; the gateway resolves the key to upstream credentials the consumer never sees, checks scope and remaining budget, looks for an answer it can legitimately reuse, and picks a provider.
- routing by declared capability, context window, projected cost of this request and health;
- virtual keys with scopes and immediate revocation, over credentials the consumer never sees;
- request, token and budget quotas, reserved before the call and settled after it;
- fallback chains with three dispositions, and a recorded attempt for every candidate, so the response says which provider served the request and why each earlier one did not;
- a circuit breaker with a single-call half-open probe;
- a semantic cache with an exact fast path and a discriminator guard against near-duplicates;
- cost accounting in integer micro-dollars over a replaceable price table;
- 951 tests, no network, no keys and no HTTP client anywhere in the repository.
QUESTIONS.md
answers the eleven questions this code invites, including what breaks first if it is deployed.
Standalone runnable repository. python benchmark.py reproduces every number in
reports/results.md.
- Pydantic chunk and search contracts;
- BM25 implemented from scratch alongside TF-IDF, so the lexical baseline is explainable;
- reciprocal-rank fusion for hybrid search, as a testable free function;
- cross-encoder reranking behind an optional interface;
- Recall@k, MRR, retrieval failure taxonomy;
- McNemar and paired bootstrap intervals, so a delta is separated from noise;
- an annotation protocol for the golden set, with inter-annotator agreement;
- interchangeable vector stores (in-memory, pgvector, Qdrant) and model providers;
- 166 tests, no network and no keys, green in GitHub Actions on Python 3.11 and 3.12.
A shortened copy of this material also lives in this repository.
Sanitized architecture reference for a private production system used by 3,000+ users across 7 nodes and 5 client channels.
Relevant controls
Health-aware routing Automatic node exclusion Versioned API contracts Staging Smoke tests Rolling deployment Rollback Prometheus Sentry Audit logs Incident response
This background shapes how I design LLM systems: agents need the same discipline around contracts, failure isolation, permissions, observability, rollout safety, and recovery as any other distributed component.
Python FastAPI Pydantic OpenAI Responses API OpenAI Agents SDK Claude Agent SDK LangGraph Temporal MCP Tool Calling Structured Outputs JSON Schema Human-in-the-loop Context Management Bounded Retries Fallbacks Stop Conditions
Chunking Embeddings Lexical Search Vector Search Hybrid Retrieval Reranking Patterns Recall@k MRR Golden Sets Regression Scenarios Tracing Error Taxonomy Prompt/Model Versioning Langfuse Familiarity
PostgreSQL Redis SQL REST APIs Webhooks Docker Linux GitHub Actions CI/CD Prometheus Sentry Health Checks Staged Rollout Rollback Incident Response
Claude Code Codex Cursor OpenCode Git Worktrees AgentOps Automated Tests Security Review Documentation
Samara National Research University
5th-year Specialist student, Information Security of Automated Systems (10.05.03)
Expected graduation: 2027 · English: C1



