A production-tested pattern for scaling LLM agents past the tool-overload wall. Instead of giving the agent all its tools at once, show it only the tools relevant to the current user intent, and let the LLM navigate between scopes itself by calling a
load_contexttool.
Give an LLM agent 20+ tools and accuracy starts dropping. The model hallucinates tool names, picks semantically similar wrong tools, and forgets which user journey it's in. Research is consistent: tool-selection accuracy degrades as tool count rises, and tool descriptions compete with the user's message for the model's attention budget. A 2025 paper on RAG-MCP found that in typical MCP deployments, roughly 72% of the agent's context window is consumed by tool definitions before any user work begins.
You don't fix this by making the model smarter. You fix it by managing what it attends to. That's attention scoping.
One agent. All tools in the registry. But on any given turn, the model only sees the tools relevant to the current context, and only the prompt text that applies to that scope.
A middleware layer runs before every LLM call and filters the tool list based
on an active_context field in agent state. The seven lines that do the real
work are in middleware.py:
def _scope_tools(self, request, state):
active = state.get("active_context")
if active and active in TOOL_SCOPES:
allowed = TOOL_SCOPES[active]
tools = [t for t in (request.tools or []) if t.name in allowed]
return request.override(tools=tools)
return requestTOOL_SCOPES is a dict mapping context names to sets of tool names. Three
tools are always in every scope: load_context, get_knowledge, and logout.
The hidden tools are not in the function-calling schema at all, so there is
literally nothing for the model to hallucinate against.
Every model call gets three layers of prompt text, assembled fresh:
- Capability index (always loaded, ~5 lines): a short list of everything the agent could do across every context. Tells the model the universe of possibilities without drowning it in detail.
- Always-on core (always loaded): reasoning framework, voice, boundaries, security rules. Who the agent is, independent of what it's doing.
- Switchable module (one, hot-swapped per turn): the detailed playbook for the current context. Booking gets flight codes. Translation gets language tables. Neither knows about the other.
The capability index matters more than it looks. Without it, an agent stuck
in one scope has no idea it could also help with something from another scope,
so it never calls load_context when it should. The index is a tiny attention
cost that unlocks the whole navigation system.
load_context is NOT a code-side router. It is a tool the LLM calls. When
the user in math context asks for a translation, the model decides on its
own to call load_context("translation"). The tool returns a state update
that sets active_context. On the next turn, the middleware sees the new
context, swaps the prompt module, and filters the tool list.
From the user's side: nothing happened. They asked for a translation, they got one. From your side: you wrote zero routing logic. The model navigates between rooms on its own.
| File | What it is |
|---|---|
middleware.py |
Tool-scoping middleware (30 lines, 7 of which do the real work) |
prompts.py |
Three-layer prompt assembly: capability index, core behavior, context modules |
example_agent.py |
Runnable demo with three scoped contexts (math, translation, weather) |
REPLICATE.md |
A complete prompt you can give Claude Code, Cursor, or ChatGPT to replicate this pattern in your own codebase |
requirements.txt |
langchain, langgraph, langchain-openai |
The example uses OpenRouter so you can try any
model (Claude, GPT, Llama, Gemini, etc.) with a single API key. The default
model is anthropic/claude-haiku-4-5 — cheap, fast, and good at tool calling.
pip install -r requirements.txt
export OPENROUTER_API_KEY=sk-or-...
# optional: export OPENROUTER_MODEL=openai/gpt-4o-mini
python example_agent.pyThen try these turns in sequence:
You: What is 47 times 83?
[active_context = math]
You: Translate that answer to French
[active_context = translation — model called load_context("translation")]
You: What's the weather in Paris?
[active_context = weather — model called load_context("weather")]
Watch the log output: each turn you will see the filtered tool list shrink
to only the active context's tools, plus the always-available navigation
tools (load_context, get_knowledge, logout).
Two paths:
- Copy the pattern manually: read
middleware.pyandprompts.py, adapt to your agent's tool list, and define your ownTOOL_SCOPESdict. - Let an AI do it: open
REPLICATE.md, paste the prompt into Claude Code / Cursor / ChatGPT along with your existing agent code. The AI will inventory your tools, propose a context map, and wire up the middleware.
Full writeup of why this pattern exists, the two architectures that failed first (a plain ReAct loop with 53 tools, then a distributed state graph), and the research behind it:
Scaling an AI agent to 53 tools without making it dumber
MIT. Take this pattern and ship it.
