Skip to content

[Feature]: On-demand Skill and MCP discovery to reduce context overhead #263

Description

@TheBaby5

Product or interface

CLI - interactive TUI

Use case and problem

As MCode grows, enabled Skills and MCP tools can become a large capability catalog that is carried into model turns even when most of it is unrelated to the current task.

MCode currently renders the available Skills catalog into the system prompt, and its runtime assembles the active tool/MCP catalog for the model. With many Skills and MCP servers, this can add unnecessary input tokens, cache writes, context pressure, model attention, and repeated processing.

The goal is to keep MCode’s normal behavior and safety model, while making capability discovery demand-driven: keep a small stable base, discover what is relevant for the current task, and load only what is actually needed..

Desired behavior

Add an optional native on-demand discovery layer for both Skills and MCP tools.

user task → lightweight capability discovery → load only relevant Skill(s)/MCP tool(s) → main model does the work

instead of:

user task → send the whole Skill/MCP catalog → main model filters everything every turn

Keep direct Skill/MCP selection authoritative, preserve essential built-in tools, keep all permissions/safety unchanged, and safely fall back to the current full catalog when discovery is uncertain or unavailable.

For Skills, load only the relevant SKILL.md instructions. For MCPs, discover the relevant server/tools and expose only the needed tool schemas, lazily connecting where practical.

The discovery layer should work without requiring Jev or another paid API; a local/deterministic matcher can be the default, with optional pluggable fast rankers.

Acceptance ideas

Large Skill/MCP setups send materially fewer irrelevant catalog/schema tokens.

Explicitly named Skills/MCP tools always remain directly usable.

Discovery failure safely falls back instead of hiding capabilities.

Permissions and tool behavior remain unchanged.

MCode shows what was discovered/loaded.

Benchmarks measure token/context reduction, latency, and cache effects.

Platform

Multiple platforms

Alternatives and additional context

The safzanpirani/pi-jev-skill-picker experiment demonstrates the general idea for Skills: remove the repeated Skills catalog and use a small ranking step to load relevant Skills only when needed. Its benchmark numbers are specific to that setup and should not be treated as MCode results.

The same idea seems even more useful if MCP discovery is included as part of the same capability-routing problem.

This also complements #262:

This issue → reduce what enters context in the first place
#262 → preserve reusable cache when runtime changes later

Together:

smaller stable context → discover only what is needed → preserve cache whenever possible

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestneeds-triageAwaiting maintainer assessmenttuiInteractive terminal UI (TUI)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions