Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 9 additions & 7 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,13 +39,15 @@ scripts/validate-full.sh # validates this repo
scripts/validate-full.sh <path> # or another bundle repo
```

It builds a throwaway `uv` venv with `hatchling` + `amplifier-foundation` + `amplifier-core` +
`pyyaml`, puts it first on `PATH`, and runs `validate-bundle-repo` so its `python3` resolves to an
interpreter that has the deps → `validation_mode: full`.

**Last full run: ✅ PASS** — 10/10 bundles clean, all hygiene/structure/placement/freshness gates
green, the lone mode "error" confirmed a false positive (name collision). Only the build *dry-run*
is skipped (no `pip wheel` in the venv); the wheels build cleanly under `uv build`.
It builds a fresh `uv` venv containing the pinned public CLI/Foundation, a prebuilt Core wheel,
`hatchling`, `pyyaml`, and `pip`, then runs **that venv's CLI**. PATH alone is insufficient:
the CLI supplies its own interpreter to recipe shell steps. The venv is removed on exit.
Set `CI_VALIDATE_RECIPE` to an explicit recipe path if multiple Foundation caches exist;
`CI_VALIDATE_VENV`, if supplied, must be a new directory. No validator findings are suppressed.

Full mode includes the actual `pip wheel` build check. Record each run's mode, build result,
and findings; a successful recipe exit is not itself a validation PASS. Review the documented
mode-name false positive above without suppressing other findings.

## Testing & what "done" looks like

Expand Down
7 changes: 5 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,8 +62,11 @@ Before opening a PR that touches bundle structure, run the repo's **full** valid
scripts/validate-full.sh
```

It builds a throwaway `uv` venv with `hatchling` + `amplifier-foundation` +
`amplifier-core` so the validator runs at `validation_mode: full`. The lone
It runs a pinned public CLI from a fresh `uv` venv with Foundation, a prebuilt
Core wheel, `hatchling`, `pyyaml`, and `pip`, so the validator runs at
`validation_mode: full` without compiling Core. The venv is removed on exit.
If multiple Foundation recipes are cached, select one with `CI_VALIDATE_RECIPE`.
The lone
mode-advertising **ERROR** it reports is a **documented FALSE POSITIVE** (a name
collision — see `AGENTS.md`); **do not "fix" it** by advertising the internal mode or
deleting path/skill references.
Expand Down
16 changes: 14 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,20 @@ Two agents are included for querying session data:
The `context-intelligence-navigation` behavior and the standalone
`context-intelligence-transcript` behavior mount a user-invocable `/transcript` skill
and the `session_transcript` tool. With no arguments, `/transcript` replays the current
session's native user/assistant capture. Pass an intent to use that transcript as source
material, or pass `--session ID[,ID...]` to target other session captures. The reader
session's native user/assistant capture. Use `/transcript abcdef12` for a unique
short ID (at least eight characters), or `/transcript FULL-ID` for an exact ID.
Pass an intent to use the current transcript as source material, or pass
`--session ID[,ID...]` to target multiple captures. Prefixes are resolved across
the mounted hook's local capture store; ambiguous matches name the character
position where the candidates diverge and ask for at least that many, never
guess. Exact IDs take precedence, including over the derived `<id>_<agent>`
sub-agent captures that every session ID prefixes: a session whose capture is
still initialising is reported, never answered with one of its own sub-agents.
Use the returned full ID when paging.
Embedding hosts with a custom capture resolver own resolution in their storage.
The shared library exports `resolve_session_id(reference, candidates)` so hosts
and other tools can reuse the same exact-first, unique-prefix lookup without
reading transcripts or contacting a graph server. The reader
preserves stored message strings, paginates only between messages, and does not call the
graph or parse provider-raw payloads. Other hosts can provide their own capture resolver;
the reusable library itself takes explicit event and metadata paths. Stored captures are
Expand Down
8 changes: 8 additions & 0 deletions context_intelligence/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,11 @@
sessions_dir_for_project,
workspace_slug,
)
from context_intelligence.session_ids import (
SessionResolutionError,
is_safe_session_id,
resolve_session_id,
)

__all__ = [
"AsyncCIClient",
Expand All @@ -68,4 +73,7 @@
"discover_sessions",
"workspace_slug",
"sessions_dir_for_project",
"SessionResolutionError",
"is_safe_session_id",
"resolve_session_id",
]
87 changes: 87 additions & 0 deletions context_intelligence/session_ids.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
"""Resolve session references without knowing a host's storage layout.

Hosts supply IDs from their authorized scope. Lookup never reads transcripts,
queries a server, guesses a match, or expands that scope.
"""

from __future__ import annotations

from collections.abc import Iterable


class SessionResolutionError(ValueError):
"""A reference is invalid, absent, or ambiguous within the supplied scope."""

def __init__(self, code: str, message: str, candidates: tuple[str, ...] = ()) -> None:
super().__init__(message)
self.code = code
self.candidates = candidates


def is_safe_session_id(value: object) -> bool:
"""Accept opaque IDs, never paths, whitespace, or glob expressions."""
return (
isinstance(value, str)
and bool(value)
and value not in {".", ".."}
and not any(character.isspace() for character in value)
and not any(character in value for character in "/\\\0*?[]")
)


def resolve_session_id(reference: str, candidates: Iterable[str]) -> str:
"""Prefer an exact ID; otherwise require one unique prefix of 8+ characters.

Exact opaque IDs may be shorter than eight characters. Duplicate IDs denote
one identity here; the host must still reject ambiguous capture locations.
"""
if not is_safe_session_id(reference):
raise SessionResolutionError("invalid_session_id", "session reference must be a safe ID")
matches: set[str] = set()
for candidate in candidates:
if candidate == reference:
return candidate
if candidate.startswith(reference):
matches.add(candidate)
if len(reference) < 8:
raise SessionResolutionError(
"invalid_session_id", "use an exact session ID or a prefix of at least 8 characters"
)
if not matches:
raise SessionResolutionError("session_not_found", f"no session matches {reference!r}")
if len(matches) > 1:
shown = tuple(sorted(matches)[:5])
raise SessionResolutionError(
"ambiguous_session",
f"{reference!r} matches {len(matches)} sessions; {_disambiguation_hint(matches)} "
f"Candidates (up to 5): {', '.join(shown)}",
shown,
)
return matches.pop()


def _shared_prefix_length(matches: set[str]) -> int:
"""Count the leading characters every match has in common."""
shortest = min(matches, key=len)
for index, character in enumerate(shortest):
if any(match[index] != character for match in matches):
return index
return len(shortest)


def _disambiguation_hint(matches: set[str]) -> str:
"""Say how much more is needed, not merely that more is needed.

A flat "use a longer prefix" is unactionable when the matches share a long
run of leading characters -- a spawned sub-agent ID begins with a
zero-padded parent block, so the separating character can sit well past the
eight-character minimum. Report where the matches actually diverge.
"""
shared = _shared_prefix_length(matches)
if shared == len(min(matches, key=len)):
# One match is a prefix of another, so no longer prefix can separate them.
return "one is a prefix of another, so only an exact full ID can select it."
return (
f"all of them share their first {shared} characters, "
f"so supply at least {shared + 1} characters or an exact full ID."
)
2 changes: 1 addition & 1 deletion modules/tool-context-intelligence-query/uv.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Original file line number Diff line number Diff line change
Expand Up @@ -13,12 +13,20 @@
read_native_transcript,
render_native_transcript,
)
from context_intelligence.session_ids import (
SessionResolutionError,
is_safe_session_id,
resolve_session_id,
)

_CAPTURE_RESOLVER_CAPABILITY = "context_intelligence.capture_resolver"
_HOOK_RESOLVER_CAPABILITY = "context_intelligence.hook_config_resolver"
_MAX_SESSIONS_PER_REQUEST = 3
_MAX_TOTAL_CONTENT_CHARS = 100_000
_GLOB_METACHARACTERS = frozenset("*?[]")
# A spawned sub-agent's capture directory is named `{parent_session_id}_{agent}`,
# so every session ID is a strict prefix of every capture it spawned. Prefix
# lookup must never let a derived capture stand in for the session that names it.
_DERIVED_CAPTURE_MARK = "_"


class SessionTranscriptTool:
Expand Down Expand Up @@ -52,7 +60,10 @@ def input_schema(self) -> dict[str, Any]:
"type": "array",
"items": {"type": "string"},
"maxItems": _MAX_SESSIONS_PER_REQUEST,
"description": "Optional session IDs. Omit to retrieve the calling session.",
"description": (
"Optional full session IDs or unique prefixes (at least 8 characters). "
"Omit to retrieve the calling session. Use returned full IDs for pagination."
),
},
"after_event_line": {"type": "integer", "minimum": 0, "default": 0},
"after_event_lines": {
Expand Down Expand Up @@ -97,19 +108,43 @@ def _coerce_locator(value: Any) -> CaptureLocator:

@staticmethod
def _find_capture_metadata(base_path: Path, session_id: str) -> list[Path]:
"""Find literal session-directory matches without treating the ID as a glob."""
matches: list[Path] = []
"""Match directory names only; do not open every capture to resolve a prefix."""
candidates: dict[str, list[Path]] = {}
matched_names: set[str] = set()
for project_dir in base_path.iterdir():
sessions_dir = project_dir / "sessions"
if not sessions_dir.is_dir():
continue
for session_dir in sessions_dir.iterdir():
if session_dir.name != session_id:
if not session_dir.name.startswith(session_id):
continue
matched_names.add(session_dir.name)
metadata_path = session_dir / "context-intelligence" / "metadata.json"
if metadata_path.is_file():
matches.append(metadata_path)
return matches
candidates.setdefault(session_dir.name, []).append(metadata_path)

# An exactly-named capture directory IS the requested session, even before
# its metadata.json exists -- the logging handler writes that file lazily,
# on the first event. Reporting the gap keeps a session whose capture is
# still initialising from being answered with one of its own sub-agents.
if session_id in matched_names and session_id not in candidates:
raise NativeTranscriptError(
"capture_unavailable",
f"session {session_id!r} has a capture directory but no readable "
"metadata.json; its capture is incomplete or still initialising",
)
derived = sorted(
name for name in candidates if name.startswith(session_id + _DERIVED_CAPTURE_MARK)
)
if derived and len(derived) == len(candidates):
raise NativeTranscriptError(
"session_not_found",
f"session {session_id!r} has no capture of its own; only sub-agent "
f"sessions derived from it were captured. Candidates (up to 5): "
f"{', '.join(derived[:5])}",
)
canonical_id = resolve_session_id(session_id, candidates)
return candidates[canonical_id]

def _resolve_locator(self, session_id: str) -> CaptureLocator:
"""Resolve through an embedding host first, then the mounted CI hook."""
Expand All @@ -127,11 +162,16 @@ def _resolve_locator(self, session_id: str) -> CaptureLocator:
directory = session_dir(session_id)
if isinstance(directory, (Path, str)):
locator = CaptureLocator.from_session_dir(directory)
if locator.metadata_path.is_file() and locator.events_path.is_file():
return locator
base_path = getattr(self._hook_resolver, "base_path", None)
if isinstance(base_path, Path):
matches = self._find_capture_metadata(base_path, session_id)
if isinstance(base_path, (Path, str)):
try:
matches = self._find_capture_metadata(Path(base_path), session_id)
except SessionResolutionError as exc:
raise NativeTranscriptError(exc.code, str(exc)) from exc
except OSError as exc:
raise NativeTranscriptError(
"capture_unavailable", f"cannot search local captures: {exc}"
) from exc
if len(matches) == 1:
return CaptureLocator(
events_path=matches[0].with_name("events.jsonl"),
Expand All @@ -156,16 +196,7 @@ def _resolve_locator(self, session_id: str) -> CaptureLocator:

@staticmethod
def _is_safe_session_id(value: object) -> bool:
"""Accept opaque IDs but never path components or glob expressions."""
return (
isinstance(value, str)
and bool(value)
and value not in {".", ".."}
and "/" not in value
and "\\" not in value
and "\0" not in value
and not _GLOB_METACHARACTERS.intersection(value)
)
return is_safe_session_id(value)

def _requested_sessions(self, input_data: dict[str, Any]) -> list[tuple[str, int]]:
raw_ids = input_data.get("session_ids")
Expand Down
4 changes: 2 additions & 2 deletions modules/tool-context-intelligence-transcript/pyproject.toml
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
[project]
name = "amplifier-module-tool-context-intelligence-transcript"
version = "0.1.0"
version = "0.1.1"
description = "Bounded native Context Intelligence transcript retrieval"
requires-python = ">=3.11"
license = "MIT"

dependencies = [
"amplifier-bundle-context-intelligence @ git+https://github.com/microsoft/amplifier-bundle-context-intelligence@3e4887fdc58ac2299986f08afdd140865e9aa063",
"amplifier-bundle-context-intelligence @ git+https://github.com/microsoft/amplifier-bundle-context-intelligence@2818e0519f1c59c71d6a5b53eee50592853f8d7e",
]

[project.entry-points."amplifier.modules"]
Expand Down
Loading
Loading