Skip to content

fix: restore native count request contract and safe diagnostics - #50

Merged
Brian Krabach (bkrabach) merged 2 commits into
mainfrom
lane/gcr-provider
Sep 16, 2026
Merged

Brian Krabach (bkrabach) merged 2 commits into
mainfrom
lane/gcr-provider

Conversation

@bkrabach

@bkrabach Brian Krabach (bkrabach) commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

What changed

  • Restores the required generateContentRequest.model field for native countTokens requests as models/<effective model>, derived from the same request plan as the endpoint while preserving the projected generation request.
  • Adds a best-effort provider:request_budget_unavailable diagnostic with fixed reason codes and an optional HTTP status. It never publishes request/response content, URLs, exception text, credentials, or environment values; model telemetry is restricted to documented exact model IDs.
  • Disables SDK automatic function calling for tool-less as well as tool-carrying requests before existing explicit extra_request_params overrides are applied.

Why

Native counting needs the complete documented request envelope. When counting is unavailable, callers need a safe explanation without changing the established None-or-budget decision contract or leaking request data.

Validation

  • Offline provider suite: 395 passed; 2 live-marked tests deselected.
  • Executed command: pytest -q -p no:cacheprovider -m 'not live' tests, in an isolated environment with Python 3.12.3, Core 1.6.1, google-genai 2.23.0, and httpx 0.28.1.
  • Seven coupled integration checks passed, and changed-file Ruff checks reported no diagnostics. Six privacy regression cases failed against the previous candidate and passed with the exact-model telemetry restriction.
  • The live Developer API contract control returned HTTP 400 INVALID_ARGUMENT for the old body, and HTTP 200 with 236 native tokens after adding only the required nested model.

The same reviewed provider and Loop #59 were validated together with the existing Context implementation:

Mode Native before Native after compaction Next native count Next SDK prompt count
Nonstreaming 879,659 443,422 443,643 443,945
Streaming 879,659 443,422 443,642 443,642

Each first SDK prompt count was 443,422. The nonstreaming next-request difference was 302 tokens (0.0681%), within the predeclared 1% tolerance; the other three comparisons were exact. The effective policy budget was 999,200 and its compaction target 499,600.

The final live run made 10 count requests and four generations, all HTTP 200. It verified normalized count/generation payload parity, durable next-request reduction, retained system/first-user/overlay content and final five tool pairs, and no unavailable-count fallback or AFC warning. The 64 synthetic tool results were individually below the unchanged 131,072-byte ingress limit.

Scope: gemini-3.8-flash, Developer API, SDK 2.23.0, synthetic history, declared tools with mode NONE. This is not an endurance test or live post-compaction tool-execution test. Earlier harness failures were corrected without product changes or weakening assertions; these figures describe the final successful run, not all attempts.

Independent read-only review found no ship blockers. The initial commits' statements that tests were unexecuted describe their creation time; the results above were executed afterward against the final candidate.

Compatibility

request_budget() retains its existing public result behavior. Explicit supported AFC overrides remain possible; overriding AFC can now produce the existing override warning on tool-less requests too. No dependency changes or breaking changes are intended.

…sable AFC

Three scoped corrections to the Gemini Developer API preflight counter and
request config.

1. _count_request_payload keeps its documented generateContentRequest
   envelope and now carries the REST-REQUIRED nested `model` field as
   `models/<selected model>`. The endpoint path and the nested field are both
   built from the same plan.model, so they cannot name different models.
   Eligibility still admits only bare `gemini-` ids, so no double-prefix is
   possible.

2. request_budget() still returns None when native counting is unavailable,
   but the reason is now visible through a best-effort
   provider:request_budget_unavailable event, registered on the existing
   observability.events contribution. Payload carries provider, method, a
   fixed reason code, the selected model when it is a plain bounded string,
   and an integer http_status when one exists. Reason codes:
   unsupported_model, unsupported_route, request_plan_unavailable,
   request_projection_unavailable, invalid_output_limit, http_error,
   invalid_response. _count_request_tokens returns (count, reason,
   http_status) and never emits; request_budget is the single reporting
   site, so one failed call produces exactly one report. No exception text,
   response body, request body, URL, credential or environment value enters
   the payload or the matching log line; exc_info dumps removed from both
   count paths. A failing subscriber cannot convert a successful or None
   count into a product failure (ordinary Exception caught, BaseException
   and cancellation propagate). With no hooks channel the same fixed reason
   is logged as a warning.

3. AutomaticFunctionCallingConfig(disable=True) is now set for every
   GenerateContentConfig, not only tool-carrying requests. It is applied
   before extra_request_params, so a supported explicit override still wins,
   and it remains an SDK-local field excluded from the count projection.

Tests added for both items are WRITTEN BUT UNEXECUTED: this source-only lane
is barred from running pytest. See the lane DONE.json for the exact DTU
commands and the specific assumptions that need a real run to confirm.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Publish a model identifier in unavailable-count diagnostics only when it exactly matches a documented limit-table key.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach
Brian Krabach (bkrabach) merged commit e191990 into main Sep 16, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants