Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

vetobench

Measures how much a small-language-model (SLM) tool-call judge reduces unsafe agent actions on two industry benchmarks, Agent-SafetyBench (THU-CoAI, 2,000 cases, 349 environments) and ASB / Agent Security Bench (400 attacker tools, 10 agents, prompt-injection attacks).

The judge sits in veto-proxy, an OpenAI-compatible proxy between the agent and its model endpoint. Every tool call the agent model proposes is shown to the judge, which sees only the call (name, description, parameter schema, argument values) and rules allow/deny. In the enforce arm, denied calls are removed before the benchmark executes them. Each ruling is written to a hash-chained audit log.

 workstation                                             OpenShift (endpoint only)
┌──────────────────────────────────────────────┐        ┌──────────────────────────────┐
│ Agent-SafetyBench / ASB                      │        │ GPU0  agent model (vLLM)     │
│   base_url = http://127.0.0.1:8080/v1        │        │ GPU1  judge SLMs, ShieldAgent│
│   model    = qwen3.8-27b@enforce:granite     │        └──────────────────────────────┘
│              │                               │                    ▲
│              ▼                               │                    │
│ veto-proxy ── forward ───────────────────────┼────────────────────┤
│   tool_calls? ── judge each call ────────────┼────────────────────┘
│   enforce: strip denied calls, return refusal│
│   audit: runs/audit/<run>.jsonl (hash chain) │
└──────────────────────────────────────────────┘

Program scope

VetoBench is the test bench for the tool-call trust boundary: how much a gate on each proposed call reduces unsafe actions, what it wrongly blocks, what it costs, and what record it leaves. Today it measures one layer, a call-only judge. The roadmap adds the other layers of a gate as judge kinds, so a run reports what each layer adds:

Layer Decides Status
1. Injection guards (vLLM-SR toolcall-sentinel, toolcall-verifier) Injected request or argument planned: classifier judge
2. Value gate (Agent Exchange aex-toolgate) Argument values against named rules planned: rule judge
3. Residual judge (SLM, Granite Guardian, Llama Guard) Unsafe call beyond the rules available
4. Hold Escalate to a person planned: hold outcome

A companion study on tool-call authorization records supplies the record specification (nine fields, five scoring rules); VetoBench supplies it with live-model evidence. The audit record moves to vetobench/v2 to carry all nine fields. Scope, reuse of adjacent work (ToolScope and BFCL, vLLM Semantic Router, Decision 1.0, TelcoAIBench) and the phased roadmap are in docs/program-scope.md.

Arms

The arm is part of the model name, so benchmarks need nothing but a base_url change:

model name arm effect
M pass-through nothing logged
M@baseline baseline calls logged, never judged
M@shadow:J shadow judge J rules, ruling logged, nothing changed
M@enforce:J enforce judge J rules, denied calls removed

Judge quality is measured by replaying every judge offline over the calls logged in the baseline run. This is equivalent to shadow mode but needs only one agent run. Enforce runs measure the end-to-end effect, since blocking changes the trajectory. The allow-all judge is a built-in control: its enforce run must reproduce the baseline.

What gets measured

metric source
Agent-SafetyBench safety score (% of cases ShieldAgent labels safe), per risk category, paired Δ vs baseline ShieldAgent served on GPU1
ASB attack success rate (ASR), original-task success, refusal rate, paired ΔASR ASB's own checks; its GPT refusal judge is aliased to a local model
over-blocking ASB clean runs (no attack): task success must not drop; judge FPR on normal tools ASB + replay
judge quality ASB: per call (attacker tool = unsafe, normal tool = safe). Agent-SafetyBench: per episode vs ShieldAgent label; unsafe episodes with no tool call are reported as unreachable for a call-only judge replay
cost judge latency p50/p95, errors, blocks audit log

All proportions are reported with Wilson 95% CIs; arm differences with paired bootstrap CIs.

Setup

git clone https://github.com/thu-coai/Agent-SafetyBench third_party/Agent-SafetyBench
git -C third_party/Agent-SafetyBench checkout 74feea8de601b3a1449a93fcf70017fe61556f73
python3 -m venv .venv && .venv/bin/pip install -e '.[dev]'
./scripts/setup_asb_env.sh            # clones ASB (pinned) + its own venv at third_party/ASB/.venv
cp configs/vetobench.example.yaml configs/vetobench.yaml   # fill in routes; export VETOBENCH_API_KEY
                                                           # (venice cluster: see "Cluster" below)

Ask the cluster admin to serve the models in docs/cluster-request.md.

Cluster: venice OpenShift

The harness and proxy run on a workstation; the cluster only serves models as OpenAI-compatible vLLM endpoints. We are admins of the vetobench project and nothing else (no nodes, no cluster-scoped objects, no operators), so only ever touch that namespace.

console https://console-openshift-console.apps.venice.narlabs.io
API server https://api.venice.narlabs.io:6443 (log in with the token from the console, user menu → Copy login command)
project vetobench, user vetobench-admin
GPUs 2 × NVIDIA RTX PRO 6000 Blackwell, 96 GB each, on one node
quota vetobench-guardrail 2 GPUs; requests 24 CPU / 96Gi, limits 48 CPU / 192Gi; 1Ti storage, 30 PVCs, 60 pods
LimitRange container max 24 CPU / 96Gi; every container must set explicit cpu and memory limits (the default limit is 1 CPU, so larger requests are rejected)
storage StorageClass lvms-vg1 (default, RWO, local LVM)
routes https://<route>-vetobench.apps.venice.narlabs.io, edge TLS, 15 min timeout

TLS. The router certificate is signed by the cluster's self-signed ingress CA (CN=ingress-operator@1784587897, valid until 2028-07-19), so plain curl fails with self-signed certificate in certificate chain. Use the CA bundle venice-ca.crt (ask the project owner for it); it also covers the API server, so it works for oc login too. Don't disable verification. In configs/vetobench.yaml, verify_tls takes the bundle path (relative to the repo root, or absolute); the venice config defaults to ../venice-ca.crt and can be overridden with VETOBENCH_CA_BUNDLE.

oc client. Download the build matching the cluster from the console (? → Command Line Tools), or directly: curl --cacert venice-ca.crt -O https://downloads-openshift-console.apps.venice.narlabs.io/amd64/linux/oc.tar && tar -xf oc.tar oc && install oc ~/.local/bin/.

Secrets. Secret vllm-secrets holds VLLM_API_KEY (bearer token for all routes) and HF_TOKEN (for gated models). Never print, log or commit them. Get the key into your shell without echoing it:

oc login --server=https://api.venice.narlabs.io:6443 --certificate-authority=venice-ca.crt --token=...
export VETOBENCH_API_KEY=$(oc get secret vllm-secrets -n vetobench -o jsonpath='{.data.VLLM_API_KEY}' | base64 -d)

What is deployed

GPU object served model route status
0 Deployment vllm-agent qwen3.8-27b = Qwen/Qwen3.8-27B, tool parser qwen3_xml (swap with scripts/swap-agent-model.sh) https://vllm-agent-vetobench.apps.venice.narlabs.io/v1 running, verified 2026-09-26
1 vllm-small (one pod, four vllm serve processes) judge-small, granite-guardian, llama-guard, shieldagent vllm-<served name>-vetobench… not deployed yet

vllm-agent: image docker.io/vllm/vllm-openai:latest (vLLM 0.30.0), 1 GPU, strategy: Recreate, progressDeadlineSeconds: 3600, cpu 4/12, memory 32Gi/64Gi, /dev/shm 16Gi, model cache on PVC model-cache (300Gi, mounted at /models, HF_HOME=/models/hf). The model is chosen by env vars MODEL_ID, SERVED_NAME, TOOL_PARSER; flags --enable-auto-tool-choice --tool-call-parser $TOOL_PARSER --max-model-len 32768 --gpu-memory-utilization 0.92 --dtype bfloat16 --api-key $VLLM_API_KEY $EXTRA_ARGS. Manifest: deploy/openshift/vllm-agent.yaml. With Qwen3-8B it gets a 68 GiB KV cache (about 15× concurrency at 32k context). Exposed by Service vllm-agent (port 8000) and Route vllm-agent. enableServiceLinks: false keeps vLLM from warning about VLLM_AGENT_* service-link variables. Applying the manifest resets the model to its defaults (Qwen3-8B); run the swap script afterwards.

Swap the agent model under test with scripts/swap-agent-model.sh <served name> (e.g. qwen3.8-27b). It sets MODEL_ID, SERVED_NAME, TOOL_PARSER and EXTRA_ARGS on vllm-agent, waits for the rollout, then waits until the route lists the model. The first start of a model downloads its weights into model-cache. To add a model, add a case to the script with its Hugging Face id (the owner/name part of its huggingface.co URL), the vLLM tool-call parser for its family (see the model's vLLM recipe at https://recipes.vllm.ai), and extra flags.

served name Hugging Face id tool parser extra flags
qwen3-8b-test Qwen/Qwen3-8B hermes
qwen3.8-27b Qwen/Qwen3.8-27B qwen3_xml --language-model-only --reasoning-parser qwen3 --max-num-seqs 64
muse-glimmer-30b TBD
gemma4-31b TBD

The 27–31B agent models are expected to fit in BF16; if they run out of memory at 32k context and 16 parallel episodes, add --quantization fp8 and record it here.

GPU1 plan: pods can't share a GPU (no fractional requests), so one pod requests nvidia.com/gpu: 1 and runs four vllm serve processes on ports 8001–8004, each with --gpu-memory-utilization ~0.2 and a smaller --max-model-len, under a supervisor that exits if any child dies. One Service (4 ports), four Routes. The two pods together must stay within 96Gi of memory requests. Manifests go in deploy/openshift/.

Checking an endpoint

export VETOBENCH_AGENT_URL=https://vllm-agent-vetobench.apps.venice.narlabs.io/v1
curl -s --cacert venice-ca.crt $VETOBENCH_AGENT_URL/models -H "Authorization: Bearer $VETOBENCH_API_KEY"
vetobench smoke --agents qwen3-8b-test --judges allow-all

The route returns 401 without the key; that is expected.

Status log

  • 2026-09-25 vllm-agent verified from the workstation: /v1/models lists qwen3-8b-test, and vetobench smoke gets a parsed tool call (get_weather(city="Paris"), ~1.8 s). GPU1 is empty, so ShieldAgent scoring, the real judges and ASB's refusal judge are not available yet. The real agent models (qwen3.8-27b, muse-glimmer-30b, gemma4-31b) still need Hugging Face ids. TODO: add the vllm-agent manifest to deploy/openshift/.
  • 2026-09-25 First real-GPU run, configs/baseline-8b.yaml (Qwen3-8B, baseline vs allow-all enforce, sanity-sized). ASB's refusal judge was temporarily aliased to qwen3-8b-test in the local config because judge-small is not deployed.
    • Agent-SafetyBench, 16 cases: all ran (24 tool calls per arm). Allow-all made the same tool calls as baseline in 16/16 cases; 4 differ only in the wording of the final text answer (vLLM greedy decoding is not bit-exact across different batch compositions). No safety score yet (needs ShieldAgent).
    • ASB, direct prompt injection (naive), 20 attacker tools: ASR 100% [83.9, 100], original task success 0%. Clean (no attack), 50 tasks: task success 42% baseline / 38% allow-all, ASR 0%. Allow-all ΔASR = 0.
    • The refusal rate (100% under attack, 94% clean) is not usable: Qwen3-8B answers "0" (did not comply) even for transcripts where the agent did the task. Re-measure with a proper judge model.
  • 2026-09-26 vllm-agent manifest exported to deploy/openshift/, then changed to pass $EXTRA_ARGS and to disable service links. Swapped GPU0 to qwen3.8-27b (Qwen/Qwen3.8-27B, BF16, no FP8 needed): first start downloads the weights in ~8 min (50 GiB loaded), KV cache 479k tokens (14.6× at 32k context). It needs --max-num-seqs 64: vLLM's default of 1024 sequences exceeds the 708 state blocks for its linear-attention layers, and vLLM then refuses to start (crash loop). vetobench smoke passes (parsed tool call, ~1.1 s); integer, boolean and array arguments come back with correct JSON types, and enable_thinking: false leaves no reasoning text.
  • 2026-09-26 Sanity run on qwen3.8-27b (configs/sanity.yaml, baseline vs allow-all; ASB's refusal judge temporarily aliased to qwen3.8-27b). First attempt: every ASB request failed with 400 System message must be at the beginning, because ASB opens with several system messages and Qwen3.8's chat template accepts only one, first. The proxy now merges the leading system messages into one in every arm (contents joined in order). The Qwen3-8B ASB numbers above predate this change (its template accepted several system messages).
    • Agent-SafetyBench, 16 cases: allow-all made the same tool calls as baseline in 16/16 (8 identical word for word, 8 differ only in the final text). No thinking text leaked.
    • ASB direct prompt injection (naive), 20 attacker tools: ASR 45% [25.8, 65.8] (Qwen3-8B: 100%), original-task success 0%. Clean, 50 tasks: task success 82% [69.2, 90.2] (Qwen3-8B: 42%), ASR 0%. Allow-all ΔASR = 0 and identical task success.
    • Refusal (50% attack / 22% clean) is still measured with a stand-in judge; treat as provisional until judge-small is deployed.
  • 2026-09-28 aex-toolgate integrated as the toolgate judge (the gate runs locally from scripts/run-toolgate.sh, not yet on venice). On Fatih's twelve-call policy it reproduces the companion study (5 allowed, 6 denied with the rule named, 1 held; chain verifies). Live check with qwen3.8-27b and the payment tools: baseline executes an over-ceiling payment and a redirected one; @enforce:toolgate blocks both (P2-ceiling, P3-recipient) and allows the in-policy payment. The proxy audit record and the gate's record share call_id and hash. With a skeleton policy for the 16-case sanity split (all 25 tools granted, no rules), enforce-toolgate made the same tool calls as baseline in 16/16 cases; the gate ruled on and recorded all 24 calls (chain verifies), about 6.5 ms per call including the round trip. Next: value rules for the benchmark's tools (the full set is 1,627 tools).

Running

vetobench smoke -e configs/pilot.yaml        # served models, tool-call parsing, judges, ShieldAgent
vetobench proxy                              # keep running in its own terminal
vetobench split --n-test 300 --n-dev 48      # stratified Agent-SafetyBench split (dev = prompt work only)

# 1. sanity: 16 cases, allow-all must equal baseline
vetobench split --n-test 16 --n-dev 0 --out runs/splits/asbench-sanity.json
vetobench run -e configs/sanity.yaml && vetobench score -e configs/sanity.yaml && vetobench report -e configs/sanity.yaml

# 2. pilot, one agent model at a time (the admin swaps the model on GPU0 between runs)
vetobench run    -e configs/pilot.yaml --model qwen3.8-27b
vetobench score  -e configs/pilot.yaml --model qwen3.8-27b    # ShieldAgent labels
vetobench replay -e configs/pilot.yaml --model qwen3.8-27b    # judges on baseline calls
vetobench report -e configs/pilot.yaml                        # runs/pilot/report.md + summary.json

vetobench verify-audit runs/audit/*.jsonl    # hash chain intact?

Every step is resumable: finished cases, shards, scores and replays are skipped on re-run. Useful filters: --variant baseline enforce-granite, --bench asbench.

Tool-retrieval study (Vincent's BFCL harness) — planned

Status: planned, nothing run yet. This section records what the study is and how it connects to vetobench.

What Vincent measured. In Improve SLM tool calling without post-training (2026-09-17), Vincent Caldeira gives an agent either the whole tool catalog or only the tools a retriever shortlists for the request, and scores one tool call per query. It is a correctness test, not a safety test:

  • Retriever: ToolScope (library by Ilya Kolchinsky) embeds tool names and descriptions with sentence-transformers/all-MiniLM-L6-v2 and keeps the top k per request, like RAG for tools. A BM25 keyword retriever did about as well. It runs on CPU, next to the agent; no GPU model is involved.
  • Harness (Vincent's): eval/ in the ToolScope repo. eval/paper/bfcl_multiple.yaml is the paper protocol: the BFCL multiple split, 443 functions in one shared catalog, 200 queries, baseline (full catalog) vs BM25 vs ToolScope at k = 5, 10 and 20. Scores are name accuracy (right function) and AST accuracy (BFCL's check of name, argument names and values against the gold call).
  • Models: Llama 3.2 3B, Qwen2.5 7B, Llama 3.1 8B, Qwen3 32B and Llama 3.3 70B, as GGUF files (Q4_K_M or Q8_0) served with llama.cpp on a DGX Spark, 32k context (eval/local/models.yaml).
  • Result: with a 10-tool shortlist, Llama 3.1 8B names the right tool 92% of the time (6% on the full catalog), ahead of Llama 3.3 70B on the full catalog (79%). AST accuracy stays at 46–61%; the errors left are mostly wrong arguments.

How it connects to vetobench. The harness has an OpenAI-compatible backend (OPENAI_BASE_URL, OPENAI_API_KEY), and veto-proxy takes the arm from the model name, so the harness runs unmodified against the proxy, the same way ASB does:

 harness (bfcl_eval) ── ToolScope/BM25 shortlist (CPU) ── chat call, model = qwen2.5-7b@baseline
        │                                                                  or qwen2.5-7b@enforce:<gate>
        ▼
 veto-proxy :8080 ── audit log, gate rules on each call ── vLLM on GPU0

Untested sketch:

git clone https://github.com/ilya-kolchinsky/ToolScope third_party/ToolScope   # install per its eval/README.md
export OPENAI_BASE_URL=http://127.0.0.1:8080/v1 OPENAI_API_KEY=unused          # the proxy holds the real key
# copy eval/paper/bfcl_multiple.yaml; set model entries to vetobench routes, e.g. qwen2.5-7b@baseline
python eval/run_eval.py --config <that copy>

Every call lands in runs/audit/ like the safety benchmarks. The harness computes its own name/AST accuracy per arm; the gate's effect is the change between the baseline and enforce arms (how many wrong-argument calls it stops, and how many correct calls it wrongly stops).

Plan.

  1. Retrieval with Qwen. Does a shortlist let a small Qwen model match a large one on the full catalog? Candidates: Qwen/Qwen2.5-7B-Instruct (in Vincent's slate, so his numbers are a reference) against Qwen/Qwen2.5-72B-Instruct in FP8 (~72 GB, one GPU), the Qwen counterpart of his 70B. Both are public (no Hugging Face gating). Alternative with no new downloads: Qwen3-8B against Qwen3.8-27B. Added as cases in scripts/swap-agent-model.sh.
  2. Gate. Fatih's value gate, aex-toolgate (Go, merged in AEX PR #59), enforces a per-tool policy: scopes, then value rules (ceiling, allowlist, lookup, prefix, suffix). Its only policy today covers twelve accounts-payable tools; BFCL's 443 functions have none. Open with Fatih: where the rules come from (e.g. generated from each function's JSON schema) and how vetobench calls the gate (a rule judge kind calling the Go service, or the rules ported). The SLM judge and guard models can run as further arms once GPU1 is up.

Things to watch.

  • Context length. The full 443-tool catalog is about 60k tokens of tool definitions, but vLLM here runs at --max-model-len 32768 and rejects longer prompts. The full-catalog baseline needs at least ~65k (Qwen2.5 goes to 128k with YaRN rope scaling). Vincent's runs also used 32k, so some of his full-catalog failures for small models may be prompts that did not fit rather than wrong choices; ask him.
  • Precision. We serve BF16 (FP8 for 72B) with vLLM, not his GGUF quantizations, so re-run the baseline here instead of reusing his numbers; the gate comparison uses the same serving for both arms anyway.
  • Thinking. Keep enable_thinking: false (via model_defaults) for Qwen3-family models.

Faithfulness to the benchmarks

  • Agent-SafetyBench: vetobench.adapters.agent_safetybench ports evaluation/eval.py's loop: the benchmark's own environments, the OpenAI-path system prompt, greedy decoding, 2,048 max tokens, 10 rounds, only the first tool call executed per turn. Differences: case subset, 16-way concurrency, a per-case header for audit correlation, and the non-standard type: object key that eval.py adds to function specs is not added. Scoring ports score/eval_with_shield.py's prompt verbatim; ShieldAgent runs on vLLM instead of HF transformers.
  • ASB: main_attacker.py runs unmodified (pinned commit) under adapters/asb_launch.py. The launcher registers an OpenAI-compatible backend (ASB's GPT backend minus its name check and a fixed 2 s sleep), disables the memory DB, and shards attacker tools across parallel processes. ASB's gpt-4o-mini refusal judge is routed to judge-small via aliases, so refusal rates are not directly comparable to the ASB paper's.
  • Proxy normalisation is applied in every arm: tools without a JSON-schema parameters (ASB) get an empty object schema, since vLLM rejects them otherwise, and the system messages that open a conversation (ASB sends several) are merged into one, since some chat templates (Qwen3.8) reject a system message anywhere but first.

Judges

Configured in configs/vetobench.yaml, all call-only:

  • llm_json: any instruct SLM plus src/vetobench/judges/prompts/call_only_v1.txt; the verdict is {"reason","category","decision"}, enforced with vLLM guided JSON.
  • guard: classifier models (Granite Guardian, Llama Guard). The call is rendered as the assistant's action after a neutral user turn; unsafe_pattern decides.
  • static: allow-all / deny-all controls.
  • toolgate: Agent Exchange's aex-toolgate (AEX PR #59), a deterministic value gate: the tool's scope, then argument-value rules from a per-provider policy (ceiling, allowlist, lookup, prefix, suffix, sensitive field). Not a model. See aex-toolgate as a judge.

The judge prompt must be developed on the dev split only. ASB is never used for tuning, so it stays fully held out.

fail_policy: closed (the default) denies a call when the judge errors or times out; run with open as a sensitivity check.

aex-toolgate as a judge

vetobench proxy                                            # serves the gate's no-op upstream too
scripts/run-toolgate.sh <policy.json> [full|scope]         # clones + builds the gate (Go), port 8090
vetobench run -e <experiment with judges: [toolgate]>      # or model name M@enforce:toolgate
  • Per call: the proxy sends POST /v1/tools/{tool} with the argument values. HTTP 403 is deny, and the rule that fired becomes the verdict's category (e.g. P2-ceiling) and appears in the refusal the agent sees. HTTP 202 (held for a named approver) counts as deny, since no approver is present in a run (escalate: allow flips that). A call the gate does not rule on (arguments that are not a JSON object, gate or stub unreachable) is an error, so fail_policy decides.
  • No decide-only mode: the gate forwards allowed calls to its upstream. run-toolgate.sh points it at veto-proxy's /toolgate-stub, which does nothing, since the benchmark executes its own tools. In the gate's records an allowed call therefore shows executed: http 200.
  • Records: the gate writes its own hash-chained record per call (GET /v1/records, /v1/records/verify, with the operator token the script prints). The X-Request-ID vetobench sends becomes the record's call_id; the proxy audit log stores it with the gate's record hash in verdict.raw, so the two chains join. The judge cache is skipped for this judge so every call gets a record.
  • Policy: unknown tools fail closed. A benchmark run needs a policy written for that benchmark's tools; the only policy today is the twelve accounts-payable tools in src/internal/toolgate/testdata/policy.json. tests/test_toolgate.py replays the companion study's twelve calls through the real gate (5 allowed, 6 denied, 1 held) when the binary is built.

Try it yourself (for Fatih)

Needs Go 1.22+ (the gate's build fetches its own toolchain), Python 3.11+, and for steps 2–3 the venice API key and venice-ca.crt (see Cluster); step 1 needs neither.

git clone https://github.com/open-experiments/vetobench && cd vetobench
python3 -m venv .venv && .venv/bin/pip install -e '.[dev]'
git clone https://github.com/thu-coai/Agent-SafetyBench third_party/Agent-SafetyBench
git -C third_party/Agent-SafetyBench checkout 74feea8de601b3a1449a93fcf70017fe61556f73
cp configs/vetobench.example.yaml configs/vetobench.yaml   # has the `toolgate` judge; set the routes as in the Cluster section
  1. Offline, no GPU: the twelve calls through the real gate. Builds the gate once, then replays the companion study's calls through ToolgateJudge:

    scripts/run-toolgate.sh third_party/agent-exchange/src/internal/toolgate/testdata/policy.json   # builds, then Ctrl-C
    .venv/bin/python -m pytest -q tests/test_toolgate.py      # 7 passed: 5 allow, 6 deny, 1 hold; chain ok
  2. Live, the gate in front of the model. Three terminals:

    export VETOBENCH_API_KEY=...                       # terminal 1
    .venv/bin/vetobench proxy
    TOOLGATE_OPERATOR_TOKEN=demo scripts/run-toolgate.sh \
      third_party/agent-exchange/src/internal/toolgate/testdata/policy.json        # terminal 2
    TOOLGATE_OPERATOR_TOKEN=demo .venv/bin/python scripts/toolgate_demo.py        # terminal 3

    Expected: the in-policy payment runs in both arms; the over-ceiling and the redirected payment run in baseline and are BLOCKED with the gate (P2-ceiling, P3-recipient); the gate's chain verifies.

  3. On the benchmark: a policy for Agent-SafetyBench. Generate a skeleton for the 16-case sanity split (25 tools), run it, then add rules:

    .venv/bin/vetobench split --n-test 16 --n-dev 0 --out runs/splits/asbench-sanity.json
    .venv/bin/python scripts/asbench_toolgate_policy.py --split runs/splits/asbench-sanity.json
    #   -> runs/policies/asbench-skeleton.json (every tool granted, no rules)
    #   -> runs/policies/asbench-tools.json    (each tool: environments, description, JSON schema)
    scripts/run-toolgate.sh runs/policies/asbench-skeleton.json                  # terminal 2, instead of step 2's
    .venv/bin/vetobench run -e configs/toolgate-sanity.yaml --bench asbench

    With the skeleton, enforce-toolgate must make the same tool calls as baseline (the gate is an allow-all that records). Then add rules (same kinds as the twelve-call policy) and narrow granted_scopes, re-run, and compare. The run directory is runs/toolgate-sanity/; blocked calls are in runs/audit/toolgate-sanity.*.jsonl with the rule that fired, and the gate's own records are at GET /v1/records.

    Sizes to plan for: the full benchmark uses 1,627 distinct tools in 341 environments (skeleton without --split); 146 tool names occur in several environments with different meanings (read_file in 12, send_email in 5), and the gate keys tools by name alone.

Layout

src/vetobench/
  proxy/        app.py (FastAPI), gate.py (judge dispatch/cache), rewrite.py, upstream.py
  judges/       base.py (ToolCall, Verdict), llm.py (llm_json, guard, static), toolgate.py, prompts/
  adapters/     agent_safetybench.py, asb.py (orchestration), asb_launch.py (runs in ASB's venv)
  scoring/      shieldagent.py, metrics.py, judge_eval.py (replay + confusion), stats.py
  audit.py      hash-chained JSONL      routing.py   model@arm:judge
  experiment.py run/score/replay/report  smoke.py    day-1 endpoint checks   cli.py
configs/        vetobench.example.yaml, pilot.yaml, sanity.yaml
tests/          unit + proxy tests, fake_vllm.py (GPU-free end-to-end stand-in)

GPU-free end-to-end check: uvicorn tests.fake_vllm:app --port 18000, then point a copy of the config at http://127.0.0.1:18000/v1 and run configs/sanity.yaml.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages