Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions docs/api/core-protocols.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,16 @@ Protocols and ABCs that define RAMPART's extension points. Implement these to co
- register_default_handler_factory
- clear_default_handler_factory

## Trace Execution

::: rampart.core.trace
options:
members:
- EvaluationRecord
- TraceRun
- run_trace_async
- evaluate_final_trace_async

## Errors

::: rampart.core.errors
Expand Down
1 change: 0 additions & 1 deletion docs/api/core-types.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,6 @@ available from `rampart.core`; established result types remain importable from
- resolve_attack_verdict
- resolve_probe_verdict
- resolve_as_attack
- resolve_as_probe

## Configuration

Expand Down
15 changes: 10 additions & 5 deletions docs/concepts/overview.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ A single test run flows from your pytest test, through a RAMPART attack or probe

*Request / response cycle for a single test run.*

Under the hood, every execution follows a common lifecycle owned by [`BaseExecution`][rampart.core.execution.BaseExecution], which drives the per-turn loop between the strategy, your adapter, and the evaluator:
Under the hood, every execution follows a common lifecycle owned by [`BaseExecution`][rampart.core.execution.BaseExecution]. The strategy drives requests through your adapter. Probes evaluate the completed trace once unless an explicit online stop condition is configured; attacks still use prefix evaluation pending their cadence migration.

```mermaid
sequenceDiagram
Expand All @@ -76,11 +76,16 @@ sequenceDiagram
Strat->>Strat: driver.next_prompt_async(history)
Strat->>Adapter: session.send_async(request)
Adapter-->>Strat: Response
Strat->>Eval: evaluate_async(context)
Eval-->>Strat: EvalResult
Note over Strat: Early stop if detected
opt Explicit online stop condition
Strat->>Eval: evaluate_async(prefix context)
Eval-->>Strat: stop EvalResult
Note over Strat: Stop if detected
end
end

Strat->>Eval: evaluate_async(final trace context)
Eval-->>Strat: final EvalResult

Strat-->>Exec: Result
Exec->>Exec: fire ON_POST_EXECUTE
Exec-->>Test: Result
Expand Down Expand Up @@ -112,7 +117,7 @@ Evaluators are **polarity-free**. They answer "did X happen?" — not "is X good
- In an **attack**, detection means the attack objective was achieved → **UNSAFE**
- In a **probe**, detection means the expected behavior is present → **SAFE**

The [`Attacks`][rampart.attacks.Attacks] and [`Probes`][rampart.probes.Probes] factories handle this mapping automatically via [`resolve_as_attack`][rampart.core.result.resolve_as_attack] and [`resolve_as_probe`][rampart.core.result.resolve_as_probe].
The [`Attacks`][rampart.attacks.Attacks] and [`Probes`][rampart.probes.Probes] factories handle this mapping automatically. Probes use [`resolve_probe_verdict`][rampart.core.result.resolve_probe_verdict] over one final-trace evaluation; attacks retain [`resolve_as_attack`][rampart.core.result.resolve_as_attack] until their cadence migration.

You can reuse the same evaluator in both contexts. A [`ToolCalled`][rampart.evaluators.tool_called.ToolCalled] evaluator detects whether a tool was called — whether that's good or bad depends on whether you're attacking or probing.

Expand Down
16 changes: 10 additions & 6 deletions docs/concepts/probes.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,9 @@ Probes use the inverse mapping from evaluator outcomes:
| `NOT_DETECTED` | `UNSAFE` | The expected behavior is missing — a regression |
| `UNDETERMINED` | `UNDETERMINED` | The evaluator could not determine whether the behavior is present |

Precedence: `NOT_DETECTED` > `UNDETERMINED` > `DETECTED`. If any turn failed to detect the expected behavior, the agent is non-compliant.

This logic lives in [`resolve_as_probe`][rampart.core.result.resolve_as_probe].
The evaluator runs once over the completed trace, and the outcome maps directly
to the verdict. This logic lives in
[`resolve_probe_verdict`][rampart.core.result.resolve_probe_verdict].

---

Expand All @@ -26,9 +26,10 @@ Probe executions are simpler than attacks — no injection phase:

1. **Create session** — Open a fresh session with the agent
2. **Send prompts** — Drive the conversation via the prompt driver
3. **Evaluate** — Check whether the expected behavior is present
4. **Clean up** — Close the session
5. **Report** — Produce a [`Result`][rampart.core.result.Result]
3. **Stop (optional)** — Check an explicit online `stop_when` condition
4. **Evaluate** — Check the completed trace once for expected behavior
5. **Clean up** — Close the session
6. **Report** — Produce a [`Result`][rampart.core.result.Result]

---

Expand All @@ -51,6 +52,9 @@ assert result, result.summary

Provide exactly one of `prompt`, `prompts`, or `driver`.

Probes run the full prompt sequence by default. Pass `stop_when=` only when an
online condition should intentionally end the trace early.

---

## Available Probes
Expand Down
5 changes: 3 additions & 2 deletions docs/concepts/trace-schema.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,8 +61,9 @@ stops, including when the turn budget is reached.
why that online evaluation ran. A non-null purpose requires an evaluation on the
same turn. `Result.trace_end_reason` records why turn production stopped. These
provenance fields are optional: missing or null means the producer did not record
them, not that the last online evaluation is the terminal one. The codec never
infers terminal evidence or a stop reason from the result status or turns.
them, not that the last online evaluation is the final-trace evaluation. The
codec never infers final-trace evidence or a stop reason from the result status
or turns.
Both placements of `EvalResult` receive the same strict type, finite-confidence,
Unicode-scalar, and closed-enum validation.

Expand Down
2 changes: 1 addition & 1 deletion docs/contributing/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ When adding a new attack or probe, you add a static factory method — not a new
Evaluators are **polarity-free**. They report whether a condition was detected, not whether it's good or bad. The attack/probe factory applies the correct polarity:

- `resolve_as_attack`: detected → UNSAFE
- `resolve_as_probe`: detected → SAFE
- `resolve_probe_verdict`: detected → SAFE

This allows the same evaluator (e.g., `ToolCalled`) to be used in both attack and probe contexts.

Expand Down
47 changes: 33 additions & 14 deletions docs/contributing/extending-rampart.md
Original file line number Diff line number Diff line change
Expand Up @@ -181,29 +181,48 @@ The process mirrors the [Attack](#attack) walkthrough. The differences are summa
|---|---|---|
| **Location** | `rampart/attacks/_name.py` | `rampart/probes/_name.py` |
| **Factory class** | `Attacks` | `Probes` |
| **Resolution function** | `resolve_as_attack` | `resolve_as_probe` |
| **Resolution function** | `resolve_as_attack` (pending cadence migration) | `resolve_probe_verdict` |
| **Detected means** | UNSAFE | SAFE |
| **Injection phase** | Often yes | No |

### 1. Create the Execution Class

The file structure mirrors the [Attack walkthrough](#1-create-the-execution-class) — same imports, `__init__`, and `_execute_async` loop. The diff from `MyAttackExecution` is:
Probe strategies drive the full trace first, then evaluate it once while the
session is still active:

```diff
-from rampart.core import (..., resolve_as_attack)
+from rampart.core import (..., resolve_as_probe)

-class MyAttackExecution(BaseExecution):
+class MyProbeExecution(BaseExecution):
```python
from rampart.core import (
SafetyStatus,
evaluate_final_trace_async,
resolve_probe_verdict,
run_trace_async,
)

- return "my_attack"
+ return "my_probe"
async with await adapter.create_session_async() as session:
run = await run_trace_async(
session=session,
driver=self._driver,
max_turns=self._max_turns,
observability_level=adapter.observability_profile,
stop_when=self._stop_when,
manifest=adapter.manifest,
)
evaluation = await evaluate_final_trace_async(
evaluator=self._evaluator,
run=run,
)

- status = resolve_as_attack(eval_results=eval_results)
+ status = resolve_as_probe(eval_results=eval_results)
status = (
SafetyStatus.ERROR
if evaluation is None
else resolve_probe_verdict(evaluation=evaluation)
)
```

Place the file in `rampart/probes/` (e.g. `_my_probe.py`). Most probes skip the injection phase — just session creation, prompt driving, and evaluation. For a complete working reference, see [`rampart/probes/_single_turn.py`](https://github.com/microsoft/RAMPART/blob/main/rampart/probes/_single_turn.py).
Store `final_trace_evaluation`, `run.turns`, and `run.trace_end_reason` on the returned
`Result`. Most probes skip the injection phase. For a complete working
reference, see
[`rampart/probes/_single_turn.py`](https://github.com/microsoft/RAMPART/blob/main/rampart/probes/_single_turn.py).

### 2. Add a Factory Method to `Probes`

Expand All @@ -214,7 +233,7 @@ Add a static method to the `Probes` class in `rampart/probes/__init__.py`, mirro
Probe tests have the same surface as attack tests, with two differences:

- **No injection phase** to test.
- **Result resolution** uses `resolve_as_probe` semantics (detected → SAFE, not detected → UNSAFE).
- **Result resolution** uses `resolve_probe_verdict` semantics (detected → SAFE, not detected → UNSAFE).


## Evaluator
Expand Down
2 changes: 1 addition & 1 deletion docs/contributing/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,7 +147,7 @@ When adding a new attack, test:
Similar to attacks, but:

1. No injection phase to test
2. Result resolution uses `resolve_as_probe` (detected → SAFE, not detected → UNSAFE)
2. Result resolution uses `resolve_probe_verdict` over one final-trace evaluation (detected → SAFE, not detected → UNSAFE)

### Testing a New Evaluator

Expand Down
29 changes: 26 additions & 3 deletions docs/probes/behavioral.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,10 @@ Use behavioral probes for regression testing: ensure your agent still does the r

1. **Create session** — Open a fresh session with the agent
2. **Send prompts** — Drive the conversation via a prompt driver
3. **Evaluate** — Check each turn for the expected behavior. Early-stops on detection.
4. **Clean up** — Close the session
5. **Result** — Produce a [`Result`][rampart.core.result.Result] via `resolve_as_probe` semantics
3. **Stop (optional)** — Evaluate `stop_when` after each response and stop when detected
4. **Evaluate** — Check the expected behavior once over the completed trace
5. **Clean up** — Close the session
6. **Result** — Map the final evaluation using probe semantics

No injection phase.

Expand Down Expand Up @@ -80,6 +81,27 @@ result = await Probes.behavior(
be ignored. Scope applies only to turns in the evaluator context; it does
not force an execution to produce every planned turn.

Probes do not stop early unless `stop_when` is configured. The verdict
evaluator therefore receives the completed trace, and `ALL_TURNS` or
negated `ANY_TURN` applies to every response that was produced.

!!! note "Driver budgets"
An adaptive driver such as `LLMDriver` does not stop itself. Without
`stop_when`, it runs until `max_turns` and then evaluates that completed
trace once. Set an intentional budget, and add an explicit stop condition
when earlier termination is part of the scenario.

!!! note "Upgrading from per-turn probe verdicts"
Earlier releases evaluated a probe after each response, stopped at the
first detection, and combined the per-turn results. Probes now evaluate the
completed trace once. Single-prompt probes with deterministic evaluators
keep the same verdicts. Multi-turn probes can resolve differently because
the evaluator's scope now applies to the full trace, which runs up to
`max_turns` unless `stop_when` is set. A stochastic evaluator, such as an
LLM judge, is sampled once per run instead of once per turn, so trial pass
rates can shift. Replace `resolve_as_probe(eval_results=...)` with
`resolve_probe_verdict(evaluation=...)`.

---

## Parameters
Expand All @@ -92,6 +114,7 @@ See [`Probes.behavior()`][rampart.probes.Probes.behavior] for the full API refer
| `prompts` | `list[str] \| None` | `None` | A list of prompt strings. |
| `driver` | [`PromptDriver`][rampart.core.prompt_driver.PromptDriver] `\| None` | `None` | A pre-built prompt driver. |
| `evaluator` | [`Evaluator`][rampart.core.evaluator.Evaluator] | required | What behavior to detect. |
| `stop_when` | [`Evaluator`][rampart.core.evaluator.Evaluator] `\| None` | `None` | Optional online condition that stops the trace when detected. |
| `max_turns` | `int` | `25` | Maximum exchanges; reaching the limit resolves the trace normally. |

!!! warning
Expand Down
4 changes: 4 additions & 0 deletions docs/usage/authoring-tests.md
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,10 @@ example `Pattern found on turn(s): 0, 2`. `CURRENT_TURN` uses the same format
with only the latest turn number. A failed `ALL_TURNS` match identifies the
missing turns with `Pattern missing on turn(s): ...`.

Probes evaluate their verdict once over the completed trace unless an
explicit `stop_when` ends the scenario. Attack cadence is documented in
the attack guide.

#### How Each Evaluator Sees the Transcript

Built-in evaluators reach their temporal behavior in two ways. Quantifying
Expand Down
28 changes: 17 additions & 11 deletions docs/usage/results-and-reporting.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,16 +60,20 @@ stops, not simply the last online evaluation. It is an
input to the final status, not a duplicate status: execution policy can still
adjust the verdict, and `result.status` remains authoritative.

This layer makes terminal provenance durable before changing execution
cadence. Existing prefix-evaluated strategies leave these fields as `None`
until their follow-up migration; manually constructed and error results may do
the same intentionally.
Behavioral probes evaluate the completed trace once by default. Their
`Result.final_trace_evaluation` contains the verdict evidence, while
`Result.turn_evaluations` is normally empty. Configure `stop_when` only when online stop evidence is
intentionally needed.

Strategies that do not record a final-trace verdict leave
`final_trace_evaluation` and `trace_end_reason` as `None`; manually constructed
and error results may do the same intentionally.

Online evaluations attached to turns are available as
`result.turn_evaluations`; this list excludes the terminal evaluation.
`result.turn_evaluations`; this list excludes the final-trace evaluation.
The former `result.eval_results` property has been removed. Use
`result.turn_evaluations` for online evidence and `result.final_trace_evaluation`
for terminal verdict evidence.
for final-trace verdict evidence.

`TraceEndReason.MAX_TURNS_REACHED` records budget truncation. It does not by
itself claim that the scenario reached semantic completion; each execution
Expand All @@ -80,19 +84,21 @@ within that size, and a finite threshold from 0.0 through 1.0.

### Observability Gaps on a Passing Run

A run can resolve `SAFE` while part of the evaluation was never observable. Such a run is graded as a pass: `result.safe` is `True`, the result line reads `PASS`, an execution population counts it toward the pass rate, and pytest exits zero. `result.summary` names the gap, and `turn.eval_result.undetermined_operands` carries it one reason at a time, so a caller that wants to fail on it has to say so:
A run can resolve `SAFE` while part of the evaluation was never observable. Such a run is graded as a pass: `result.safe` is `True`, the result line reads `PASS`, an execution population counts it toward the pass rate, and pytest exits zero. `result.summary` names the gap, and each evaluation's `undetermined_operands` carries it one reason at a time. Inspect final-trace evidence as well as any online evaluations when choosing to fail on a gap:

```python
evaluations = result.turn_evaluations
if result.final_trace_evaluation is not None:
evaluations.append(result.final_trace_evaluation)
gaps = [
reason
for turn in result.turns
if turn.eval_result is not None
for reason in turn.eval_result.undetermined_operands
for evaluation in evaluations
for reason in evaluation.undetermined_operands
]
assert result and not gaps, result.summary
```

`JsonFileReportSink` writes the same list as `eval_undetermined_operands` on each turn that has one, and omits the key otherwise. A failing run can carry the key too, so read it alongside `status`: together they tell a fully observed pass from one reached with a gap. No counter makes that distinction, because a qualified pass lands in `safe_count` like any other.
`JsonFileReportSink` writes final-trace gaps as `final_trace_evaluation.undetermined_operands` and online gaps as `eval_undetermined_operands` on each turn. Empty gap lists are omitted. A failing run can carry these keys too, so read them alongside `status`: together they tell a fully observed pass from one reached with a gap. No counter makes that distinction, because a qualified pass lands in `safe_count` like any other.

XPIA applies one further rule of its own to `RESPONSE_ONLY` adapters, which does move the verdict. See [Observability Adjustment](../attacks/xpia.md#observability-adjustment).

Expand Down
2 changes: 0 additions & 2 deletions rampart/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,6 @@
Result,
SafetyStatus,
resolve_as_attack,
resolve_as_probe,
)
from rampart.core.types import (
EvalContext,
Expand Down Expand Up @@ -111,7 +110,6 @@
"execute_trials_async",
"record_result",
"resolve_as_attack",
"resolve_as_probe",
]


Expand Down
12 changes: 10 additions & 2 deletions rampart/core/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,10 +32,15 @@
Result,
SafetyStatus,
resolve_as_attack,
resolve_as_probe,
resolve_attack_verdict,
resolve_probe_verdict,
)
from rampart.core.trace import (
EvaluationRecord,
TraceRun,
evaluate_final_trace_async,
run_trace_async,
)
from rampart.core.types import (
EvalContext,
EvalOutcome,
Expand Down Expand Up @@ -63,6 +68,7 @@
"EvalOutcome",
"EvalResult",
"EvaluationPurpose",
"EvaluationRecord",
"Evaluator",
"ExecutionEvent",
"ExecutionEventData",
Expand Down Expand Up @@ -92,11 +98,13 @@
"ToolCall",
"ToolDeclaration",
"TraceEndReason",
"TraceRun",
"Turn",
"evaluate_final_trace_async",
"evaluate_turn_async",
"execute_trials_async",
"resolve_as_attack",
"resolve_as_probe",
"resolve_attack_verdict",
"resolve_probe_verdict",
"run_trace_async",
]
Loading
Loading