Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
51 commits
Select commit Hold shift + click to select a range
cea1113
Execute Level 2 probes with raw evidence and basic graders
burtenshaw Sep 16, 2026
a74c97b
Exercise runtime faults and reference Echo with complete evidence
burtenshaw Sep 21, 2026
623cce7
fix: bound runtime transport and worker inputs
burtenshaw Sep 21, 2026
7282482
fix: honor runtime timeout and UTF-8 schema input
cursoragent Sep 21, 2026
b9ae0de
fix: address runtime validation reviews
burtenshaw Sep 21, 2026
29c5801
fix: retain findings when teardown fails
burtenshaw Sep 21, 2026
ccca191
fix: preserve runtime failure reports
burtenshaw Sep 22, 2026
7e9b888
refactor: simplify runtime validation
burtenshaw Sep 22, 2026
bea5115
Merge main (5b1982ac) into the Level 2 runtime probes branch
cursoragent Sep 23, 2026
65be5dd
fix: report initial source digest failures
cursoragent Sep 23, 2026
4c746aa
fix: stop validation after source digest failure
cursoragent Sep 23, 2026
16240f1
Merge main into ben/rfc008-l2-03-runtime for Zach review
cursoragent Sep 23, 2026
d94f021
fix: preserve runtime evidence on close interrupt
cursoragent Sep 23, 2026
f84f8a3
Merge remote-tracking branch 'origin/main' into ben/rfc008-l2-04-tele…
burtenshaw Sep 24, 2026
434bb1d
feat: add session validation telemetry
burtenshaw Sep 24, 2026
f1575e9
fix: keep process evidence free of bytecode
burtenshaw Sep 24, 2026
297b578
Fix source digest portability on Windows
cursoragent Sep 24, 2026
776dace
Merge remote-tracking branch 'origin/ben/rfc008-l2-03-runtime' into b…
burtenshaw Sep 24, 2026
317983c
fix: preserve runtime evidence on telemetry errors
burtenshaw Sep 24, 2026
d365379
Merge remote-tracking branch 'origin/main' into HEAD
burtenshaw Oct 1, 2026
ee9094e
Merge commit 'd365379d1a4109b90e1f58efaad917d56c4fe5dd' into HEAD
burtenshaw Oct 1, 2026
237d11f
fix: drop unsafe malformed telemetry frames
burtenshaw Oct 1, 2026
8937647
Merge commit '237d11f5c883a2c12684465c064eedb82c49f3b0' into HEAD
burtenshaw Oct 1, 2026
9cdaff0
fix: runtime validation diagnostics and deadlines
burtenshaw Oct 1, 2026
b31d6ee
Merge commit '9cdaff04' into HEAD
burtenshaw Oct 1, 2026
d3d153b
fix: integrate runtime review fixes
burtenshaw Oct 1, 2026
aa85ab8
fix: hash sources portably without following swaps
burtenshaw Oct 1, 2026
c7722bd
Merge commit 'aa85ab80' into HEAD
burtenshaw Oct 1, 2026
94494cc
Merge commit 'c7722bd1509bf963053a16ab8a15ed3583059e8a' into HEAD
burtenshaw Oct 1, 2026
74df6cc
fix: enforce HTTP collection deadlines
burtenshaw Oct 1, 2026
4e75e52
fix: enforce runtime HTTP deadlines
burtenshaw Oct 1, 2026
0cf6760
test: bound socket cancellation regression
burtenshaw Oct 1, 2026
4ebd9f0
fix: inherit bounded HTTP transport
burtenshaw Oct 1, 2026
80afc19
test: inherit bounded cancellation check
burtenshaw Oct 1, 2026
36149d7
test: bound socket cancellation regression
burtenshaw Oct 1, 2026
ca2d293
fix: preserve startup errors
burtenshaw Oct 2, 2026
b795b40
fix: align reset and unscored reward contracts
burtenshaw Oct 2, 2026
46d678a
fix: integrate reset contracts
burtenshaw Oct 2, 2026
c7c4101
fix: protect reset schema evidence
burtenshaw Oct 2, 2026
2c9a486
docs: explain runtime contracts
burtenshaw Oct 2, 2026
680a0aa
docs: sync runtime contracts
burtenshaw Oct 2, 2026
2704615
fix: locate echo in isolated protocol tests
burtenshaw Oct 2, 2026
d98155f
test: sync installed-wheel regression
burtenshaw Oct 2, 2026
f60fa9e
fix: isolate factory test setup
burtenshaw Oct 2, 2026
5473641
docs: clarify failure handling
burtenshaw Oct 2, 2026
d27048e
fix: derive runtime checks from registry
burtenshaw Oct 5, 2026
b9e3e29
fix: enforce level two contracts
burtenshaw Oct 5, 2026
f97c0ec
fix: sync validation decisions
burtenshaw Oct 5, 2026
1d7a968
Merge branch 'main' into ben/rfc008-l2-03-runtime
burtenshaw Oct 5, 2026
eb1bb16
fix: sync current validation base
burtenshaw Oct 5, 2026
b712048
fix: reconcile merged validation parent
burtenshaw Oct 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 33 additions & 4 deletions rfcs/008-environment-auto-validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -197,9 +197,10 @@ configuration (model, version, params) in the manifest; the oracle check becomes
bit-exact. RFC 004 rubrics are **leveraged, not required**: the contract stays spec-neutral
(graders read the manifest), but for the served OpenEnv format the rubric tree is the native
satisfaction path — `LLMJudge` is the in-repo `llm_judged` implementation, and the
introspectability and reward-attribution graders read `named_rubrics()` / `state_dict()` /
per-child scores. Judge pinning stays a *manifest* declaration because the rubric object does not
serialize model/version/params today.
introspectability and reward-attribution graders read `named_rubrics()`, explicit
`validation_config()` and fresh per-child scores. `state_dict()` is never serialized
as validation configuration. Judge pinning stays a *manifest* declaration because
the rubric object does not serialize model/version/params today.

Tolerances, margins, and variance bounds are author-declared in the manifest, **bounded by the
versioned severity policy**, and carried verbatim in reports so hubs can apply stricter ceilings.
Expand Down Expand Up @@ -624,7 +625,35 @@ attached to the **same** replay connection, reject unauthorized/cross-session
reads and never expose telemetry as agent MCP tools. A second WebSocket creates
another environment and cannot supply evidence for the measured instance.
A validator transcript alone cannot pass subject-emitted trajectory recording.
These authorization requirements do not add new public wire messages in this slice.
The initial contracts slice added no public wire messages. PR4 implements the
following opt-in protocol.

#### Session telemetry protocol (PR4)

The server opts in only when `OPENENV_VALIDATION_TOKEN` is provisioned explicitly
by the validation supervisor. On its existing simulation `/ws` connection, the
collector sends `validation_open` with `data: {schema_version: 1, token: ...}`.
The reply returns a random capability bound to that connection. Subsequent
`validation_read` messages carry that capability and return a `validation`
snapshot. Disabled, unauthorized and cross-connection requests fail without
echoing credentials. Closing the connection destroys the capability. MCP and
production endpoints never expose these operations. Authentication exchanges
are excluded from persisted evidence.

Snapshots identify requested and actually forwarded seed arguments; successful
reset alone is not acceptance. They include a named rubric tree rooted at `root`,
explicit safe configuration, and per-step attribution with operation-local
evaluation flags. Unevaluated gated children never reuse an earlier score.
Stock container aggregation is named explicitly; custom rubrics may supply
`validation_config()` to expose public JSON configuration. Arbitrary attributes
and `state_dict()` are never serialized as configuration.

The subject server also emits a bounded record of the reset/step/state request
and response envelopes it executed. The validator captures the wire separately
and later compares the two. The record is bound to the authenticated session,
limited to 100 steps/202 operations and 8 MiB, and marks truncation explicitly.
Missing configuration, attribution or records cannot be inferred from other
successful operations. This transport introduces no new passing grader by itself.

Applicability predicates must distinguish empty declared sets from absent
capabilities. Missing subject features, missing provider support and checks whose
Expand Down
108 changes: 108 additions & 0 deletions src/openenv/core/env_server/http_server.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@
import json
import logging
import os
import secrets
import time
import uuid
from concurrent.futures import ThreadPoolExecutor
Expand Down Expand Up @@ -61,6 +62,17 @@
)
from .route_config import GetEndpointConfig, register_get_endpoints
from .serialization import deserialize_action, serialize_observation
from .session_telemetry import (
rubric_counts,
rubric_snapshot,
SeedAcceptance,
SessionTelemetry,
ValidationOpenedData,
ValidationOpenedResponse,
ValidationOpenMessage,
ValidationReadMessage,
ValidationResponse,
)
from .types import (
Action,
ConcurrencyConfig,
Expand Down Expand Up @@ -879,6 +891,12 @@ def register_routes(
f"Invalid mode: '{mode}'. Must be one of: {valid_modes}"
)

# Only explicitly provisioned simulation servers accept validation controls.
validation_token = os.environ.get("OPENENV_VALIDATION_TOKEN", "")
validation_enabled = (
mode == ServerMode.SIMULATION and 32 <= len(validation_token) <= 256
)

# Wire up idle-session reaper lifecycle via app events
server_ref = self

Expand Down Expand Up @@ -1753,6 +1771,8 @@ async def websocket_endpoint(websocket: WebSocket):
session_env = None
owns_session = False
attached_session = False
telemetry = None
operations_started = False

try:
requested_session_id = websocket.query_params.get("session_id")
Expand Down Expand Up @@ -1817,6 +1837,64 @@ async def websocket_endpoint(websocket: WebSocket):

msg_type = message_dict.get("type", "")

if msg_type in {"validation_open", "validation_read"}:
# Do not return Pydantic input/error details: these messages
# contain credentials and must never echo or log them.
try:
if not validation_enabled or not owns_session:
raise ValueError("Validation unavailable")
if msg_type == "validation_open":
auth = ValidationOpenMessage.model_validate(
message_dict
)
if telemetry is not None or operations_started:
raise ValueError("Validation already started")
if not secrets.compare_digest(
validation_token.encode(),
auth.data.token.get_secret_value().encode(),
):
raise ValueError("Unauthorized")
telemetry = SessionTelemetry()
response = ValidationOpenedResponse(
data=ValidationOpenedData(
capability=telemetry.capability
)
)
else:
auth = ValidationReadMessage.model_validate(
message_dict
)
if telemetry is None or not telemetry.authorized(
auth.data.capability.get_secret_value()
):
raise ValueError("Unauthorized")
response = ValidationResponse(
data=telemetry.snapshot
)
except Exception:
response = WSErrorResponse(
data={
"message": "Validation unavailable or unauthorized",
"code": WSErrorCode.VALIDATION_ERROR,
}
)
await websocket.send_text(response.model_dump_json())
continue

seed_acceptance = None
before_scores = None
if msg_type in {"reset", "step", "state"}:
operations_started = True
if telemetry is not None and msg_type == "step":
try:
before_scores = rubric_counts(
getattr(session_env, "rubric", None)
)
except Exception:
telemetry.snapshot.rubric_error = (
"Rubric introspection unavailable"
)

try:
match msg_type:
case "reset":
Expand Down Expand Up @@ -1848,6 +1926,12 @@ async def websocket_endpoint(websocket: WebSocket):
)
)

if telemetry is not None:
seed_acceptance = SeedAcceptance(
requested="seed" in msg.data,
value=msg.data.get("seed"),
accepted="seed" in valid_kwargs,
)
self._update_session_activity(session_id)

response = WSObservationResponse(
Expand Down Expand Up @@ -1925,6 +2009,30 @@ async def websocket_endpoint(websocket: WebSocket):
}
)

if telemetry is not None and msg_type in {
"reset",
"step",
"state",
}:
nodes = None
try:
if msg_type in {"reset", "step"}:
nodes = rubric_snapshot(
getattr(session_env, "rubric", None),
before_scores,
)
except Exception:
nodes = []
telemetry.snapshot.rubric_error = (
"Rubric introspection unavailable"
)
telemetry.append(
msg_type,
message_dict,
response.model_dump(mode="json"),
seed=seed_acceptance,
rubric=nodes,
)
await websocket.send_text(response.model_dump_json())

except ValidationError as e:
Expand Down
Loading
Loading