diff --git a/docs/source/environments/harbor.md b/docs/source/environments/harbor.md index d2a28e62d..e79bdae73 100644 --- a/docs/source/environments/harbor.md +++ b/docs/source/environments/harbor.md @@ -240,7 +240,9 @@ This path involves no env server, which makes it the one to reach for when somet openenv harbor serve --llm-url $LLM --dataset org/train,org/eval ``` -You get a Task API for discovery, one long-running `run_rollout` MCP tool, and a UI at `/web`. +You get a Task API for discovery, one long-running `run_rollout` MCP tool, and a UI at `/web` for +reading tasks, running agents on them and reading what they did (see [The web UI](#the-web-ui)). +`--llm-url` is optional: without it, whoever uses the UI connects a model of their own. ```python from harbor_env import HarborEnv @@ -257,6 +259,96 @@ with HarborEnv(base_url="http://localhost:8000") as env: `harness` and `sandbox` are per call, so consecutive rollouts against the same server can use different agents and different backends. +## The web UI + +`serve` (and a Space made with `push`) serves a UI at `/web`, in four tabs. + +**Tasks.** Every task of every served dataset, as cards you can search and filter by dataset, +category, difficulty and tag. Opening one shows exactly what the agent will receive, the task's files +(with a full-window viewer), the environment and verifier settings from `task.toml`, and the runs of +that task. Metadata fields that hold the answer (`gold_answer`, `solution`, ...) are left out of the +summary. **Add from the Hub** lists public datasets tagged `harbor`, each with its task count and +size, and adds one in the background with its progress shown; a dataset added this way can be +removed again. Only datasets in Harbor's `tasks//` layout (flat or grouped) can be added. + +**Running a task.** The run card on a task picks the model, the agent and the sandbox: + +- **This server**: the endpoint `serve` was started with. Its key stays in the capture proxy. +- **Hugging Face**: a model on Inference Providers, chosen from a searchable list with providers, + context length and price, reached with a Hugging Face token (locally, this machine's token; on a + Space, optionally by signing in with Hugging Face). +- **Your endpoint**: any OpenAI-compatible URL (vLLM, SGLang, ...) or an Anthropic endpoint, probed + before use. Training capture needs vLLM with `--return-tokens-as-token-ids`. + +A rollout runs on the server whether or not the page stays open. Model usage is billed to the +account or token selected on the card; sandbox compute is billed to the server operator, including +Hugging Face Sandbox on a Space. + +**Runs.** Every run you may see, filterable by status. A run shows what the agent did as a timeline +(its prompt, thinking, each tool call with its input and output, terminal sessions, the final +answer), the result and reward, the trace checks, and downloads of the result JSON and, for a +trainable rollout, the training contract. Tick two to four runs to compare them. + +**Setup.** The server's endpoint, which sandboxes have working credentials, the deployment settings +below, the datasets and the agents. + +### Deployment settings + +The same UI runs on a laptop and as a public Space, and what a visitor may do differs. The defaults +depend on where the server runs: **local** is `serve --host 127.0.0.1`; **network** is any other bind +address, including the default `0.0.0.0`, since everyone who can reach it is then a visitor; a +**Space** is detected from `SPACE_ID`. Each setting is an environment variable (a Space variable on +a Space), and some have a flag on `serve` and `push` (`--private-urls` is on `serve` only). + +| setting | variable | flag | local | network | Space | +|---|---|---|---|---|---| +| Visitors may start rollouts | `OPENENV_HARBOR_UI_ROLLOUTS` | `--rollouts` | on | on | on | +| Visitors may use the server's endpoint | `OPENENV_HARBOR_UI_SERVER_ENDPOINT` | `--share-endpoint` | on | on | off | +| Visitors may connect their own | `OPENENV_HARBOR_UI_VISITOR_ENDPOINTS` | `--visitor-endpoints` | on | on | on | +| A visitor's URL may be private or local | `OPENENV_HARBOR_UI_PRIVATE_URLS` | `--private-urls` | on | off | off | +| Offer this machine's HF token | `OPENENV_HARBOR_UI_LOCAL_TOKEN` | | on | off | never | +| Add and remove Hub datasets from the page | `OPENENV_HARBOR_UI_ADD_DATASETS` | `--add-datasets` | on | off | off | +| Who sees runs (`all` or `own`) | `OPENENV_HARBOR_RUN_VISIBILITY` | `--run-visibility` | all | all | own | +| Keep finished runs across restarts | `OPENENV_HARBOR_RUN_HISTORY` | `--run-history` | on | on | off | +| Headline reward of a task with several | `OPENENV_HARBOR_REWARD_KEY` | `--reward-key` | unset | unset | unset | +| Rollouts at once | `OPENENV_HARBOR_UI_MAX_RUNS` | | 4 | 4 | 4 | +| Rollouts at once per visitor | `OPENENV_HARBOR_UI_MAX_RUNS_PER_VISITOR` | | 4 | 4 | 2 | + +`own` visibility ties runs to a random id kept in the visitor's browser; a run stores only a digest +of it. That keeps visitors' runs apart, but it is not sign-in: a visitor who clears the browser's +storage loses their runs, and it is no substitute for access control over traces you consider +private. The per-visitor limit counts a signed-in visitor's Hugging Face account, and otherwise that +browser id, so for anonymous visitors it stops one page from taking every slot, not someone set on +it; `OPENENV_HARBOR_UI_MAX_RUNS` bounds the total. Run history goes to `OPENENV_HARBOR_RUNS_DIR`, by default `~/.cache/openenv/harbor/runs`, or +`/data/harbor-runs` on a Space with the bucket mounted. `OPENENV_HARBOR_UI_MAX_ADD_GB` (default `5`) +caps the size of a dataset added from the page. + +What the UI guarantees whatever the settings: + +- A key or token typed into the page is sent to that endpoint by the server, held in server memory + for that page only, and never written to disk, to run history or back to the page. +- With private URLs off, a visitor's URL, and every redirect it answers with, must resolve to public + addresses, and is checked again before each rollout. +- Every dataset, task and file the page asks for is one the server serves or that was added from + the page. Every file it reads, for a card, the task view, the file viewer or the environment + check, must resolve inside the dataset's own folder (a link to that folder itself is followed), + so a task reached through a link to anywhere else shows nothing and cannot run on a visitor's + model. The Task API's instruction preview follows the same rule, and a registry dataset is + anchored on Harbor's cache. A dataset added from the page may not contain a symbolic link at all. +- A task that reads the server's environment variables (`${VAR}` in `task.toml` or a compose file, + or a bare name under a compose `environment:`), which is where the server's keys are, or its files + (a compose `env_file`, `include`, `extends`, or a host path in a mount, secret, build context, + cache, watch rule or device), or asks for more of the host than a folder (`privileged`, + `cap_add`, the host's namespaces, the Docker socket, another container's volumes, a named volume + or network with settings, the build's SSH agent) runs only on the server's own endpoint, never on + a model a visitor connects: that model does what the visitor asks, printing the sandbox's + environment included, into a trace the visitor reads. A dataset added from the page may not do + either at all. Both files are checked as parsed, and one that doesn't parse counts as reading. +- A request that changes something in the UI (a rollout, an added dataset) is refused when a + browser sends it from another site, so another page can't act through a visitor's browser. Behind + a proxy that rewrites `Host`, list the public host in `OPENENV_HARBOR_UI_HOSTS` (comma-separated). +- Everything a model, a task or a tool produced is escaped before it is shown. + ## CLI reference Four commands. Every flag below is the complete set, with its type and default. @@ -340,16 +432,26 @@ Start the env server: Task API for discovery, one long-running `run_rollout` MCP | flag | type | default | meaning | |---|---|---|---| -| `--llm-url` | str | **required** | OpenAI-spec endpoint | +| `--llm-url` | str | `""` | OpenAI-spec endpoint. Optional: without one, UI visitors connect their own | | `--dataset` | str | none | Dataset specs to serve as splits. Repeatable | | `--model` | str | `""` | Served model id | -| `--host` | str | `0.0.0.0` | Bind address | +| `--host` | str | `0.0.0.0` | Bind address. `127.0.0.1` gives the UI its local defaults, see [Deployment settings](#deployment-settings) | | `--port` | int | `8000` | Env server port. Faces the trainer and the browser | | `--capture-port` | int | `8100` | Capture proxy port. Faces the sandbox | | `--expose` | str | `gradio` | How the sandbox reaches the proxy | | `--env-file` | path | `""` | dotenv with provider credentials | | `--api-key` | str | `$OPENENV_LLM_API_KEY` | Credential for the endpoint, for a hosted provider | | `--auth-header` | str | `Authorization` | Header to send it under, e.g. `x-api-key` | +| `--share-endpoint/--no-share-endpoint` | flag | unset | UI visitors may use this endpoint and its key | +| `--visitor-endpoints/--no-visitor-endpoints` | flag | unset | UI visitors may connect their own model | +| `--run-visibility` | str | unset | `all` or `own`: who sees runs in the UI | +| `--run-history/--no-run-history` | flag | unset | Keep finished UI runs across restarts | +| `--add-datasets/--no-add-datasets` | flag | unset | UI visitors may add and remove Hub datasets | +| `--rollouts/--no-rollouts` | flag | unset | UI visitors may start rollouts | +| `--private-urls/--no-private-urls` | flag | unset | UI visitors may connect an endpoint on a private or local address | +| `--reward-key` | str | unset | The reward UI runs report for tasks with several and none named `reward`. Unset, such a run lists each one | + +An unset UI flag keeps the default for where the server runs. Refuses to start only if the endpoint is unreachable. One that cannot return token ids starts as an eval deployment and says so. @@ -359,17 +461,25 @@ Deploy the same server to a Hugging Face Space. | flag | type | default | meaning | |---|---|---|---| -| `--llm-url` | str | **required** | Endpoint the deployed Space will use | +| `--llm-url` | str | `""` | Endpoint the deployed Space will use. Optional: without one, visitors connect their own | | `--repo-id` | str | **required** | Target Space, e.g. `you/harbor-env` | | `--dataset` | str | none | Dataset specs. Repeatable | | `--model` | str | `""` | Served model id | | `--bucket` | str | Space name | Storage bucket holding the task suites. `none` disables the mount and downloads instead | +| `--public-bucket/--private-bucket` | flag | private | Visibility of a new bucket. Given for an existing bucket, it changes it; not given, an existing bucket keeps its visibility | +| `--hf-login/--no-hf-login` | flag | on | Let visitors sign in with Hugging Face to use their own account for Inference Providers | | `--hardware` | str | `""` | Space hardware, e.g. `cpu-basic` | -| `--private` | flag | off | Create it private. Rollouts then cannot work, see below | +| `--private` | flag | off | Create the Space private. Rollouts then cannot work, see below | | `--recreate` | flag | off | Delete the Space first, then deploy fresh | | `--dry-run` | flag | off | Print exactly what would be sent and stop | | `--env-file` | path | `""` | dotenv whose provider keys become Space **secrets** | +`push` also takes the UI flags of `serve` (`--share-endpoint`, `--visitor-endpoints`, +`--run-visibility`, `--run-history`, `--add-datasets`, `--rollouts`, `--reward-key`) and sets them as Space +variables. Unlike `serve`, it doesn't share the endpoint with UI visitors unless `--share-endpoint` is +given. Each push reconciles these UI variables: an omitted flag removes an earlier override and +returns to the Space default, so a one-time `--share-endpoint` does not stay enabled forever. + ```bash openenv harbor push --llm-url $LLM --dataset org/train,org/eval \ --repo-id you/harbor-env --env-file .env --dry-run @@ -379,6 +489,18 @@ openenv harbor push --llm-url $LLM --dataset org/train,org/eval \ at `/capture`, and a private Space requires an auth header the sandboxed agent does not send. Use it only to park a deployment. +### Migration notes for the new UI + +- `serve` binds `0.0.0.0` by default and therefore uses the network-safe UI defaults. Visitor + endpoints such as `http://localhost:8000/v1` are refused. Bind the whole server to + `--host 127.0.0.1`, or pass `--private-urls` when the server still needs a remote trainer. +- An existing Space re-pushed with `--llm-url` no longer shares that endpoint with UI visitors + unless `--share-endpoint` is passed. `--hf-login` also writes `hf_oauth: true` and the + `inference-api` OAuth scope into the Space README. +- Capture `/health` still returns status and aggregate fields publicly, but its per-session + `upstreams` list is empty when an admin key is configured. Set + `OPENENV_CAPTURE_ADMIN_KEY` and send it as a bearer token to inspect that field. + ## Supported harnesses 16 of the 30 known agents are validated end to end. "Validated" means a real rollout produced token @@ -506,7 +628,25 @@ Two details that matter: **Task suites are mounted, not downloaded.** A Harbor suite is thousands of small files and Space disk is ephemeral, so a download is re-paid on every restart. `push` syncs the suites into a storage bucket named after the Space and mounts it at `/data`. The copy is server side, and re-running -`push` copies only what is new. +`push` copies only what is new. The bucket is created private (`--public-bucket` to change that), +since with run history on it also holds every visitor's runs. With `--add-datasets`, the `tasks/` +folder of a dataset added from the UI is copied into the same bucket, server side, and read through +the mount, so it survives restarts; removing it deletes it from the bucket. + +**Visitors can sign in with Hugging Face.** `push` turns on OAuth for the Space (`hf_oauth: true`, +scope `inference-api`), and the UI offers "sign in with Hugging Face" as a way to use Inference +Providers on the visitor's own account instead of pasting a token. It is optional and not needed +for anything else; `--no-hf-login` leaves it out. + +**Visitors bring their own model.** On a Space, a visitor runs on a model they connect: signing in +with Hugging Face, a token, or their own endpoint. `--share-endpoint` lets them use the Space's own +endpoint and key too, at your cost; the Task API and MCP use it either way. On that endpoint a +served task that reads the Space's environment does run, and the visitor who starts it reads its +trace, so don't combine `--share-endpoint` with such tasks on a public Space. + +**Rollouts on a Space use the Space's credentials for sandboxes.** An `hf-sandbox` rollout is billed +to the `HF_TOKEN` the Space holds, whoever started it. `--no-rollouts` makes a public deployment a +read-only task browser. **The Space must be public.** The capture proxy is served at `/capture`, and a private Space requires an auth header that the agent inside the sandbox does not send. This is safe because @@ -542,6 +682,14 @@ unaffected. **Many roots for one rollout.** Normal for agents that run subagents or auxiliary calls. Each root is a separate conversation, and only agent conversations are counted as trainable. +**A dataset is greyed out in "Add from the Hub".** It has no `tasks//` folder, which is how +Harbor datasets are laid out, or it is over `OPENENV_HARBOR_UI_MAX_ADD_GB`. + +**A URL is refused as "private or local".** The server does not call private addresses for visitors +unless `OPENENV_HARBOR_UI_PRIVATE_URLS` is on, which by default it is only for `--host 127.0.0.1`; +`serve --private-urls` turns it on for any bind address. To reach a vLLM on your own network from a +Space, run the UI yourself. + **Exit code 137.** The agent was killed inside the sandbox, almost always by the OOM killer on a large input. That is a task failure, not a capture failure. diff --git a/envs/harbor_env/README.md b/envs/harbor_env/README.md index f743e6d05..c5765cd2c 100644 --- a/envs/harbor_env/README.md +++ b/envs/harbor_env/README.md @@ -248,7 +248,9 @@ This path involves no env server, which makes it the one to reach for when somet openenv harbor serve --llm-url $LLM --dataset org/train,org/eval ``` -You get a Task API for discovery, one long-running `run_rollout` MCP tool, and a UI at `/web`. +You get a Task API for discovery, one long-running `run_rollout` MCP tool, and a UI at `/web` for +reading tasks, running agents on them and reading what they did (see [The web UI](#the-web-ui)). +`--llm-url` is optional: without it, whoever uses the UI connects a model of their own. ```python from harbor_env import HarborEnv @@ -265,6 +267,96 @@ with HarborEnv(base_url="http://localhost:8000") as env: `harness` and `sandbox` are per call, so consecutive rollouts against the same server can use different agents and different backends. +## The web UI + +`serve` (and a Space made with `push`) serves a UI at `/web`, in four tabs. + +**Tasks.** Every task of every served dataset, as cards you can search and filter by dataset, +category, difficulty and tag. Opening one shows exactly what the agent will receive, the task's files +(with a full-window viewer), the environment and verifier settings from `task.toml`, and the runs of +that task. Metadata fields that hold the answer (`gold_answer`, `solution`, ...) are left out of the +summary. **Add from the Hub** lists public datasets tagged `harbor`, each with its task count and +size, and adds one in the background with its progress shown; a dataset added this way can be +removed again. Only datasets in Harbor's `tasks//` layout (flat or grouped) can be added. + +**Running a task.** The run card on a task picks the model, the agent and the sandbox: + +- **This server**: the endpoint `serve` was started with. Its key stays in the capture proxy. +- **Hugging Face**: a model on Inference Providers, chosen from a searchable list with providers, + context length and price, reached with a Hugging Face token (locally, this machine's token; on a + Space, optionally by signing in with Hugging Face). +- **Your endpoint**: any OpenAI-compatible URL (vLLM, SGLang, ...) or an Anthropic endpoint, probed + before use. Training capture needs vLLM with `--return-tokens-as-token-ids`. + +A rollout runs on the server whether or not the page stays open. Model usage is billed to the +account or token selected on the card; sandbox compute is billed to the server operator, including +Hugging Face Sandbox on a Space. + +**Runs.** Every run you may see, filterable by status. A run shows what the agent did as a timeline +(its prompt, thinking, each tool call with its input and output, terminal sessions, the final +answer), the result and reward, the trace checks, and downloads of the result JSON and, for a +trainable rollout, the training contract. Tick two to four runs to compare them. + +**Setup.** The server's endpoint, which sandboxes have working credentials, the deployment settings +below, the datasets and the agents. + +### Deployment settings + +The same UI runs on a laptop and as a public Space, and what a visitor may do differs. The defaults +depend on where the server runs: **local** is `serve --host 127.0.0.1`; **network** is any other bind +address, including the default `0.0.0.0`, since everyone who can reach it is then a visitor; a +**Space** is detected from `SPACE_ID`. Each setting is an environment variable (a Space variable on +a Space), and some have a flag on `serve` and `push` (`--private-urls` is on `serve` only). + +| setting | variable | flag | local | network | Space | +|---|---|---|---|---|---| +| Visitors may start rollouts | `OPENENV_HARBOR_UI_ROLLOUTS` | `--rollouts` | on | on | on | +| Visitors may use the server's endpoint | `OPENENV_HARBOR_UI_SERVER_ENDPOINT` | `--share-endpoint` | on | on | off | +| Visitors may connect their own | `OPENENV_HARBOR_UI_VISITOR_ENDPOINTS` | `--visitor-endpoints` | on | on | on | +| A visitor's URL may be private or local | `OPENENV_HARBOR_UI_PRIVATE_URLS` | `--private-urls` | on | off | off | +| Offer this machine's HF token | `OPENENV_HARBOR_UI_LOCAL_TOKEN` | | on | off | never | +| Add and remove Hub datasets from the page | `OPENENV_HARBOR_UI_ADD_DATASETS` | `--add-datasets` | on | off | off | +| Who sees runs (`all` or `own`) | `OPENENV_HARBOR_RUN_VISIBILITY` | `--run-visibility` | all | all | own | +| Keep finished runs across restarts | `OPENENV_HARBOR_RUN_HISTORY` | `--run-history` | on | on | off | +| Headline reward of a task with several | `OPENENV_HARBOR_REWARD_KEY` | `--reward-key` | unset | unset | unset | +| Rollouts at once | `OPENENV_HARBOR_UI_MAX_RUNS` | | 4 | 4 | 4 | +| Rollouts at once per visitor | `OPENENV_HARBOR_UI_MAX_RUNS_PER_VISITOR` | | 4 | 4 | 2 | + +`own` visibility ties runs to a random id kept in the visitor's browser; a run stores only a digest +of it. That keeps visitors' runs apart, but it is not sign-in: a visitor who clears the browser's +storage loses their runs, and it is no substitute for access control over traces you consider +private. The per-visitor limit counts a signed-in visitor's Hugging Face account, and otherwise that +browser id, so for anonymous visitors it stops one page from taking every slot, not someone set on +it; `OPENENV_HARBOR_UI_MAX_RUNS` bounds the total. Run history goes to `OPENENV_HARBOR_RUNS_DIR`, by default `~/.cache/openenv/harbor/runs`, or +`/data/harbor-runs` on a Space with the bucket mounted. `OPENENV_HARBOR_UI_MAX_ADD_GB` (default `5`) +caps the size of a dataset added from the page. + +What the UI guarantees whatever the settings: + +- A key or token typed into the page is sent to that endpoint by the server, held in server memory + for that page only, and never written to disk, to run history or back to the page. +- With private URLs off, a visitor's URL, and every redirect it answers with, must resolve to public + addresses, and is checked again before each rollout. +- Every dataset, task and file the page asks for is one the server serves or that was added from + the page. Every file it reads, for a card, the task view, the file viewer or the environment + check, must resolve inside the dataset's own folder (a link to that folder itself is followed), + so a task reached through a link to anywhere else shows nothing and cannot run on a visitor's + model. The Task API's instruction preview follows the same rule, and a registry dataset is + anchored on Harbor's cache. A dataset added from the page may not contain a symbolic link at all. +- A task that reads the server's environment variables (`${VAR}` in `task.toml` or a compose file, + or a bare name under a compose `environment:`), which is where the server's keys are, or its files + (a compose `env_file`, `include`, `extends`, or a host path in a mount, secret, build context, + cache, watch rule or device), or asks for more of the host than a folder (`privileged`, + `cap_add`, the host's namespaces, the Docker socket, another container's volumes, a named volume + or network with settings, the build's SSH agent) runs only on the server's own endpoint, never on + a model a visitor connects: that model does what the visitor asks, printing the sandbox's + environment included, into a trace the visitor reads. A dataset added from the page may not do + either at all. Both files are checked as parsed, and one that doesn't parse counts as reading. +- A request that changes something in the UI (a rollout, an added dataset) is refused when a + browser sends it from another site, so another page can't act through a visitor's browser. Behind + a proxy that rewrites `Host`, list the public host in `OPENENV_HARBOR_UI_HOSTS` (comma-separated). +- Everything a model, a task or a tool produced is escaped before it is shown. + ## CLI reference Four commands. Every flag below is the complete set, with its type and default. @@ -348,16 +440,26 @@ Start the env server: Task API for discovery, one long-running `run_rollout` MCP | flag | type | default | meaning | |---|---|---|---| -| `--llm-url` | str | **required** | OpenAI-spec endpoint | +| `--llm-url` | str | `""` | OpenAI-spec endpoint. Optional: without one, UI visitors connect their own | | `--dataset` | str | none | Dataset specs to serve as splits. Repeatable | | `--model` | str | `""` | Served model id | -| `--host` | str | `0.0.0.0` | Bind address | +| `--host` | str | `0.0.0.0` | Bind address. `127.0.0.1` gives the UI its local defaults, see [Deployment settings](#deployment-settings) | | `--port` | int | `8000` | Env server port. Faces the trainer and the browser | | `--capture-port` | int | `8100` | Capture proxy port. Faces the sandbox | | `--expose` | str | `gradio` | How the sandbox reaches the proxy | | `--env-file` | path | `""` | dotenv with provider credentials | | `--api-key` | str | `$OPENENV_LLM_API_KEY` | Credential for the endpoint, for a hosted provider | | `--auth-header` | str | `Authorization` | Header to send it under, e.g. `x-api-key` | +| `--share-endpoint/--no-share-endpoint` | flag | unset | UI visitors may use this endpoint and its key | +| `--visitor-endpoints/--no-visitor-endpoints` | flag | unset | UI visitors may connect their own model | +| `--run-visibility` | str | unset | `all` or `own`: who sees runs in the UI | +| `--run-history/--no-run-history` | flag | unset | Keep finished UI runs across restarts | +| `--add-datasets/--no-add-datasets` | flag | unset | UI visitors may add and remove Hub datasets | +| `--rollouts/--no-rollouts` | flag | unset | UI visitors may start rollouts | +| `--private-urls/--no-private-urls` | flag | unset | UI visitors may connect an endpoint on a private or local address | +| `--reward-key` | str | unset | The reward UI runs report for tasks with several and none named `reward`. Unset, such a run lists each one | + +An unset UI flag keeps the default for where the server runs. Refuses to start only if the endpoint is unreachable. One that cannot return token ids starts as an eval deployment and says so. @@ -367,17 +469,25 @@ Deploy the same server to a Hugging Face Space. | flag | type | default | meaning | |---|---|---|---| -| `--llm-url` | str | **required** | Endpoint the deployed Space will use | +| `--llm-url` | str | `""` | Endpoint the deployed Space will use. Optional: without one, visitors connect their own | | `--repo-id` | str | **required** | Target Space, e.g. `you/harbor-env` | | `--dataset` | str | none | Dataset specs. Repeatable | | `--model` | str | `""` | Served model id | | `--bucket` | str | Space name | Storage bucket holding the task suites. `none` disables the mount and downloads instead | +| `--public-bucket/--private-bucket` | flag | private | Visibility of a new bucket. Given for an existing bucket, it changes it; not given, an existing bucket keeps its visibility | +| `--hf-login/--no-hf-login` | flag | on | Let visitors sign in with Hugging Face to use their own account for Inference Providers | | `--hardware` | str | `""` | Space hardware, e.g. `cpu-basic` | -| `--private` | flag | off | Create it private. Rollouts then cannot work, see below | +| `--private` | flag | off | Create the Space private. Rollouts then cannot work, see below | | `--recreate` | flag | off | Delete the Space first, then deploy fresh | | `--dry-run` | flag | off | Print exactly what would be sent and stop | | `--env-file` | path | `""` | dotenv whose provider keys become Space **secrets** | +`push` also takes the UI flags of `serve` (`--share-endpoint`, `--visitor-endpoints`, +`--run-visibility`, `--run-history`, `--add-datasets`, `--rollouts`, `--reward-key`) and sets them as Space +variables. Unlike `serve`, it doesn't share the endpoint with UI visitors unless `--share-endpoint` is +given. Each push reconciles these UI variables: an omitted flag removes an earlier override and +returns to the Space default, so a one-time `--share-endpoint` does not stay enabled forever. + ```bash openenv harbor push --llm-url $LLM --dataset org/train,org/eval \ --repo-id you/harbor-env --env-file .env --dry-run @@ -387,6 +497,18 @@ openenv harbor push --llm-url $LLM --dataset org/train,org/eval \ at `/capture`, and a private Space requires an auth header the sandboxed agent does not send. Use it only to park a deployment. +### Migration notes for the new UI + +- `serve` binds `0.0.0.0` by default and therefore uses the network-safe UI defaults. Visitor + endpoints such as `http://localhost:8000/v1` are refused. Bind the whole server to + `--host 127.0.0.1`, or pass `--private-urls` when the server still needs a remote trainer. +- An existing Space re-pushed with `--llm-url` no longer shares that endpoint with UI visitors + unless `--share-endpoint` is passed. `--hf-login` also writes `hf_oauth: true` and the + `inference-api` OAuth scope into the Space README. +- Capture `/health` still returns status and aggregate fields publicly, but its per-session + `upstreams` list is empty when an admin key is configured. Set + `OPENENV_CAPTURE_ADMIN_KEY` and send it as a bearer token to inspect that field. + ## Supported harnesses 16 of the 30 known agents are validated end to end. "Validated" means a real rollout produced token @@ -514,7 +636,25 @@ Two details that matter: **Task suites are mounted, not downloaded.** A Harbor suite is thousands of small files and Space disk is ephemeral, so a download is re-paid on every restart. `push` syncs the suites into a storage bucket named after the Space and mounts it at `/data`. The copy is server side, and re-running -`push` copies only what is new. +`push` copies only what is new. The bucket is created private (`--public-bucket` to change that), +since with run history on it also holds every visitor's runs. With `--add-datasets`, the `tasks/` +folder of a dataset added from the UI is copied into the same bucket, server side, and read through +the mount, so it survives restarts; removing it deletes it from the bucket. + +**Visitors can sign in with Hugging Face.** `push` turns on OAuth for the Space (`hf_oauth: true`, +scope `inference-api`), and the UI offers "sign in with Hugging Face" as a way to use Inference +Providers on the visitor's own account instead of pasting a token. It is optional and not needed +for anything else; `--no-hf-login` leaves it out. + +**Visitors bring their own model.** On a Space, a visitor runs on a model they connect: signing in +with Hugging Face, a token, or their own endpoint. `--share-endpoint` lets them use the Space's own +endpoint and key too, at your cost; the Task API and MCP use it either way. On that endpoint a +served task that reads the Space's environment does run, and the visitor who starts it reads its +trace, so don't combine `--share-endpoint` with such tasks on a public Space. + +**Rollouts on a Space use the Space's credentials for sandboxes.** An `hf-sandbox` rollout is billed +to the `HF_TOKEN` the Space holds, whoever started it. `--no-rollouts` makes a public deployment a +read-only task browser. **The Space must be public.** The capture proxy is served at `/capture`, and a private Space requires an auth header that the agent inside the sandbox does not send. This is safe because @@ -550,6 +690,14 @@ unaffected. **Many roots for one rollout.** Normal for agents that run subagents or auxiliary calls. Each root is a separate conversation, and only agent conversations are counted as trainable. +**A dataset is greyed out in "Add from the Hub".** It has no `tasks//` folder, which is how +Harbor datasets are laid out, or it is over `OPENENV_HARBOR_UI_MAX_ADD_GB`. + +**A URL is refused as "private or local".** The server does not call private addresses for visitors +unless `OPENENV_HARBOR_UI_PRIVATE_URLS` is on, which by default it is only for `--host 127.0.0.1`; +`serve --private-urls` turns it on for any bind address. To reach a vLLM on your own network from a +Space, run the UI yourself. + **Exit code 137.** The agent was killed inside the sandbox, almost always by the OOM killer on a large input. That is a task failure, not a capture failure. diff --git a/envs/harbor_env/pyproject.toml b/envs/harbor_env/pyproject.toml index 3c843640e..69d6275cc 100644 --- a/envs/harbor_env/pyproject.toml +++ b/envs/harbor_env/pyproject.toml @@ -11,11 +11,14 @@ dependencies = [ # reports the pair as unsatisfiable and the image build fails. Neither is a sandbox backend # we offer, so both are dropped and everything else kept. Re-check on a Harbor upgrade. "harbor[e2b,modal,daytona,gke,ec2,runloop,novita,blaxel,beam,islo,opensandbox,cwsandbox,use-computer,cua]>=0.22.0", - "huggingface_hub>=1.12", + "huggingface_hub>=1.29.0", "fastapi>=0.104", "uvicorn[standard]>=0.24", "httpx>=0.27", "gradio>=5", + # "Sign in with Hugging Face" on a Space (Gradio's OAuth: the `gradio[oauth]` extra's two). + "authlib>=1.3", + "itsdangerous>=2.1", ] [project.scripts] diff --git a/envs/harbor_env/uv.lock b/envs/harbor_env/uv.lock index 6c41f39d7..70c70164b 100644 --- a/envs/harbor_env/uv.lock +++ b/envs/harbor_env/uv.lock @@ -1918,6 +1918,15 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/c0/b5/a8ecdea027ccc20f5ab44042c8629cd22588ca7c33cc199c5f103f92ead4/islo-0.3.19-py3-none-any.whl", hash = "sha256:ab7e9bab5678e486e0bdfa8ed25e40fe2eef95c5748c3d666b5c6b1226354126", size = 372082, upload-time = "2026-08-27T12:48:18.582Z" }, ] +[[package]] +name = "itsdangerous" +version = "2.2.0" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/9c/cb/8ac0172223afbccb63986cc25049b154ecfb5e85932587206f42317be31d/itsdangerous-2.2.0.tar.gz", hash = "sha256:e0050c0b7da1eea53ffaf149c0cfbb5c6e2e2b69c4bef22c81fa6eb73e5f6173", size = 54410, upload-time = "2024-04-16T21:28:15.614Z" } +wheels = [ + { url = "https://files.pythonhosted.org/packages/04/96/92447566d16df59b2a776c0fb82dbc4d9e07cd95062562af01e408583fc4/itsdangerous-2.2.0-py3-none-any.whl", hash = "sha256:c6242fc49e35958c8b15141343aa660db5fc54d4f13a1db01a3f5891b98700ef", size = 16234, upload-time = "2024-04-16T21:28:14.499Z" }, +] + [[package]] name = "jaraco-classes" version = "3.4.0" @@ -2692,22 +2701,26 @@ name = "openenv-harbor-env" version = "0.1.0" source = { editable = "." } dependencies = [ + { name = "authlib" }, { name = "fastapi" }, { name = "gradio" }, { name = "harbor", extra = ["beam", "blaxel", "cua", "cwsandbox", "daytona", "e2b", "ec2", "gke", "islo", "modal", "novita", "opensandbox", "runloop", "use-computer"] }, { name = "httpx" }, { name = "huggingface-hub" }, + { name = "itsdangerous" }, { name = "openenv" }, { name = "uvicorn", extra = ["standard"] }, ] [package.metadata] requires-dist = [ + { name = "authlib", specifier = ">=1.3" }, { name = "fastapi", specifier = ">=0.104" }, { name = "gradio", specifier = ">=5" }, { name = "harbor", extras = ["e2b", "modal", "daytona", "gke", "ec2", "runloop", "novita", "blaxel", "beam", "islo", "opensandbox", "cwsandbox", "use-computer", "cua"], specifier = ">=0.22.0" }, { name = "httpx", specifier = ">=0.27" }, - { name = "huggingface-hub", specifier = ">=1.12" }, + { name = "huggingface-hub", specifier = ">=1.29.0" }, + { name = "itsdangerous", specifier = ">=2.1" }, { name = "openenv" }, { name = "uvicorn", extras = ["standard"], specifier = ">=0.24" }, ] @@ -4022,8 +4035,8 @@ name = "secretstorage" version = "3.5.0" source = { registry = "https://pypi.org/simple" } dependencies = [ - { name = "cryptography" }, - { name = "jeepney" }, + { name = "cryptography", marker = "sys_platform != 'emscripten' and sys_platform != 'win32'" }, + { name = "jeepney", marker = "sys_platform != 'emscripten' and sys_platform != 'win32'" }, ] sdist = { url = "https://files.pythonhosted.org/packages/1c/03/e834bcd866f2f8a49a85eaff47340affa3bfa391ee9912a952a1faa68c7b/secretstorage-3.5.0.tar.gz", hash = "sha256:f04b8e4689cbce351744d5537bf6b1329c6fc68f91fa666f60a380edddcd11be", size = 19884, upload-time = "2025-11-23T19:02:53.191Z" } wheels = [ diff --git a/pyproject.toml b/pyproject.toml index d69d73be4..739085ad5 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -24,7 +24,7 @@ dependencies = [ "typer>=0.9.0", "rich>=13.0.0", "pyyaml>=6.0", - "huggingface_hub>=0.20.0", + "huggingface_hub>=1.29.0", "openai>=2.7.2", "tomli>=2.3.0", "tomli-w>=1.2.0", @@ -83,6 +83,7 @@ include-package-data = true "openenv.cli.importers" = ["templates/**/*"] "openenv.validation" = ["policies/*.json", "schemas/*.json"] "openenv.discovery" = ["schemas/**/*.json"] +"openenv.harbor" = ["ui_assets/*"] [tool.setuptools.exclude-package-data] "*" = ["*.pyc", "*.pyo", "__pycache__/*"] diff --git a/src/openenv/cli/commands/harbor.py b/src/openenv/cli/commands/harbor.py index b9e586414..8d509d4f6 100644 --- a/src/openenv/cli/commands/harbor.py +++ b/src/openenv/cli/commands/harbor.py @@ -56,6 +56,10 @@ "OpenAI-spec inference endpoint. Optional here: without it, `info` " "still reports sandboxes, datasets and harnesses." ) +_LLM_HELP_PUSH = ( + "The Space's own inference endpoint. Optional: without one, visitors connect their own model " + "in the UI (see --visitor-endpoints)." +) _KEY_HELP = ( "Credential for the inference endpoint, for a hosted provider (OpenAI, Anthropic, HF " "Inference Providers). Defaults to $OPENENV_LLM_API_KEY. This is NOT the key the agent " @@ -68,6 +72,105 @@ ) +_SHARE_HELP = ( + "Let visitors of the UI run on this server's endpoint and key. --no-share-endpoint makes each " + "visitor connect their own model. Default: shared." +) +_SHARE_HELP_PUSH = ( + "Let visitors of the Space's UI run on its endpoint and key, at your cost. Default: not shared, " + "so each visitor connects their own model (signing in with Hugging Face, a token, or an endpoint)." +) +_VISITOR_HELP = ( + "Let visitors connect their own model in the UI: a Hugging Face token and model, or any " + "OpenAI-compatible URL such as vLLM. Default: allowed." +) +_VISIBILITY_HELP = ( + "Who sees runs in the UI: `all` (every visitor sees every run) or `own` (each browser sees the " + "runs it started). Default: all locally, own on a Space." +) +_HISTORY_HELP = "Keep finished UI runs on disk across restarts. Default: on locally, off on a Space." +_ROLLOUTS_HELP = ( + "Let visitors start rollouts from the UI. --no-rollouts makes it a read-only task browser. " + "Default: on." +) +_PRIVATE_URLS_HELP = ( + "Let visitors connect an endpoint on a private or local address, such as http://localhost:8000/v1. " + "Default: allowed only when the server listens on 127.0.0.1, since --host 0.0.0.0 makes this " + "server call those addresses for anyone who can reach it." +) +_ADD_HELP = ( + "Let visitors add and remove Hub datasets from the UI. On a Space with a bucket they are copied " + "into it; otherwise downloaded. Default: on only for a server on 127.0.0.1." +) +_UI_VARIABLES = ( + "OPENENV_HARBOR_UI_SERVER_ENDPOINT", + "OPENENV_HARBOR_UI_VISITOR_ENDPOINTS", + "OPENENV_HARBOR_RUN_HISTORY", + "OPENENV_HARBOR_UI_ADD_DATASETS", + "OPENENV_HARBOR_UI_ROLLOUTS", + "OPENENV_HARBOR_RUN_VISIBILITY", + "OPENENV_HARBOR_REWARD_KEY", +) +_REWARD_KEY_HELP = ( + "Which reward UI runs report for tasks with several and none named `reward`, or a " + "comma-separated preference order. Unset, such a run shows each of them instead." +) + + +def _ui_env( + share_endpoint: Optional[bool], + visitor_endpoints: Optional[bool], + run_visibility: str, + run_history: Optional[bool], + add_datasets: Optional[bool] = None, + rollouts: Optional[bool] = None, + private_urls: Optional[bool] = None, + reward_key: str = "", +) -> dict[str, str]: + """The UI's deployment settings as the variables `openenv.harbor.ui_settings` reads. + + Only flags that were given become variables, so an unset one keeps the default for where the + server runs (a laptop or a Space). + """ + if run_visibility and run_visibility not in ("all", "own"): + raise typer.BadParameter("--run-visibility is `all` or `own`") + out = {} + for name, value in ( + ("OPENENV_HARBOR_UI_SERVER_ENDPOINT", share_endpoint), + ("OPENENV_HARBOR_UI_VISITOR_ENDPOINTS", visitor_endpoints), + ("OPENENV_HARBOR_RUN_HISTORY", run_history), + ("OPENENV_HARBOR_UI_ADD_DATASETS", add_datasets), + ("OPENENV_HARBOR_UI_ROLLOUTS", rollouts), + # `serve` only, so not one of `_UI_VARIABLES`: on a Space these addresses are its own network + ("OPENENV_HARBOR_UI_PRIVATE_URLS", private_urls), + ): + if value is not None: + out[name] = "1" if value else "0" + if run_visibility: + out["OPENENV_HARBOR_RUN_VISIBILITY"] = run_visibility + if reward_key.strip(): + out["OPENENV_HARBOR_REWARD_KEY"] = reward_key.strip() + return out + + +def _remove_omitted_ui_variables(repo_id: str, supplied: dict[str, str]) -> None: + """Reset omitted UI flags to deployment defaults on an incremental Space push. + + Space variables survive a push. Without this reconciliation, one old `--share-endpoint` remains + enabled forever even when later pushes omit it and the CLI says the Space default is not shared. + Only variables owned by this command are removed; unrelated operator configuration and every + secret are left alone. + """ + from huggingface_hub import HfApi + + api = HfApi() + existing = api.get_space_variables(repo_id) + for key in _UI_VARIABLES: + if key in existing and key not in supplied: + api.delete_space_variable(repo_id=repo_id, key=key) + print(f"variable {key} reset to its Space default") + + def _split(values: Optional[list[str]]) -> list[str]: """Accept both `--dataset a --dataset b` and `--dataset a,b`.""" out: list[str] = [] @@ -76,6 +179,14 @@ def _split(values: Optional[list[str]]) -> list[str]: return out +def _hub_not_found(exc: Exception) -> bool: + """Whether a Hub operation failed specifically because its resource does not exist.""" + if isinstance(exc, FileNotFoundError): + return True + response = getattr(exc, "response", None) + return getattr(response, "status_code", None) == 404 + + @app.command("info") def info( llm_url: Annotated[str, typer.Option("--llm-url", help=_LLM_HELP_OPTIONAL)] = "", @@ -254,6 +365,36 @@ def serve( str, typer.Option("--auth-header", help=_AUTH_HEADER_HELP) ] = "Authorization", env_file: Annotated[str, typer.Option("--env-file")] = "", + share_endpoint: Annotated[ + Optional[bool], + typer.Option("--share-endpoint/--no-share-endpoint", help=_SHARE_HELP), + ] = None, + visitor_endpoints: Annotated[ + Optional[bool], + typer.Option("--visitor-endpoints/--no-visitor-endpoints", help=_VISITOR_HELP), + ] = None, + run_visibility: Annotated[ + str, typer.Option("--run-visibility", help=_VISIBILITY_HELP) + ] = "", + run_history: Annotated[ + Optional[bool], + typer.Option("--run-history/--no-run-history", help=_HISTORY_HELP), + ] = None, + add_datasets: Annotated[ + Optional[bool], + typer.Option("--add-datasets/--no-add-datasets", help=_ADD_HELP), + ] = None, + rollouts: Annotated[ + Optional[bool], + typer.Option("--rollouts/--no-rollouts", help=_ROLLOUTS_HELP), + ] = None, + private_urls: Annotated[ + Optional[bool], + typer.Option("--private-urls/--no-private-urls", help=_PRIVATE_URLS_HELP), + ] = None, + reward_key: Annotated[ + str, typer.Option("--reward-key", help=_REWARD_KEY_HELP) + ] = "", ) -> None: """Serve Harbor tasks over the OpenEnv Task API and MCP. @@ -263,6 +404,20 @@ def serve( """ from openenv.harbor.serving import serve_harbor + os.environ.update( + _ui_env( + share_endpoint, + visitor_endpoints, + run_visibility, + run_history, + add_datasets, + rollouts, + private_urls, + reward_key, + ) + ) + # The UI's laptop defaults (this machine's token, private URLs) are for a loopback-only server. + os.environ["OPENENV_HARBOR_UI_HOST"] = host serve_harbor( llm_url=llm_url, model=model or None, @@ -280,7 +435,7 @@ def serve( @app.command("push") def push( - llm_url: Annotated[str, typer.Option("--llm-url", help=_LLM_HELP)], + llm_url: Annotated[str, typer.Option("--llm-url", help=_LLM_HELP_PUSH)] = "", repo_id: Annotated[ str, typer.Option("--repo-id", help="Target, e.g. your-org/harbor-env.") ] = "", @@ -315,6 +470,23 @@ def push( help="Storage bucket holding the task suites. Defaults to a bucket named after the Space. Pass `none` to skip the bucket and let the Space download datasets instead.", ), ] = "", + public_bucket: Annotated[ + Optional[bool], + typer.Option( + "--public-bucket/--private-bucket", + help="Visibility of the Space's bucket. Private unless --public-bucket: it holds the task " + "suites and, with run history on, every visitor's runs. Given for an existing bucket, it " + "changes that bucket's visibility; not given, an existing bucket is left as it is.", + ), + ] = None, + hf_login: Annotated[ + bool, + typer.Option( + "--hf-login/--no-hf-login", + help="Let visitors sign in with Hugging Face to use their own account for Inference " + "Providers (the `inference-api` scope), instead of pasting a token.", + ), + ] = True, recreate: Annotated[ bool, typer.Option( @@ -325,6 +497,32 @@ def push( dry_run: Annotated[ bool, typer.Option("--dry-run", help="Show what would be pushed and stop.") ] = False, + share_endpoint: Annotated[ + Optional[bool], + typer.Option("--share-endpoint/--no-share-endpoint", help=_SHARE_HELP_PUSH), + ] = None, + visitor_endpoints: Annotated[ + Optional[bool], + typer.Option("--visitor-endpoints/--no-visitor-endpoints", help=_VISITOR_HELP), + ] = None, + run_visibility: Annotated[ + str, typer.Option("--run-visibility", help=_VISIBILITY_HELP) + ] = "", + run_history: Annotated[ + Optional[bool], + typer.Option("--run-history/--no-run-history", help=_HISTORY_HELP), + ] = None, + add_datasets: Annotated[ + Optional[bool], + typer.Option("--add-datasets/--no-add-datasets", help=_ADD_HELP), + ] = None, + rollouts: Annotated[ + Optional[bool], + typer.Option("--rollouts/--no-rollouts", help=_ROLLOUTS_HELP), + ] = None, + reward_key: Annotated[ + str, typer.Option("--reward-key", help=_REWARD_KEY_HELP) + ] = "", ) -> None: """Deploy this environment to a Hugging Face Space. @@ -341,10 +539,36 @@ def push( raise typer.BadParameter("--repo-id is required, e.g. your-org/harbor-env") datasets = _split(dataset) - if not llm_url: + ui_variables = _ui_env( + share_endpoint, + visitor_endpoints, + run_visibility, + run_history, + add_datasets, + rollouts, + reward_key=reward_key, + ) + if not share_endpoint and visitor_endpoints is False: + # On a Space the endpoint is shared only when asked, so unset counts as not shared. + raise typer.BadParameter( + "--no-visitor-endpoints needs --share-endpoint: on a Space visitors use its endpoint " + "only when it is shared, so they would have no model to run with." + ) + if not llm_url and visitor_endpoints is False: raise typer.BadParameter( - "--llm-url is required: a Space with no engine cannot run anything, and finding " - "that out after deploying is worse than finding it out now." + "--llm-url is required with --no-visitor-endpoints: a Space with no engine of its own " + "that accepts none from visitors cannot run anything, and finding that out after " + "deploying is worse than finding it out now." + ) + if not llm_url: + print( + "NOTE: no --llm-url. Visitors run rollouts on a model they connect themselves (a " + "Hugging Face token, or their own endpoint)." + ) + elif share_endpoint is None: + print( + "NOTE: UI visitors connect their own model; the endpoint serves the Task API and MCP. " + "Pass --share-endpoint to let them run on it too, at your cost." ) if private: @@ -372,7 +596,16 @@ def push( mounts = {spec: f"{_MOUNT_ROOT}/{spec.replace('/', '__')}" for spec in hf_specs} # Non-secret configuration travels as plain Space variables. - variables = {"OPENENV_LLM_URL": llm_url, "ENABLE_WEB_INTERFACE": "true"} + variables = { + "OPENENV_LLM_URL": llm_url, + "ENABLE_WEB_INTERFACE": "true", + **ui_variables, + } + if bucket: + # Where the UI puts datasets added from the page: a server-side copy into this bucket, + # read back through the mount, as for the datasets pushed here. + variables["OPENENV_HARBOR_BUCKET"] = bucket + variables["OPENENV_HARBOR_BUCKET_MOUNT"] = _MOUNT_ROOT if datasets: variables["OPENENV_DATASETS"] = ",".join(datasets) if model: @@ -447,8 +680,9 @@ def push( if recreate: _delete_space(repo_id) - if bucket and hf_specs: - _fill_bucket(bucket, hf_specs) + if bucket: + # Created even with no datasets to copy: datasets added from the page go into it too. + _fill_bucket(bucket, hf_specs, public=public_bucket) with tempfile.TemporaryDirectory(prefix="openenv-harbor-push-") as tmp: staged = Path(tmp) / "env" @@ -470,6 +704,8 @@ def push( "__pycache__", "*.pyc", "forwarding.py", "cli" ), ) + if hf_login: + _enable_hf_login(staged / "README.md") _prune_removed_files(repo_id, staged) _push( directory=str(staged), @@ -479,6 +715,7 @@ def push( env_vars=[f"{k}={v}" for k, v in variables.items()], secrets=[f"{k}={v}" for k, v in secrets.items()], ) + _remove_omitted_ui_variables(repo_id, ui_variables) # After the push, because volumes attach to a Space that already exists and `--recreate` has # just deleted it. Setting them triggers one more rebuild, which is why this is last. @@ -501,6 +738,33 @@ def push( print("mount OPENENV_DATASETS switched to mount paths") +# What visitors may sign in for: Inference Providers, billed to their own account. Nothing else. +_HF_LOGIN = ( + "hf_oauth: true", + "hf_oauth_scopes:", + " - inference-api", + "hf_oauth_expiration_minutes: 480", +) + + +def _enable_hf_login(readme: Path) -> None: + """Turn on "Sign in with Hugging Face" for the Space, in its README front matter. + + The Hub then provides the OAuth app (`OAUTH_CLIENT_ID` and friends), and the UI offers sign-in + as a way to use Inference Providers with the visitor's own account. + """ + text = readme.read_text() if readme.is_file() else "" + if text.startswith("---"): + end = text.index("\n---", 3) + # Only the front matter counts: the README's own text may well mention `hf_oauth:`. + if any(line.startswith("hf_oauth:") for line in text[:end].splitlines()): + return + text = text[:end] + "\n" + "\n".join(_HF_LOGIN) + text[end:] + else: + text = "---\n" + "\n".join(_HF_LOGIN) + "\n---\n\n" + text + readme.write_text(text) + + def _prune_removed_files(repo_id: str, staged: Path) -> None: """Delete files on the Space that this push no longer produces. @@ -636,9 +900,12 @@ def _delete_space(repo_id: str) -> None: print(f"recreate nothing to delete ({type(exc).__name__})") -def _fill_bucket(bucket: str, specs: list[str]) -> None: +def _fill_bucket(bucket: str, specs: list[str], public: bool | None = None) -> None: """Create `bucket` if missing and copy each task suite into it, server side. + A new bucket is private unless `public` is `True`. An existing one keeps its visibility unless + `public` says otherwise: changing who can read a bucket is not something to do by default. + `copy_files` copies by xet hash: the Hub moves the references, nothing is downloaded here and nothing is re-uploaded. That is the difference between seconds and the ~47k-file upload a local sync performs, and it is why the bucket is filled before the Space exists rather than after. @@ -649,7 +916,25 @@ def _fill_bucket(bucket: str, specs: list[str]) -> None: from huggingface_hub import HfApi api = HfApi() - api.create_bucket(bucket, private=False, exist_ok=True) + try: + existing = api.bucket_info(bucket) + except Exception as exc: # noqa: BLE001 - Hub exception classes vary across client versions + if not _hub_not_found(exc): + raise + existing = None + if existing is None: + api.create_bucket(bucket, private=not public, exist_ok=True) + print(f"bucket {bucket} created ({'public' if public else 'private'})") + else: + private = bool(existing.private) + if public is not None and private == public: + api.update_bucket_settings(bucket, private=not public) + print(f"bucket {bucket} is now {'public' if public else 'private'}") + elif not private and public is None: + print( + f"bucket {bucket} is PUBLIC. Anyone can read it, including run history if that is " + "on. Pass --private-bucket to make it private." + ) try: present = { diff --git a/src/openenv/core/harness/capture/server.py b/src/openenv/core/harness/capture/server.py index 8aa30ca9e..32b9dad48 100644 --- a/src/openenv/core/harness/capture/server.py +++ b/src/openenv/core/harness/capture/server.py @@ -424,6 +424,16 @@ def known(self) -> list[dict[str, Any]]: ) in self._by_engine.items() ] + def forget(self, upstream: Upstream) -> None: + """Drop the cached client for `upstream`, and with it the credential it holds. + + Every proxied call resolves its session's engine here, so the next call naming this engine + probes it again. Call it only once no session uses the engine: for a caller whose key should + not outlive their use of it. + """ + with self._lock: + self._by_engine.pop(upstream.cache_key, None) + async def resolve(self, upstream: Upstream) -> tuple[InferenceClient, str]: """Client and measured capture level for `upstream`, probing once per engine.""" key = upstream.cache_key @@ -655,7 +665,7 @@ def _level_of(session) -> str: return session.capture_level or app.state.capture_level @app.get("/health") - async def health() -> dict[str, Any]: + async def health(request: Request) -> dict[str, Any]: return { "status": "ok", "instance": app.state.instance_id, @@ -669,9 +679,10 @@ async def health() -> dict[str, Any]: # readable from an endpoint that, on a Space, is public. "capture_level": app.state.capture_level, "rollout_type": "train" if app.state.capture_level == "tokens" else "eval", - # Engines named per session and already probed. A caller can see what this server - # measured without minting a session to find out. - "upstreams": app.state.upstreams.known(), + # Engines named per session and already probed, so a caller can see what this server + # measured without minting a session. Behind the admin key when one is set: on a public + # Space these are other callers' endpoints, private tunnel URLs among them. + "upstreams": app.state.upstreams.known() if _admin_ok(request) else [], "upstream_auth": bool(app.state.inference and app.state.inference.api_key), "param_fixes": ( [str(f) for f in app.state.inference.param_fixes] diff --git a/src/openenv/harbor/capabilities.py b/src/openenv/harbor/capabilities.py index 6fb2c156c..2c0981d7b 100644 --- a/src/openenv/harbor/capabilities.py +++ b/src/openenv/harbor/capabilities.py @@ -28,7 +28,7 @@ # Backends worth advertising. Harbor registers 23; these are the ones with a credential story we # check and have exercised. Others still work via `--sandbox `, just unadvertised. -KNOWN_SANDBOXES = ("docker", "e2b", "modal", "daytona") +KNOWN_SANDBOXES = ("docker", "e2b", "modal", "daytona", "hf-sandbox") @dataclass diff --git a/src/openenv/harbor/rollout.py b/src/openenv/harbor/rollout.py index afedf0ed4..77422b828 100644 --- a/src/openenv/harbor/rollout.py +++ b/src/openenv/harbor/rollout.py @@ -286,6 +286,7 @@ async def run_rollout( trials_dir: Path | str, dataset: str = "", reward_key: str = "", + require_reward: bool = True, keep_sandbox: bool = False, agent_timeout_sec: float | None = None, agent_step_limit: int | None = None, @@ -319,6 +320,9 @@ async def run_rollout( Where Harbor writes trial artifacts. reward_key (`str`, *optional*): Which reward key is the training signal, for multi-reward tasks. + require_reward (`bool`, *optional*, defaults to `True`): + Fail the rollout when the verifier's rewards leave no headline one. `False` keeps it, + with every reward in `rewards` and a warning in `findings`: for a rollout to look at. keep_sandbox (`bool`, *optional*, defaults to `False`): Leave the sandbox alive after the run, for debugging. capture_level (`str`, *optional*, defaults to `"tokens"`): @@ -543,8 +547,12 @@ async def run_rollout( result.rewards, reward_key ) except ValueError as exc: - result.ok = False - result.error = str(exc) + if require_reward: + result.ok = False + result.error = str(exc) + else: + # A rollout to look at, not to train on: every reward stays in `rewards`. + result.findings.append(f"[WARN] no headline reward: {exc}") for step in getattr(trial_result, "step_results", None) or []: step_rewards = dict( getattr(getattr(step, "verifier_result", None), "rewards", None) @@ -715,7 +723,12 @@ async def run_rollout( # A rollout that produced no reward is not a zero: the verifier never ran. Keeping the two # distinct is what stops a dead sandbox being scored as a wrong answer. - if result.ok and result.reward is None and trial_result is not None: + if ( + result.ok + and result.reward is None + and not result.rewards # several, none the headline, is a warning of its own above + and trial_result is not None + ): result.findings.append( "[WARN] ungraded: the verifier produced no reward for this trial" ) diff --git a/src/openenv/harbor/serving.py b/src/openenv/harbor/serving.py index 6c9cb2145..9a15a36ba 100644 --- a/src/openenv/harbor/serving.py +++ b/src/openenv/harbor/serving.py @@ -86,6 +86,7 @@ def __init__( self.llm_url = llm_url self.model = model self.datasets = datasets + self.provider = provider self.capture_level = "text" if provider == "anthropic" else capture_level self.capture = CaptureServer( llm_url=llm_url, @@ -297,10 +298,12 @@ def gradio_builder( ) -> Any: """OpenEnv calls this positionally with six web-interface arguments. - Only the title is useful here: the Harbor UI drives rollouts through its own handlers rather - than the generic action-field form, because a rollout is one long tool call, not a step. + None of them is used: the Harbor UI drives rollouts through its own handlers rather than the + generic action-field form, because a rollout is one long tool call, not a step. Even the + title is its own, since OpenEnv's ("OpenEnv Agentic Environment: harbor_env") names the + env class rather than what the page is. """ - return harbor_gradio_builder(datasets=datasets, title=display_title or "Harbor") + return harbor_gradio_builder(datasets=datasets, title="OpenEnv × Harbor") app = create_app( HarborEnvironment, @@ -321,4 +324,99 @@ def gradio_builder( if service is not None and service.mounted: app.mount(CAPTURE_MOUNT, service.capture.app) + _attach_hf_login(app) + # Every UI handler that changes something (a rollout, an added dataset) is a POST under /web, + # and Gradio accepts those from any origin. Without this, any page a visitor opens could make + # their browser start rollouts on this server's endpoint, or reach a server on their own machine + # that the page itself cannot. Nothing a real visitor does is cross-site. + app.add_middleware(SameOrigin) return app + + +def _attach_hf_login(app: Any) -> bool: + """ "Sign in with Hugging Face", where the Hub has set it up for the Space (`hf_oauth: true`). + + Attached to this app rather than to the Gradio UI mounted at `/web`: Gradio's OAuth routes and + the callback URL it registers assume the site root. Gradio then hands the signed-in visitor's + token to any UI handler that asks for a `gr.OAuthToken`, through the session cookie set here. + + A cookie that identifies the visitor makes another website's request count as theirs; the + same-origin check `build_app` installs for the UI covers that too. + """ + from .ui_settings import load + + if not load().hf_login: + return False + # Gradio tells a Space from a laptop by `SYSTEM=spaces`, which Docker Spaces do not set. Without + # it, `attach_oauth` installs its local stand-in, which signs every visitor in as the account of + # the token the server holds (the operator's), with a token that calls nothing. `hf_login` is + # only true on a Space that has an OAuth app, so this is that Space. It stays process-wide on + # purpose: `attach_oauth` reads it once (its routes then use `SPACE_HOST`), and whatever else in + # Gradio asks `gradio.utils.get_space()` afterwards should get the answer the login got. + os.environ.setdefault("SYSTEM", "spaces") + try: + from gradio.oauth import attach_oauth + from gradio.utils import get_space + + if get_space() is None: + print("hf login off: Gradio does not see a Space here") + return False + attach_oauth(app) + except (ImportError, ValueError) as exc: + print(f"hf login off: {exc}") + return False + print("hf login on (Inference Providers with the visitor's own account)") + return True + + +class SameOrigin: + """Refuse a state-changing request to the UI that a browser made from another site. + + Browsers label every request they make (`Sec-Fetch-Site`, `Origin`); servers calling this app, + such as a sandbox reaching the capture proxy or a trainer on the Task API, send neither and pass. + Only the UI (`/web`) is guarded: the Task API and MCP carry no visitor state, and browser tools + such as the MCP Inspector call them from another origin on purpose. + """ + + def __init__(self, app: Any, prefix: str = "/web") -> None: + self.app = app + self.prefix = prefix + + async def __call__(self, scope: dict, receive: Any, send: Any) -> None: + path = scope.get("path") or "" + if ( + scope.get("type") == "http" + and scope.get("method") not in ("GET", "HEAD", "OPTIONS") + and (path == self.prefix or path.startswith(self.prefix + "/")) + ): + headers = { + k.decode().lower(): v.decode() for k, v in scope.get("headers") or [] + } + site = headers.get("sec-fetch-site", "") + origin = headers.get("origin", "") + if site == "cross-site" or ( + origin and not _same_host(origin, headers.get("host", "")) + ): + from starlette.responses import PlainTextResponse + + await PlainTextResponse("cross-site request refused", status_code=403)( + scope, receive, send + ) + return + await self.app(scope, receive, send) + + +def _same_host(origin: str, host: str) -> bool: + """Whether a page at `origin` is this server's own. + + Its own hosts are the `Host` it was reached at, a Space's `SPACE_HOST`, and any listed in + `OPENENV_HARBOR_UI_HOSTS` (for a proxy that rewrites `Host`). Never `X-Forwarded-Host`: a + client sets that itself, so it would let any origin vouch for itself. + """ + from urllib.parse import urlparse + + listed = ",".join( + os.environ.get(k, "") for k in ("SPACE_HOST", "OPENENV_HARBOR_UI_HOSTS") + ) + allowed = {host, *[h.strip() for h in listed.split(",") if h.strip()]} + return urlparse(origin).netloc in allowed diff --git a/src/openenv/harbor/tasks.py b/src/openenv/harbor/tasks.py index f29b4e4bc..a7145937f 100644 --- a/src/openenv/harbor/tasks.py +++ b/src/openenv/harbor/tasks.py @@ -61,6 +61,31 @@ def _is_hf_repo(spec: str) -> bool: ) +_MAX_TASK_DEPTH = 6 + + +def _nested_task_dirs(base: Path) -> list[Path]: + """Every folder under `base` holding a `task.toml`, by relative path; a task's own subfolders + are not searched. + + A symlinked folder is not followed: it can point outside the dataset, or back into it, which + would recurse until the thread dies. Neither is anything deeper than `_MAX_TASK_DEPTH` levels. + """ + found: list[Path] = [] + + def walk(folder: Path, depth: int) -> None: + for p in sorted(folder.iterdir()): + if p.name.startswith(".") or p.is_symlink() or not p.is_dir(): + continue + if (p / "task.toml").is_file(): + found.append(p) + elif depth < _MAX_TASK_DEPTH: + walk(p, depth + 1) + + walk(base, 1) + return found + + def _task_dirs_from_directory( root: Path, *, validate: bool | None = None ) -> list[Path]: @@ -78,6 +103,11 @@ def _task_dirs_from_directory( candidates = sorted( p for p in base.iterdir() if p.is_dir() and not p.name.startswith(".") ) + if candidates and not any((p / "task.toml").is_file() for p in candidates): + # Grouped: `tasks////` (terminal-bench-science). Only when no top + # level folder is a task, so a flat dataset keeps exactly the order, and so the indexes, it + # always had. + candidates = _nested_task_dirs(base) if not (_VALIDATE_TASKS if validate is None else validate): return candidates try: @@ -87,9 +117,13 @@ def _task_dirs_from_directory( return [p for p in candidates if Task.is_valid_dir(p, disable_verification=True)] -def resolve_task_dirs(spec: str, *, refresh: bool = False) -> list[Path]: +def resolve_task_dirs( + spec: str, *, refresh: bool = False, tqdm_class: Any = None +) -> list[Path]: """Resolve a dataset spec to an ordered list of Harbor task directories. + `tqdm_class` is handed to the Hub download, for a caller that shows its progress. + Order is stable (sorted by directory name) because a task's *index* is its identity everywhere downstream — a trainer's dataset row, a `run_rollout` argument, a result. An unstable order would silently change which task an index refers to between runs. @@ -102,7 +136,9 @@ def resolve_task_dirs(spec: str, *, refresh: bool = False) -> list[Path]: if path.is_dir(): dirs = _task_dirs_from_directory(path) elif _is_hf_repo(spec): - dirs = _task_dirs_from_directory(_materialise_hf_dataset(spec)) + dirs = _task_dirs_from_directory( + _materialise_hf_dataset(spec, tqdm_class=tqdm_class) + ) else: dirs = _registry_task_dirs(spec) @@ -145,7 +181,7 @@ def resolve_task_dirs(spec: str, *, refresh: bool = False) -> list[Path]: _DOWNLOAD_WORKERS = int(os.environ.get("OPENENV_DATASET_WORKERS", "32")) -def _materialise_hf_dataset(spec: str) -> Path: +def _materialise_hf_dataset(spec: str, *, tqdm_class: Any = None) -> Path: """Download an HF dataset as real files and return its local root. Mounting beats downloading where it is available: a deployed Space can attach the dataset repo @@ -162,6 +198,7 @@ def _materialise_hf_dataset(spec: str) -> Path: allow_patterns=["tasks/**"], local_dir=str(target), max_workers=_DOWNLOAD_WORKERS, + tqdm_class=tqdm_class, ) return target @@ -184,13 +221,62 @@ def _registry_task_dirs(spec: str) -> list[Path]: return [Path(str(t.get_local_path())) for t in task_configs] -def read_instruction(task_dir: Path, *, limit: int = 4000) -> str: +def _dataset_folder(spec: str, task_dir: Path) -> Path | None: + """The folder a dataset's files must stay in, with a link to that folder itself followed (an + operator may serve `~/datasets/current`): the local folder, the Hub download, or Harbor's cache + for a registry task. `None` for a registry task Harbor keeps elsewhere (a local path its + registry names), which Harbor has already resolved.""" + path = Path(spec).expanduser() + if path.is_dir(): + return path.resolve() + if _is_hf_repo(spec): + return (_DATASET_ROOT / spec.replace("/", "__")).resolve() + try: + from harbor.constants import CACHE_DIR + except ImportError: + return None + if task_dir.absolute().is_relative_to(CACHE_DIR.absolute()): + return CACHE_DIR.resolve() + return None + + +def task_root(spec: str | None, task_dir: Path) -> Path | None: + """The one folder a task's files are read from to be shown (the Task API's instruction, and + everything the UI shows): `task_dir` resolved, or `None` when a link (the task folder, `tasks/`, + anything between) takes it outside its dataset, and then nothing in it is read. Without a + dataset to anchor on, the task folder may not be a link. Rollouts are not affected: Harbor + reads a task for itself.""" + real = task_dir.resolve() + folder = _dataset_folder(spec, task_dir) if spec else None + if folder is None: + return None if task_dir.is_symlink() else real + return real if real.is_relative_to(folder) else None + + +def own_file(root: Path | None, name: str) -> Path: + """`root / name` when it resolves inside the task's root (see `task_root`), else `OSError`, + which readers already treat as a missing file.""" + if root is None: + raise OSError("this task's folder is outside its dataset") + path = root / name + if not path.resolve().is_relative_to(root): + raise OSError(f"{name} points outside its task") + return path + + +def read_instruction( + task_dir: Path, *, limit: int = 4000, spec: str | None = None +) -> str: """The task's prompt, for previewing in discovery. Truncated: this is not the authoritative copy. The sandbox gets the real instruction from Harbor at run time. Serving a huge prompt over the - Task API for every listed task would make `list_tasks` enormous for no benefit. + Task API for every listed task would make `list_tasks` enormous for no benefit. Read only + inside the task's dataset (`task_root`), since the Task API is public on a Space. """ - path = task_dir / "instruction.md" + try: + path = own_file(task_root(spec, task_dir), "instruction.md") + except OSError: + return "" if not path.is_file(): return "" text = path.read_text(errors="replace").strip() @@ -298,5 +384,5 @@ def _ref(spec: str, index: int, task_dir: Path) -> HarborTaskRef: task_id=str(task_dir), task_name=task_dir.name, dataset=spec, - instruction=read_instruction(task_dir), + instruction=read_instruction(task_dir, spec=spec), ) diff --git a/src/openenv/harbor/ui.py b/src/openenv/harbor/ui.py index b86d64d4f..037b1ed59 100644 --- a/src/openenv/harbor/ui.py +++ b/src/openenv/harbor/ui.py @@ -1,793 +1,734 @@ """Human-facing UI for a Harbor env server. -Two columns: the LLM on the left, the task on the right. Validate, pick, run. - -Status text is deliberately terse. The long explanations belong in docs — what a person needs on -screen is whether it will work, what got rewritten, and which sandboxes are usable. - -Validation is a gate, not a hint: an LLM endpoint without token-id capture answers every request -normally and returns nothing trainable, so a rollout looks perfect and is worthless. - -Rich output (the rollout graph, per-turn tokens) is rendered as HTML rather than Gradio widgets, -because a conversation tree with branches and discarded retries is a shape, and a dataframe cannot -show a shape. +Three tabs. **Tasks**: pick a dataset and a task, read it (instruction, files, settings), and run an +agent on it. **Runs**: every rollout this server has run, live or finished, one at a time or two to +four side by side. **Setup**: what this machine can run, and why anything it cannot. + +It is a Gradio app, like every OpenEnv UI, with custom HTML components where Gradio has no widget +for the job: the task list (thousands of rows, filtered as you type), the task viewer (a file tree +read on demand), the run card, the run list and the trajectory. Gradio supplies the tabs, the state +and the event wiring. + +A rollout uses the endpoint the server was started with, or one the visitor connects in the page: a +Hugging Face token and a model on Inference Providers, or any OpenAI-compatible URL such as vLLM. +Connecting is a gate, not a hint: an endpoint without token-id capture answers every request +normally and returns nothing trainable, so a rollout looks perfect and is worthless for training. +What visitors may do (use the server's endpoint, bring their own, see each other's runs) is set per +deployment; see `ui_settings`. + +Every argument that names a dataset is checked against the datasets this server serves or that were +added from the Hub in this process. The UI's handlers are callable by anyone who can load the page, +and a dataset spec that is a local path would otherwise let a browser list and read the server's +own files. """ from __future__ import annotations import html import json +import os import re +import threading +import time +from collections import OrderedDict +from importlib import resources +from pathlib import Path from typing import Any +from urllib.parse import urlparse import gradio as gr -_UNVALIDATED = "_Enter your LLM URL and press Validate._" - -_CSS = """ -.hb-wrap { max-width: 1400px; margin: 0 auto; } -.hb-card { border: 1px solid var(--border-color-primary); border-radius: 10px; padding: 14px 16px; } -.hb-dim { opacity: .6; } -.hb-kv { display: flex; gap: 22px; flex-wrap: wrap; margin: 4px 0 2px; } -.hb-kv b { font-variant-numeric: tabular-nums; } - -/* The two panels read as one undifferentiated wall of controls without a boundary; the border is - what makes "pick a model" and "pick a task" look like two separate decisions. */ -.hb-cell { border: 1px solid var(--border-color-primary); border-radius: 10px; - padding: 14px 16px; } -.hb-panel { border: 1px solid var(--border-color-primary) !important; - border-radius: 10px !important; padding: 16px !important; - background: var(--block-background-fill); } -.hb-cell { min-width: 0 !important; } -.hb-wrap .hb-panel + .hb-panel { margin-top: 12px; } -.hb-tx, .hb-card { overflow-wrap: anywhere; } -@media (max-width: 700px) { - .hb-cell, .hb-panel { padding: 12px !important; } - .hb-hero { flex-wrap: wrap; } - .hb-kv { gap: 12px; } -} -@media (prefers-reduced-motion: reduce) { - .hb-pulse, .hb-step.now .hb-dot { animation: none; } -} - -/* Live conversation. Roles are colour-coded down the left edge so the shape of the loop - (assistant calls a tool, tool answers, assistant calls again) is readable at a glance. */ -/* No max-height here. A fixed-height scroll box nests a second scroller inside the page: - the wheel gets captured while the pointer is over the conversation, and the page stops - growing so there is nothing left to scroll to. Let it run at natural height and let the - page do the scrolling. Length is bounded by the message cap, not by CSS. */ -.hb-tx { margin-top: 10px; } -.hb-msg { border-left: 3px solid var(--border-color-primary); padding: 6px 0 6px 10px; - margin: 8px 0; font-size: 13px; line-height: 1.45; } -.hb-msg pre { white-space: pre-wrap; word-break: break-word; margin: 4px 0 0; - font-size: 12px; opacity: .85; } -.hb-role { display: inline-block; font-size: 11px; text-transform: uppercase; - letter-spacing: .04em; opacity: .65; margin-bottom: 2px; } -.hb-assistant { border-left-color: #22c55e; } -.hb-tool { border-left-color: #38bdf8; } -.hb-user { border-left-color: #a78bfa; } -.hb-system { border-left-color: #94a3b8; opacity: .75; } -.hb-tc { margin-top: 4px; padding: 4px 8px; border-radius: 6px; - background: var(--background-fill-secondary); } -.hb-tc { display: block; } -.hb-tc b { font-family: ui-monospace, SFMono-Regular, Menlo, monospace; font-size: 12px; } -.hb-arrow { opacity: .5; margin-right: 6px; } -.hb-tr { margin-top: 4px; padding: 4px 8px; border-radius: 6px; border-left: 2px solid #38bdf8; - background: var(--background-fill-secondary); } -/* No inner scroller here either, for the same reason as the conversation above, and the previous - version of this rule was the bug: `overscroll-behavior: contain` does not stop a box from - swallowing the page scroll, it is what *prevents* the wheel from chaining to the page once the - box reaches its own end. Tool output is clipped to 500 characters server side, but 500 - characters of shell output is 25 short lines, which overflowed the 220px cap and left the page - feeling frozen wherever the pointer happened to be. Length is bounded by the clip, not by CSS. */ -.hb-tr pre{ margin: 0; font-size: 11.5px; opacity: .8; } - -/* A run in flight should look like one. */ -.hb-live { display: flex; align-items: center; gap: 10px; margin-bottom: 6px; } -.hb-pulse { width: 8px; height: 8px; border-radius: 50%; background: #22c55e; - animation: hb-blink 1.2s ease-in-out infinite; } -@keyframes hb-blink { 0%, 100% { opacity: 1; } 50% { opacity: .25; } } -.hb-drop-msg { opacity: .5; border-left-color: #ef4444; } - -/* Verdict. The outcome should be legible from across the room; the numbers behind it should not - compete with it for attention. */ -.hb-verdict { border-left-width: 4px; } -.hb-head { font-size: 17px; font-weight: 650; margin-bottom: 8px; } -.hb-good { border-left-color: #22c55e; } -.hb-warn { border-left-color: #f59e0b; } -.hb-bad { border-left-color: #ef4444; } -.hb-err { white-space: pre-wrap; word-break: break-word; font-size: 12px; margin: 10px 0 0; - padding: 8px 10px; border-radius: 6px; background: var(--background-fill-secondary); } - -/* A qualifier on the result: true, load-bearing, and not an error. Bordered rather than coloured - like a finding, so "this rollout is eval-only" does not read as "this rollout failed". */ -.hb-note { font-size: 12.5px; line-height: 1.5; margin: 10px 0 0; padding: 8px 11px; - border-radius: 6px; border: 1px solid var(--border-color-primary); - background: var(--background-fill-secondary); } -.hb-note code { font-size: 11.5px; } - -/* Hover explanations. `data-tip` rather than `title=` for the two long ones: the native tooltip - truncates, takes a second to appear, and cannot wrap a paragraph. Short hints use Gradio's own - `info=`, which renders under the label and needs no hover at all. */ -.hb-i { display: inline-flex; align-items: center; justify-content: center; cursor: help; - width: 15px; height: 15px; margin-left: 6px; border-radius: 50%; font-size: 10px; - font-weight: 700; font-style: normal; vertical-align: 1px; - border: 1px solid var(--border-color-primary); opacity: .75; position: relative; } -.hb-i:hover { opacity: 1; } -.hb-i::after { content: attr(data-tip); position: absolute; left: 50%; bottom: 130%; - transform: translateX(-50%); width: max-content; max-width: 320px; padding: 8px 10px; - border-radius: 6px; border: 1px solid var(--border-color-primary); - background: var(--background-fill-primary); color: var(--body-text-color); - font-size: 11.5px; font-weight: 400; line-height: 1.5; text-align: left; - white-space: pre-line; opacity: 0; visibility: hidden; transition: opacity .12s; - z-index: 40; box-shadow: 0 4px 14px rgba(0,0,0,.18); } -.hb-i:hover::after { opacity: 1; visibility: visible; } -/* The label row the icon sits on, so the icon lines up with a Gradio label rather than floating. */ -.hb-lbl { display: flex; align-items: center; font-size: 13px; font-weight: 600; - margin: 2px 0 -6px; } - -/* Findings carry severity: a FATAL means unusable, a WARN means read before training on it. */ -.hb-find { font-size: 12.5px; margin: 5px 0; line-height: 1.45; } -.hb-tag { display: inline-block; min-width: 46px; margin-right: 8px; padding: 1px 6px; - border-radius: 4px; font-size: 10px; font-weight: 700; letter-spacing: .04em; - text-align: center; vertical-align: 1px; } -.hb-fatal .hb-tag { background: #ef4444; color: #fff; } -.hb-warn2 .hb-tag { background: #f59e0b; color: #1f2937; } -.hb-info .hb-tag { background: var(--background-fill-secondary); opacity: .7; } -.hb-info { opacity: .7; } - -/* Turn table: dense, aligned, and the numbers read as numbers. */ -.hb-tbl { width: 100%; border-collapse: collapse; margin-top: 8px; font-size: 13px; } -.hb-tbl th{ text-align: left; font-weight: 600; font-size: 11px; text-transform: uppercase; - letter-spacing: .04em; opacity: .55; padding: 4px 10px 6px 0; - border-bottom: 1px solid var(--border-color-primary); } -.hb-tbl td{ padding: 7px 10px 7px 0; border-bottom: 1px solid var(--border-color-primary); - vertical-align: top; } -.hb-tbl code { font-size: 12px; padding: 1px 6px; border-radius: 4px; - background: var(--background-fill-secondary); } -.hb-num { font-variant-numeric: tabular-nums; text-align: right; white-space: nowrap; - padding-right: 14px !important; } -.hb-prev { margin-top: 3px; font-size: 12px; } -.hb-drop-row { opacity: .45; } -.hb-drop-tag { background: #ef4444; color: #fff; } -.hb-conf { display: inline-block; width: 76px; height: 7px; border-radius: 4px; - background: var(--background-fill-secondary); overflow: hidden; vertical-align: middle; } -.hb-conf span { display: block; height: 100%; } - -/* Each conversation folds away; the main one starts open. */ -.hb-convo { margin-top: 10px; border-top: 1px solid var(--border-color-primary); padding-top: 8px; } -.hb-convo summary { cursor: pointer; padding: 4px 0; } - -/* Setup, before the agent has said anything. */ -.hb-steps { margin: 8px 0 0; } -.hb-step { display: flex; align-items: center; gap: 9px; padding: 3px 0; font-size: 13px; } -.hb-dot { width: 7px; height: 7px; border-radius: 50%; background: var(--border-color-primary); } -.hb-step.done .hb-dot { background: #22c55e; } -.hb-step.now .hb-dot { background: #f59e0b; animation: hb-blink 1.2s ease-in-out infinite; } -.hb-step.todo { opacity: .45; } - -/* The outcome, at a glance. */ -.hb-hero { display: flex; align-items: center; justify-content: space-between; gap: 20px; - padding-bottom: 12px; margin-bottom: 4px; - border-bottom: 1px solid var(--border-color-primary); } -.hb-badge { display: inline-flex; align-items: center; gap: 9px; font-size: 19px; - font-weight: 700; letter-spacing: -.01em; } -.hb-mark { display: inline-flex; align-items: center; justify-content: center; - width: 30px; height: 30px; border-radius: 50%; font-size: 15px; color: #fff; } -.hb-b-good .hb-mark { background: #22c55e; } -.hb-b-warn .hb-mark { background: #f59e0b; } -.hb-b-bad .hb-mark { background: #ef4444; } -.hb-score { text-align: right; line-height: 1.05; } -.hb-score-v { font-size: 42px; font-weight: 700; font-variant-numeric: tabular-nums; - letter-spacing: -.02em; } -.hb-score-c { font-size: 11px; text-transform: uppercase; letter-spacing: .06em; opacity: .55; } -.hb-kv-big span { font-size: 11px; text-transform: uppercase; letter-spacing: .04em; - opacity: .55; } -.hb-kv-big b { display: block; font-size: 19px; margin-top: 3px; text-transform: none; - letter-spacing: normal; opacity: 1; } -.hb-kv-big .hb-key b { color: var(--body-text-color); } -.hb-kv-big .hb-key { opacity: .85; } - -footer { display: none !important; } -""" +from . import ui_data, ui_icons, ui_pages, ui_runs, ui_settings +from .ui_icons import js_prelude +# Hub datasets added from the page in this process, on top of the ones the server was started with. +_ADDED: list[str] = [] +_ADDED_LOCK = threading.RLock() +_REMOVING: set[str] = set() +_HUB_ID = re.compile(r"^[A-Za-z0-9][\w.-]*/[\w.-]+$") -def _labelled(label: str, tip: str) -> str: - """A field label with a hover-explained `i` beside it. - For the explanations too long to sit under a Gradio label as `info=` text — which is where every - one-liner belongs instead, since it needs no hover to be seen. - """ - return ( - f'
{html.escape(label)}' - f'i
' - ) +def _asset(name: str) -> str: + return resources.files("openenv.harbor").joinpath("ui_assets", name).read_text() -_KEY_TIP = ( - "Only needed for a hosted endpoint: OpenAI, Anthropic, HF Inference Providers.\n\n" - "It is sent to the inference endpoint by this server and nothing else. It is NOT the key the " - "agent receives — that one is a capture session id, minted per rollout, which is how one proxy " - "serves many rollouts and how an unregistered caller is rejected.\n\n" - "Leave empty for a local vLLM or SGLang." -) +def _can_add_datasets() -> bool: + """Whether the page may download datasets from the Hub. Off by default on a Space: a public page + that downloads any dataset a visitor names is a disk-filling button.""" + return ui_settings.load().add_datasets -_LEVEL_TIP = ( - "There are two kinds of rollout, and the endpoint decides which you get.\n\n" - "TRAIN needs the engine to return token ids and per-token logprobs: vLLM started with " - "--return-tokens-as-token-ids --logprobs-mode processed_logprobs, or SGLang built from git " - "main. You get the reward, the trace, and the exact tokens and logprobs to train on.\n\n" - "EVAL is everything else, including a vLLM started without those flags. You get the reward and " - "the full trace; there are no token ids, so nothing is trainable. Logprobs alone do not help — " - "with no ids to pair them with there is nothing to align them to." -) +def _allowed(spec: str, served: list[str]) -> bool: + with _ADDED_LOCK: + return bool(spec) and (spec in served or spec in _ADDED) -def _clip(text: Any, limit: int = 400) -> str: - """Escape and shorten a value for display, keeping the head where the meaning usually is.""" - body = text if isinstance(text, str) else json.dumps(text, default=str) - body = body.strip() - return html.escape(body[:limit]) + ("…" if len(body) > limit else "") +_MAX_ADDED = 20 +_MAX_INSPECTED = 128 +_INSPECT_TTL = 300 +_INSPECTED: OrderedDict[str, tuple[float, dict[str, Any]]] = OrderedDict() +_INSPECTED_LOCK = threading.Lock() -def _tool_calls(message: dict[str, Any]) -> list[dict[str, Any]]: - """Tool calls on a message, normalised across all four dialects. - Chat-completions puts them in `tool_calls`; Anthropic puts them in the content block list as - `tool_use`. Reading only the former shows claude-code as a stream of text with no visible - actions, which is exactly the case the live view exists to make visible. - """ - out: list[dict[str, Any]] = [] - for call in message.get("tool_calls") or []: - function = call.get("function") or {} - name = function.get("name") or call.get("name") - if name: - out.append( - { - "name": str(name), - "arguments": function.get("arguments", call.get("arguments", "")), - } - ) - content = message.get("content") - if isinstance(content, list): - for block in content: - if ( - isinstance(block, dict) - and block.get("type") == "tool_use" - and block.get("name") - ): - out.append( - {"name": str(block["name"]), "arguments": block.get("input", "")} - ) - return out - - -def _message_text(message: dict[str, Any]) -> str: - """Readable text of a message, ignoring tool-call and tool-result blocks.""" - content = message.get("content") - if isinstance(content, str): - return content - if isinstance(content, list): - parts = [] - for block in content: - if not isinstance(block, dict): - continue - if block.get("type") in ("tool_use", "tool_result"): - continue - if block.get("text"): - parts.append(str(block["text"])) - return " ".join(parts) - return "" - - -def _tool_results(message: dict[str, Any]) -> list[str]: - """What came back from a tool, in either the chat-completions or the Anthropic shape.""" - if message.get("role") == "tool": - return [ - _message_text(message) or json.dumps(message.get("content"), default=str) - ] - content = message.get("content") - if not isinstance(content, list): - return [] - out = [] - for block in content: - if isinstance(block, dict) and block.get("type") == "tool_result": - body = block.get("content") - if isinstance(body, list): - body = " ".join(b.get("text", "") for b in body if isinstance(b, dict)) - out.append(str(body if body is not None else "")) - return out - - -def _render_calls(calls: list[dict[str, Any]]) -> str: - return "".join( - f'
▸' - f"{html.escape(str(c.get('name', 'tool')))}" - f"
{_clip(c.get('arguments', ''), 600)}
" - for c in calls - ) +def _added_file(settings: ui_settings.UISettings) -> Path | None: + """Where datasets added from the page are listed, so they come back after a restart: next to + the run history, whose runs may refer to them. `None` when history is off.""" + if settings.bucket and settings.bucket_mount: + return None # the bucket's own folders are the list (`ui_data.added_in_bucket`) + return settings.runs_dir / ".added-datasets.json" if settings.runs_dir else None -def _render_message(message: dict[str, Any], *, label: str = "") -> str: - """One row of the conversation: who spoke, what they said, what they invoked or returned.""" - role = str(message.get("role", "?")) - calls = _tool_calls(message) - results = _tool_results(message) - text = _message_text(message) - - # A user message carrying only tool results is the tool speaking, not the user; labelling it - # "user" makes the agent look like it is being prompted between every action. - shown_role = "tool" if results and role != "assistant" else role - # For a `role: tool` message the content IS the result, so rendering both duplicates it. - if shown_role == "tool": - text = "" - body = _clip(text, 700 if shown_role in ("user", "system") else 450) if text else "" - blocks = "".join( - f'
{_clip(r, 500)}
' for r in results - ) - if not body and not blocks and not calls: - return "" - return ( - f'
' - f'{html.escape(label or shown_role)}' - + (f"
{body}
" if body else "") - + _render_calls(calls) - + blocks - + "
" - ) - +def _save_added(settings: ui_settings.UISettings) -> None: + path = _added_file(settings) + if path is None: + return + with _ADDED_LOCK: + saved = sorted(set(_ADDED)) + try: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(saved, indent=0)) + except OSError: + pass -def _transcript_html(session: Any) -> str: - """The conversation as it stands right now: what the agent said, called, and got back. - Counters answer "is it alive"; this answers "is it doing the right thing", which is the question - worth asking while a rollout is still running. The newest turn's `request_messages` already holds - the whole conversation the harness assembled, tool results included, so rendering that plus the - latest response needs no reconstruction from deltas. +def _load_added(settings: ui_settings.UISettings, served: list[str]) -> None: + path = _added_file(settings) + if path is None or not settings.add_datasets: + return + try: + listed = json.loads(path.read_text()) + except (OSError, ValueError): + return + with _ADDED_LOCK: + for spec in listed if isinstance(listed, list) else []: + # Checked again without the network: the file is ours, but a path must never slip in. + spec = str(spec) + parts = spec.split("/") + if ( + _HUB_ID.match(spec) + and not any(p in (".", "..") or p.startswith(".") for p in parts) + and spec not in served + and spec not in _ADDED + and len(_ADDED) < _MAX_ADDED + ): + _ADDED.append(spec) + + +def _inspect_hub(spec: str) -> dict[str, Any]: + """Inspect one Hub dataset, caching only successful answers for a bounded time.""" + now = time.monotonic() + with _INSPECTED_LOCK: + cached = _INSPECTED.get(spec) + if cached is not None and now - cached[0] < _INSPECT_TTL: + _INSPECTED.move_to_end(spec) + return cached[1] + _INSPECTED.pop(spec, None) + summary = ui_data.hub_summary(spec) + with _INSPECTED_LOCK: + _INSPECTED[spec] = (time.monotonic(), summary) + _INSPECTED.move_to_end(spec) + while len(_INSPECTED) > _MAX_INSPECTED: + _INSPECTED.popitem(last=False) + return summary + + +def _remove_added( + spec: str, served: list[str], settings: ui_settings.UISettings +) -> dict[str, Any]: + """Serialise removal so two requests cannot delete or mutate the same dataset.""" + with _ADDED_LOCK: + if spec in served or spec not in _ADDED or spec in _REMOVING: + return {"error": "Only datasets added from this page can be removed."} + _REMOVING.add(spec) + try: + ui_data.remove_added(spec, settings) + except Exception as exc: # noqa: BLE001 + try: + return { + "error": f"Could not remove it: {type(exc).__name__}: {str(exc)[:200]}" + } + finally: + with _ADDED_LOCK: + _REMOVING.discard(spec) + with _ADDED_LOCK: + _ADDED.remove(spec) + _REMOVING.discard(spec) + _save_added(settings) + return {"ok": True, "spec": spec} + + +def _hub_problem(spec: str) -> str | None: + """Why a dataset typed into the page may not be added, or `None`. + + A public Hub dataset only. Not a local path (the loader tries one first, so `src/..` would serve + this server's own files), and not a private or gated one, which the server's token might open + for a visitor who has no access of their own. """ - nodes = sorted(session.graph.nodes(), key=lambda n: n.index) - if not nodes: - return "" - latest = nodes[-1] - - rows = [ - row - for row in (_render_message(m) for m in (latest.request_messages or [])) - if row - ] - - response = latest.response_message or {} - tail = _render_message( - {**response, "role": "assistant"}, - label=f"assistant · completed call {latest.index + 1}", - ) - if tail: - rows.append(tail) - - # Only the tail is ever new, so cap from the front and say what was dropped. - shown = rows[-18:] - elided = ( - f'
… {len(rows) - len(shown)} earlier message(s)
' - if len(rows) > len(shown) - else "" - ) - # Count across the conversation, not just the response messages: Anthropic carries tool use in - # the assistant content blocks the harness replays back, so a response-only tally reads 0. - calls_so_far = sum( - len(_tool_calls(m)) for m in (latest.request_messages or []) - ) + len(_tool_calls(latest.response_message or {})) - return ( - f'
' - f'Live conversation' - f'turn {latest.index} · {calls_so_far} tool call(s) so far · ' - f"{latest.n_tools} tool(s) offered
{elided}{''.join(shown)}
" - ) - - -# Capture-session creation precedes sandbox allocation; it is not evidence that -# setup has finished. Only a completed captured call proves the agent is running. -_SETUP_STEPS = ( - "preparing the sandbox, task and agent", - "waiting for a completed model call", + parts = spec.split("/") + if ( + not _HUB_ID.match(spec) + or any(p in (".", "..") or p.startswith(".") for p in parts) + or Path(spec).expanduser().exists() + ): + return "Enter a Hugging Face dataset id, like org/name." + try: + from huggingface_hub import HfApi + + info = HfApi(token=False).dataset_info(spec) + except Exception: # noqa: BLE001 - private, missing, or the Hub is down: all the same answer + return f"{spec} is not a public dataset on the Hub." + if info.private or info.gated: + return f"{spec} is private or gated. Only public datasets can be added from the page." + return None + + +from .ui_trace import ( # noqa: F401 - re-exported: tests and callers read them from `ui` + _conversation_html, + _findings_html, + _markdown, + _result_html, + _transcript_html, + _turns_html, + count_actions, ) -def _steps_html(stage: int) -> str: - """The setup sequence, with the current stage marked.""" - rows = [] - for i, label in enumerate(_SETUP_STEPS): - cls = "done" if i < stage else ("now" if i == stage else "todo") - rows.append( - f'
' - f"{html.escape(label)}
" - ) - return f'
{"".join(rows)}
' - - -def _live_html( - harness: str, - sandbox: str, - phase: str, - elapsed: float, - stats: dict[str, Any] | None, - stage: int = -1, -) -> str: - """The running header: what is running, how far in, and what it has produced so far.""" - bits = [ - f'
' - f'
' - f"Running {html.escape(harness)} on " - f"{html.escape(sandbox)}" - f'{html.escape(phase)} · {elapsed:.0f}s
' - ] - # Before the first call there are no numbers worth showing, so show progress instead. A row of - # zeros for a minute reads as "stuck" when the sandbox is simply still booting. - if stage >= 0: - bits.append(_steps_html(stage)) - if stats: - bits.append( - '
' - + "".join(f"{k}
{v}
" for k, v in stats.items()) - + "
" - ) - bits.append("
") - return "".join(bits) - - -# `warn` is already a verdict tone; the finding variant needs its own class name. -_FINDING_CLASS = {"FATAL": "fatal", "WARN": "warn2", "INFO": "info"} +def _contract(r: dict[str, Any]) -> dict[str, Any] | None: + """The training contract of a rollout: exactly what a trainer consumes, nothing else. + Per turn, `(prompt_token_ids, completion_token_ids, per_token_logps)` plus the reward. The + logprobs are the load-bearing part and the reason this is a separate download: they are the + behaviour policy's, recorded at sampling time, and cannot be recovered afterwards by re-running + the prompt. Discarded turns are kept but flagged, because they were generated and billed and a + trainer must be able to see them in order to exclude them deliberately. -def _findings_html(findings: list[str]) -> str: - """Findings, grouped by how much they should worry you. + `None` for an eval rollout: a download named `contract.json` whose every `prompt_token_ids` is + `[]` would look like a contract and contain none. - They were previously all rendered the same dim grey and truncated to 220 characters, which put - "the intercept saw no model calls" and "3 roots across 7 turns" at equal weight. A FATAL means - the rollout is unusable; a WARN means read it before training on it. + Raises: + `ValueError`: when the result has FATAL findings or an invalid mask; the exporter refuses it. """ - if not findings: - return "" - buckets: dict[str, list[str]] = {"FATAL": [], "WARN": [], "INFO": []} - for raw in findings: - level = ( - "FATAL" - if raw.startswith("[FATAL") - else "WARN" - if raw.startswith("[WARN") - else "INFO" - ) - buckets[level].append( - raw.split("]", 1)[-1].strip() if raw.startswith("[") else raw - ) - - out = [] - for level, items in buckets.items(): - for item in items: - out.append( - f'
' - f'{level}{html.escape(item[:400])}
' - ) - return "".join(out) - + turns = r.get("turns") or [] + if not turns or r.get("rollout_type", "train") == "eval": + return None + from .contract import export_training_contract + from .models import HarborRolloutResult -def _result_html(r: dict[str, Any]) -> str: - """The verdict, the numbers behind it, and anything that qualifies it. + return export_training_contract(HarborRolloutResult.model_validate(r)) - The outcome is the one thing every reader wants first, so the reward is set at display size and - the supporting counts are deliberately quieter. Getting that hierarchy wrong is how a failed - rollout reads as a successful one at a glance. - """ - reward = r.get("reward") - if not r.get("ok"): - tone, mark, label = "bad", "✕", "Failed" - value, caption = "—", str(r.get("exception_type") or "error") - elif reward is None: - # Not a zero. The verifier never ran, so this says nothing about the model. - tone, mark, label = "warn", "!", "Not graded" - value, caption = "—", "the verifier never ran" - elif reward > 0: - tone, mark, label = "good", "✓", "Solved" - value, caption = f"{reward:.2f}", "reward" - else: - tone, mark, label = "warn", "○", "Not solved" - value, caption = f"{reward:.2f}", "reward" - turns = r.get("turns") or [] - generated = sum(len(t.get("completion_token_ids") or []) for t in turns) - dropped = sum( - len(t.get("completion_token_ids") or []) for t in turns if t.get("discarded") - ) - tools = sum(len(t.get("tool_calls") or []) for t in turns) - atif = r.get("atif", "none") - - # `key` marks the figures that decide whether this rollout is usable, as opposed to describing it. - # The initial prompt: task instruction plus the harness's system prompt and tool manifest. - # Constant across turns, so it is a property of the rollout rather than a per-row column. - context = len((turns[0].get("prompt_token_ids") or [])) if turns else 0 - - is_eval = r.get("rollout_type", "train") == "eval" - kv = [ - # "trainable tokens: 0" on an eval rollout reads as a capture failure. It is not one, so the - # slot says what kind of rollout this is instead of reporting a zero that means nothing here. - ("rollout", f"EVAL · {r.get('capture_level', '?')}", True) - if is_eval - else ("trainable tokens", f"{r.get('n_trainable_tokens', 0):,}", True), - ("context", f"{context:,}", False), - ("trace check", atif, atif != "match"), - ("model calls", r.get("n_turns", 0), False), - ("tool calls", tools, False), - ("conversations", r.get("n_roots", 0), False), - ( - "generated", - f"{generated:,}" + (f" · {dropped:,} discarded" if dropped else ""), - False, - ), - ("wall", f"{r.get('wall_s', 0):.0f}s", False), - ] +# ── the page's own pieces ───────────────────────────────────────────────────────────────────────── - out = [ - f'
', - '
', - f'
{mark}' - f"{html.escape(label)}
", - f'
{html.escape(value)}
' - f'
{html.escape(caption)}
', - "
", - '
' - + "".join( - f'{k}
{v}
' - for k, v, key in kv - ) - + "
", - ] - if is_eval: - out.append( - '
This is an eval rollout. The endpoint returned ' - f"{'logprobs but no token ids' if r.get('capture_level') == 'logprobs' else 'no token ids and no logprobs'}, " - "so you get the reward and the full trace below, but nothing trainable — there is no " - "contract.json and no per-token logprobs. Point the server at vLLM " - "(--return-tokens-as-token-ids --logprobs-mode processed_logprobs) or " - "SGLang built from main for trainable rollouts.
" - ) - for fix in r.get("param_fixes") or []: - out.append( - f'
upstream compatibility: {html.escape(fix)} — ' - "the request differs from what the harness asked for.
" - ) +# Capabilities are credential checks and SDK imports. Worth doing once, not per page load, and +# worth redoing on request, because adding a key is the usual fix for an unusable sandbox. +_CAPS: dict[str, Any] = {"value": None} +_CAPS_LOCK = threading.Lock() - rewards = r.get("rewards") or {} - if len(rewards) > 1: - chosen = r.get("reward_key", "") - parts = [ - f"{html.escape(k)} {v:.3f}" + (" ←" if k == chosen else "") - for k, v in sorted(rewards.items()) - ] - out.append(f'
{"   ".join(parts)}
') - for step in r.get("step_results") or []: - vals = ", ".join(f"{k}={v:.2f}" for k, v in (step.get("rewards") or {}).items()) - out.append( - f'
step {html.escape(step.get("name", ""))} {vals}
' - ) +def _capabilities(datasets: list[str], *, refresh: bool = False) -> Any: + from .capabilities import capabilities + from .serving import HarborService - if r.get("error"): - out.append(f'
{html.escape(str(r["error"])[:1200])}
') - if r.get("agent_log_tail"): - out.append( - '
agent log' - f"
{html.escape(str(r['agent_log_tail'])[:4000])}
" - ) + with _CAPS_LOCK: + if _CAPS["value"] is None or refresh: + service = HarborService.current() + llm = ( + { + "url": service.llm_url, + "model": service.model, + "capture_level": service.capture_level, + "reachable": True, + "ok": service.capture_level == "tokens", + } + if service is not None and service.llm_url + else {} + ) + _CAPS["value"] = capabilities(datasets=datasets or None, llm=llm) + return _CAPS["value"] - out.append(_findings_html(r.get("findings") or [])) - out.append( - '
Capture quality and reward are ' - "independent: a perfectly captured rollout can still score 0 because the model was " - "wrong, and reward — means the verifier never ran at all.
" - ) - return "".join(out) +def _agent_choices( + caps: Any, *, purpose: str, provider: str, include_experimental: bool +) -> tuple[list[tuple[str, str]], dict[str, str], int]: + """Which agents to offer, per the qualification evidence, and the profile each was qualified with. -def _conversation_html(r: dict[str, Any]) -> str: - """The whole conversation as it was actually sent: system prompt, tools, results, replies. + Stable agents are offered by default and experimental ones only on request, as the provider + qualification guide specifies. Without a report every agent is experimental, which is why the + count of hidden ones is returned: an empty list with no reason reads as a broken page. - Rebuilt from the result rather than the live session, so it survives the run. Several are - possible: each root is a separate conversation, and an auxiliary one (a next-speaker check, a - summariser) is labelled as such so it is not mistaken for the agent working on the task. + Returns: + `tuple` of `(label, name)` choices, `{harness: profile}`, and how many experimental agents + the filter hides. """ - conversations = r.get("conversations") or [] - if not conversations: - return "" + from .qualification import harness_maturity_rows + from .seams import get as get_seam - agents = [c for c in conversations if c.get("role", "agent") == "agent"] - blocks = [] - seen_agents = 0 - for i, convo in enumerate(conversations): - role = convo.get("role", "agent") - if role == "agent": - seen_agents += 1 - # Numbered when there is more than one, so two blocks are never both "main". - badge = ( - "main conversation" - if len(agents) == 1 - else f"conversation {seen_agents} of {len(agents)}" + report_path = os.environ.get("OPENENV_HARBOR_QUALIFICATION_REPORT", "") + try: + evidence = json.loads(Path(report_path).read_text()) if report_path else None + tiers = { + name: tier + for name, tier, _ in harness_maturity_rows( + [h.name for h in caps.harnesses], evidence ) - else: - badge = { - "auxiliary": "auxiliary call", - "discarded": "discarded branch", - }.get(role, role) - rows = [ - row - for row in (_render_message(m) for m in convo.get("messages") or []) - if row - ] - if not rows: + } + except (OSError, ValueError, TypeError): + tiers = {h.name: "experimental" for h in caps.harnesses} + evidence = None + profile_provider = "vllm" if purpose == "train" else provider + profiles: dict[str, str] = {} + unavailable: set[str] = set() + for cell in (evidence or {}).get("cells", []): + if cell.get("provider") != profile_provider: continue - blocks.append( - f'
' - f"{html.escape(badge)} " - f'{convo.get("n_turns", 0)} model call(s), ' - f"{len(rows)} message(s){''.join(rows)}
" - ) - if not blocks: - return "" - return ( - f'
Conversation ' - f'everything the model saw and produced' - f"{''.join(blocks)}
" + config = cell.get("configuration") or {} + profile = config.get("acp_profile") or config.get("nemo_profile") + if profile: + name = cell["harness"] + profiles[name] = profile + try: + get_seam(name, profile=profile) + except (ValueError, KeyError): + unavailable.add(name) + # Agents that passed a development run first, so the likely choices lead the list. + ordered = sorted( + (h for h in caps.harnesses if h.name not in unavailable), + key=lambda h: (h.status != "validated", h.name), ) + choices, hidden = [], 0 + for h in ordered: + tier = tiers.get(h.name, "experimental") + if tier == "stable" or (include_experimental and tier == "experimental"): + label = f"{h.name} · {h.dialect} · {tier}" + if h.name in profiles: + label += f" · profile: {profiles[h.name]}" + choices.append((label, h.name)) + elif tier == "experimental": + hidden += 1 + return choices, profiles, hidden + + +def _server_engine( + caps: Any, + include_experimental: bool = False, + settings: ui_settings.UISettings | None = None, +) -> dict[str, Any]: + """The endpoint the server was started with, as the page's default engine. + + It carries no key: the capture proxy already holds the server's credential and uses it for any + rollout that does not name a different endpoint. A deployment that does not share it with + visitors (`OPENENV_HARBOR_UI_SERVER_ENDPOINT=0`) starts every visitor with no engine. + """ - -def _confidence(mean_logp: float) -> str: - """A bar for mean logprob. Closer to 0 is more confident; -1.0 is the practical floor here.""" - pct = max(0.0, min(1.0, 1.0 + mean_logp)) # -0 -> 1.0, -1 -> 0.0 - hue = 8 + int(112 * pct) # red through amber to green - return ( - f'' - f'' + settings = settings or ui_settings.load() + service = ui_settings.shared_endpoint(settings) + if service is None: + return { + "ok": False, + "reason": "Connect a model to run rollouts.", + "include_experimental": include_experimental, + } + level = service.capture_level + purpose = "train" if level == "tokens" else "eval" + provider = str(service.provider or "openai") + choices, profiles, hidden = _agent_choices( + caps, + purpose=purpose, + provider=provider, + include_experimental=include_experimental, ) + return { + "ok": True, + "server_default": True, + "source": "server", + "url": "", + "model": service.model, + "host": urlparse(service.llm_url).netloc or service.llm_url, + "capture_level": level, + "trainable": level == "tokens", + "purpose": purpose, + "provider": provider, + "allowed_harnesses": [v for _, v in choices], + "harness_profiles": profiles, + "choices": choices, + "hidden_agents": hidden, + "include_experimental": include_experimental, + "notes": [], + } + + +def validate_endpoint( + url: str, + model: str, + api_key: str, + provider: str = "openai", + purpose: str = "eval", + include_experimental: bool = False, + datasets: list[str] | None = None, + private_urls: bool = True, + kind: str = "custom URL", +) -> dict[str, Any]: + """Probe an endpoint typed into the page and describe it as an engine for this browser session. + A rollout reaches it through the capture proxy's pool, so the tier comes from a real probe of + this endpoint rather than from what the server booted with. -def _turns_html(r: dict[str, Any]) -> str: - """Turn by turn: what it did, how much it wrote, how sure it was. + Args: + url (`str`): + OpenAI-spec endpoint. + model (`str`): + Served model id; read from the endpoint when it serves exactly one. + api_key (`str`): + Credential for a hosted endpoint, or `""`. + provider (`str`, *optional*, defaults to `"openai"`): + Upstream API family. + purpose (`str`, *optional*, defaults to `"eval"`): + `"eval"`, or `"train"` to require exact token capture. + include_experimental (`bool`, *optional*, defaults to `False`): + Offer agents the qualification evidence marks experimental. + datasets (`list[str]`, *optional*): + Served datasets, for the capability report. + private_urls (`bool`, *optional*, defaults to `True`): + Whether `url` may resolve to a loopback or private address. The server makes the call, so + on a public deployment this is off (see `ui_settings.url_problem`). + kind (`str`, *optional*, defaults to `"custom URL"`): + How a run from this engine is labelled in the run list. - Replaces a table whose most prominent column was "tools", meaning the number of tools *offered* - to the model. That number is a property of the harness, identical on every row, and told nobody - anything. What varies per turn, and is worth reading, is the action taken, the tokens spent on - it, and the model's confidence while producing them. + Returns: + `dict`: the engine, with `ok` and either `reason` or the agent choices. It holds `api_key` + for the server-side session state; the page is only ever sent `_card` of it. """ - turns = r.get("turns") or [] - if not turns: - return '
No model calls were captured.
' - - used: dict[str, int] = {} - for t in turns: - for call in t.get("tool_calls") or []: - name = str(call.get("name", "?")) - used[name] = used.get(name, 0) + 1 - - rows = [] - for t in turns: - lp = t.get("per_token_logps") or [] - mean = sum(lp) / len(lp) if lp else 0.0 - gen = len(t.get("completion_token_ids") or []) - calls = t.get("tool_calls") or [] - if calls: - action = " ".join( - f"{html.escape(str(c.get('name', 'tool')))}" for c in calls - ) - elif t.get("finish_reason") == "stop": - action = 'final answer' - else: - action = 'text only' - note = ( - ' discarded' - if t.get("discarded") - else "" - ) - preview = _clip(t.get("text") or "", 160) - rows.append( - f'' - f'{t.get("turn")}' - f"{action}{note}" - + (f'
{preview}
' if preview else "") - + f'{gen:,}' - f"{_confidence(mean) if lp else ''}" - f'{html.escape(str(t.get("finish_reason") or ""))}' - ) - - histogram = "" - if used: - top = sorted(used.items(), key=lambda kv: -kv[1]) - histogram = ( - '
tools used: ' - + "   ".join(f"{html.escape(k)}×{v}" for k, v in top) - + "
" + from openenv.core.harness.capture.validate_llm import list_models, validate_llm + + from .capabilities import capabilities + from .seams import agent_facing_model + from .serving import HarborService + + url = (url or "").strip().rstrip("/") + api_key = (api_key or "").strip() or None + base = { + "include_experimental": include_experimental, + "custom": { + "url": url, + "model": model or "", + "provider": provider, + "purpose": purpose, + }, + } + if not url: + return { + **base, + "ok": False, + "reason": "Enter the endpoint's URL.", + } + problem = ui_settings.url_problem(url, private_ok=private_urls) + if problem: + return {**base, "ok": False, "reason": problem} + if not model: + served_models = list_models(url, timeout=15, api_key=api_key) + if len(served_models) != 1: + return { + **base, + "ok": False, + "reason": "Pick a model: this endpoint serves " + + ", ".join(served_models[:12]) + if served_models + else "Nothing reachable at that URL. Check it, and the API key if it needs one.", + } + model = served_models[0] + # Bounded: the page waits on this, and an endpoint that never answers must not hold it for minutes. + report = validate_llm(url, model, api_key=api_key, provider=provider, timeout=45) + if not report.reachable or (purpose == "train" and not report.trainable): + why = "; ".join(report.findings) or ( + "exact engine tokens are required for training" + if report.reachable + else "unreachable" ) - - return ( - '
Turn by turn' - '' - "" - + "".join(rows) - + "
#actiontokensconfidencestopped because
" - + histogram - + '
Confidence is the mean logprob of the ' - "sampled tokens: full bar means the model was near-certain, short means it was " - "guessing. Discarded turns were generated and billed but lead nowhere, so they are " - "excluded from training paths.
" + return { + **base, + "ok": False, + "reason": f"Not usable: {why}. Training needs vLLM with --return-tokens-as-token-ids " + "--logprobs-mode processed_logprobs, or SGLang from git main; eval works with any " + "reachable endpoint.", + } + caps = capabilities( + datasets=list(datasets or []) or None, + llm={ + "url": url, + "model": model, + "ok": report.ok, + "capture_level": report.capture_level, + "reachable": True, + "authenticated": bool(api_key), + }, ) + choices, profiles, hidden = _agent_choices( + caps, + purpose=purpose, + provider=provider, + include_experimental=include_experimental, + ) + notes = [] + leaf = agent_facing_model(model) + if leaf != model: + notes.append(f"Sent to agents as {leaf}, rewritten back on the way out.") + notes += [f"Upstream compatibility: {fix}" for fix in report.param_fixes] + notes += [ + f.split(": ", 2)[-1] + for f in report.findings + if "behaviour_changed" in f or "tool_call" in f + ] + service = HarborService.current() + server_url = service.llm_url if service is not None else "" + if server_url and server_url.rstrip("/") != url: + notes.append( + "Rollouts from this page use this endpoint, not the server's default." + ) + return { + **base, + "ok": True, + "kind": kind, + "url": url, + "model": model, + "host": urlparse(url).netloc or url, + "capture_level": report.capture_level, + "trainable": report.trainable, + # Held in server-side session state so Run can reach a token-gated endpoint. Never sent + # back to the page: `_card` leaves it out. + "api_key": api_key or "", + "provider": provider, + "purpose": purpose, + "allowed_harnesses": [v for _, v in choices], + "harness_profiles": profiles, + "choices": choices, + "hidden_agents": hidden, + "sandboxes": list(caps.available_sandboxes), + "notes": notes, + } + + +def _local_token() -> str: + try: + from huggingface_hub import get_token + return get_token() or "" + except Exception: # noqa: BLE001 - no hub client, or an unreadable token file: just none + return "" -def _write_contract(r: dict[str, Any]) -> str | None: - """Write `contract.json`: exactly what a trainer consumes, nothing else. - - Per turn, `(prompt_token_ids, completion_token_ids, per_token_logps)` plus the reward. The - logprobs are the load-bearing part and the reason this is a separate file: they are the - behaviour policy's, recorded at sampling time, and cannot be recovered afterwards by re-running - the prompt. Discarded turns are kept but flagged, because they were generated and billed and a - trainer must be able to see them in order to exclude them deliberately. - - Returns `None` for an eval rollout. Writing a file whose every `prompt_token_ids` is `[]` would - hand someone a download named `contract.json` containing no contract, and a file on disk is far - more convincing than an empty list in a JSON blob. - """ - import tempfile - from pathlib import Path as _Path - - turns = r.get("turns") or [] - if not turns or r.get("rollout_type", "train") == "eval": - return None - from .contract import export_training_contract - from .models import HarborRolloutResult - contract = export_training_contract(HarborRolloutResult.model_validate(r)) - name = re.sub(r"[^A-Za-z0-9_.-]", "_", str(r.get("task_name") or "rollout")) - target = ( - _Path(tempfile.mkdtemp(prefix="harbor-contract-")) / f"{name}.contract.json" - ) - target.write_text(json.dumps(contract, indent=2)) - return str(target) +def connect_endpoint( + form: dict[str, Any], + *, + include_experimental: bool = False, + datasets: list[str] | None = None, + settings: ui_settings.UISettings | None = None, + account_token: str | None = None, +) -> dict[str, Any]: + """The run card's endpoint form, validated into an engine. + Two sources. `hf`: a model on Hugging Face Inference Providers, reached through the router with + the visitor's token (or, run locally, this machine's), optionally pinned to one provider or to the + router's `fastest`/`cheapest` policy. `url`: any OpenAI-compatible or Anthropic endpoint, such as a + vLLM the visitor runs. Both end in `validate_endpoint`, which probes the endpoint for real. -def _summary_json(r: dict[str, Any]) -> str: - """The result with the token arrays summarised, which is the part anyone actually reads. + Args: + form (`dict`): + `source`, and for `hf`: `model`, `route`, `api_key`, `local_token`, `use_account`; for + `url`: `url`, `model`, `api_key`, `api` (`openai` or `anthropic`), `purpose`. + account_token (`str`, *optional*): + The signed-in visitor's Hugging Face token (`inference-api` scope), used when the form + asks for `use_account`. Gradio reads it from the session; it is never in the form. - The full document stays available below; printing 8000 integers first buries the fields that - carry meaning. + Returns: + `dict`: the engine, as `validate_endpoint` returns it, with the form (minus the key) under + `custom` so the card can show what was connected. """ - compact = {k: v for k, v in r.items() if k not in ("turns", "conversations")} - compact["turns"] = [ - { - "turn": t.get("turn"), - "action": [c.get("name") for c in (t.get("tool_calls") or [])] or "text", - "prompt_token_ids": f"<{len(t.get('prompt_token_ids') or [])} ids>", - "completion_token_ids": f"<{len(t.get('completion_token_ids') or [])} ids>", - "per_token_logps": f"<{len(t.get('per_token_logps') or [])} floats>", - "finish_reason": t.get("finish_reason"), - "discarded": t.get("discarded"), - "text": (t.get("text") or "")[:200], + settings = settings or ui_settings.load() + source = "hf" if form.get("source") == "hf" else "url" + key = str(form.get("api_key") or "").strip() + remembered = { + k: form.get(k) + for k in ( + "source", + "model", + "route", + "url", + "api", + "purpose", + "local_token", + "use_account", + ) + if k in form + } + remembered["source"] = source + if not settings.visitor_endpoints: + return { + "ok": False, + "custom": remembered, + "reason": "This server runs rollouts on its own endpoint only.", } - for t in (r.get("turns") or [])[:200] - ] - compact["conversations"] = [ + common = {"include_experimental": include_experimental, "datasets": datasets} + if source == "hf": + model = str(form.get("model") or "").strip() + route = str(form.get("route") or "").strip() + if not model: + return {"ok": False, "custom": remembered, "reason": "Pick a model."} + if not key and form.get("use_account") and settings.hf_login: + key = account_token or "" + if not key: + return { + "ok": False, + "custom": remembered, + "reason": "Your Hugging Face sign-in has expired. Sign in again, or paste a token.", + } + if not key and form.get("local_token") and settings.local_token: + key = _local_token() + if not key: + return { + "ok": False, + "custom": remembered, + "reason": "Enter a Hugging Face token with the Inference Providers permission.", + } + engine = validate_endpoint( + ui_data.HF_ROUTER, + f"{model}:{route}" if route else model, + key, + provider="hf", + purpose="eval", # the router returns no token ids, so nothing from it is trainable + kind="Hugging Face", + **common, + ) + else: + engine = validate_endpoint( + str(form.get("url") or ""), + str(form.get("model") or ""), + key, + provider="anthropic" if form.get("api") == "anthropic" else "openai", + purpose="train" if form.get("purpose") == "train" else "eval", + private_urls=settings.private_urls, + kind="custom URL", + **common, + ) + engine["source"] = source + engine["custom"] = remembered + return engine + + +def _card( + engine: dict[str, Any], + selection: dict[str, Any] | None, + caps: Any, + message: tuple[str, str] | None = None, + settings: ui_settings.UISettings | None = None, + profile: Any = None, +) -> dict[str, Any]: + """What the run card shows. Everything the browser receives about the engine is chosen here, so + the API key held in the engine never reaches the page. `profile` is the signed-in visitor's + Hugging Face profile, where sign-in is set up and they used it.""" + from .serving import HarborService + + settings = settings or ui_settings.load() + service = HarborService.current() + shared = ui_settings.shared_endpoint(settings) + server = ( { - "role": c.get("role"), - "n_turns": c.get("n_turns"), - "messages": f"<{len(c.get('messages') or [])} messages>", + "model": shared.model, + "host": urlparse(shared.llm_url).netloc or shared.llm_url, + "level_text": ui_pages.LEVEL_TEXT.get(shared.capture_level, ""), + "train": shared.capture_level == "tokens", } - for c in (r.get("conversations") or []) + if shared is not None + else None + ) + sources = (["server"] if server else []) + ( + ["hf", "url"] if settings.visitor_endpoints else [] + ) + public = str(service.public_url or "") if service is not None else "" + host = urlparse(public).hostname or "" + harnesses = {h.name: h for h in caps.harnesses} + agents = [] + for label, name in engine.get("choices") or []: + h = harnesses.get(name) + agents.append( + { + "value": name, + "label": label, + "host_side": h is not None and h.kind == "base", + } + ) + values = [a["value"] for a in agents] + hidden = engine.get("hidden_agents") or 0 + sandboxes = [ + {"name": s.name, "available": bool(s.available), "detail": s.detail} + for s in caps.sandboxes ] - return json.dumps(compact, indent=2)[:200_000] - - -def _read(path: Any, limit: int = 20000) -> str: - try: - text = path.read_text(errors="replace") - except Exception: # noqa: BLE001 - return "" - return text if len(text) <= limit else text[:limit] + "\n…truncated…" + available = [s["name"] for s in sandboxes if s["available"]] + empty = ( + "No agent is qualified as stable for this model yet." + if hidden + else "No agent is available for this endpoint." + ) + return { + "stamp": time.time(), + "task": ( + {k: selection.get(k) for k in ("dataset", "index", "title")} + if selection and selection.get("spec") + else None + ), + "rollouts": settings.rollouts, + "sources": sources, + "server": server, + "local_token": settings.local_token and bool(_local_token()), + "hf_login": { + "on": settings.hf_login, + "user": (profile.username or profile.name) if profile is not None else None, + }, + "private_urls": settings.private_urls, + "engine": { + "ok": bool(engine.get("ok")), + "model": engine.get("model"), + "host": engine.get("host"), + "source": engine.get("source") + or ("server" if engine.get("server_default") else ""), + "train": engine.get("purpose") == "train", + "level_text": ui_pages.LEVEL_TEXT.get( + engine.get("capture_level") or "", "" + ), + "notes": engine.get("notes") or [], + "reason": engine.get("reason"), + }, + "custom": engine.get("custom"), + "agents": agents, + "agent": "opencode" + if "opencode" in values + else (values[0] if values else None), + "agents_empty": empty, + "hidden_agents": hidden, + "include_experimental": bool(engine.get("include_experimental")), + "sandboxes": sandboxes, + "sandbox": "e2b" + if "e2b" in available + else (available[0] if available else None), + # A capture proxy on loopback cannot be reached from a remote sandbox, only by host-side agents. + "proxy_local": host in ("127.0.0.1", "localhost", "0.0.0.0", "::1"), + "message": {"tone": message[0], "text": message[1]} if message else None, + } + + +def _payload(evt: Any) -> dict[str, Any]: + data = getattr(evt, "_data", None) + return data if isinstance(data, dict) else {} + + +def _signature(listing: list[dict[str, Any]]) -> str: + """What changes when a run changes state; ticks that change nothing re-render nothing.""" + return json.dumps([(r.get("id"), r.get("status")) for r in listing]) + + +def _js(name: str) -> str: + """A component's script, with the shared icon set in front of it.""" + return js_prelude() + _asset(name) def harbor_gradio_builder( @@ -799,713 +740,774 @@ def harbor_gradio_builder( Args: datasets (`list[str]`, *optional*): - Dataset specs served by this server; each becomes a selectable split. + Dataset specs served by this server. Each is browsable; Hub datasets can be added from the + page when `OPENENV_HARBOR_UI_ADD_DATASETS` allows it. + title (`str`, *optional*): + Page title. Defaults to `"OpenEnv × Harbor"`. Returns: `gr.Blocks`: The interface. """ - from .tasks import HarborTaskProvider, resolve_task_dirs - - datasets = list(datasets or []) - - def on_validate( - url: str, - model: str, - api_key: str, - provider: str = "openai", - purpose: str = "eval", - include_experimental: bool = False, - ): - from openenv.core.harness.capture.validate_llm import list_models, validate_llm + from .serving import HarborService + + served = list(datasets or []) + title = title or "OpenEnv × Harbor" + settings = ui_settings.load() + runs = ui_runs.manager() + if not settings.private_urls: + ui_settings.guard_redirects() + _load_added(settings, served) + with _ADDED_LOCK: + for spec in ui_data.added_in_bucket(settings, served): + if spec not in _ADDED: + _ADDED.append(spec) + + def viewer(visitor: str | None) -> str | None: + """Whose runs this visitor may see: everyone's (`None`), or their own.""" + return ( + None if settings.run_visibility == "all" else ui_settings.owner_of(visitor) + ) - from .capabilities import capabilities - from .seams import agent_facing_model, get as get_seam - from .serving import HarborService + def listing(visitor: str | None) -> list[dict[str, Any]]: + return runs.list(viewer(visitor)) - url = (url or "").strip().rstrip("/") - api_key = (api_key or "").strip() or None - if not url: - return ( - _UNVALIDATED, - gr.update(), - gr.update(), - {}, - gr.update(interactive=False), + # History is any `*.json` in the runs folder, so a record can be old or hand-edited. One that + # cannot be drawn is reported as that, rather than breaking the tab it is on. + def runs_page( + visitor: str | None, selected: str = "", compare: list[str] | None = None + ) -> str: + try: + return ui_pages.runs_html( + listing(visitor), + selected, + compare, + history=runs.store is not None, + own=settings.run_visibility == "own", ) + except Exception as exc: # noqa: BLE001 + return ui_pages.unreadable("Runs", exc) - if not model: - served = list_models(url, api_key=api_key) - if len(served) != 1: - hint = ( - f"`{', '.join(served[:12])}`" - if served - else "nothing reachable — check the URL, and the API key if it needs one" - ) - return ( - f"**Pick a model** — this endpoint serves {hint}.", - gr.update(), - gr.update(), - {}, - gr.update(interactive=False), - ) - model = served[0] - - report = validate_llm(url, model, api_key=api_key, provider=provider) - if not report.reachable or (purpose == "train" and not report.trainable): - why = "; ".join(report.findings) or ( - "Exact engine tokens are required for training" - if report.reachable - else "unreachable" - ) - return ( - f"**Not usable** — {why}\n\n" - "Needs vLLM with `--return-tokens-as-token-ids --logprobs-mode " - "processed_logprobs`, SGLang built from git main, or any reachable OpenAI-spec " - "endpoint (with an API key) for eval rollouts.", - gr.update(), - gr.update(), - {}, - gr.update(interactive=False), + def run_page(run_id: str, visitor: str | None) -> str: + owner = viewer(visitor) + try: + return ui_pages.run_html( + runs.get(run_id, owner), + runs.live(run_id, owner), + HarborService.current(), ) + except Exception as exc: # noqa: BLE001 + return ui_pages.unreadable("Runs", exc, run=True) + + # ── server functions: called from the components' JavaScript ────────────────────────────── + # Gradio passes a server function one value: nothing becomes `[]`, one argument arrives as is, + # several arrive as a list. So each takes exactly one parameter. + def hb_datasets(_: Any = None) -> dict[str, Any]: + from .tasks import resolve_task_dirs + + out = [] + with _ADDED_LOCK: + added = list(_ADDED) + for spec in served + [s for s in added if s not in served]: + row: dict[str, Any] = { + "spec": spec, + "label": ui_data.hub_id(spec), + "added": spec not in served, + # Removing is adding's undo, so it takes the same permission. + "removable": spec in added + and spec not in served + and _can_add_datasets(), + } + try: + row["num_tasks"] = len(resolve_task_dirs(spec)) + except Exception as exc: # noqa: BLE001 - one broken dataset must not hide the others + row["error"] = f"{type(exc).__name__}: {str(exc)[:160]}" + out.append(row) + return {"datasets": out, "can_add": _can_add_datasets()} + + def hb_tasks(spec: str) -> dict[str, Any]: + if not _allowed(spec, served): + return {"error": "That dataset is not served here."} + try: + return {"rows": ui_data.task_rows(spec)} + except Exception as exc: # noqa: BLE001 + return {"error": f"{type(exc).__name__}: {str(exc)[:300]}"} - caps = capabilities( - datasets=datasets, - llm={ - "url": url, - "model": model, - "ok": report.ok, - "capture_level": report.capture_level, - "reachable": True, - "authenticated": bool(api_key), - }, - ) - sandboxes = caps.available_sandboxes - from .qualification import harness_maturity_rows - + def hb_hub(query: str) -> list[dict[str, Any]]: + if not _can_add_datasets(): + return [] try: - evidence = ( - json.loads(Path(report_path).read_text()) if report_path else None - ) - tiers = { - name: tier - for name, tier, _ in harness_maturity_rows( - [h.name for h in caps.harnesses], evidence - ) + return ui_data.search_hub(query) + except Exception: # noqa: BLE001 - the Hub being unreachable just means no suggestions + return [] + + def hb_add(spec: str) -> dict[str, Any]: + """Start adding a Hub dataset; the page follows it with `hb_add_status`.""" + spec = (spec or "").strip() + if not _can_add_datasets(): + return { + "state": "error", + "error": "Adding datasets is turned off on this server.", } - except (OSError, ValueError, TypeError): - tiers = {h.name: "experimental" for h in caps.harnesses} - evidence = None - profile_provider = "vllm" if purpose == "train" else provider - harness_profiles = {} - unavailable_profiles = set() - for cell in (evidence or {}).get("cells", []): - if cell.get("provider") != profile_provider: - continue - config = cell.get("configuration") or {} - profile = config.get("acp_profile") or config.get("nemo_profile") - if profile: - name = cell["harness"] - harness_profiles[name] = profile - try: - get_seam(name, profile=profile) - except (ValueError, KeyError): - unavailable_profiles.add(name) - choices = [ - ( - f"{h.name} ({h.dialect}; {tiers[h.name]}" - + ( - f"; profile: {harness_profiles[h.name]}" - if h.name in harness_profiles - else "" - ) - + ")", - h.name, - ) - for h in sorted(caps.harnesses, key=lambda h: h.name) - if h.name not in unavailable_profiles - and ( - tiers[h.name] == "stable" - or (include_experimental and tiers[h.name] == "experimental") - ) - ] - values = [v for _, v in choices] + target = ui_data.added_spec(spec, settings) + with _ADDED_LOCK: + if target in served or target in _ADDED or spec in served: + return {"spec": spec, "state": "done", "target": target} + full = len(_ADDED) >= _MAX_ADDED + problem = _hub_problem(spec) + if problem: + return {"spec": spec, "state": "error", "error": problem} + if full: + return { + "spec": spec, + "state": "error", + "error": f"This server already holds {_MAX_ADDED} added datasets.", + } + return ui_data.start_add(spec, _remember_added, settings) + + def hb_remove(spec: str) -> dict[str, Any]: + """Remove a dataset added from the page, files and all: its bucket folder, or its download. + One the server was started with is not the page's to remove.""" + spec = str(spec or "") + if not _can_add_datasets(): + return { + "error": "Adding and removing datasets is turned off on this server." + } + return _remove_added(spec, served, settings) - leaf = agent_facing_model(model) - if purpose == "train": - lines = [ - f"**Endpoint ready — TRAINING CAPTURE** · `{model}` · token ids + logprobs ✓" - ] - else: - detail = ( - "token capture available; training export disabled for this eval" - if report.capture_level == "tokens" - else "logprobs, no token ids" - if report.capture_level == "logprobs" - else "no token ids, no logprobs" - ) - lines = [ - f"**Endpoint ready — EVAL** · `{model}` · {detail}", - "Rollouts carry the reward and the full trace, but nothing trainable.", - ] - if leaf != model: - lines.append(f"Sent to agents as `{leaf}`, rewritten back on the way out.") - if not values: - lines.append( - "No agents match the support filter. Load qualification evidence or explicitly include experimental adapters." - ) - for fix in report.param_fixes: - lines.append(f"upstream compat: {fix}") - # The one thing a user cannot discover by reading the endpoint's own docs: whether a model - # will actually sustain an agent loop here. Shown at Validate rather than after a rollout, - # because a rollout costs a sandbox and several minutes to learn the same thing. - for finding in report.findings: - if "behaviour_changed" in finding or "tool_call" in finding: - detail = finding.split(": ", 2)[-1] - lines.append(f"⚠️ {detail}") - - # Run uses the endpoint typed above. The engine is a per-rollout argument, so a browser can - # point this server at any reachable OpenAI-spec endpoint without restarting it — which is - # the whole point of validating a URL here. Say which one will be used, because a server may - # also have been booted with a default and the two can differ. - service = HarborService.current() - if ( - service is not None - and service.llm_url - and service.llm_url.rstrip("/") != url - ): - lines.append( - f"Rollouts will use **this** endpoint, not the server's default " - f"(`{service.llm_url}`)." - ) - lines.append( - f"Sandboxes: {', '.join(f'`{s}`' for s in sandboxes) or '**none usable**'}" - ) - blocked = [s.name for s in caps.sandboxes if not s.available] - if blocked: - lines.append( - f"unavailable: {', '.join(blocked)}" - ) + def hb_add_status(spec: str) -> dict[str, Any]: + return ui_data.add_status(str(spec or "")) or {"spec": spec, "state": "unknown"} - return ( - " \n".join(lines), - gr.update( - choices=choices, - value="opencode" - if "opencode" in values - else (values[0] if values else None), - ), - gr.update(choices=sandboxes, value=sandboxes[0] if sandboxes else None), - # `ok` gates the Run button and now means "reachable", not "trainable": an eval endpoint - # is a perfectly good thing to press Run against. - { - "url": url, - "model": model, - "ok": True, - "capture_level": report.capture_level, - "trainable": report.trainable, - # Carried so Run can reach a token-gated endpoint. Without it, validating a hosted - # provider succeeded and pressing Run then failed to authenticate against the same - # URL. `gr.State` is held server-side and this is never rendered back into the page, - # which is the same rule the API key box itself follows. - "api_key": api_key or "", - "provider": provider, - "purpose": purpose, - "allowed_harnesses": values, - "harness_profiles": harness_profiles, - }, - gr.update( - interactive=bool(sandboxes and values), - value="Run training capture" - if purpose == "train" - else "Run eval rollout", - ), - ) + def hb_inspect(spec: str) -> dict[str, Any]: + """A Hub dataset's task count (null when not in Harbor's layout) and size, before adding it.""" + spec = str(spec or "").strip() + if not _can_add_datasets() or not _HUB_ID.match(spec): + return {"spec": spec, "tasks": None, "bytes": None} + try: + return {"spec": spec, **_inspect_hub(spec)} + except Exception as exc: # noqa: BLE001 - distinguish Hub failure from an empty dataset + return { + "spec": spec, + "tasks": None, + "bytes": None, + "error": f"Could not inspect it: {type(exc).__name__}: {str(exc)[:200]}", + } + + def _remember_added(spec: str) -> None: + with _ADDED_LOCK: + if spec not in _ADDED: + _ADDED.append(spec) + _save_added(settings) - def on_dataset(spec: str): - if not spec: - return gr.update(), "" + def hb_file(args: list[Any]) -> dict[str, Any]: + spec, index, path = (list(args or []) + ["", 0, ""])[:3] + if not _allowed(str(spec), served): + return {"error": "That dataset is not served here."} try: - n = len(resolve_task_dirs(spec)) + return ui_data.read_task_file(spec, int(index), str(path)) except Exception as exc: # noqa: BLE001 - return gr.update(value=0), f"Cannot load `{spec}` — {exc}" - return gr.update(value=0), f"**{n}** tasks · 0–{n - 1}" + return {"error": f"{type(exc).__name__}: {str(exc)[:200]}"} - def on_task(spec: str, index: int): - if not spec: - return "", "", "", "", "" + def hb_models(_: Any = None) -> dict[str, Any]: + if not settings.visitor_endpoints: + return {"models": []} try: - task_dir = HarborTaskProvider([spec]).task_dir(spec, int(index)) + return {"models": ui_data.hf_models()} + except Exception as exc: # noqa: BLE001 - no list means the visitor types a model id instead + return {"models": [], "error": f"Could not list models: {str(exc)[:160]}"} + + def hb_served(args: list[Any]) -> dict[str, Any]: + """The models an endpoint serves, for the card's model field. Guarded like a connect.""" + from openenv.core.harness.capture.validate_llm import list_models + + url, key = (list(args or []) + ["", ""])[:2] + url = str(url or "").strip().rstrip("/") + if not settings.visitor_endpoints: + return {"error": "This server runs rollouts on its own endpoint only."} + problem = ui_settings.url_problem(url, private_ok=settings.private_urls) + if problem: + return {"error": problem} + models = list_models(url, timeout=10, api_key=str(key or "").strip() or None) + if not models: + return { + "error": "Nothing listed at that URL. Check it, and the key if it needs one." + } + return {"models": models[:200]} + + def hb_download(args: list[Any]) -> dict[str, Any]: + token, kind = (list(args or []) + ["", ""])[:2] + run_id = ui_pages.granted(str(token)) + rec = runs.get(run_id) if run_id else None + result = (rec or {}).get("result") + if not result: + return {"error": "This link has expired. Open the run again."} + name = re.sub(r"[^A-Za-z0-9_.-]", "_", str(rec.get("id") or "rollout")) + if kind == "contract": + try: + contract = _contract(result) + except ( + ValueError + ) as exc: # the exporter refuses a rollout with FATAL findings + return {"error": str(exc)[:200]} + if contract is None: + return { + "error": "An eval rollout has nothing to train on, so it has no contract." + } + return { + "name": f"{name}.contract.json", + "text": json.dumps(contract, indent=2), + } + return { + "name": f"{name}.json", + "text": json.dumps(result, indent=2, default=str), + } + + # ── handlers ───────────────────────────────────────────────────────────────────────────────── + # Each tab shows a list or one item, never both. Which one is decided by the stylesheet from what + # is rendered (a task head means a task is open; a run page means a run is), not by toggling + # visibility, so the list keeps its filters and scroll, and no late update can show both. + + def on_load(visitor: str | None): + # Not the run card or its engine: those belong to the task page and are written when a task + # opens. A `#task=` link opens one while this is still running, and writing the card here too + # would replace that task's card with an empty one whenever this finished second. + visitor = visitor or ui_settings.new_visitor() + caps = _capabilities(served) + return ( + ui_pages.header_html(title, served, caps, settings), + ui_pages.setup_html(caps, served, settings), + runs_page(visitor), + _signature(listing(visitor)), + visitor, + ) + + def open_task( + engine: dict, visitor: str | None, spec: str, index: int, profile: Any = None + ): + """The task page for one task: (selection, head, body, card, engine).""" + caps = _capabilities(served) + # The first task a page opens starts from the server's endpoint. + engine = engine or _server_engine(caps, settings=settings) + if not _allowed(spec, served): + return (gr.skip(),) * 5 + try: + detail = ui_data.task_detail(spec, index) except Exception as exc: # noqa: BLE001 - return f"_{exc}_", "", "", "", "" - env_dir, tests_dir = task_dir / "environment", task_dir / "tests" + return ( + {}, + ui_pages.task_error_head(spec, index), + f'
{ui_pages.empty("alert", "Could not open this task", html.escape(str(exc)[:300]))}
', + _card(engine, None, caps, settings=settings, profile=profile), + engine, + ) + selection = { + "spec": spec, + "dataset": spec, + "index": index, + "name": detail["name"], + "title": detail["title"], + } + mine = [ + r + for r in listing(visitor) + if r.get("dataset") == spec and r.get("task_index") == index + ] return ( - f"`{task_dir.name}`", - _read(task_dir / "instruction.md"), - _read(env_dir / "Dockerfile"), - _read(task_dir / "task.toml"), - _read(tests_dir / "test.sh"), + selection, + ui_pages.task_head_html(detail, ui_data.common_tags(spec)), + ui_pages.task_html(detail, mine), + _card(engine, selection, caps, settings=settings, profile=profile), + engine, + ) + + def on_open_task( + engine: dict, + visitor: str | None, + evt: gr.EventData, + profile: gr.OAuthProfile | None = None, + ): + data = _payload(evt) + return open_task( + engine, + visitor, + str(data.get("spec") or ""), + int(data.get("index") or 0), + profile, + ) + + def on_back_to_tasks(): + return "" + + def on_tab_again(sel: dict, visitor: str | None, evt: gr.EventData): + """Tasks or Runs clicked while already showing: back to that tab's list.""" + skip = gr.skip() + if _payload(evt).get("back") == "tasks": + return "", skip, skip, skip + state = {**sel, "id": ""} + return skip, state, "", runs_page(visitor, "", sel.get("compare", [])) + + def show_run(run_id: str, visitor: str | None, sel: dict): + return {"id": run_id, "compare": sel.get("compare", [])}, run_page( + run_id, visitor ) - def on_run(engine: dict, spec: str, index: int, harness: str, sandbox: str): - """Stream progress while the rollout runs, then the result and its graph.""" - import asyncio - import queue - import threading - import time - from pathlib import Path + def on_open_run(sel: dict, visitor: str | None, evt: gr.EventData): + return show_run(str(_payload(evt).get("id") or ""), visitor, sel) + + def on_pick_runs(sel: dict, evt: gr.EventData): + ids = [str(i) for i in (_payload(evt).get("ids") or [])][:4] + return {**sel, "compare": ids} + + def on_compare(sel: dict, visitor: str | None, evt: gr.EventData): + ids = [str(i) for i in (_payload(evt).get("ids") or [])][:4] + owner = viewer(visitor) + return {"id": "", "compare": ids}, ui_pages.compare_html( + [runs.get(i, owner) for i in ids] + ) - from .rollout import run_rollout as _run - from .serving import HarborService + def on_task_page(sel: dict, visitor: str | None, evt: gr.EventData): + """A run listed on a task page was clicked: show it on the Runs tab.""" + run_id = str(_payload(evt).get("run") or "") + if not run_id: + return (gr.skip(),) * 4 + state, view = show_run(run_id, visitor, sel) + return gr.Tabs(selected="runs"), state, view, runs_page(visitor, run_id) + + def on_run_page( + engine: dict, + sel: dict, + visitor: str | None, + evt: gr.EventData, + profile: gr.OAuthProfile | None = None, + ): + """The run page's own links: back to the list, another run, or the task it ran.""" + data = _payload(evt) + skip = gr.skip() + if data.get("run"): + return (*show_run(str(data["run"]), visitor, sel), *(skip,) * 7) + if data.get("task") is not None: + task = open_task( + engine, + visitor, + str(data.get("task") or ""), + int(data.get("index") or 0), + profile, + ) + return (skip, skip, gr.Tabs(selected="tasks"), *task, skip) + listing_now = runs_page(visitor, "", sel.get("compare", [])) + return ({**sel, "id": ""}, "", skip, *(skip,) * 5, listing_now) + + def on_experimental( + engine: dict, + selection: dict, + evt: gr.EventData, + profile: gr.OAuthProfile | None = None, + ): + include = bool(_payload(evt).get("include_experimental")) + caps = _capabilities(served) + if engine.get("server_default") or not engine: + engine = _server_engine( + caps, include_experimental=include, settings=settings + ) + elif not engine.get("ok"): + engine = {**engine, "include_experimental": include} + else: + choices, profiles, hidden_n = _agent_choices( + caps, + purpose=engine.get("purpose", "eval"), + provider=engine.get("provider", "openai"), + include_experimental=include, + ) + engine = { + **engine, + "choices": choices, + "allowed_harnesses": [v for _, v in choices], + "harness_profiles": profiles, + "hidden_agents": hidden_n, + "include_experimental": include, + } + return _card( + engine, selection, caps, settings=settings, profile=profile + ), engine + + def on_connect( + engine: dict, + selection: dict, + evt: gr.EventData, + profile: gr.OAuthProfile | None = None, + token: gr.OAuthToken | None = None, + ): + new = connect_endpoint( + _payload(evt), + include_experimental=bool(engine.get("include_experimental")), + datasets=served, + settings=settings, + account_token=token.token if token is not None else None, + ) + caps = _capabilities(served) + if not new.get("ok"): + # A failed check leaves the working engine exactly as it was, `custom` included: the card + # compares its form with `custom` to know whether Run still means what it shows. + return _card( + engine, + selection, + caps, + ("bad", str(new.get("reason"))), + settings, + profile=profile, + ), engine + return _card( + new, + selection, + caps, + ("ok", f"Connected to {new['model']}."), + settings, + profile=profile, + ), new + + def on_server_default( + engine: dict, selection: dict, profile: gr.OAuthProfile | None = None + ): + caps = _capabilities(served) + new = _server_engine( + caps, + include_experimental=bool(engine.get("include_experimental")), + settings=settings, + ) + return _card(new, selection, caps, settings=settings, profile=profile), new + + def on_run( + engine: dict, + selection: dict, + visitor: str | None, + evt: gr.EventData, + profile: gr.OAuthProfile | None = None, + ): + """Start a rollout and show it on the Runs tab. It keeps running if the page is closed.""" + data = _payload(evt) + harness, sandbox = str(data.get("agent") or ""), str(data.get("sandbox") or "") + caps = _capabilities(served) + keep = (gr.skip(),) * 5 + def say(text: str): + return ( + _card( + engine, selection, caps, ("bad", text), settings, profile=profile + ), + *keep, + ) + + if not settings.rollouts: + return say("Rollouts are turned off on this server.") if not engine.get("ok"): - yield _UNVALIDATED, "", "", "{}", None, gr.update(interactive=True) - return + return say("Connect a model first.") + if engine.get("server_default") and not settings.server_endpoint: + return say("This server's endpoint is not shared. Connect your own model.") + if not engine.get("server_default") and not settings.visitor_endpoints: + return say("This server runs rollouts on its own endpoint only.") if harness not in engine.get("allowed_harnesses", []): - yield ( - "Selected harness is outside the validated support filter. Validate again.", - "", - "", - "{}", - None, - gr.update(interactive=False), + return say( + "That agent is outside the qualified set for this endpoint. Pick another." + ) + spec = (selection or {}).get("spec", "") + if not _allowed(spec, served): + return say("Pick a task first.") + # The card only offers these, but the event is the browser's to send. + if sandbox not in caps.available_sandboxes: + return say("That sandbox is not available on this server.") + # Harbor fills `${VAR}` in a task's settings from this server's environment, where its keys + # are. Even a served task cannot receive them when a visitor controls the model: its trace is + # visible to that visitor, so the model can print a secret it finds in the sandbox. + try: + reads_env = ui_data.reads_environment(spec, int(selection.get("index", 0))) + except Exception as exc: # noqa: BLE001 - a task that cannot be read cannot be run + return say( + f"Could not read this task: {type(exc).__name__}: {str(exc)[:200]}" + ) + if reads_env and spec not in served: + return say( + "This task reads environment variables or files from the server, which only " + "datasets the server was started with may do." + ) + if reads_env and not engine.get("server_default"): + return say( + "This task passes the server's environment variables or files into the sandbox, so " + "it runs only on the server's own endpoint, never on a model you connect." ) - return service = HarborService.current() if service is None: - yield ( - "Server not initialised — no capture proxy running.", - "", - "", - "{}", - None, - gr.update(interactive=True), + return say( + "This server has no capture proxy running, so it cannot run rollouts." ) - return + visitor = visitor or ui_settings.new_visitor() try: - task_dir = HarborTaskProvider([spec]).task_dir(spec, int(index)) - except Exception as exc: # noqa: BLE001 - yield ( - f"Bad task — {html.escape(str(exc))}", - "", - "", - "{}", - None, - gr.update(interactive=True), - ) - return - - done: queue.Queue = queue.Queue(maxsize=1) - live_sessions: queue.Queue[str] = queue.Queue(maxsize=1) - - async def _run_with_engine(): - """Resolve the engine the user validated, then run against it. - - The engine is per rollout, so the URL in the box is the one used. Resolving it through the - capture server's pool means the tier comes from a real probe of that endpoint rather than - from whatever the server happened to boot with — and the probe is cached, so pressing Run - repeatedly costs nothing after the first time. - """ - from openenv.core.harness.capture.sessions import Upstream - - pool = service.capture.app.state.upstreams - typed_url = str((engine or {}).get("url") or "").strip() - if typed_url: - upstream = Upstream( - llm_url=typed_url, - model=str((engine or {}).get("model") or ""), - api_key=str((engine or {}).get("api_key") or "") or None, - provider=str((engine or {}).get("provider") or "openai"), - ) - client, level = await pool.resolve(upstream) - served = client.served_model or upstream.model - else: - # Nothing validated in the box: fall back to the server's default, which is what a - # server booted with --llm-url provides. With neither, the rollout reports the - # missing engine rather than silently producing an untrainable result. - upstream, (client, level) = None, pool.default - level = getattr(service, "capture_level", "text") - served = service.model - return await _run( - task_dir=task_dir, + run_id = runs.start( + engine=engine, + spec=spec, + index=int(selection.get("index", 0)), harness=harness, - harness_profile=engine.get("harness_profiles", {}).get(harness), sandbox=sandbox, - registry=service.capture.registry, - intercept_url=service.public_url, - model=served, - trials_dir=Path("/tmp/openenv-harbor-trials"), - dataset=spec, - capture_level=level, - purpose=str((engine or {}).get("purpose") or "eval"), - upstream=upstream, - inference=client, - on_session_created=live_sessions.put_nowait, + service=service, + owner=ui_settings.owner_of(visitor), + title=str(selection.get("title") or ""), + per_owner=settings.max_runs_per_visitor, + private_urls=settings.private_urls, + # a browser id is free to replace, so a signed-in visitor's cap follows the account + quota=f"hf:{profile.username}" + if profile is not None and profile.username + else "", ) + except (RuntimeError, IndexError, ValueError) as exc: + return say(str(exc)) + except Exception as exc: # noqa: BLE001 - the card must never be left on "Starting…" + return say(f"Could not start: {type(exc).__name__}: {str(exc)[:200]}") + return ( + _card( + engine, + selection, + caps, + ("ok", f"Started {harness} on {sandbox}."), + settings, + profile=profile, + ), + gr.Tabs(selected="runs"), + {"id": run_id, "compare": []}, + runs_page(visitor, run_id), + run_page(run_id, visitor), + _signature(listing(visitor)), + ) - def worker() -> None: - try: - res = asyncio.run(_run_with_engine()) - done.put(("ok", res.model_dump())) - except Exception as exc: # noqa: BLE001 - show it, never take the server down - done.put(("err", f"{type(exc).__name__}: {exc}")) - - thread = threading.Thread(target=worker, daemon=True) - started = time.monotonic() - thread.start() - session_id = None - - while thread.is_alive(): - if session_id is None: - try: - session_id = live_sessions.get_nowait() - except queue.Empty: - pass - stats, phase, stage = None, "starting up", 0 - if session_id: - session = service.capture.registry.get(session_id) - if session is not None: - st = session.graph.stats() - # n_trainable_tokens only exists after export; mid-run we can count only what - # has been sampled, before masking and discards. - sampled = sum( - len(n.sampled_ids or []) for n in session.graph.nodes() - ) - turns = st.get("n_turns", 0) - stats = { - "calls": turns, - "roots": st.get("n_roots", 0), - "sampled tokens": sampled, - "discarded": st.get("n_discarded", 0), - } - if turns: - stage = -1 # past setup; the numbers mean something now - phase = "agent working" - # Only meaningful once a call has landed. Before that `idle_seconds` counts - # from session creation, which renders as a stall during a normal boot. - stats["since last call"] = f"{session.idle_seconds:.0f}s" - else: - stage = 0 - phase = "preparing the run or waiting for its first response" - # The transcript rides in the graph slot: it is empty until the run finishes anyway, - # and the two answer the same question at different times. - transcript = "" - if session_id: - live = service.capture.registry.get(session_id) - if live is not None: - transcript = _transcript_html(live) - yield ( - _live_html( - harness, sandbox, phase, time.monotonic() - started, stats, stage - ), - transcript, - "", - "{}", - None, - gr.update(interactive=False), - ) - time.sleep(2.0) - - kind, payload = done.get() - if kind == "err": - yield ( - f'
Run failed
' - f'
{html.escape(payload)}
', - "", - "", - "{}", - None, - gr.update(interactive=True), - ) - return - contract = None - contract_error = "" - try: - contract = _write_contract(payload) - except (ValueError, TypeError) as exc: - contract_error = ( - "

Training export rejected: " + html.escape(str(exc)) + "

" - ) - yield ( - _result_html(payload) + contract_error, - _conversation_html(payload), - _turns_html(payload), - _summary_json(payload), - contract, - gr.update(interactive=True), + def on_tick(sel: dict, sig: str, visitor: str | None): + """Refresh the run list when a run changes state, and the open run while it is live.""" + items = listing(visitor) + new_sig = _signature(items) + run_id = sel.get("id") or "" + live = runs.live(run_id, viewer(visitor)) if run_id else None + changed = new_sig != sig + list_out = ( + runs_page(visitor, run_id, sel.get("compare", [])) if changed else gr.skip() + ) + # live: redraw as it grows; just finished: draw the result once + view = ( + run_page(run_id, visitor) + if run_id and (live is not None or changed) + else gr.skip() ) + return list_out, view, new_sig - with gr.Blocks(title=title or "Harbor") as app: - gr.HTML(f"") - state = gr.State({}) + def on_refresh_setup(): + caps = _capabilities(served, refresh=True) + return ui_pages.setup_html(caps, served, settings), ui_pages.header_html( + title, served, caps, settings + ) - with gr.Column(elem_classes="hb-wrap"): - gr.Markdown( - "## Harbor task playground\nChoose a model, validate the connection, then run an agent " - "on a task. Follow its tool calls and results below. " - "Evaluation works with supported hosted providers; training capture requires " - "verified engine token IDs and log probabilities." - ) + # ── layout ─────────────────────────────────────────────────────────────────────────────────── + html_opts = {"padding": False, "apply_default_css": False, "elem_classes": "hb"} + with gr.Blocks(title=title, fill_width=True) as app: + # apply_default_css=False on every HTML block: Gradio's default wraps it in `.prose`, and the + # OpenEnv theme strips the border and background from anything directly inside `.prose`. + gr.HTML( + f"{ui_icons.json_tag()}", + padding=False, + apply_default_css=False, + ) + header = gr.HTML( + ui_pages.header_html(title, served, settings=settings), + js_on_load=_js("header.js"), + **html_opts, + ) + engine_state = gr.State({}) + selection = gr.State({}) + run_sel = gr.State({"id": "", "compare": []}) + runs_sig = gr.State("") + # Which runs are this browser's: a random id kept in the browser, stored on each run as a digest. + visitor = gr.BrowserState( + None, + storage_key="openenv-harbor-visitor", + secret=ui_settings.visitor_secret(settings), + ) - with gr.Row(equal_height=False): - # left — the model - with gr.Column(scale=1, elem_classes="hb-cell"): - with gr.Column(elem_classes="hb-panel"): - gr.Markdown("### 1 · Connect a model") - # Deliberately empty. Prefilling meant the box already held whatever URL the - # server was started with, so Validate confirmed a value nobody chose and a - # stale endpoint could be used without anyone noticing it was stale. - provider_in = gr.Dropdown( - label="Upstream provider", - choices=[ - ("OpenAI-compatible", "openai"), - ("Anthropic native", "anthropic"), - ("Hugging Face Inference Providers", "hf"), - ("vLLM", "vllm"), - ], - value="openai", - info="Select the upstream API. Exact training tokens are verified separately.", - ) - purpose_in = gr.Dropdown( - label="Use", - choices=[ - ("Evaluation", "eval"), - ("Training capture", "train"), - ], - value="eval", - ) - url_in = gr.Textbox( - label="LLM URL", - placeholder="https://your-endpoint/v1", - info="vLLM, SGLang, OpenAI, Anthropic, HF Inference Providers. " - "Accepts a bare root or one ending in /v1.", - ) - gr.HTML(_labelled("API key (optional)", _KEY_TIP)) - key_in = gr.Textbox( - label="", - type="password", - placeholder="only for a hosted provider", - show_label=False, - ) - model_in = gr.Textbox( - label="Model (optional)", - placeholder="read from the endpoint", - info="Required when the endpoint serves more than one model.", - ) - validate_btn = gr.Button( - "Validate connection", variant="secondary" - ) - gr.HTML(_labelled("Capture level", _LEVEL_TIP)) - engine_md = gr.Markdown(_UNVALIDATED) - - # right — task preview and agent. Implementation files stay folded away. - with gr.Column(scale=1, elem_classes="hb-cell"): - gr.Markdown("### 2 · Choose a task") - with gr.Row(): - ds_in = gr.Dropdown( - label="Dataset", - choices=datasets, - value=datasets[0] if datasets else None, - scale=3, - ) - idx_in = gr.Number( - label="Index", value=0, precision=0, minimum=0, scale=1 - ) - count_md = gr.Markdown() - task_md = gr.Markdown() - instruction_box = gr.Code( - label="Task instruction", language="markdown", lines=10 + with gr.Tabs(selected="tasks", elem_classes="hb-tabs") as tabs: + with gr.Tab("Tasks", id="tasks"): + with gr.Column(elem_classes="hb-page hb-browse"): + browser = gr.HTML( + html_template=_asset("task_browser.html"), + js_on_load=_js("task_browser.js"), + server_functions=[ + hb_datasets, + hb_tasks, + hb_hub, + hb_add, + hb_add_status, + hb_inspect, + hb_remove, + ], + **html_opts, ) - with gr.Accordion("Task files and grader", open=False): - with gr.Accordion("Dockerfile", open=False): - dockerfile_box = gr.Code( - label="", language="dockerfile", lines=12 + with gr.Column(elem_classes="hb-page hb-taskpage"): + task_head = gr.HTML("", js_on_load=_js("task_head.js"), **html_opts) + with gr.Row(elem_classes="hb-row", equal_height=False): + with gr.Column(elem_classes="hb-main", min_width=0): + task_view = gr.HTML( + "", + js_on_load=_js("task_view.js"), + server_functions=[hb_file], + **html_opts, + ) + with gr.Column(elem_classes="hb-side", min_width=0): + card = gr.HTML( + {}, + html_template="", + js_on_load=_js("run_card.js"), + server_functions=[hb_models, hb_served], + **html_opts, ) - with gr.Accordion("task.toml", open=False): - toml_box = gr.Code(label="", language="python", lines=12) - with gr.Accordion("Grader", open=False): - tests_box = gr.Code(label="", language="shell", lines=12) - - with gr.Column(elem_classes="hb-panel"): - gr.Markdown("### 3 · Choose an agent") - experimental_in = gr.Checkbox( - label="Include experimental adapters", - value=False, - info="Stable adapters are shown by default. Unstable adapters remain excluded.", - ) - harness_in = gr.Dropdown( - label="Agent", - choices=[], - info="The coding agent to run. Its dialect is shown in brackets; the " - "capture proxy connects it to your selected provider.", - ) - sandbox_in = gr.Dropdown( - label="Sandbox", - choices=[], - info="Where the agent executes. Harbor's backends, not OpenEnv's " - "container providers — only those with working credentials are listed.", - ) - - # Full width, under both columns: the action belongs to the pair, not to either one. - run_btn = gr.Button( - "Run rollout", variant="primary", interactive=False, scale=1 - ) - with gr.Column(elem_classes="hb-panel"): - gr.Markdown("### Run status") - result_html = gr.HTML( - '

Validate your model and choose a task to begin.

' - ) - with gr.Column(elem_classes="hb-panel"): - gr.Markdown( - "### Live trace\nAgent messages, tool calls and tool results appear as " - "model calls complete. A request in progress may take a moment." - ) - convo_html = gr.HTML() - with gr.Accordion("Token details and training export", open=False): - analysis_html = gr.HTML() - contract_file = gr.File( - label="Training contract — captured tokens, log probabilities and reward", - interactive=False, - visible=True, - ) - with gr.Accordion("Harness/provider qualification evidence", open=False): - import os - from pathlib import Path - - from .qualification import ( - harness_maturity_rows, - qualification_details, - qualification_rows, - ) - from .seams import SEAMS - - report_path = os.environ.get("OPENENV_HARBOR_QUALIFICATION_REPORT", "") - - def read_qualification_evidence(): - try: - evidence = ( - json.loads(Path(report_path).read_text()) - if report_path - else None - ) - return ( - qualification_rows(list(SEAMS), evidence), - qualification_details(evidence), - harness_maturity_rows(list(SEAMS), evidence), - "Loaded recorded evidence." - if evidence - else "No qualification report configured.", - ) - except (OSError, ValueError, TypeError) as exc: - return ( - qualification_rows(list(SEAMS)), - [], - harness_maturity_rows(list(SEAMS)), - "Invalid qualification report: " + str(exc), - ) - - evidence_rows, evidence_details, maturity_rows, evidence_status = ( - read_qualification_evidence() - ) - gr.Markdown( - "Recorded results apply to the listed model, harness version, and captures. " - "They do not certify the endpoint currently selected above. " - "Capture/reader passes exclude optimizer validation; optimizer details state " - "whether the test used diagnostic replay and whether it covered weight sync." - ) - evidence_status_md = gr.Markdown(evidence_status) - gr.Markdown( - "Stable means all four recorded profiles passed, including optimizer replay. " - "It is limited to this test coverage, not a production-scale guarantee. " - "Experimental adapters have partial or pending support; unstable adapters " - "have no passing profile in the recorded matrix." - ) - maturity_table = gr.Dataframe( - headers=["Harness", "Maturity", "Qualification scope"], - value=maturity_rows, - interactive=False, - ) - evidence_table = gr.Dataframe( - headers=[ - "Harness", - "OpenAI eval", - "Anthropic eval", - "HF eval", - "vLLM training", - ], - value=evidence_rows, - interactive=False, - ) - evidence_detail_table = gr.Dataframe( - headers=[ - "Harness", - "Provider", - "Status", - "Model", - "Harness version", - "Tasks", - "Workflow profile", - "Optimizer scope", - "Optimizer model revision", - "Capture evidence", - "Reason", - ], - value=evidence_details, - interactive=False, - ) - refresh_evidence = gr.Button("Refresh recorded evidence") - refresh_evidence.click( - read_qualification_evidence, - [], - [ - evidence_table, - evidence_detail_table, - maturity_table, - evidence_status_md, - ], - ) - with gr.Accordion("Result JSON", open=False): - raw_json = gr.Code(language="json", lines=22) - - validate_btn.click( - on_validate, - [url_in, model_in, key_in, provider_in, purpose_in, experimental_in], - [engine_md, harness_in, sandbox_in, state, run_btn], + with gr.Tab("Runs", id="runs"): + with gr.Column(elem_classes="hb-page hb-runlist"): + runs_list = gr.HTML( + runs_page(None), js_on_load=_js("run_list.js"), **html_opts + ) + with gr.Column(elem_classes="hb-page hb-runpage"): + run_view = gr.HTML( + "", + js_on_load=_js("run_view.js"), + server_functions=[hb_download], + **html_opts, + ) + + with gr.Tab("Setup", id="setup"): + setup_view = gr.HTML("", **html_opts) + with gr.Row(elem_classes="hb-actions"): + refresh_btn = gr.Button("Check sandboxes again", size="sm") + _qualification_panel() + + timer = gr.Timer(2.0) + + # ── wiring ─────────────────────────────────────────────────────────────────────────────── + task_out = [selection, task_head, task_view, card, engine_state] + run_out = [run_sel, run_view] + app.load( + on_load, + [visitor], + [header, setup_view, runs_list, runs_sig, visitor], ) - for setting in ( - url_in, - model_in, - key_in, - provider_in, - purpose_in, - experimental_in, - ): - setting.change( - lambda: ({}, gr.update(interactive=False), _UNVALIDATED), - outputs=[state, run_btn, engine_md], - ) - ds_in.change(on_dataset, [ds_in], [idx_in, count_md]) - ds_in.change( - on_task, - [ds_in, idx_in], - [task_md, instruction_box, dockerfile_box, toml_box, tests_box], + header.select( + on_tab_again, [run_sel, visitor], [task_head, run_sel, run_view, runs_list] ) - idx_in.change( - on_task, - [ds_in, idx_in], - [task_md, instruction_box, dockerfile_box, toml_box, tests_box], + browser.select(on_open_task, [engine_state, visitor], task_out) + task_head.select(on_back_to_tasks, None, [task_head]) + task_view.select(on_task_page, [run_sel, visitor], [tabs, *run_out, runs_list]) + # Not `.change`: Gradio fires that on every value update from Python too, which would loop. + card.select(on_experimental, [engine_state, selection], [card, engine_state]) + # Gradio runs one call of each event at a time by default. A connect waits on someone's + # endpoint and a tick fires every two seconds per open page, so neither may queue the rest. + card.input( + on_connect, + [engine_state, selection], + [card, engine_state], + concurrency_limit=8, ) - run_btn.click( + card.clear(on_server_default, [engine_state, selection], [card, engine_state]) + card.submit( on_run, - [state, ds_in, idx_in, harness_in, sandbox_in], - [result_html, convo_html, analysis_html, raw_json, contract_file, run_btn], + [engine_state, selection, visitor], + [card, tabs, run_sel, runs_list, run_view, runs_sig], + concurrency_limit=8, ) + runs_list.select(on_open_run, [run_sel, visitor], run_out) + runs_list.input(on_pick_runs, [run_sel], [run_sel]) + runs_list.submit(on_compare, [run_sel, visitor], run_out) + run_view.select( + on_run_page, + [engine_state, run_sel, visitor], + [*run_out, tabs, *task_out, runs_list], + ) + timer.tick( + on_tick, + [run_sel, runs_sig, visitor], + [runs_list, run_view, runs_sig], + show_progress="hidden", + concurrency_limit=None, + ) + refresh_btn.click(on_refresh_setup, None, [setup_view, header]) + return app + - if datasets: - app.load(on_dataset, [ds_in], [idx_in, count_md]) - app.load( - on_task, - [ds_in, idx_in], - [task_md, instruction_box, dockerfile_box, toml_box, tests_box], +def _qualification_panel() -> None: + """The recorded harness/provider evidence: what was qualified, where, and with what.""" + from .qualification import ( + harness_maturity_rows, + qualification_details, + qualification_rows, + ) + from .seams import SEAMS + + report_path = os.environ.get("OPENENV_HARBOR_QUALIFICATION_REPORT", "") + + def read_evidence(): + try: + evidence = ( + json.loads(Path(report_path).read_text()) if report_path else None ) - return app + return ( + qualification_rows(list(SEAMS), evidence), + qualification_details(evidence), + harness_maturity_rows(list(SEAMS), evidence), + "Loaded recorded evidence." + if evidence + else "No qualification report configured (OPENENV_HARBOR_QUALIFICATION_REPORT).", + ) + except (OSError, ValueError, TypeError) as exc: + return ( + qualification_rows(list(SEAMS)), + [], + harness_maturity_rows(list(SEAMS)), + f"Invalid qualification report: {exc}", + ) + + rows, details, maturity, status = read_evidence() + with gr.Accordion("Harness and provider qualification evidence", open=False): + gr.Markdown( + "Recorded results apply to the listed model, harness version and captures. They do not " + "certify the endpoint selected on the Tasks tab. Stable means all four recorded profiles " + "passed, including optimizer replay; experimental adapters have partial or pending " + "support; unstable ones have no passing profile." + ) + status_md = gr.Markdown(status) + maturity_table = gr.Dataframe( + headers=["Harness", "Maturity", "Qualification scope"], + value=maturity, + interactive=False, + ) + evidence_table = gr.Dataframe( + headers=[ + "Harness", + "OpenAI eval", + "Anthropic eval", + "HF eval", + "vLLM training", + ], + value=rows, + interactive=False, + ) + detail_table = gr.Dataframe( + headers=[ + "Harness", + "Provider", + "Status", + "Model", + "Harness version", + "Tasks", + "Workflow profile", + "Optimizer scope", + "Optimizer model revision", + "Capture evidence", + "Reason", + ], + value=details, + interactive=False, + ) + refresh = gr.Button("Refresh recorded evidence", size="sm") + refresh.click( + read_evidence, [], [evidence_table, detail_table, maturity_table, status_md] + ) diff --git a/src/openenv/harbor/ui_assets/harbor.css b/src/openenv/harbor/ui_assets/harbor.css new file mode 100644 index 000000000..7404eaf09 --- /dev/null +++ b/src/openenv/harbor/ui_assets/harbor.css @@ -0,0 +1,659 @@ +@import url("https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700&family=JetBrains+Mono:wght@400;500&display=swap"); + +/* Harbor UI. One stylesheet: tokens, base, shared parts, then one block per view. + Prefixes: hb- shared, tb- task browser, tp- task page, rc- run card, rs- run list, rp- run page, + ev-/tl- trajectory, cmp- comparison, su- setup. + + Corners. OpenEnv's theme squares every element with a zero `border-radius` marked `!important`. Rather than + restating every radius with an `!important` of its own, a component sets `--hb-radius`, and the one + rule below applies it. The property is registered as not inherited, so a rounded panel does not + round everything inside it. + + Specificity. Gradio resets every button and input (`.gradio-container-X [type="button"]` sets padding, + margin, font and background), which outranks a single class. So every rule below the tokens is + written under `.gradio-container .hb`, the wrapper of each component on this page. */ + +@property --hb-radius { syntax: "*"; inherits: false; } +body .gradio-container .contain .hb, body .gradio-container .contain .hb * { border-radius: var(--hb-radius, 0) !important; } + +/* ── tokens ───────────────────────────────────────────────────────────── */ +.gradio-container { + --bg: #fafafa; --surface: #ffffff; --surface-2: #f4f4f5; --surface-3: #e9e9ec; + --border: #e6e6ea; --border-strong: #d2d2d8; + --text: #18181b; --muted: #55555e; --faint: #8b8b94; + --primary: #18181b; --on-primary: #ffffff; --focus: #2563eb; + --ok: #15803d; --warn: #b45309; --bad: #c2410c; --err: #dc2626; --live: #2563eb; --think: #7c4ee4; + --term-bg: #0f1115; --term-text: #d4d4d8; --term-faint: #71717a; --term-prompt: #4ade80; --term-border: #1f2127; + --shadow-sm: 0 1px 2px rgba(16, 16, 24, .05); + --shadow: 0 1px 2px rgba(16, 16, 24, .04), 0 6px 20px -6px rgba(16, 16, 24, .10); + --shadow-lg: 0 24px 64px -12px rgba(16, 16, 24, .28), 0 2px 6px rgba(16, 16, 24, .06); + --r-sm: 6px; --r: 10px; --r-lg: 14px; + --font: "Inter", ui-sans-serif, system-ui, -apple-system, "Segoe UI", sans-serif; + --mono: "JetBrains Mono", ui-monospace, SFMono-Regular, Menlo, monospace; +} +.dark .gradio-container { + --bg: #0b0b0d; --surface: #131316; --surface-2: #1a1a1e; --surface-3: #232328; + --border: #25252b; --border-strong: #34343c; + --text: #ececef; --muted: #a3a3ad; --faint: #72727c; + --primary: #ececef; --on-primary: #111113; --focus: #60a5fa; + --ok: #4ade80; --warn: #fbbf24; --bad: #fb923c; --err: #f87171; --live: #60a5fa; --think: #a58bfa; + --term-bg: #0c0c0e; --term-border: #25252b; + --shadow-sm: 0 1px 2px rgba(0, 0, 0, .4); + --shadow: 0 1px 2px rgba(0, 0, 0, .35), 0 8px 24px -8px rgba(0, 0, 0, .6); + --shadow-lg: 0 24px 64px -12px rgba(0, 0, 0, .75); +} + +/* ── the page around the components ─────────────────────────────────── */ +/* One page colour edge to edge: Gradio's paints the OpenEnv theme's own around the page. */ +html:has(.hb) { scrollbar-gutter: stable; } /* a tab that does not scroll keeps the width of one that does */ +body:has(.hb), body:has(.hb) gradio-app { background: #fafafa !important; } +body.dark:has(.hb), body.dark:has(.hb) gradio-app, .dark body:has(.hb) gradio-app { background: #0b0b0d !important; } +/* Always the full width (to 1600px): Gradio sizes its container to the content, so a tab with a + narrower table would otherwise shrink the whole page. */ +.gradio-container { background: var(--bg) !important; width: 100% !important; max-width: 1600px !important; margin: 0 auto !important; } +.gradio-container .main, .gradio-container .wrap, .gradio-container .contain { background: transparent !important; } +.gradio-container .hb-row { gap: 24px !important; align-items: flex-start !important; flex-wrap: nowrap !important; } +.gradio-container .hb-side { flex: 0 0 360px !important; min-width: 0 !important; max-width: 360px; position: sticky; top: 12px; } +.gradio-container .hb-main { min-width: 0 !important; flex: 1 1 0 !important; } +.gradio-container .hb-page { gap: 0 !important; } +/* A tab shows its list, or one item from it: switched on what is rendered, see `ui.py`. */ +.gradio-container:has(.tp-head) .hb-browse, .gradio-container:not(:has(.tp-head)) .hb-taskpage, .gradio-container:has(.rp) .hb-runlist, .gradio-container:not(:has(.rp)) .hb-runpage { display: none !important; } + +/* Gradio's tab bar, as a plain nav */ +.gradio-container .hb-tabs { background: transparent !important; border: 0 !important; padding: 0 !important; } +.gradio-container .hb-tabs .tab-wrapper { border-bottom: 1px solid var(--border) !important; margin-bottom: 20px; padding: 0 !important; } +.gradio-container .hb-tabs .tab-container { gap: 2px !important; } +.gradio-container .hb-tabs [role="tab"] { font: 500 14px var(--font) !important; color: var(--muted) !important; background: transparent !important; + border: 0 !important; padding: 10px 12px !important; position: relative; } +.gradio-container .hb-tabs [role="tab"]:hover { color: var(--text) !important; } +.gradio-container .hb-tabs [role="tab"][aria-selected="true"] { color: var(--text) !important; } +.gradio-container .hb-tabs [role="tab"][aria-selected="true"]::after { content: ""; position: absolute; left: 12px; right: 12px; bottom: -1px; height: 2px; + background: var(--text) !important; } +.gradio-container .hb-tabs .tabitem { background: transparent !important; border: 0 !important; padding: 0 !important; } +.gradio-container .hb-tabs [role="tab"]:focus { outline: none !important; box-shadow: none !important; } +.gradio-container .hb-tabs [role="tab"]:focus-visible { outline: 2px solid var(--focus) !important; outline-offset: -4px; } +.gradio-container .hb-docs { display: inline-flex; align-items: center; gap: 5px; padding: 10px 12px; font: 500 14px var(--font); color: var(--muted) !important; + text-decoration: none !important; white-space: nowrap; } +.gradio-container .hb-docs:hover { color: var(--text) !important; } +.gradio-container .hb-docs svg { opacity: .7; } + +/* ── base ───────────────────────────────────────────────────────────── */ +.gradio-container .hb { font: 14px/1.55 var(--font); color: var(--text); font-feature-settings: "cv11", "ss01", "ss03"; -webkit-font-smoothing: antialiased; } +.gradio-container .hb *, .gradio-container .hb *::before, .gradio-container .hb *::after { box-sizing: border-box; } +.gradio-container .hb h1, .gradio-container .hb h2, .gradio-container .hb h3, .gradio-container .hb h4, .gradio-container .hb h5 { margin: 0; font-weight: 600; letter-spacing: -.011em; color: var(--text); } +.gradio-container .hb p { margin: 0; } +.gradio-container .hb a { color: inherit; text-decoration: none; } +.gradio-container .hb code, .gradio-container .hb kbd { font-family: var(--mono); font-size: .88em; } +.gradio-container .hb code, .gradio-container .hb pre, .gradio-container .hb kbd { font-variant-ligatures: none; font-feature-settings: "liga" 0, "calt" 0; } +.gradio-container .hb button, .gradio-container .hb input, .gradio-container .hb select, .gradio-container .hb textarea { font: inherit; color: inherit; } +.gradio-container .hb [hidden] { display: none !important; } +.gradio-container .hb :focus-visible { outline: 2px solid var(--focus); outline-offset: 2px; } +.gradio-container .hb .ic { flex: none; display: inline-block; vertical-align: -3px; } +.gradio-container .hb .num { font-variant-numeric: tabular-nums; } +.gradio-container .hb ::-webkit-scrollbar { width: 8px; height: 8px; } +.gradio-container .hb ::-webkit-scrollbar-thumb { background: var(--border-strong); --hb-radius: 99px; } +.gradio-container .hb ::-webkit-scrollbar-track { background: transparent; } + +/* ── header ─────────────────────────────────────────────────────────── */ +.gradio-container .hb .hb-top { display: flex; align-items: center; gap: 12px 24px; flex-wrap: wrap; padding: 16px 0 6px; } +.gradio-container .hb .hb-brand { display: flex; align-items: center; gap: 10px; min-width: 0; } +.gradio-container .hb .hb-brand b { font-size: 15px; font-weight: 600; letter-spacing: -.01em; } +.gradio-container .hb .hb-brand span { color: var(--faint); font-size: 13px; } +.gradio-container .hb .hb-top .hb-facts { margin-left: auto; } + +/* ── buttons, inputs ────────────────────────────────────────────────── */ +.gradio-container .hb .hb-btn { display: inline-flex; align-items: center; justify-content: center; gap: 7px; height: 34px; padding: 0 13px; --hb-radius: 8px; white-space: nowrap; + border: 1px solid var(--border-strong); background: var(--surface); color: var(--text); font-weight: 500; font-size: 13.5px; cursor: pointer; + box-shadow: var(--shadow-sm); transition: background .12s, border-color .12s, opacity .12s; } +.gradio-container .hb .hb-btn:hover { background: var(--surface-2); } +.gradio-container .hb .hb-btn:disabled { opacity: .5; cursor: not-allowed; } +.gradio-container .hb .hb-btn.primary { background: var(--primary); color: var(--on-primary); border-color: var(--primary); } +.gradio-container .hb .hb-btn.primary:hover { opacity: .88; background: var(--primary); } +.gradio-container .hb .hb-btn.ghost { background: transparent; border-color: transparent; box-shadow: none; color: var(--muted); } +.gradio-container .hb .hb-btn.ghost:hover { background: var(--surface-2); color: var(--text); } +.gradio-container .hb .hb-btn.sm { height: 28px; padding: 0 10px; font-size: 12.5px; --hb-radius: 7px; gap: 6px; } +.gradio-container .hb .hb-btn.lg { height: 40px; font-size: 14px; --hb-radius: 9px; } +.gradio-container .hb .hb-btn.block { width: 100%; } +.gradio-container .hb .hb-link { background: none; border: 0; padding: 0; color: var(--muted); cursor: pointer; font-size: 12.5px; display: inline-flex; align-items: center; gap: 5px; } +.gradio-container .hb .hb-link:hover { color: var(--text); } +.gradio-container .hb .hb-input, .gradio-container .hb select.hb-input, .gradio-container .hb input.hb-input[type] { width: 100%; height: 36px; padding: 0 11px; --hb-radius: 8px; border: 1px solid var(--border-strong); + background: var(--surface); color: var(--text); font-size: 13.5px; outline: none; box-shadow: none; min-width: 0; margin: 0; } +.gradio-container .hb .hb-input:focus, .gradio-container .hb input.hb-input[type]:focus { border-color: var(--focus); box-shadow: 0 0 0 3px color-mix(in srgb, var(--focus) 16%, transparent); } +.gradio-container .hb .hb-input::placeholder { color: var(--faint); } +.gradio-container .hb select.hb-input { appearance: none; padding-right: 30px; cursor: pointer; + background-image: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='12' height='12' viewBox='0 0 24 24' fill='none' stroke='%238b8b94' stroke-width='2'%3E%3Cpath d='m6 9 6 6 6-6'/%3E%3C/svg%3E"); + background-repeat: no-repeat; background-position: right 10px center; } +.gradio-container .hb input[type="checkbox"] { appearance: auto !important; -webkit-appearance: checkbox !important; width: 14px; height: 14px; margin: 0; accent-color: var(--primary); + cursor: pointer; flex: none; background: none; border: 0; padding: 0; } +.gradio-container .hb .hb-field { display: grid; gap: 5px; min-width: 0; } +.gradio-container .hb .hb-field > span { font-size: 12.5px; font-weight: 600; display: flex; align-items: baseline; gap: 6px; } +.gradio-container .hb .hb-field > span em { font-style: normal; font-weight: 400; color: var(--faint); font-size: 12px; } +.gradio-container .hb .hb-field > span .hb-link { margin-left: auto; font-weight: 500; } +.gradio-container .hb .hb-pair { display: grid; grid-template-columns: 1fr 1fr; gap: 8px; } +.gradio-container .hb .hb-seg { display: inline-flex; gap: 2px; padding: 3px; --hb-radius: 9px; background: var(--surface-2); border: 1px solid var(--border); max-width: 100%; overflow-x: auto; } +.gradio-container .hb .hb-seg button { display: inline-flex; align-items: center; justify-content: center; gap: 6px; border: 0; background: none; --hb-radius: 6px; padding: 5px 12px; + font-size: 13px; font-weight: 500; color: var(--muted); cursor: pointer; white-space: nowrap; } +.gradio-container .hb .hb-seg button span { color: var(--faint); font-size: 12px; font-variant-numeric: tabular-nums; } +.gradio-container .hb .hb-seg button:hover { color: var(--text); } +.gradio-container .hb .hb-seg button[aria-pressed="true"] { background: var(--surface); color: var(--text); box-shadow: var(--shadow-sm); } +.gradio-container .hb .hb-seg.fill { display: grid; grid-auto-flow: column; grid-auto-columns: minmax(0, 1fr); width: 100%; } +.gradio-container .hb .hb-seg.fill button { padding: 5px 6px; } + +/* ── small parts ────────────────────────────────────────────────────── */ +.gradio-container .hb .hb-chip { display: inline-flex; align-items: center; gap: 5px; height: 22px; padding: 0 8px; --hb-radius: 6px; font-size: 12px; color: var(--muted); + background: var(--surface-2); white-space: nowrap; max-width: 100%; overflow: hidden; text-overflow: ellipsis; } +.gradio-container .hb .hb-chip.outline { background: transparent; border: 1px solid var(--border); } +.gradio-container .hb .hb-chips { display: flex; flex-wrap: wrap; gap: 5px; } +.gradio-container .hb .hb-status { display: inline-flex; align-items: center; gap: 6px; height: 22px; padding: 0 8px; --hb-radius: 99px; font-size: 12px; font-weight: 600; white-space: nowrap; + color: var(--muted); background: var(--surface-2); } +.gradio-container .hb .hb-status.live { color: var(--live); background: color-mix(in srgb, var(--live) 11%, transparent); } +.gradio-container .hb .hb-status.ok { color: var(--ok); background: color-mix(in srgb, var(--ok) 11%, transparent); } +.gradio-container .hb .hb-status.bad { color: var(--err); background: color-mix(in srgb, var(--err) 10%, transparent); } +.gradio-container .hb .hb-status.warn { color: var(--warn); background: color-mix(in srgb, var(--warn) 12%, transparent); } +.gradio-container .hb .hb-reward { display: inline-flex; align-items: center; justify-content: center; min-width: 38px; height: 22px; padding: 0 7px; --hb-radius: 6px; + font-weight: 700; font-size: 12.5px; font-variant-numeric: tabular-nums; } +.gradio-container .hb .hb-reward.full { color: var(--ok); background: color-mix(in srgb, var(--ok) 12%, transparent); } +.gradio-container .hb .hb-reward.part { color: var(--warn); background: color-mix(in srgb, var(--warn) 13%, transparent); } +.gradio-container .hb .hb-reward.zero { color: var(--err); background: color-mix(in srgb, var(--err) 10%, transparent); } +.gradio-container .hb .hb-reward.none { color: var(--faint); font-weight: 500; font-size: 12px; } +.gradio-container .hb .hb-dot { width: 7px; height: 7px; --hb-radius: 50%; background: var(--faint); display: inline-block; flex: none; } +.gradio-container .hb .hb-dot.ok { background: var(--ok); } .gradio-container .hb .hb-dot.warn { background: var(--warn); } .gradio-container .hb .hb-dot.bad { background: var(--err); } .gradio-container .hb .hb-dot.live { background: var(--live); } +.gradio-container .hb .hb-kv { display: grid; grid-template-columns: max-content minmax(0, 1fr); gap: 8px 20px; margin: 0; font-size: 13px; } +.gradio-container .hb .hb-kv dt { color: var(--faint); } .gradio-container .hb .hb-kv dd { margin: 0; overflow-wrap: anywhere; min-width: 0; } +.gradio-container .hb .hb-kv code { font-size: 12.5px; } +.gradio-container .hb .hb-facts { display: flex; flex-wrap: wrap; gap: 6px 18px; align-items: center; font-size: 13px; color: var(--muted); } +.gradio-container .hb .hb-facts > span, .gradio-container .hb .hb-facts > button { display: inline-flex; align-items: center; gap: 6px; min-width: 0; } +.gradio-container .hb .hb-facts .ic { color: var(--faint); } +.gradio-container .hb .hb-facts b { color: var(--text); font-weight: 600; } +.gradio-container .hb .hb-crumbs { display: flex; align-items: center; gap: 6px; font-size: 13px; color: var(--faint); min-width: 0; } +.gradio-container .hb .hb-crumbs button { color: var(--muted); background: none; border: 0; padding: 0; cursor: pointer; font-size: 13px; } +.gradio-container .hb .hb-crumbs button:hover { color: var(--text); text-decoration: underline; text-underline-offset: 3px; } +.gradio-container .hb .hb-crumbs button.back { display: inline-flex; align-items: center; gap: 5px; color: var(--text); font-weight: 500; } +.gradio-container .hb .hb-crumbs .ic { color: var(--border-strong); } +.gradio-container .hb .hb-crumbs span:last-child { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; font-family: var(--mono); font-size: 12px; } +.gradio-container .hb .hb-empty { display: grid; justify-items: center; text-align: center; gap: 6px; padding: 48px 16px; color: var(--muted); font-size: 13.5px; } +.gradio-container .hb .hb-empty .ic { color: var(--faint); margin-bottom: 4px; } +.gradio-container .hb .hb-empty h3 { font-size: 15px; } +.gradio-container .hb .hb-empty p { max-width: 440px; } +.gradio-container .hb .hb-note { display: flex; gap: 9px; align-items: flex-start; padding: 9px 12px; --hb-radius: var(--r); font-size: 13px; line-height: 1.5; + background: var(--surface-2); color: var(--muted); min-width: 0; overflow-wrap: anywhere; } +.gradio-container .hb .hb-note .ic { margin-top: 1px; color: var(--faint); } +.gradio-container .hb .hb-note b { color: var(--text); font-weight: 600; } +.gradio-container .hb .hb-note code { font-size: 12px; } +.gradio-container .hb .hb-note.warn { background: color-mix(in srgb, var(--warn) 9%, var(--surface)); color: var(--text); } +.gradio-container .hb .hb-note.warn .ic { color: var(--warn); } +.gradio-container .hb .hb-note.err, .gradio-container .hb .hb-note.bad { background: color-mix(in srgb, var(--err) 8%, var(--surface)); color: var(--text); } +.gradio-container .hb .hb-note.err .ic, .gradio-container .hb .hb-note.bad .ic { color: var(--err); } +.gradio-container .hb .hb-note.ok { background: color-mix(in srgb, var(--ok) 8%, var(--surface)); color: var(--text); } +.gradio-container .hb .hb-note.ok .ic { color: var(--ok); } +.gradio-container .hb .hb-fine { font-size: 12px; color: var(--faint); line-height: 1.5; } +.gradio-container .hb .hb-fine code { font-size: 11.5px; } +@keyframes hb-spin { to { transform: rotate(360deg); } } +@keyframes hb-pulse { 0%, 100% { opacity: 1; } 50% { opacity: .35; } } +@keyframes hb-shimmer { from { background-position: 200% 0; } to { background-position: -200% 0; } } +@keyframes hb-rise { from { opacity: 0; transform: translateY(4px); } } +.gradio-container .hb .hb-spinner { width: 14px; height: 14px; --hb-radius: 50%; border: 2px solid var(--border-strong); border-top-color: var(--text); animation: hb-spin .7s linear infinite; + display: inline-block; flex: none; } +.gradio-container .hb .hb-btn.primary .hb-spinner { border-color: color-mix(in srgb, var(--on-primary) 30%, transparent); border-top-color: var(--on-primary); } +.gradio-container .hb .hb-pulse { width: 7px; height: 7px; --hb-radius: 50%; background: var(--live); display: inline-block; animation: hb-pulse 1.3s ease-in-out infinite; flex: none; } +.gradio-container .hb .hb-sk { display: block; --hb-radius: 6px; background: linear-gradient(90deg, var(--surface-2) 25%, var(--surface-3) 37%, var(--surface-2) 63%); + background-size: 400% 100%; animation: hb-shimmer 1.4s ease infinite; } +.gradio-container .hb details > summary { cursor: pointer; list-style: none; } +.gradio-container .hb details > summary::-webkit-details-marker { display: none; } +.gradio-container .hb details > summary .ic.chev { color: var(--faint); transition: transform .15s; } +.gradio-container .hb details[open] > summary .ic.chev { transform: rotate(90deg); } +.gradio-container .hb .hb-disclose { display: flex; align-items: center; gap: 6px; font-size: 13px; font-weight: 500; color: var(--muted); padding: 6px 0; } +.gradio-container .hb .hb-disclose:hover { color: var(--text); } + +/* panels */ +.gradio-container .hb .hb-panel { background: var(--surface); border: 1px solid var(--border); --hb-radius: var(--r-lg); min-width: 0; } +.gradio-container .hb .hb-panel-h { display: flex; align-items: center; gap: 10px; padding: 13px 18px; border-bottom: 1px solid var(--border); min-width: 0; } +.gradio-container .hb .hb-panel-h h2, .gradio-container .hb .hb-panel-h h3 { font-size: 14px; display: flex; align-items: center; gap: 8px; min-width: 0; } +.gradio-container .hb .hb-panel-h h2 .ic, .gradio-container .hb .hb-panel-h h3 .ic { color: var(--faint); } +.gradio-container .hb .hb-panel-h .aside { margin-left: auto; font-size: 12.5px; color: var(--faint); text-align: right; white-space: nowrap; overflow: hidden; text-overflow: ellipsis; min-width: 0; } +.gradio-container .hb .hb-panel-b { padding: 16px 18px; min-width: 0; } +.gradio-container .hb .hb-panel-b > :first-child { margin-top: 0; } +.gradio-container .hb .hb-label { font-size: 12.5px; font-weight: 600; margin: 16px 0 8px; color: var(--text); display: flex; align-items: center; gap: 8px; } +.gradio-container .hb .hb-label:first-child { margin-top: 0; } +.gradio-container .hb .hb-label span { font-weight: 400; color: var(--faint); margin-left: auto; } + +/* tables */ +.gradio-container .hb .hb-tbl { overflow: auto; border: 1px solid var(--border); --hb-radius: var(--r); background: var(--surface); } +.gradio-container .hb .hb-tbl table { border-collapse: collapse; font-size: 13px; width: 100%; margin: 0; } +.gradio-container .hb .hb-tbl th, .gradio-container .hb .hb-tbl td { padding: 8px 12px; border: 0; border-bottom: 1px solid var(--border); text-align: left; vertical-align: top; overflow-wrap: anywhere; background: transparent; } +.gradio-container .hb .hb-tbl tr:last-child td { border-bottom: 0; } +.gradio-container .hb .hb-tbl th { background: var(--surface-2); font-weight: 600; color: var(--muted); font-size: 12.5px; white-space: nowrap; } +.gradio-container .hb .hb-tbl td.num, .gradio-container .hb .hb-tbl th.num { text-align: right; font-variant-numeric: tabular-nums; } +.gradio-container .hb .hb-tbl code { font-size: 12px; } + +/* code and prose */ +.gradio-container .hb pre { margin: 0; } +.gradio-container .hb .hb-code { background: var(--surface-2); border: 1px solid var(--border); --hb-radius: 8px; padding: 10px 12px; overflow: auto; + font: 12.5px/1.55 var(--mono); max-height: 480px; tab-size: 2; white-space: pre-wrap; overflow-wrap: anywhere; color: var(--text); margin: 0; } +.gradio-container .hb .hb-prose { font-size: 14px; line-height: 1.7; overflow-wrap: anywhere; color: var(--text); } +.gradio-container .hb .hb-prose > :first-child { margin-top: 0; } .gradio-container .hb .hb-prose > :last-child { margin-bottom: 0; } +.gradio-container .hb .hb-prose p { margin: .7em 0; } +.gradio-container .hb .hb-prose h1, .gradio-container .hb .hb-prose h2, .gradio-container .hb .hb-prose h3, .gradio-container .hb .hb-prose h4 { font-size: 14.5px; margin: 1.2em 0 .4em; } +.gradio-container .hb .hb-prose ul, .gradio-container .hb .hb-prose ol { padding-left: 1.3em; margin: .6em 0; } +.gradio-container .hb .hb-prose li { margin: .2em 0; } +.gradio-container .hb .hb-prose code { background: var(--surface-2); padding: 1px 5px; --hb-radius: 4px; } +.gradio-container .hb .hb-prose pre { background: var(--surface-2); border: 1px solid var(--border); --hb-radius: 8px; padding: 10px 12px; overflow: auto; font: 12.5px/1.55 var(--mono); margin: .8em 0; } +.gradio-container .hb .hb-prose pre code { background: none; padding: 0; } +.gradio-container .hb .hb-prose table { border-collapse: collapse; font-size: 13px; margin: .8em 0; display: block; overflow: auto; } +.gradio-container .hb .hb-prose th, .gradio-container .hb .hb-prose td { border: 1px solid var(--border); padding: 5px 10px; text-align: left; } +.gradio-container .hb .hb-prose th { background: var(--surface-2); } +.gradio-container .hb .hb-prose a { text-decoration: underline; text-decoration-color: var(--border-strong); text-underline-offset: 2px; } +.gradio-container .hb .hb-prose blockquote { margin: .7em 0; padding-left: 12px; border-left: 2px solid var(--border-strong); color: var(--muted); } + +/* ═════════════════════════ TASK BROWSER ═══════════════════════════════ */ +.gradio-container .hb .tb-head { display: flex; align-items: flex-end; justify-content: space-between; gap: 16px 24px; flex-wrap: wrap; padding: 4px 0 16px; } +.gradio-container .hb .tb-head h1 { font-size: 22px; letter-spacing: -.02em; } +.gradio-container .hb .tb-head p { color: var(--muted); margin-top: 4px; font-size: 13.5px; max-width: 720px; } +.gradio-container .hb .tb-stats { display: flex; gap: 28px; } +.gradio-container .hb .tb-stats div { display: grid; } +.gradio-container .hb .tb-stats b { font-size: 20px; font-weight: 650; letter-spacing: -.02em; font-variant-numeric: tabular-nums; } +.gradio-container .hb .tb-stats span { font-size: 12px; color: var(--faint); } +.gradio-container .hb .tb-toolbar { display: flex; gap: 8px; align-items: center; margin-bottom: 16px; } +.gradio-container .hb .tb-search { position: relative; flex: 1; min-width: 0; } +.gradio-container .hb .tb-search > .ic { position: absolute; left: 12px; top: 50%; transform: translateY(-50%); color: var(--faint); pointer-events: none; z-index: 1; } +.gradio-container .hb .tb-search input.hb-input[type] { height: 40px; padding-left: 38px; font-size: 14px; --hb-radius: 9px; box-shadow: var(--shadow-sm); } +.gradio-container .hb .tb-toolbar .hb-btn { height: 40px; } +.gradio-container .hb .tb-layout { display: grid; grid-template-columns: 236px minmax(0, 1fr); gap: 28px; } +.gradio-container .hb .tb-filters { position: sticky; top: 12px; align-self: start; max-height: calc(100vh - 24px); overflow: hidden auto; padding: 0 4px 12px 0; } +.gradio-container .hb .tb-filters-h { display: flex; justify-content: space-between; align-items: center; font-weight: 600; font-size: 13px; margin: 2px 0 8px; padding: 0 6px; } +.gradio-container .hb .tb-facet { border-top: 1px solid var(--border); padding: 12px 0 8px; } +.gradio-container .hb .tb-facet > summary { display: flex; align-items: center; gap: 8px; padding: 2px 6px 6px; } +.gradio-container .hb .tb-facet > summary h3 { font-size: 12.5px; font-weight: 600; color: var(--muted); } +.gradio-container .hb .tb-facet > summary .sel { font-size: 12px; color: var(--focus); font-weight: 500; } +.gradio-container .hb .tb-facet > summary .ic.chev { margin-left: auto; } +.gradio-container .hb .tb-opt { display: flex; align-items: center; gap: 9px; width: 100%; padding: 4px 6px; margin: 0; --hb-radius: 6px; background: none; border: 0; + cursor: pointer; text-align: left; font-size: 13px; color: var(--text); } +.gradio-container .hb .tb-opt:hover { background: var(--surface-2); } +.gradio-container .hb .tb-opt .box { width: 15px; height: 15px; --hb-radius: 4px; border: 1.5px solid var(--border-strong); flex: none; display: grid; place-items: center; background: var(--surface); color: transparent; } +.gradio-container .hb .tb-opt[aria-pressed="true"] .box { background: var(--primary); border-color: var(--primary); color: var(--on-primary); } +.gradio-container .hb .tb-opt .lab { flex: 1; min-width: 0; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } +.gradio-container .hb .tb-opt .c { color: var(--faint); font-size: 12px; font-variant-numeric: tabular-nums; } +.gradio-container .hb .tb-opt.zero { opacity: .45; } +.gradio-container .hb .tb-facet .hb-link { margin: 4px 6px 0; } +.gradio-container .hb .tb-results-h { display: flex; align-items: center; gap: 10px; flex-wrap: wrap; min-height: 30px; margin: 0 0 10px; } +.gradio-container .hb .tb-results-h .count { font-size: 13px; font-weight: 600; font-variant-numeric: tabular-nums; } +.gradio-container .hb .tb-active { display: flex; flex-wrap: wrap; gap: 6px; } +.gradio-container .hb .tb-pill { display: inline-flex; align-items: center; gap: 6px; height: 26px; padding: 0 6px 0 10px; --hb-radius: 99px; background: var(--surface); + border: 1px solid var(--border-strong); font-size: 12.5px; cursor: pointer; max-width: 100%; color: var(--text); } +.gradio-container .hb .tb-pill:hover { background: var(--surface-2); } +.gradio-container .hb .tb-pill span { color: var(--faint); } +.gradio-container .hb .tb-pill .ic { color: var(--faint); } +.gradio-container .hb .tb-list { list-style: none; margin: 0; padding: 0; display: grid; grid-template-columns: repeat(auto-fill, minmax(min(100%, 420px), 1fr)); gap: 10px; } +.gradio-container .hb .tb-list > li { min-width: 0; } +.gradio-container .hb .tb-card { display: flex; flex-direction: column; width: 100%; height: 100%; padding: 14px 16px; --hb-radius: var(--r-lg); background: var(--surface); border: 1px solid var(--border); + transition: border-color .12s, box-shadow .12s; min-width: 0; text-align: left; cursor: pointer; color: var(--text); margin: 0; } +.gradio-container .hb .tb-card:hover { border-color: var(--border-strong); box-shadow: var(--shadow); } +.gradio-container .hb .tb-card .r1 { display: flex; align-items: center; gap: 7px; font-size: 12px; color: var(--muted); margin-bottom: 6px; min-width: 0; width: 100%; } +.gradio-container .hb .tb-card .r1 .ds { font-weight: 600; color: var(--text); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; flex: 0 1 auto; } +.gradio-container .hb .tb-card .r1 .sep { color: var(--border-strong); } +.gradio-container .hb .tb-card .r1 .cat { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; flex: 0 1 auto; } +.gradio-container .hb .tb-card .r1 .idx { margin-left: auto; color: var(--faint); font-variant-numeric: tabular-nums; font-family: var(--mono); font-size: 11.5px; flex: none; } +.gradio-container .hb .tb-card .t { font-size: 14.5px; font-weight: 600; line-height: 1.4; letter-spacing: -.006em; overflow-wrap: anywhere; + display: -webkit-box; -webkit-line-clamp: 2; -webkit-box-orient: vertical; overflow: hidden; } +.gradio-container .hb .tb-card .s { font-size: 13px; color: var(--muted); margin-top: 4px; overflow-wrap: anywhere; line-height: 1.5; + display: -webkit-box; -webkit-line-clamp: 2; -webkit-box-orient: vertical; overflow: hidden; } +.gradio-container .hb .tb-card .hb-chips { margin-top: auto; padding-top: 10px; } +.gradio-container .hb mark { background: color-mix(in srgb, #f5c542 42%, transparent); color: inherit; --hb-radius: 3px; padding: 0 1px; } +.gradio-container .hb .tb-more { display: flex; flex-direction: column; align-items: center; gap: 10px; padding: 24px 0 8px; font-size: 12.5px; color: var(--faint); } + +.gradio-container .hb .tb-optrow { display: flex; align-items: center; min-width: 0; } +.gradio-container .hb .tb-optrow .tb-opt { flex: 1; min-width: 0; } +.gradio-container .hb .tb-rm { display: grid; place-items: center; width: 22px; height: 22px; flex: none; border: 0; background: none; color: var(--faint); cursor: pointer; + --hb-radius: 5px; opacity: 0; transition: opacity .12s; } +.gradio-container .hb .tb-optrow:hover .tb-rm, .gradio-container .hb .tb-rm:focus-visible { opacity: 1; } +.gradio-container .hb .tb-rm:hover { color: var(--err); background: var(--surface-2); } +.gradio-container .hb .tb-confirm { margin: 4px 0; padding: 8px 10px; border: 1px solid color-mix(in srgb, var(--err) 30%, var(--border)); --hb-radius: 8px; font-size: 12.5px; + background: color-mix(in srgb, var(--err) 5%, var(--surface)); display: grid; gap: 8px; } +.gradio-container .hb .tb-confirm p { margin: 0; color: var(--muted); line-height: 1.45; overflow-wrap: anywhere; } +.gradio-container .hb .tb-confirm p b { color: var(--text); font-weight: 500; } +.gradio-container .hb .tb-confirm p.err { color: var(--err); } +.gradio-container .hb .tb-confirm div { display: flex; gap: 6px; } +.gradio-container .hb .tb-confirm [data-rm-yes] { color: var(--err); } +.gradio-container .hb .tb-rmrow { display: flex; align-items: center; gap: 9px; padding: 4px 6px; font-size: 13px; color: var(--muted); } +@media (hover: none) { .gradio-container .hb .tb-rm { opacity: 1; } } +.gradio-container .hb .tb-add { margin-bottom: 16px; } +.gradio-container .hb .tb-add .hb-panel-h code { font-size: 11.5px; } +.gradio-container .hb .tb-add .hb-panel-h .hb-btn { margin-left: 4px; width: 28px; padding: 0; } +.gradio-container .hb .tb-add .hb-panel-b { display: grid; gap: 10px; padding-top: 12px; } +.gradio-container .hb .tb-add-list { display: grid; max-height: 420px; overflow: auto; border: 1px solid var(--border); --hb-radius: var(--r); } +.gradio-container .hb .tb-add-row { display: grid; grid-template-columns: minmax(0, 1fr) 150px minmax(220px, auto); gap: 14px; align-items: center; padding: 9px 12px; border-top: 1px solid var(--border); min-width: 0; } +.gradio-container .hb .tb-add-row:first-child { border-top: 0; } +.gradio-container .hb .tb-add-row .nm { display: grid; min-width: 0; } +.gradio-container .hb .tb-add-row .nm b { font: 500 13px var(--mono); overflow-wrap: anywhere; color: var(--text); } +.gradio-container .hb .tb-add-row .nm b span { color: var(--faint); font-weight: 400; } +.gradio-container .hb .tb-add-row .nm em { font-style: normal; font-size: 12px; color: var(--faint); } +.gradio-container .hb .tb-add-row .tk { font-size: 12.5px; color: var(--muted); font-variant-numeric: tabular-nums; } +.gradio-container .hb .tb-add-row .tk .big { color: var(--warn); } +.gradio-container .hb .tb-add-row .ac { display: flex; align-items: center; justify-content: flex-end; gap: 8px; min-width: 0; font-size: 12.5px; } +.gradio-container .hb .tb-add-row .ac .ok { display: inline-flex; align-items: center; gap: 5px; color: var(--ok); font-weight: 500; } +.gradio-container .hb .tb-add-row .ac .err { display: inline-flex; align-items: center; gap: 5px; color: var(--err); min-width: 0; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; max-width: 260px; } +.gradio-container .hb .tb-add-row .prog { display: grid; gap: 5px; width: 240px; } +.gradio-container .hb .tb-add-row .prog .w { display: inline-flex; align-items: center; gap: 7px; color: var(--muted); font-variant-numeric: tabular-nums; white-space: nowrap; } +.gradio-container .hb .tb-add-row .prog .bar { height: 4px; --hb-radius: 99px; background: var(--surface-3); overflow: hidden; } +.gradio-container .hb .tb-add-row .prog .bar i { display: block; height: 100%; --hb-radius: 99px; background: var(--live); transition: width .4s ease; } +.gradio-container .hb .tb-add-empty { display: flex; align-items: center; gap: 8px; padding: 18px 12px; color: var(--muted); font-size: 13px; } +.gradio-container .hb .tb-opt .hb-spinner { width: 12px; height: 12px; border-width: 1.5px; } +@media (max-width: 760px) { + .gradio-container .hb .tb-add-row { grid-template-columns: minmax(0, 1fr) auto; } + .gradio-container .hb .tb-add-row .tk { grid-column: 1; } + .gradio-container .hb .tb-add-row .ac { grid-column: 1 / -1; justify-content: flex-start; } + .gradio-container .hb .tb-add-row .prog { width: 100%; } +} + +/* ═════════════════════════ TASK PAGE ══════════════════════════════════ */ +.gradio-container .hb .tp-head { display: grid; gap: 10px; margin-bottom: 20px; min-width: 0; } +.gradio-container .hb .tp-head .kick { display: flex; flex-wrap: wrap; gap: 6px; align-items: center; } +.gradio-container .hb .tp-head h1 { font-size: 24px; line-height: 1.3; letter-spacing: -.02em; max-width: 1040px; overflow-wrap: anywhere; } +.gradio-container .hb .tp-grid { display: grid; grid-template-columns: 164px minmax(0, 1fr); gap: 24px; align-items: start; } +.gradio-container .hb .tp-toc { position: sticky; top: 12px; display: grid; gap: 1px; font-size: 13px; } +.gradio-container .hb .tp-toc p { font-size: 12px; font-weight: 600; color: var(--faint); margin: 0 0 6px 10px; } +.gradio-container .hb .tp-toc a { display: flex; align-items: center; justify-content: space-between; gap: 8px; padding: 6px 10px; --hb-radius: 6px; color: var(--muted); cursor: pointer; } +.gradio-container .hb .tp-toc a:hover { color: var(--text); background: var(--surface-2); } +.gradio-container .hb .tp-toc a.on { color: var(--text); background: var(--surface-2); font-weight: 500; } +.gradio-container .hb .tp-toc a em { font-style: normal; font-size: 11.5px; color: var(--faint); font-variant-numeric: tabular-nums; } +.gradio-container .hb .tp-body { display: grid; gap: 16px; min-width: 0; } +.gradio-container .hb .tp-sec { scroll-margin-top: 16px; } +.gradio-container .hb .tp-meta { margin-top: 16px; padding-top: 14px; border-top: 1px solid var(--border); } +.gradio-container .hb .tp-groups { display: grid; grid-template-columns: repeat(auto-fit, minmax(min(100%, 260px), 1fr)); gap: 18px 28px; } +.gradio-container .hb .tp-groups h4 { font-size: 12.5px; font-weight: 600; margin: 0 0 8px; color: var(--muted); } +.gradio-container .hb .tp-files { display: grid; grid-template-columns: minmax(180px, 260px) minmax(0, 1fr); border: 1px solid var(--border); --hb-radius: var(--r); overflow: hidden; min-height: 360px; } +.gradio-container .hb .tp-tree { border-right: 1px solid var(--border); overflow: auto; max-height: 560px; padding: 6px; background: var(--surface); } +.gradio-container .hb .tp-tree button { display: flex; align-items: center; gap: 7px; width: 100%; padding: 4px 8px; border: 0; --hb-radius: 6px; background: none; color: var(--text); + font: 12.5px var(--mono); cursor: pointer; text-align: left; min-width: 0; } +.gradio-container .hb .tp-tree button:hover { background: var(--surface-2); } +.gradio-container .hb .tp-tree button[aria-current="true"] { background: var(--surface-2); font-weight: 500; } +.gradio-container .hb .tp-tree button .ic { color: var(--faint); } +.gradio-container .hb .tp-tree button span { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; flex: 1; } +.gradio-container .hb .tp-tree button em { font-style: normal; color: var(--faint); font-size: 11px; font-family: var(--font); } +.gradio-container .hb .tp-tree .dir { padding: 8px 8px 3px; font: 600 11.5px var(--font); color: var(--faint); display: flex; align-items: center; gap: 6px; } +.gradio-container .hb .tp-view { min-width: 0; display: flex; flex-direction: column; background: var(--surface-2); } +.gradio-container .hb .tp-view-h { display: flex; align-items: center; gap: 8px; padding: 4px 6px 4px 12px; min-height: 38px; border-bottom: 1px solid var(--border); font: 12.5px var(--mono); background: var(--surface); min-width: 0; } +.gradio-container .hb .tp-view-h .p { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; } +.gradio-container .hb .tp-view-h em { font: 11.5px var(--font); color: var(--faint); font-style: normal; white-space: nowrap; } +.gradio-container .hb .tp-view-h .tools { margin-left: auto; display: flex; gap: 2px; } +.gradio-container .hb .tp-view-h .tools .hb-btn { width: 28px; padding: 0; } +.gradio-container .hb .tp-view-h .tools .hb-btn[aria-pressed="true"] { background: var(--surface-2); color: var(--text); } +.gradio-container .hb .tp-code { flex: 1; margin: 0; padding: 10px 0; overflow: auto; max-height: 560px; font: 12.5px/1.6 var(--mono); white-space: pre; tab-size: 4; color: var(--text); counter-reset: ln; } +.gradio-container .hb .tp-code .l { display: block; padding: 0 16px 0 58px; position: relative; min-height: 1.6em; } +.gradio-container .hb .tp-code .l::before { counter-increment: ln; content: counter(ln); position: absolute; left: 0; width: 42px; text-align: right; color: var(--faint); user-select: none; } +.gradio-container .hb .tp-code .l:hover { background: color-mix(in srgb, var(--text) 4%, transparent); } +.gradio-container .hb .tp-code.wrap .l { white-space: pre-wrap; overflow-wrap: anywhere; } +.gradio-container .hb .tp-code .msg { padding: 0 16px; color: var(--faint); font-family: var(--font); } +/* full view: the file browser over the whole window, like an editor */ +.gradio-container .hb .tp-files.full { position: fixed; inset: 12px; z-index: 1000; min-height: 0; --hb-radius: var(--r-lg); box-shadow: var(--shadow-lg); background: var(--surface); + grid-template-columns: minmax(220px, 300px) minmax(0, 1fr); } +.gradio-container .hb .tp-files.full .tp-tree, .gradio-container .hb .tp-files.full .tp-code { max-height: none; } +.gradio-container .hb .tp-files.full .tp-tree { height: 100%; } +.gradio-container .hb .tp-files.full .tp-view { height: 100%; min-height: 0; } +body:has(.tp-files.full) { overflow: hidden !important; } +.gradio-container .hb .hb-scrim { position: fixed; inset: 0; background: rgba(9, 9, 12, .5); z-index: 999; } +.gradio-container .hb .tp-view .hb-empty { padding: 36px 16px; } +.gradio-container .hb .tp-runs { display: grid; } +.gradio-container .hb .tp-run { display: grid; grid-template-columns: auto auto minmax(0, 1fr) auto; gap: 12px; align-items: center; padding: 9px 2px; font-size: 13px; + border: 0; border-bottom: 1px solid var(--border); background: none; cursor: pointer; text-align: left; width: 100%; color: var(--text); } +.gradio-container .hb .tp-run:last-child { border-bottom: 0; } +.gradio-container .hb .tp-run:hover .m { color: var(--text); } +.gradio-container .hb .tp-run .m { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; color: var(--muted); } +.gradio-container .hb .tp-run .w { color: var(--faint); font-size: 12.5px; font-variant-numeric: tabular-nums; } + +/* ═════════════════════════ RUN CARD ═══════════════════════════════════ */ +.gradio-container .hb .rc .hb-panel-b { display: grid; gap: 16px; } +.gradio-container .hb .rc-intro { font-size: 13px; color: var(--muted); line-height: 1.5; } +.gradio-container .hb .rc-sec { display: grid; gap: 8px; min-width: 0; } +.gradio-container .hb .rc-sec > .hb-label { margin: 0; } +.gradio-container .hb .rc-model { display: grid; grid-template-columns: minmax(0, 1fr) auto; column-gap: 10px; align-items: center; padding: 8px 12px; --hb-radius: 9px; + border: 1px solid var(--border); background: var(--surface-2); min-width: 0; } +.gradio-container .hb .rc-model b { font-size: 13.5px; font-weight: 600; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } +.gradio-container .hb .rc-model span { font-size: 11.5px; color: var(--faint); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; grid-column: 1; } +.gradio-container .hb .rc-model em { grid-row: 1 / span 2; grid-column: 2; font-style: normal; font-size: 12px; color: var(--muted); white-space: nowrap; } +.gradio-container .hb .rc-conn { display: flex; align-items: center; gap: 7px; font-size: 12.5px; color: var(--muted); min-width: 0; } +.gradio-container .hb .rc-conn b { color: var(--text); font-weight: 500; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; } +.gradio-container .hb .rc-conn .ic.ok { color: var(--ok); } +.gradio-container .hb .rc-form { display: grid; gap: 12px; } +.gradio-container .hb .rc-boxes { display: grid; grid-template-columns: repeat(auto-fill, minmax(96px, 1fr)); gap: 6px; } +.gradio-container .hb .rc-boxes button { height: 32px; --hb-radius: 8px; border: 1px solid var(--border); background: var(--surface); color: var(--text); font: 12.5px var(--mono); cursor: pointer; } +.gradio-container .hb .rc-boxes button:hover:not(:disabled) { border-color: var(--border-strong); background: var(--surface-2); } +.gradio-container .hb .rc-boxes button[aria-pressed="true"] { border-color: var(--text); box-shadow: inset 0 0 0 1px var(--text); } +.gradio-container .hb .rc-boxes button:disabled { color: var(--faint); text-decoration: line-through; cursor: not-allowed; background: transparent; } +.gradio-container .hb .rc-acct { display: flex; align-items: center; justify-content: space-between; gap: 8px; flex-wrap: wrap; font-size: 12.5px; } +.gradio-container .hb .rc-acct .hb-link { font-size: 12px; } +.gradio-container .hb .rc-check { display: flex; align-items: center; gap: 8px; font-size: 12.5px; color: var(--muted); cursor: pointer; } +.gradio-container .hb .rc-go { display: grid; gap: 8px; } + +/* model picker */ +.gradio-container .hb .pk { position: relative; } +.gradio-container .hb .pk-btn { width: 100%; display: grid; grid-template-columns: minmax(0, 1fr) auto auto; grid-template-rows: auto auto; column-gap: 10px; align-items: center; text-align: left; + padding: 8px 12px; --hb-radius: 9px; border: 1px solid var(--border-strong); background: var(--surface); cursor: pointer; box-shadow: var(--shadow-sm); color: var(--text); } +.gradio-container .hb .pk-btn:hover, .gradio-container .hb .pk-btn[aria-expanded="true"] { border-color: var(--text); } +.gradio-container .hb .pk-name { font-size: 13.5px; font-weight: 600; white-space: nowrap; overflow: hidden; text-overflow: ellipsis; } +.gradio-container .hb .pk-org { font-size: 11.5px; color: var(--faint); white-space: nowrap; overflow: hidden; text-overflow: ellipsis; } +.gradio-container .hb .pk-btn .pk-name { grid-column: 1; grid-row: 1; } +.gradio-container .hb .pk-btn .pk-org { grid-column: 1; grid-row: 2; } +.gradio-container .hb .pk-price { font-size: 12px; color: var(--muted); font-variant-numeric: tabular-nums; white-space: nowrap; } +.gradio-container .hb .pk-btn .pk-price { grid-column: 2; grid-row: 1 / span 2; } +.gradio-container .hb .pk-btn > .ic { grid-column: 3; grid-row: 1 / span 2; } +.gradio-container .hb .pk-btn > .ic { color: var(--faint); } +.gradio-container .hb .pk-pop { position: absolute; z-index: 40; right: 0; width: 420px; max-width: calc(100vw - 32px); top: calc(100% + 6px); background: var(--surface); + border: 1px solid var(--border-strong); --hb-radius: var(--r-lg); box-shadow: var(--shadow-lg); animation: hb-rise .14s ease; overflow: hidden; } +.gradio-container .hb .pk-search { position: relative; border-bottom: 1px solid var(--border); display: flex; align-items: center; } +.gradio-container .hb .pk-search > .ic { position: absolute; left: 12px; color: var(--faint); } +.gradio-container .hb .pk-q[type] { width: 100%; height: 42px; border: 0; background: transparent; color: var(--text); padding: 0 12px 0 36px; outline: none; font-size: 13.5px; box-shadow: none; } +.gradio-container .hb .pk-search label { display: flex; align-items: center; gap: 6px; padding: 0 12px; font-size: 12px; color: var(--muted); white-space: nowrap; cursor: pointer; } +.gradio-container .hb .pk-list { max-height: 340px; overflow: auto; padding: 4px 6px 6px; } +.gradio-container .hb .pk-row { width: 100%; display: grid; grid-template-columns: minmax(0, 1fr) auto; column-gap: 10px; text-align: left; padding: 7px 8px; border: 0; --hb-radius: 7px; + background: none; cursor: pointer; color: var(--text); } +.gradio-container .hb .pk-row:hover, .gradio-container .hb .pk-row.sel { background: var(--surface-2); } +.gradio-container .hb .pk-row .pk-name { font-weight: 500; } +.gradio-container .hb .pk-row .pk-meta { font-size: 11.5px; color: var(--faint); grid-column: 1; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } +.gradio-container .hb .pk-row .pk-price { grid-row: 1 / span 2; grid-column: 2; align-self: center; } +.gradio-container .hb .pk-row .ic.tick { vertical-align: -2px; margin-left: 4px; } +.gradio-container .hb .pk-empty { padding: 18px 12px; text-align: center; color: var(--faint); font-size: 13px; } +.gradio-container .hb .pk-note { font-size: 11.5px; color: var(--faint); padding: 9px 14px; line-height: 1.45; border-top: 1px solid var(--border); background: var(--surface-2); } + +/* ═════════════════════════ RUN LIST ═══════════════════════════════════ */ +.gradio-container .hb .rs-head { display: flex; align-items: flex-end; flex-wrap: wrap; gap: 12px 24px; margin-bottom: 16px; } +.gradio-container .hb .rs-head h1 { font-size: 22px; letter-spacing: -.02em; } +.gradio-container .hb .rs-head p { color: var(--muted); font-size: 13.5px; margin-top: 4px; } +.gradio-container .hb .rs-head .tb-stats { margin-left: auto; } +.gradio-container .hb .rs-bar { display: flex; align-items: center; gap: 10px; flex-wrap: wrap; margin-bottom: 12px; } +.gradio-container .hb .rs-bar .tb-search { flex: 0 1 320px; } +.gradio-container .hb .rs-bar .tb-search input.hb-input[type] { height: 34px; font-size: 13.5px; } +.gradio-container .hb .rs-bar .right { margin-left: auto; display: flex; align-items: center; gap: 10px; font-size: 12.5px; color: var(--faint); } +.gradio-container .hb .rs-table { background: var(--surface); border: 1px solid var(--border); --hb-radius: var(--r-lg); overflow: hidden; } +.gradio-container .hb .rs-row { display: grid; grid-template-columns: 22px 96px minmax(0, 1fr) minmax(0, 150px) minmax(0, 200px) 60px 70px 92px; gap: 14px; align-items: center; + padding: 10px 16px; font-size: 13px; border-top: 1px solid var(--border); cursor: pointer; min-width: 0; } +.gradio-container .hb .rs-row.head { border-top: 0; background: var(--surface-2); font-size: 12.5px; font-weight: 600; color: var(--muted); padding-top: 8px; padding-bottom: 8px; cursor: default; } +.gradio-container .hb .rs-row:not(.head):hover, .gradio-container .hb .rs-row[aria-current="true"] { background: var(--surface-2); } +.gradio-container .hb .rs-row .t { display: grid; min-width: 0; } +.gradio-container .hb .rs-row .t b { font-weight: 500; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } +.gradio-container .hb .rs-row .t em { font-style: normal; font-size: 12px; color: var(--faint); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } +.gradio-container .hb .rs-row .m { color: var(--muted); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; min-width: 0; } +.gradio-container .hb .rs-row .m code { font-size: 12px; color: var(--text); } +.gradio-container .hb .rs-row .d, .gradio-container .hb .rs-row .w { font-size: 12.5px; color: var(--faint); text-align: right; font-variant-numeric: tabular-nums; white-space: nowrap; } +.gradio-container .hb .rs-row .rw { text-align: right; } + +/* ═════════════════════════ RUN PAGE ═══════════════════════════════════ */ +.gradio-container .hb .rp-head { display: grid; gap: 10px; margin-bottom: 18px; min-width: 0; } +.gradio-container .hb .rp-head .kick { display: flex; flex-wrap: wrap; gap: 8px; align-items: center; } +.gradio-container .hb .rp-head .kick .when { font-size: 12.5px; color: var(--faint); display: inline-flex; align-items: center; gap: 5px; } +.gradio-container .hb .rp-head h1 { font-size: 22px; line-height: 1.3; letter-spacing: -.02em; overflow-wrap: anywhere; } +.gradio-container .hb .rp-bar { display: flex; flex-wrap: wrap; align-items: center; gap: 8px 18px; } +.gradio-container .hb .rp-actions { margin-left: auto; display: flex; flex-wrap: wrap; gap: 8px; } +.gradio-container .hb .rp-stepper { display: grid; grid-template-columns: repeat(var(--n, 3), minmax(0, 1fr)); margin-bottom: 20px; background: var(--surface); border: 1px solid var(--border); + --hb-radius: var(--r-lg); overflow: hidden; } +.gradio-container .hb .rp-step { position: relative; display: grid; grid-template-columns: auto minmax(0, 1fr) auto; column-gap: 8px; align-items: center; padding: 11px 14px; min-width: 0; } +.gradio-container .hb .rp-step + .rp-step { border-left: 1px solid var(--border); } +.gradio-container .hb .rp-step .mk { display: grid; place-items: center; width: 20px; height: 20px; --hb-radius: 50%; border: 1.5px solid var(--border-strong); color: var(--faint); background: var(--surface); } +.gradio-container .hb .rp-step b { font-size: 13px; font-weight: 600; white-space: nowrap; overflow: hidden; text-overflow: ellipsis; } +.gradio-container .hb .rp-step .t { font-size: 12px; color: var(--faint); font-variant-numeric: tabular-nums; } +.gradio-container .hb .rp-step em { grid-column: 2 / -1; font-style: normal; font-size: 11.5px; color: var(--faint); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } +.gradio-container .hb .rp-step.ok .mk { background: var(--ok); border-color: var(--ok); color: var(--surface); } +.gradio-container .hb .rp-step.bad .mk { background: var(--err); border-color: var(--err); color: var(--surface); } +.gradio-container .hb .rp-step.on { background: color-mix(in srgb, var(--live) 5%, var(--surface)); } +.gradio-container .hb .rp-step.on .mk { border-color: var(--live); } +.gradio-container .hb .rp-step.on .mk::after { content: ""; width: 8px; height: 8px; --hb-radius: 50%; background: var(--live); animation: hb-pulse 1.2s ease-in-out infinite; } +.gradio-container .hb .rp-step.on::after { content: ""; position: absolute; left: 0; right: 0; bottom: 0; height: 2px; background: var(--live); } +.gradio-container .hb .rp-step.off b { color: var(--faint); font-weight: 500; } +.gradio-container .hb .rp-grid { display: grid; grid-template-columns: minmax(0, 1fr) 360px; gap: 24px; align-items: start; } +.gradio-container .hb .rp-grid > * { min-width: 0; } +.gradio-container .hb .rp-side { position: sticky; top: 12px; display: grid; gap: 12px; } +.gradio-container .hb .rp-score { display: flex; align-items: baseline; gap: 10px; } +.gradio-container .hb .rp-big { font-size: 40px; font-weight: 700; letter-spacing: -.035em; line-height: 1; font-variant-numeric: tabular-nums; } +.gradio-container .hb .rp-big.full { color: var(--ok); } .gradio-container .hb .rp-big.part { color: var(--warn); } .gradio-container .hb .rp-big.zero { color: var(--err); } .gradio-container .hb .rp-big.none { color: var(--faint); } +.gradio-container .hb .rp-score .of { font-size: 13px; color: var(--faint); } +.gradio-container .hb .rp-meter { height: 6px; --hb-radius: 99px; background: var(--surface-3); overflow: hidden; } +.gradio-container .hb .rp-meter i { display: block; height: 100%; --hb-radius: 99px; background: currentColor; } +.gradio-container .hb .rp-grade .hb-panel-b { display: grid; gap: 12px; } +.gradio-container .hb .rp-list { list-style: none; padding: 0; margin: 0; display: grid; gap: 1px; border: 1px solid var(--border); --hb-radius: var(--r); overflow: hidden; background: var(--border); } +.gradio-container .hb .rp-list li { display: grid; grid-template-columns: 16px minmax(0, 1fr) auto; gap: 8px; font-size: 13px; padding: 8px 10px; background: var(--surface); align-items: start; margin: 0; } +.gradio-container .hb .rp-list li .mk { padding-top: 1px; color: var(--faint); } +.gradio-container .hb .rp-list li.ok .mk { color: var(--ok); } .gradio-container .hb .rp-list li.bad .mk { color: var(--err); } .gradio-container .hb .rp-list li.warn .mk { color: var(--warn); } +.gradio-container .hb .rp-list li b { font-weight: 500; overflow-wrap: anywhere; } +.gradio-container .hb .rp-list li p { margin: 2px 0 0; color: var(--muted); font-size: 12px; overflow-wrap: anywhere; } +.gradio-container .hb .rp-list li .sc { font-variant-numeric: tabular-nums; color: var(--muted); font-size: 12.5px; } +.gradio-container .hb .rp-err { white-space: pre-wrap; word-break: break-word; font: 12px/1.5 var(--mono); padding: 10px 12px; --hb-radius: 8px; margin: 0; + background: color-mix(in srgb, var(--err) 7%, var(--surface)); color: var(--err); border: 1px solid color-mix(in srgb, var(--err) 25%, var(--border)); max-height: 260px; overflow: auto; } + +/* trajectory */ +.gradio-container .hb .tl-bar { display: flex; align-items: center; gap: 12px; margin-bottom: 10px; } +.gradio-container .hb .tl-bar h2 { font-size: 15px; } +.gradio-container .hb .tl-bar .n { font-size: 12.5px; color: var(--faint); } +.gradio-container .hb .tl-bar .right { margin-left: auto; display: flex; gap: 6px; align-items: center; } +.gradio-container .hb .tl-conv + .tl-conv { margin-top: 18px; } +.gradio-container .hb .tl-conv-h { display: flex; align-items: center; gap: 8px; font-size: 12.5px; color: var(--muted); margin: 0 0 6px 38px; } +.gradio-container .hb .tl-conv-h b { color: var(--text); font-weight: 600; } +.gradio-container .hb .tl-conv-h b::first-letter { text-transform: uppercase; } +.gradio-container .hb .tl { list-style: none; margin: 0; padding: 0; display: grid; grid-template-columns: minmax(0, 1fr); position: relative; } +.gradio-container .hb .tl::before { content: ""; position: absolute; left: 13px; top: 8px; bottom: 8px; width: 1px; background: var(--border); } +.gradio-container .hb .ev { position: relative; display: grid; grid-template-columns: 28px minmax(0, 1fr) 40px; column-gap: 10px; align-items: start; padding: 4px 0; margin: 0; } +.gradio-container .hb .ev > .ei { position: relative; z-index: 1; display: grid; place-items: center; width: 28px; height: 28px; --hb-radius: 8px; background: var(--surface); + border: 1px solid var(--border); color: var(--faint); } +.gradio-container .hb .ev > .ts { font-size: 11.5px; color: var(--faint); font-variant-numeric: tabular-nums; padding-top: 7px; white-space: nowrap; text-align: right; } +.gradio-container .hb .ev > .eb { min-width: 0; } +.gradio-container .hb .ev-text .msg { background: var(--surface); border: 1px solid var(--border); --hb-radius: var(--r); padding: 10px 14px; } +.gradio-container .hb .ev-text .msg.hb-prose { font-size: 14px; line-height: 1.65; } +.gradio-container .hb .ev-text > .ei { color: var(--text); } +.gradio-container .hb .ev-final > .ei { color: var(--ok); border-color: color-mix(in srgb, var(--ok) 40%, var(--border)); } +.gradio-container .hb .ev details { border: 1px solid var(--border); --hb-radius: var(--r); background: var(--surface); min-width: 0; overflow: hidden; } +.gradio-container .hb .ev details > summary { display: flex; align-items: center; gap: 8px; min-width: 0; padding: 5px 10px; min-height: 30px; font-size: 13px; } +.gradio-container .hb .ev details > summary:hover { background: var(--surface-2); } +.gradio-container .hb .ev details[open] > summary { border-bottom: 1px solid var(--border); } +.gradio-container .hb .ev details > summary b { font-size: 12.5px; font-weight: 600; white-space: nowrap; } +.gradio-container .hb .ev details > summary code { font-size: 12px; color: var(--muted); overflow: hidden; text-overflow: ellipsis; white-space: nowrap; flex: 1; min-width: 0; } +.gradio-container .hb .ev details > summary .ms { font-size: 11px; color: var(--faint); font-variant-numeric: tabular-nums; white-space: nowrap; } +.gradio-container .hb .ev details > summary > .ms:last-child { margin-left: auto; } +.gradio-container .hb .ev .peek { flex: 1; min-width: 0; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; color: var(--faint); font-size: 12.5px; font-style: italic; } +.gradio-container .hb .ev .io { padding: 10px; display: grid; gap: 10px; min-width: 0; } +.gradio-container .hb .ev .io h5 { margin: 0 0 5px; font-size: 12px; color: var(--faint); font-weight: 600; display: flex; align-items: center; gap: 8px; } +.gradio-container .hb .ev .io h5 span { font-weight: 400; margin-left: auto; } +.gradio-container .hb .ev .io .hb-code { max-height: 340px; } +.gradio-container .hb .ev .thought { padding: 10px 12px; font-size: 13px; line-height: 1.65; color: var(--muted); white-space: pre-wrap; overflow-wrap: anywhere; max-height: 420px; overflow: auto; } +.gradio-container .hb .ev .prompt { padding: 12px 14px; max-height: 440px; overflow: auto; } +.gradio-container .hb .ev .prompt.hb-prose { font-size: 13.5px; } +.gradio-container .hb .ev-tool.bad > .ei { color: var(--err); border-color: color-mix(in srgb, var(--err) 40%, var(--border)); } +.gradio-container .hb .ev-think > .ei { color: var(--think); } +.gradio-container .hb .ev-live > .ei { color: var(--live); border-color: color-mix(in srgb, var(--live) 40%, var(--border)); } +.gradio-container .hb .ev-live .eb { display: flex; align-items: center; gap: 8px; font-size: 13px; color: var(--muted); padding-top: 5px; } +.gradio-container .hb .ev-done > .ei { color: var(--ok); } +.gradio-container .hb .ev-done .eb { font-size: 13px; color: var(--muted); padding-top: 5px; } +.gradio-container .hb .tl-empty { padding: 20px 0 0 38px; color: var(--muted); font-size: 13px; } +/* terminal output: dark in both themes, because it is a terminal */ +.gradio-container .hb .term { margin: 0; padding: 10px 12px; --hb-radius: 8px; background: var(--term-bg); color: var(--term-text); border: 1px solid var(--term-border); + font: 12px/1.55 var(--mono); white-space: pre-wrap; word-break: break-word; overflow: auto; max-height: 360px; } +.gradio-container .hb .term .ps { color: var(--term-prompt); } +.gradio-container .hb .term .fold { color: var(--term-faint); } +.gradio-container .hb .term.err { border-color: color-mix(in srgb, var(--err) 45%, var(--term-border)); } +.gradio-container .hb .term-more > summary { font: 12px var(--mono); color: var(--term-faint); padding: 6px 0 0; } +.gradio-container .hb .term-more > summary:hover { color: var(--term-text); } +.gradio-container .hb .diff { font: 12px/1.55 var(--mono); white-space: pre-wrap; word-break: break-word; overflow: auto; max-height: 360px; padding: 8px 0; background: var(--surface-2); + border: 1px solid var(--border); --hb-radius: 8px; margin: 0; } +.gradio-container .hb .diff span { display: block; padding: 0 12px; } +.gradio-container .hb .diff .d { background: color-mix(in srgb, var(--err) 10%, transparent); } +.gradio-container .hb .diff .a { background: color-mix(in srgb, var(--ok) 12%, transparent); } +.gradio-container .hb .diff .d::before { content: "- "; color: var(--err); } +.gradio-container .hb .diff .a::before { content: "+ "; color: var(--ok); } +.gradio-container .hb .cmd { font: 12.5px/1.55 var(--mono); white-space: pre-wrap; word-break: break-word; padding: 8px 12px; background: var(--surface-2); border: 1px solid var(--border); + --hb-radius: 8px; margin: 0; max-height: 280px; overflow: auto; color: var(--text); } +.gradio-container .hb .cmd .ps { color: var(--faint); user-select: none; } +.gradio-container .hb .args { display: grid; grid-template-columns: max-content minmax(0, 1fr); gap: 4px 14px; margin: 0; } +.gradio-container .hb .args dt { color: var(--faint); font-family: var(--mono); font-size: 12px; } +.gradio-container .hb .args dd { margin: 0; font-family: var(--mono); font-size: 12px; overflow-wrap: anywhere; white-space: pre-wrap; } +.gradio-container .hb .rp-conf { display: inline-block; width: 64px; height: 6px; --hb-radius: 99px; background: var(--surface-3); overflow: hidden; vertical-align: middle; } +.gradio-container .hb .rp-conf i { display: block; height: 100%; background: var(--muted); } + +/* ═════════════════════════ COMPARE ════════════════════════════════════ */ +.gradio-container .hb .cmp { display: grid; gap: 16px; } +.gradio-container .hb .cmp-tag { display: inline-grid; place-items: center; width: 20px; height: 20px; --hb-radius: 5px; background: var(--primary); color: var(--on-primary); + font-size: 11px; font-weight: 700; flex: none; } +.gradio-container .hb .cmp .best { font-weight: 700; } + +/* ═════════════════════════ SETUP ══════════════════════════════════════ */ +.gradio-container .hb .su { display: grid; grid-template-columns: repeat(2, minmax(0, 1fr)); gap: 12px; align-items: start; } +.gradio-container .hb .su .wide { grid-column: 1 / -1; } +.gradio-container .hb .su .hb-tbl { border: 0; --hb-radius: 0; background: transparent; } +.gradio-container .hb .su .hb-tbl th:first-child, .gradio-container .hb .su .hb-tbl td:first-child { padding-left: 18px; } +.gradio-container .hb .su .hb-tbl th:last-child, .gradio-container .hb .su .hb-tbl td:last-child { padding-right: 18px; } +.gradio-container .hb .su-set td:nth-child(2) { white-space: nowrap; } +.gradio-container .hb .su-set td:last-child { color: var(--muted); font-size: 12.5px; } +.gradio-container .hb .su-set code { font-size: 11.5px; color: var(--muted); } +.gradio-container .hb .su .hb-tbl td:has(> code), .gradio-container .hb .su .hb-tbl code { white-space: nowrap; overflow-wrap: normal; } +.gradio-container .hb .su .hb-tbl td:first-child { white-space: nowrap; } +.gradio-container .hb-actions { justify-content: flex-start !important; gap: 8px !important; margin: 12px 0 !important; } +.gradio-container .hb-actions button { flex: 0 0 auto !important; width: auto !important; min-width: 0 !important; font: 500 12.5px var(--font) !important; + height: 30px !important; padding: 0 12px !important; --hb-radius: 7px; border: 1px solid var(--border-strong) !important; background: var(--surface) !important; + color: var(--text) !important; box-shadow: var(--shadow-sm) !important; } +.gradio-container .hb-actions button:hover { background: var(--surface-2) !important; } +.gradio-container .hb.accordion, .gradio-container .hb .accordion { --hb-radius: var(--r-lg); border: 1px solid var(--border) !important; background: var(--surface) !important; } + +/* ── responsive ─────────────────────────────────────────────────────── */ +@media (max-width: 1180px) { + .gradio-container .hb-row { flex-wrap: wrap !important; } + .gradio-container .hb-side { flex: 1 1 100% !important; max-width: none; position: static; } + .gradio-container .hb .rp-grid { grid-template-columns: minmax(0, 1fr); } + .gradio-container .hb .rp-side { position: static; } + .gradio-container .hb .tp-grid { grid-template-columns: minmax(0, 1fr); } + .gradio-container .hb .tp-toc { display: none; } +} +@media (max-width: 900px) { + .gradio-container .hb .su { grid-template-columns: minmax(0, 1fr); } + .gradio-container .hb .tb-layout { grid-template-columns: minmax(0, 1fr); gap: 12px; } + .gradio-container .hb .tb-filters { position: static; max-height: none; display: flex; gap: 8px; overflow-x: auto; padding: 0 0 4px; align-items: flex-start; } + .gradio-container .hb .tb-filters-h { display: none; } + .gradio-container .hb .tb-facet { border: 1px solid var(--border); --hb-radius: var(--r); padding: 6px 10px; min-width: 210px; flex: none; background: var(--surface); } + .gradio-container .hb .rs-row { grid-template-columns: 22px 92px minmax(0, 1fr) 60px 70px; } + .gradio-container .hb .rs-row > :nth-child(4), .gradio-container .hb .rs-row > :nth-child(5), .gradio-container .hb .rs-row > :nth-child(8) { display: none; } + .gradio-container .hb .tp-files { grid-template-columns: minmax(0, 1fr); } + .gradio-container .hb .tp-tree { border-right: 0; border-bottom: 1px solid var(--border); max-height: 220px; } + .gradio-container .hb .rp-stepper { grid-template-columns: minmax(0, 1fr); } + .gradio-container .hb .rp-step + .rp-step { border-left: 0; border-top: 1px solid var(--border); } + .gradio-container .hb .tb-stats { gap: 18px; } +} +@media (max-width: 560px) { + .gradio-container .main.fillable { padding-left: 4px !important; padding-right: 4px !important; } + .gradio-container .hb .hb-crumbs { flex-wrap: wrap; row-gap: 2px; } + .gradio-container .hb .hb-crumbs button { max-width: 62vw; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; } + .gradio-container .hb .tp-sec .hb-panel-h .aside { display: none; } + .gradio-container .hb .hb-panel-h { padding: 12px 14px; } + .gradio-container .hb .hb-panel-b { padding: 14px; } + .gradio-container .hb .hb-top .hb-facts { margin-left: 0; } + .gradio-container .hb .tp-head h1, .gradio-container .hb .rp-head h1 { font-size: 19px; } + .gradio-container .hb .tb-head h1, .gradio-container .hb .rs-head h1 { font-size: 19px; } + .gradio-container .hb .rs-row { grid-template-columns: 22px minmax(0, 1fr) 56px; gap: 10px; padding: 10px 12px; } + .gradio-container .hb .rs-row > :nth-child(2), .gradio-container .hb .rs-row > :nth-child(7) { display: none; } + .gradio-container .hb .ev { grid-template-columns: 24px minmax(0, 1fr); column-gap: 8px; } + .gradio-container .hb .ev > .ts { display: none; } + .gradio-container .hb .ev > .ei { width: 24px; height: 24px; } + .gradio-container .hb .tl::before { left: 11px; } +.gradio-container .hb .tl-conv-h, .gradio-container .hb .tl-empty { margin-left: 32px; } + .gradio-container .hb .pk-pop { position: fixed; left: 16px; right: 16px; width: auto; top: 16vh; } + .gradio-container .hb .hb-pair { grid-template-columns: minmax(0, 1fr); } + .gradio-container .hb .rp-actions { margin-left: 0; } +} +@media (prefers-reduced-motion: reduce) { + .gradio-container .hb *, .gradio-container .hb *::before, .gradio-container .hb *::after { animation: none !important; transition: none !important; } +} diff --git a/src/openenv/harbor/ui_assets/header.js b/src/openenv/harbor/ui_assets/header.js new file mode 100644 index 000000000..44c3c7809 --- /dev/null +++ b/src/openenv/harbor/ui_assets/header.js @@ -0,0 +1,50 @@ +// Header: a Docs link in the tab bar, and tabs that go back to their list when clicked again. +// +// The tab bar is Gradio's, so the link is placed into it from here and put back if Gradio redraws +// the bar. Clicking Tasks while a task is open, or Runs while a run is open, fires `select` with +// `{back: "tasks"}` or `{back: "runs"}`; Python clears that tab's open item, which shows its list. + +const DOCS = (element.querySelector("[data-docs]") || {}).dataset?.docs; +const TITLE = element.querySelector(".hb-brand b")?.textContent; +if (TITLE) document.title = TITLE; // OpenEnv's page title names the env class + +// Inside the tab list, after the last tab: Gradio moves tabs that do not fit into a "more" menu by +// measuring that list, so it must keep its own width, and a link beside it would be pushed away. +function placeDocs() { + const list = document.querySelector('.hb-tabs [role="tablist"]'); + if (!list || !DOCS) return; + const last = [...list.querySelectorAll('[role="tab"]')].pop(); + const have = list.querySelector(".hb-docs"); + if (have && (!last || have.previousElementSibling === last)) return; + if (have) have.remove(); + const a = document.createElement("a"); + a.className = "hb-docs"; + a.href = DOCS; + a.target = "_blank"; + a.rel = "noopener"; + a.innerHTML = `Docs${icon("external", 13)}`; + last ? last.after(a) : list.appendChild(a); +} +// Watched until the tab bar exists, then only the tab bar, so a live run redrawing every two seconds +// does not run this on each change. +const wide = new MutationObserver(() => { + const bar = document.querySelector(".hb-tabs .tab-wrapper"); + if (!bar) return; + placeDocs(); + wide.disconnect(); + new MutationObserver(placeDocs).observe(bar, { childList: true, subtree: true }); +}); +placeDocs(); +wide.observe(document.querySelector(".gradio-container") || document.body, { childList: true, subtree: true }); + +// Capture phase, so the tab's state is read before Gradio switches it. +document.addEventListener("click", (ev) => { + const tab = ev.target.closest('.hb-tabs [role="tab"]'); + if (!tab || tab.getAttribute("aria-selected") !== "true") return; + const name = tab.textContent.trim(); + if (name === "Tasks" && document.querySelector(".tp-head")) { + try { history.replaceState(null, "", location.pathname + location.search); } catch (_) { /* sandboxed frame */ } + trigger("select", { back: "tasks" }); + } + if (name === "Runs" && document.querySelector(".rp")) trigger("select", { back: "runs" }); +}, true); diff --git a/src/openenv/harbor/ui_assets/icons.json b/src/openenv/harbor/ui_assets/icons.json new file mode 100644 index 000000000..628937073 --- /dev/null +++ b/src/openenv/harbor/ui_assets/icons.json @@ -0,0 +1,43 @@ +{ +"search": "", +"x": "", +"check": "", +"chevronDown": "", +"chevronRight": "", +"arrowLeft": "", +"external": "", +"copy": "", +"link": "", +"file": "", +"folder": "", +"terminal": "", +"pencil": "", +"list": "", +"brain": "", +"message": "", +"play": "", +"refresh": "", +"clock": "", +"cpu": "", +"key": "", +"info": "", +"alert": "", +"shield": "", +"code": "", +"box": "", +"shuffle": "", +"database": "", +"gauge": "", +"user": "", +"columns": "", +"plus": "", +"more": "", +"target": "", +"filePlus": "", +"download": "", +"globe": "", +"tool": "", +"maximize": "", +"minimize": "", +"wrap": "" +} \ No newline at end of file diff --git a/src/openenv/harbor/ui_assets/run_card.js b/src/openenv/harbor/ui_assets/run_card.js new file mode 100644 index 000000000..d48bf14cd --- /dev/null +++ b/src/openenv/harbor/ui_assets/run_card.js @@ -0,0 +1,307 @@ +// Run card: which model, which agent, which sandbox, and the Run button. +// +// Rendered here from `props.value`, which Python sets whenever the task, the endpoint or the agent +// filter changes. Choices made inside the card live in this closure and travel with the event that +// needs them: +// submit {agent, sandbox} start a rollout +// select {include_experimental} re-filter the agents (not `change`, which Gradio +// also fires on every update from Python) +// input {source: "hf", model, route, api_key, local_token} +// {source: "url", url, model, api_key, api, purpose} connect a model +// clear {} back to the server's endpoint +// A token or key is held in this closure only, sent once with `input`, and forgotten once the +// connection works: the server keeps it in memory for this page, never on disk, never in the page. + +const ui = { + src: null, agent: null, sandbox: null, busy: "", stamp: null, msgAt: "go", quiet: false, + hf: { model: "", route: "", local: true, token: "" }, + url: { url: "", model: "", api: "openai", purpose: "eval", key: "", models: null, loading: false, error: "" }, + models: null, modelsError: "", pkOpen: false, pkQ: "", toolsOnly: true, +}; +const $ = (s) => element.querySelector(s); +const usd = (x) => `$${Number(Number(x).toPrecision(3))}`; +const price = (p) => (p && p.price_in != null ? `${usd(p.price_in)} / ${usd(p.price_out)}` : ""); +const ctx = (n) => (!n ? "" : n >= 1e6 ? `${(n / 1e6).toFixed(n % 1e6 ? 1 : 0)}M ctx` : `${Math.round(n / 1000)}k ctx`); +const cheapest = (m) => (m.providers || []).filter((p) => p.price_in != null).sort((a, b) => a.price_in + a.price_out - b.price_in - b.price_out)[0]; +const org = (id) => (id.includes("/") ? id.split("/")[0] : ""); +const leaf = (id) => (id.includes("/") ? id.split("/").slice(1).join("/") : id); + +async function loadModels() { + if (ui.models || ui.modelsLoading) return; + ui.modelsLoading = true; + try { + const res = await server.hb_models(); + ui.models = res.models || []; + ui.modelsError = res.error || ""; + } catch (e) { ui.models = []; ui.modelsError = String(e.message || e); } + ui.modelsLoading = false; + if (!ui.hf.model && ui.models.length) { + const v = props.value || {}; + const want = v.server && ui.models.find((m) => m.id === v.server.model); + ui.hf.model = (want || ui.models.find((m) => m.tools) || ui.models[0]).id; + } + render(); +} + +function pickerRows() { + const q = ui.pkQ.trim().toLowerCase(); + const list = (ui.models || []).filter((m) => (!ui.toolsOnly || m.tools) + && (!q || m.id.toLowerCase().includes(q) || (m.providers || []).some((p) => String(p.name || "").toLowerCase().includes(q)))); + if (!list.length) return `
${ui.models ? "No model matches." : "Loading models…"}
`; + return list.slice(0, 120).map((m) => { + const c = cheapest(m); + const n = (m.providers || []).length; + const meta = [`${n} provider${n === 1 ? "" : "s"}`, ctx(m.context), m.tools ? "tools" : "no tool calls"].filter(Boolean).join(" · "); + return ``; + }).join(""); +} + +function hfBody(v, e) { + const m = (ui.models || []).find((x) => x.id === ui.hf.model); + const c = m && cheapest(m); + const routes = [["", "Auto"], ["fastest", "Fastest"], ["cheapest", "Cheapest"]] + .concat((m ? m.providers : []).map((p) => [p.name, `${p.name}${p.price_in != null ? ` · ${price(p)}` : ""}${p.context ? ` · ${ctx(p.context)}` : ""}${p.tools ? "" : " · no tools"}`])); + if (!routes.some(([k]) => k === ui.hf.route)) ui.hf.route = ""; + // Signed in with Hugging Face (where the Space offers it): that account pays, and no token is typed. + const login = v.hf_login || {}; + const signedIn = !!(login.on && login.user); + if (signedIn && ui.hf.account === undefined) ui.hf.account = true; + const useAccount = signedIn && ui.hf.account; + const useLocal = !useAccount && v.local_token && ui.hf.local; + const here = encodeURIComponent(location.pathname + location.search + location.hash); + const connected = e.ok && e.source === "hf" && formMatches(e, v); + return `
+
+ + ${ui.pkOpen ? `
+
${pickerRows()}
+
Prices are per million input / output tokens, the cheapest provider's. Agents need tool calling.
` : ""} +
+ + ${signedIn ? `
+ @${esc(login.user)} · Sign out
` : ""} + ${v.local_token && !useAccount ? `` : ""} + ${useLocal || useAccount ? "" : ``} + ${login.on && !signedIn ? `

or sign in with Hugging Face to use your own account

` : ""} + ${connected ? `
${icon("check", 14, "ok")}${esc(e.model)}· ${esc(e.level_text)}
${notes(e)}` : ""} + + ${msg("model", v)} +

Billed to ${useAccount ? "your Hugging Face account" : "the token's account"}. Evaluation only: the router returns no token ids to train on.

+
`; +} + +function urlBody(v, e) { + const u = ui.url; + const connected = e.ok && e.source === "url" && formMatches(e, v); + const opts = (u.models || []).map((m) => ``).join(""); + return `
+ + + +
+ + +
+ ${connected ? `
${icon("check", 14, "ok")}${esc(e.model)}· ${esc(e.host)} · ${esc(e.level_text)}
${notes(e)}` : ""} + + ${msg("model", v)} +

Training needs exact token ids: vLLM with --return-tokens-as-token-ids. The key is sent only to that endpoint and kept in memory for this page.

+
`; +} + +// Whether the form still says what was connected. Picking another model or editing the URL after +// connecting means Run would use something the card no longer shows, so that counts as unconnected. +function formMatches(e, v) { + const c = v.custom || {}; + if (e.source === "hf") return c.model === ui.hf.model && (c.route || "") === (ui.hf.route || ""); + if (e.source === "url") { + return (c.url || "") === ui.url.url && (c.model || "") === ui.url.model + && (c.api || "openai") === ui.url.api && (c.purpose || "eval") === ui.url.purpose; + } + return true; +} + +function notes(e) { + return (e.notes || []).map((n) => `

${esc(n)}

`).join(""); +} + +// `quiet` hides the last message from the moment an action starts until Python answers it, so an +// old message never shows under the button that is now waiting. +function msg(at, v) { + return v.message && ui.msgAt === at && !ui.quiet ? `
${icon(v.message.tone === "ok" ? "check" : "alert", 15)}${esc(v.message.text)}
` : ""; +} + +function render() { + const v = props.value || {}; + const e = v.engine || {}; + const sources = v.sources || []; + if (v.stamp !== ui.stamp) { + // A new value from Python ends any pending action. The form is this card's own: it is never + // overwritten from Python, so an edit not yet connected survives a re-render. + ui.stamp = v.stamp; ui.busy = ""; ui.quiet = false; + } + // `false` from the server, not merely absent: an empty card (before its first value) is not read-only + if (v.rollouts === false) { + element.innerHTML = `
+

${icon("file", 15)}Read-only server

+
This server is a read-only task browser. You can inspect tasks and prior runs, but it does not accept model credentials or start agents and sandboxes.
+
`; + return; + } + if (!ui.src || !sources.includes(ui.src)) ui.src = e.source && sources.includes(e.source) ? e.source : sources[0] || null; + const agents = v.agents || []; + if (!agents.some((a) => a.value === ui.agent)) ui.agent = v.agent || (agents[0] && agents[0].value) || null; + const boxes = v.sandboxes || []; + if (!boxes.some((s) => s.name === ui.sandbox && s.available)) ui.sandbox = v.sandbox || (boxes.find((s) => s.available) || {}).name || null; + const agent = agents.find((a) => a.value === ui.agent); + const onEngine = e.ok && e.source === ui.src && formMatches(e, v); + const ready = !!(v.task && onEngine && ui.agent && ui.sandbox && v.rollouts) && !ui.busy; + if (ui.src === "hf") loadModels(); + + const label = { server: "This server", hf: "Hugging Face", url: "Your endpoint" }; + const seg = sources.length > 1 ? `
${sources.map((s) => ``).join("")}
` : ""; + let body = ""; + if (ui.src === "server" && v.server) { + body = `
${esc(v.server.model)}${v.server.train ? "training" : "eval only"}${esc(v.server.host)} · ${esc(v.server.level_text)}
${msg("model", v)}`; + } else if (ui.src === "hf") body = hfBody(v, e); + else if (ui.src === "url") body = urlBody(v, e); + else body = `
${icon("alert", 15)}No model is available: this server has no endpoint of its own and does not accept visitors' models.
`; + + const where = agent ? (agent.host_side ? "runs on this server" : "runs in the sandbox") : ""; + const localWarn = agent && !agent.host_side && v.proxy_local + ? `
${icon("alert", 15)}This server's capture proxy is only reachable from this machine, so an agent inside a sandbox cannot call the model. Pick one that runs on this server, like terminus-2, or start the server with --expose gradio.
` : ""; + const intro = !v.task ? "Open a task to run it." + : `A fresh ${esc(ui.sandbox || "")} sandbox from this task's image, ${ui.agent ? `${esc(ui.agent)} as the agent, ` : ""}then the task's verifier.`; + const why = !v.rollouts ? "Rollouts are turned off on this server." + : !onEngine ? (ui.src === "server" ? "" : "Connect the model first.") + : !ui.agent ? "Pick an agent." : !ui.sandbox ? "No sandbox is available." : ""; + + element.innerHTML = ` +
+

${icon("play", 15)}Run a rollout

+
+

${intro}

+
Model
${seg}${body}
+
Agent${esc(where)}
+ ${agents.length ? `` + : `
${icon("alert", 15)}${esc(v.agents_empty || "No agent is available for this model.")}
`} + +
+
Sandbox
+
${boxes.map((s) => ``).join("")}
+
+
${localWarn} + + ${why ? `

${esc(why)}

` : ""} + ${msg("go", v)} +

Model usage is billed to the endpoint or account selected above. Sandbox compute is billed to this server's operator, including Hugging Face Sandbox.

+

It keeps running if you close this page, and shows under Runs.

+
+
+
`; + // Held in this closure, never in the markup; restored so a re-render does not drop a half-typed key. + const tok = $(".rc-token"); if (tok && ui.hf.token) tok.value = ui.hf.token; + const key = $(".rc-key"); if (key && ui.url.key) key.value = ui.url.key; + if (ui.pkOpen) { const q = $(".pk-q"); if (q) { q.focus(); q.setSelectionRange(q.value.length, q.value.length); } } +} + +function connect(src) { + ui.busy = "connect"; ui.msgAt = "model"; ui.quiet = true; + let payload; + if (src === "hf") { + const v = props.value || {}; + const login = v.hf_login || {}; + const account = !!(login.on && login.user && ui.hf.account); + const local = !account && !!(v.local_token && ui.hf.local); + payload = { source: "hf", model: ui.hf.model, route: ui.hf.route, api_key: account || local ? "" : ui.hf.token, local_token: local, use_account: account }; + } else { + payload = { source: "url", url: ui.url.url, model: ui.url.model, api_key: ui.url.key, api: ui.url.api, purpose: ui.url.purpose }; + } + render(); + trigger("input", payload); +} + +element.addEventListener("click", async (ev) => { + const src = ev.target.closest("[data-src]"); + if (src) { + ui.src = src.dataset.src; ui.pkOpen = false; + const e = (props.value || {}).engine || {}; + if (ui.src === "server" && e.source !== "server") { ui.busy = "connect"; ui.msgAt = "model"; ui.quiet = true; render(); return trigger("clear", {}); } + return render(); + } + if (ev.target.closest("[data-pk]")) { ui.pkOpen = !ui.pkOpen; ui.pkQ = ""; loadModels(); return render(); } + const row = ev.target.closest("[data-model]"); + if (row) { ui.hf.model = row.dataset.model; ui.hf.route = ""; ui.pkOpen = false; return render(); } + const sb = ev.target.closest("[data-sb]"); + if (sb && !sb.disabled) { ui.sandbox = sb.dataset.sb; return render(); } + const c = ev.target.closest("[data-connect]"); + if (c && !c.disabled) return connect(c.dataset.connect); + if (ev.target.closest("[data-load]") && ui.url.url) { + ui.url.loading = true; ui.url.error = ""; render(); + try { + const res = await server.hb_served([ui.url.url, ui.url.key]); + ui.url.models = res.models || null; ui.url.error = res.error || ""; + if (res.models && res.models.length === 1 && !ui.url.model) ui.url.model = res.models[0]; + } catch (e) { ui.url.error = String(e.message || e); } + ui.url.loading = false; return render(); + } + if (ev.target.closest(".rc-run") && !ev.target.closest(".rc-run").disabled) { + ui.busy = "run"; ui.msgAt = "go"; ui.quiet = true; render(); + return trigger("submit", { agent: ui.agent, sandbox: ui.sandbox }); + } +}); +element.addEventListener("change", (ev) => { + const t = ev.target; + if (t.matches(".rc-agent")) { ui.agent = t.value; return render(); } + if (t.matches(".rc-exp")) { ui.msgAt = "go"; ui.quiet = true; return trigger("select", { include_experimental: t.checked }); } + if (t.matches("[data-hf='route']")) { ui.hf.route = t.value; return render(); } + if (t.matches("[data-hf-local]")) { ui.hf.local = t.checked; return render(); } + if (t.matches("[data-hf-account]")) { ui.hf.account = t.checked; return render(); } + if (t.matches(".pk-tools")) { ui.toolsOnly = t.checked; const l = $(".pk-list"); if (l) l.innerHTML = pickerRows(); return; } + if (t.matches("[data-u]")) { + ui.url[t.dataset.u] = t.value; + if (t.dataset.u === "url") { ui.url.models = null; ui.url.error = ""; } + render(); // `change` lands on blur, so redrawing here cannot take the caret away + } +}); +element.addEventListener("input", (ev) => { + const t = ev.target; + if (t.matches(".rc-token")) { ui.hf.token = t.value; return; } + if (t.matches(".rc-key")) { ui.url.key = t.value; return; } + if (t.matches(".pk-q")) { ui.pkQ = t.value; const l = $(".pk-list"); if (l) l.innerHTML = pickerRows(); return; } + if (t.matches("[data-u]")) { + ui.url[t.dataset.u] = t.value; + const b = $(`[data-connect="url"]`); if (b && t.dataset.u === "url") b.disabled = !t.value || ui.busy === "connect"; + const l = $("[data-load]"); if (l && t.dataset.u === "url") l.disabled = !t.value; + } +}); +element.addEventListener("keydown", (ev) => { + if (ev.key === "Escape" && ui.pkOpen) { ui.pkOpen = false; render(); } + if (ev.key === "Enter" && ev.target.matches(".pk-q")) { + const first = $(".pk-row"); if (first) { ui.hf.model = first.dataset.model; ui.hf.route = ""; ui.pkOpen = false; render(); } + } +}); +document.addEventListener("click", (ev) => { + if (ui.pkOpen && !ev.target.closest(".pk")) { ui.pkOpen = false; render(); } +}); + +// A key that worked is not kept in the page any longer than needed. +watch("value", () => { + const v = props.value || {}; + if (v.message && v.message.tone === "ok" && ui.msgAt === "model") { ui.hf.token = ""; ui.url.key = ""; } + render(); +}); +render(); diff --git a/src/openenv/harbor/ui_assets/run_list.js b/src/openenv/harbor/ui_assets/run_list.js new file mode 100644 index 000000000..23da98c16 --- /dev/null +++ b/src/openenv/harbor/ui_assets/run_list.js @@ -0,0 +1,54 @@ +// Run list: open a run, filter by status or text, tick two to four to compare. +// +// The table is rendered in Python and re-rendered when a run changes state, so the filters live here +// and are re-applied after every render. Ticked runs go to Python on each change (`input`), which +// renders them ticked again; Compare sends them with `submit`; opening a row sends `select`. + +const st = { q: "", s: "" }; + +function apply() { + const q = st.q.toLowerCase(); + let n = 0; + element.querySelectorAll(".rs-row[data-rs-id]").forEach((row) => { + const hide = (st.s && row.dataset.s !== st.s) || (q && !row.textContent.toLowerCase().includes(q)); + row.hidden = !!hide; + n += hide ? 0 : 1; + }); + element.querySelectorAll("[data-rs-seg] button").forEach((b) => b.setAttribute("aria-pressed", String(b.dataset.s === st.s))); + const box = element.querySelector(".rs-q"); + if (box && box.value !== st.q) box.value = st.q; + const ticked = [...element.querySelectorAll("[data-rs-pick]:checked")].map((c) => c.dataset.rsPick); + const btn = element.querySelector("[data-rs-compare]"); + if (btn) { + btn.disabled = ticked.length < 2; + btn.innerHTML = `${icon("columns", 14)}${ticked.length >= 2 ? `Compare ${ticked.length}` : "Compare"}`; + } + const hint = element.querySelector(".rs-n"); + if (hint) hint.textContent = ticked.length ? `${ticked.length} ticked` : (st.q || st.s ? `${n} shown` : "Tick two to four to compare"); + element.querySelectorAll("[data-rs-pick]").forEach((c) => { c.disabled = !c.checked && ticked.length >= 4; }); +} + +element.addEventListener("click", (ev) => { + if (ev.target.closest("[data-rs-pick]")) return; // ticking is not opening + const seg = ev.target.closest("[data-rs-seg] button"); + if (seg) { st.s = seg.dataset.s; return apply(); } + if (ev.target.closest("[data-rs-compare]")) { + const ids = [...element.querySelectorAll("[data-rs-pick]:checked")].map((c) => c.dataset.rsPick); + if (ids.length >= 2) trigger("submit", { ids }); + return; + } + const row = ev.target.closest(".rs-row[data-rs-id]"); + if (row) trigger("select", { id: row.dataset.rsId }); +}); +element.addEventListener("change", (ev) => { + if (!ev.target.closest("[data-rs-pick]")) return; + apply(); + trigger("input", { ids: [...element.querySelectorAll("[data-rs-pick]:checked")].map((c) => c.dataset.rsPick) }); +}); +element.addEventListener("input", (ev) => { + if (!ev.target.matches(".rs-q")) return; + st.q = ev.target.value.trim(); + apply(); +}); +watch("value", apply); +apply(); diff --git a/src/openenv/harbor/ui_assets/run_view.js b/src/openenv/harbor/ui_assets/run_view.js new file mode 100644 index 000000000..4d43c936b --- /dev/null +++ b/src/openenv/harbor/ui_assets/run_view.js @@ -0,0 +1,69 @@ +// Run page: its links, "Expand all", the downloads, and keeping what you opened open. +// +// While a rollout runs, the page is re-rendered every two seconds from Python. Without this, a row +// you unfolded would snap shut on the next refresh. So every toggle is remembered by the fold's +// position in the page (events only ever append, so positions are stable while a run grows) and +// re-applied after each render. Another run starts from its own defaults, at the top. +// Links fire `select`: `{back}` to the list, `{run}` for another run, `{task, index}` for its task. + +const opened = new Map(); +let run = null; +const folds = () => [...element.querySelectorAll(".rp-main details")]; + +// `toggle` fires after the change, including after the restore below sets `open`; recording those +// again is harmless, since it records the same value. +element.addEventListener("toggle", (ev) => { + if (ev.target.tagName !== "DETAILS") return; + const i = folds().indexOf(ev.target); + if (i >= 0) opened.set(i, ev.target.open); +}, true); + +async function download(button) { + const label = button.innerHTML; + button.disabled = true; + button.innerHTML = `${esc(button.textContent)}`; + try { + const res = await server.hb_download([button.dataset.grant, button.dataset.dl]); + if (res.error) throw new Error(res.error); + const url = URL.createObjectURL(new Blob([res.text], { type: "application/json" })); + const a = Object.assign(document.createElement("a"), { href: url, download: res.name }); + document.body.appendChild(a); a.click(); a.remove(); + setTimeout(() => URL.revokeObjectURL(url), 2000); + button.innerHTML = label; + } catch (e) { + button.innerHTML = `${icon("alert", 14)}${esc(String(e.message || e).slice(0, 60))}`; + setTimeout(() => { button.innerHTML = label; }, 3000); + } + button.disabled = false; +} + +element.addEventListener("click", (ev) => { + if (ev.target.closest("[data-back]")) return trigger("select", { back: true }); + const task = ev.target.closest("[data-task]"); + if (task) return trigger("select", { task: task.dataset.task, index: Number(task.dataset.index) }); + const other = ev.target.closest("[data-run]:not(.rp)"); + if (other) return trigger("select", { run: other.dataset.run }); + const dl = ev.target.closest("[data-dl]"); + if (dl) return download(dl); + const b = ev.target.closest("[data-expand]"); + if (!b) return; + const open = b.textContent.trim() === "Expand all"; + folds().forEach((d, i) => { d.open = open; opened.set(i, open); }); + b.textContent = open ? "Collapse all" : "Expand all"; +}); + +function restore() { + const root = element.querySelector("[data-run]"); + const id = root ? root.dataset.run : null; + if (id !== run) { + run = id; opened.clear(); + const top = element.getBoundingClientRect().top; + if (top < 0) window.scrollBy({ top: top - 12 }); + return; + } + folds().forEach((d, i) => { if (opened.has(i)) d.open = opened.get(i); }); + const b = element.querySelector("[data-expand]"); + if (b && opened.size && [...opened.values()].every(Boolean)) b.textContent = "Collapse all"; +} +watch("value", restore); +restore(); diff --git a/src/openenv/harbor/ui_assets/task_browser.html b/src/openenv/harbor/ui_assets/task_browser.html new file mode 100644 index 000000000..d010ef324 --- /dev/null +++ b/src/openenv/harbor/ui_assets/task_browser.html @@ -0,0 +1,20 @@ +
+
+

Tasks

Every task this server serves. Open one to read exactly what the agent gets, then run an agent on it.

+
+
+
+ + +
+
+ +
+
+
+
    + +
    +
    +
    diff --git a/src/openenv/harbor/ui_assets/task_browser.js b/src/openenv/harbor/ui_assets/task_browser.js new file mode 100644 index 000000000..c289da59a --- /dev/null +++ b/src/openenv/harbor/ui_assets/task_browser.js @@ -0,0 +1,387 @@ +// Task browser: every served dataset's tasks, searched and filtered as you type. +// +// Rendered here rather than in Python because a server can hold thousands of tasks, and filtering them +// on every keystroke must not cost a round trip. Rows come from `server.hb_tasks(spec)` once per +// dataset. Opening a task fires `select` with `{spec, index}`, and the Python side shows the task page. +// `props.value` is never written: that would re-render the template and wipe the list. +// Adding a Hub dataset runs on the server in the background (`hb_add`), and the add panel polls it +// (`hb_add_status`) to show the download as it happens. `esc` and `icon` come from ui_icons.py. + +const PAGE = 60; +const FACETS = [["spec", "Dataset"], ["category", "Category"], ["difficulty", "Difficulty"], ["tags", "Tags"]]; +const SHOW = 8; +const st = { datasets: [], rows: [], loading: 0, q: "", sel: {}, more: {}, shown: PAGE, canAdd: false, common: {}, + add: { open: false, q: "", hits: null, error: "", timer: null, counts: {}, jobs: {}, poll: null }, + rm: { confirm: null, busy: null, error: "" } }; +const $ = (sel) => element.querySelector(sel); +const fmt = (n) => Number(n || 0).toLocaleString(); +// A dataset on the Space's bucket is a folder like `/data/org__name`; it is shown by its Hub id. +const labels = {}; +const nameOf = (spec) => labels[spec] || spec; +const short = (spec) => String(nameOf(spec)).split("/").pop(); + +// Tags every task of a dataset carries (`terminal` on a terminal dataset) tell no two tasks apart, +// so they are left off the cards and out of the tag filter. +function tagsOf(r) { return (r.keywords || []).filter((k) => !(st.common[r.spec] || new Set()).has(k)); } +function values(r, key) { + if (key === "tags") return tagsOf(r); + const v = r[key]; + return v ? [v] : []; +} + +function matches(r, skip) { + for (const [key] of FACETS) { + if (key === skip) continue; + const want = st.sel[key]; + if (want && want.size && !values(r, key).some((v) => want.has(v))) return false; + } + if (!st.q) return true; + return st.q.split(/\s+/).every((t) => r._hay.includes(t)); +} + +// Split the raw text on the search terms and escape each piece, so a match can never land inside an +// escaped entity (searching "39" in "Don't" must not break `'`). +function mark(text) { + const raw = String(text ?? ""); + const terms = st.q.split(/\s+/).filter((t) => t.length > 1).map((t) => t.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")); + if (!terms.length) return esc(raw); + return raw.split(new RegExp(`(${terms.join("|")})`, "gi")).map((part, i) => (i % 2 ? `${esc(part)}` : esc(part))).join(""); +} + +function renderStats() { + const ok = st.datasets.filter((d) => !d.error); + $(".tb-stats").innerHTML = `
    ${fmt(ok.length)}dataset${ok.length === 1 ? "" : "s"}
    ` + + `
    ${fmt(st.rows.length)}tasks${st.loading ? " so far" : ""}
    `; +} + +function renderFilters() { + const parts = [`
    Filters${Object.values(st.sel).some((s) => s.size) ? '' : ""}
    `]; + for (const [key, label] of FACETS) { + const counts = new Map(); + for (const r of st.rows) if (matches(r, key)) for (const v of values(r, key)) counts.set(v, (counts.get(v) || 0) + 1); + const all = new Map(); + for (const r of st.rows) for (const v of values(r, key)) all.set(v, (all.get(v) || 0) + 1); + if (key === "spec") for (const d of st.datasets) if (!all.has(d.spec)) all.set(d.spec, 0); + if (all.size < (key === "spec" ? 1 : 2) && !(key === "spec" && st.canAdd)) continue; + const chosen = st.sel[key] || new Set(); + const opts = [...all.keys()].sort((a, b) => (counts.get(b) || 0) - (counts.get(a) || 0) || String(a).localeCompare(String(b))); + const limit = st.more[key] ? opts.length : SHOW; + const rows = opts.slice(0, limit).map((v) => { + const d = key === "spec" ? st.datasets.find((x) => x.spec === v) : null; + const n = d && d.error ? "failed" : d && d.loading ? '' : fmt(counts.get(v) || 0); + const opt = ``; + if (!(d && d.removable)) return opt; + if (st.rm.busy === v) return `
    Removing ${esc(short(v))}…
    `; + if (st.rm.confirm === v) { + const where = v.startsWith("/") ? "its copy in this Space's bucket" : "its download on this server"; + return `

    Remove ${esc(nameOf(v))}? This deletes ${where}. Runs of its tasks stay in history.

    + ${st.rm.error ? `

    ${esc(st.rm.error)}

    ` : ""} +
    `; + } + return `
    ${opt}
    `; + }).join(""); + const more = opts.length > SHOW ? `` : ""; + const adding = key === "spec" ? Object.values(st.add.jobs).filter((j) => !["done", "error"].includes(j.state)) : []; + const pending = adding.map((j) => `
    ${esc(short(j.spec))}${esc(jobWord(j))}
    `).join(""); + const hub = key === "spec" && st.canAdd ? `${pending}` : ""; + parts.push(`

    ${label}

    ${chosen.size ? `${chosen.size}` : ""}${icon("chevronRight", 14, "chev")}
    ${rows}${more}${hub}
    `); + } + const box = $(".tb-filters"); + // Which facets are folded survives re-renders. On a phone they start folded, so the tasks, not the + // filters, fill the first screen. + if (!st.folded) st.folded = new Set(matchMedia("(max-width: 900px)").matches ? FACETS.map(([k]) => k) : []); + box.querySelectorAll(".tb-facet").forEach((d) => { d.open ? st.folded.delete(d.dataset.facet) : st.folded.add(d.dataset.facet); }); + box.innerHTML = parts.join(""); + box.querySelectorAll(".tb-facet").forEach((d) => { if (st.folded.has(d.dataset.facet)) d.open = false; }); +} + +function card(r) { + const tags = tagsOf(r).slice(0, 4); + const bits = [r.category, r.difficulty].filter(Boolean).map(esc); + return `
  • `; +} + +function renderList() { + const hits = st.rows.filter((r) => matches(r)); + $(".tb-list").innerHTML = hits.slice(0, st.shown).map(card).join(""); + if (!hits.length) { + $(".tb-list").innerHTML = st.loading && !st.rows.length ? "" : `
  • ${icon("search", 20)}

    No tasks match

    Try fewer words, or clear the filters.

  • `; + } + if (st.loading && !st.rows.length) $(".tb-list").innerHTML = Array.from({ length: 6 }, () => '
  • ').join(""); + $(".count").textContent = st.loading && !st.rows.length ? "Loading tasks…" : `${fmt(hits.length)} task${hits.length === 1 ? "" : "s"}`; + const pills = []; + for (const [key, label] of FACETS) for (const v of st.sel[key] || []) { + pills.push(``); + } + if (st.q) pills.push(``); + $(".tb-active").innerHTML = pills.join(""); + const more = $(".tb-more"); + more.hidden = hits.length <= st.shown; + more.innerHTML = `Showing ${fmt(Math.min(st.shown, hits.length))} of ${fmt(hits.length)}`; + $(".tb-random").disabled = !hits.length; +} + +function render() { renderStats(); renderFilters(); renderList(); } + +function open(spec, index) { + try { history.replaceState(null, "", `#task=${encodeURIComponent(spec)}:${index}`); } catch (_) { /* sandboxed frame */ } + trigger("select", { spec, index: Number(index) }); +} + +function learn(spec, rows) { + const counts = new Map(); + for (const r of rows) for (const k of r.keywords || []) counts.set(k, (counts.get(k) || 0) + 1); + st.common[spec] = new Set([...counts].filter(([, n]) => rows.length > 3 && n >= rows.length * 0.9).map(([k]) => k)); + for (const r of rows) { + r.spec = spec; + r._hay = `${r.title} ${r.brief || ""} ${r.name} ${r.category} ${r.difficulty} ${(r.keywords || []).join(" ")} ${spec} #${r.index}`.toLowerCase(); + } + st.rows = st.rows.filter((r) => r.spec !== spec).concat(rows); +} + +async function load(d) { + d.loading = true; st.loading += 1; render(); + try { + const res = await server.hb_tasks(d.spec); + if (res.error) throw new Error(res.error); + d.num_tasks = res.rows.length; + learn(d.spec, res.rows); + } catch (e) { d.error = String(e.message || e); } + d.loading = false; st.loading -= 1; + render(); +} + +// ── events ────────────────────────────────────────────────────────────── +element.addEventListener("click", (ev) => { + const c = ev.target.closest(".tb-card"); + if (c) return open(c.dataset.spec, c.dataset.i); + const f = ev.target.closest("[data-f]"); + if (f) { + const set = st.sel[f.dataset.f] || (st.sel[f.dataset.f] = new Set()); + set.has(f.dataset.v) ? set.delete(f.dataset.v) : set.add(f.dataset.v); + st.shown = PAGE; return render(); + } + if (ev.target.closest("[data-more]")) { const k = ev.target.closest("[data-more]").dataset.more; st.more[k] = !st.more[k]; return renderFilters(); } + if (ev.target.closest("[data-clear]")) { st.sel = {}; st.shown = PAGE; return render(); } + if (ev.target.closest("[data-clear-q]")) { st.q = ""; $(".tb-q").value = ""; return render(); } + if (ev.target.closest("[data-page]")) { st.shown += PAGE; return renderList(); } + if (ev.target.closest(".tb-random")) { + const hits = st.rows.filter((r) => matches(r)); + if (hits.length) { const r = hits[Math.floor(Math.random() * hits.length)]; open(r.spec, r.index); } + return; + } + if (ev.target.closest("[data-add-open]")) { st.add.open = true; renderAdd(); $(".tb-addq")?.focus(); if (!st.add.hits) hubSearch(""); return; } + if (ev.target.closest("[data-add-close]")) { st.add.open = false; return renderAdd(); } + const add = ev.target.closest("[data-add]"); + if (add && !add.disabled) return addDataset(add.dataset.add); + const rm = ev.target.closest("[data-rm]"); + if (rm) { st.rm.confirm = rm.dataset.rm; st.rm.error = ""; return renderFilters(); } + if (ev.target.closest("[data-rm-no]")) { st.rm.confirm = null; return renderFilters(); } + const yes = ev.target.closest("[data-rm-yes]"); + if (yes) return removeDataset(yes.dataset.rmYes); + const show = ev.target.closest("[data-add-show]"); + if (show) { const d = st.datasets.find((x) => x.spec === show.dataset.addShow || x.label === show.dataset.addShow); st.sel = { spec: new Set([d ? d.spec : show.dataset.addShow]) }; st.add.open = false; st.shown = PAGE; render(); return renderAdd(); } +}); +let qTimer = null; +element.addEventListener("input", (ev) => { + if (ev.target.matches(".tb-q")) { + clearTimeout(qTimer); + qTimer = setTimeout(() => { st.q = ev.target.value.trim().toLowerCase(); st.shown = PAGE; renderFilters(); renderList(); }, 120); + } + if (ev.target.matches(".tb-addq")) { st.add.q = ev.target.value; hubSearch(ev.target.value); } +}); +element.addEventListener("keydown", (ev) => { + if (ev.target.matches(".tb-addq") && ev.key === "Enter" && ev.target.value.includes("/")) addDataset(ev.target.value.trim()); + if (ev.target.matches(".tb-addq") && ev.key === "Escape") { st.add.open = false; renderAdd(); } +}); +// "/" searches tasks from anywhere on the list page, as long as nothing else is being typed into. +document.addEventListener("keydown", (ev) => { + if (ev.key !== "/" || ev.metaKey || ev.ctrlKey || !element.offsetParent) return; + if (ev.target.closest && ev.target.closest("input, textarea, select, [contenteditable]")) return; + ev.preventDefault(); $(".tb-q").focus(); +}); + +// ── adding a dataset from the Hub ─────────────────────────────────────── +function bytes(n) { + return n >= 1e9 ? `${(n / 1e9).toFixed(1)} GB` : n >= 1e6 ? `${Math.round(n / 1e6)} MB` : `${Math.max(1, Math.round(n / 1e3))} KB`; +} + +function jobWord(j) { + if (j.state === "checking") return "checking"; + if (j.state === "copying") return "copying"; + if (j.state === "mounting") return "mounting"; + if (j.state === "downloading") return j.total ? `${Math.floor((100 * j.done) / j.total)}%` : "downloading"; + if (j.state === "indexing") return "reading"; + return j.state; +} + +function addRow(h) { + const spec = h.id; + const job = st.add.jobs[spec]; + const have = st.datasets.some((d) => d.spec === spec || d.label === spec); + const info = st.add.counts[spec]; + const n = info === undefined ? undefined : info.tasks; + const [org, name] = spec.includes("/") ? spec.split("/") : ["", spec]; + const meta = [h.downloads != null ? `${fmt(h.downloads)} downloads` : "", h.rl_environment ? "rl-environment" : ""].filter(Boolean).join(" · "); + const big = info && info.bytes > 1e9; + const tasks = n === undefined ? '…' + : info.error ? `couldn't check yet` + : n === null ? 'not in Harbor\'s tasks/ layout' + : `${fmt(n)} task${n === 1 ? "" : "s"}${info.bytes ? ` · ${bytes(info.bytes)}` : ""}`; + let action; + if (have && (!job || job.state === "done")) { + action = `${icon("check", 14)}Added`; + } else if (job && job.state === "error") { + action = `${icon("alert", 14)}${esc(job.error)}`; + } else if (job) { + const pct = job.state === "downloading" && job.total ? (100 * job.done) / job.total + : job.state === "indexing" ? 100 : job.state === "mounting" ? 60 : job.state === "copying" ? 30 : 4; + const words = job.state === "checking" ? "Checking the layout…" + : job.state === "copying" ? "Copying into the bucket…" + : job.state === "mounting" ? "Waiting for the bucket mount…" + : job.state === "downloading" ? (job.total ? `Downloading ${fmt(job.done)} of ${fmt(job.total)} files` : "Starting the download…") + : "Reading the tasks…"; + action = `
    ${esc(words)}
    `; + } else { + action = ``; + } + return `
    ${org ? `${esc(org)}/` : ""}${esc(name)}${esc(meta)}
    + ${tasks}
    ${action}
    `; +} + +function renderAdd() { + const box = $(".tb-addwrap"); + if (!st.add.open) { box.innerHTML = ""; return; } + const typed = st.add.q.trim(); + let hits = st.add.hits || []; + if (typed.includes("/") && !hits.some((h) => h.id === typed)) hits = [{ id: typed }, ...hits]; + const body = st.add.error ? `
    ${icon("alert", 15)}${esc(st.add.error)}
    ` + : st.add.hits === null ? '
    Searching the Hub…
    ' + : hits.length ? hits.map(addRow).join("") : '
    No Harbor datasets match. Type an id like org/name to add one directly.
    '; + const had = $(".tb-addq"); + const focus = had && document.activeElement === had; + const list = box.querySelector(".tb-add-list"); + if (had && list) { list.innerHTML = body; return; } // keep the search box, and its focus, as it is + box.innerHTML = `
    +

    ${icon("plus", 15)}Add a Harbor dataset

    public Hub datasets tagged harbor +
    +
    +
    ${body}
    `; + if (focus) $(".tb-addq")?.focus(); +} + +function hubSearch(q) { + clearTimeout(st.add.timer); + st.add.timer = setTimeout(async () => { + st.add.hits = null; st.add.error = ""; renderAdd(); + try { + st.add.hits = await server.hb_hub(q.trim()); + } catch (e) { st.add.hits = []; st.add.error = String(e.message || e); } + renderAdd(); + inspect(); + }, 250); +} + +// Task counts for what is listed, one at a time: it is what tells a Harbor dataset from one this +// server cannot load, before anything is downloaded. +async function inspect() { + const want = [...(st.add.hits || []).map((h) => h.id), ...(st.add.q.includes("/") ? [st.add.q.trim()] : [])]; + for (const spec of want.slice(0, 24)) { + if (!st.add.open) return; + if (spec in st.add.counts && !st.add.counts[spec].error) continue; // a Hub error is asked again + try { st.add.counts[spec] = await server.hb_inspect(spec); } catch (_) { st.add.counts[spec] = { tasks: null, bytes: null }; } + renderAdd(); + } +} + +async function addDataset(spec) { + try { + const job = await server.hb_add(spec); + st.add.jobs[spec] = job; + if (job.state === "done") return finished(spec, job); + renderAdd(); renderFilters(); + follow(); + } catch (e) { st.add.jobs[spec] = { spec, state: "error", error: String(e.message || e) }; renderAdd(); } +} + +function follow() { + if (st.add.poll) return; + st.add.poll = setInterval(async () => { + const live = Object.values(st.add.jobs).filter((j) => !["done", "error"].includes(j.state)); + if (!live.length) { clearInterval(st.add.poll); st.add.poll = null; return; } + for (const j of live) { + try { + const now = await server.hb_add_status(j.spec); + st.add.jobs[j.spec] = now; + if (now.state === "done") finished(j.spec, now); + } catch (_) { /* keep polling */ } + } + renderAdd(); renderFilters(); + }, 700); +} + +async function finished(spec, job) { + const target = job.target || spec; // on a Space, the dataset's folder on the bucket mount + labels[target] = spec; + let d = st.datasets.find((x) => x.spec === target); + if (!d) { d = { spec: target, label: spec, num_tasks: job.tasks, added: true, removable: true }; st.datasets.push(d); } + if (job.tasks != null) st.add.counts[spec] = { ...(st.add.counts[spec] || {}), tasks: job.tasks }; + st.sel = { spec: new Set([target]) }; + st.shown = PAGE; + renderAdd(); + await load(d); +} + +async function removeDataset(spec) { + st.rm.busy = spec; st.rm.error = ""; renderFilters(); + try { + const res = await server.hb_remove(spec); + if (res.error) throw new Error(res.error); + st.datasets = st.datasets.filter((d) => d.spec !== spec); + st.rows = st.rows.filter((r) => r.spec !== spec); + if (st.sel.spec) st.sel.spec.delete(spec); + delete st.add.jobs[labels[spec] || spec]; + st.rm.confirm = null; + } catch (e) { st.rm.error = String(e.message || e); st.rm.confirm = spec; } + st.rm.busy = null; + render(); renderAdd(); +} + +// A `#task=:` link opens that task: on load, and when one is pasted into an open page +// (the page's own URL updates use replaceState, which fires no hashchange). +function openFromHash() { + const m = /#task=([^:]+):(\d+)/.exec(location.hash || ""); + if (!m) return; + const want = decodeURIComponent(m[1]); + if (st.datasets.some((d) => d.spec === want)) trigger("select", { spec: want, index: Number(m[2]) }); +} +window.addEventListener("hashchange", openFromHash); + +// ── start ─────────────────────────────────────────────────────────────── +$(".tb-search-ic").outerHTML = icon("search", 16); +$(".tb-random").innerHTML = `${icon("shuffle", 15)} Random`; +(async () => { + st.loading += 1; renderList(); + try { + const info = await server.hb_datasets(); + st.datasets = info.datasets.map((d) => ({ ...d })); st.canAdd = info.can_add; + for (const d of st.datasets) labels[d.spec] = d.label || d.spec; + st.loading -= 1; + openFromHash(); + if (!st.datasets.length) { + render(); + $(".tb-list").innerHTML = `
  • ${icon("database", 20)}

    No datasets

    Start the server with --dataset${st.canAdd ? ", or add one from the Hub under Filters" : ""}.

  • `; + return; + } + await Promise.all(st.datasets.filter((d) => !d.error).map(load)); + } catch (e) { + st.loading = 0; + $(".count").textContent = String(e.message || e); + } +})(); diff --git a/src/openenv/harbor/ui_assets/task_head.js b/src/openenv/harbor/ui_assets/task_head.js new file mode 100644 index 000000000..23722cbfa --- /dev/null +++ b/src/openenv/harbor/ui_assets/task_head.js @@ -0,0 +1,32 @@ +// Task page head: back to the task list, and a link to this task. +// The breadcrumb fires `select` with `{back: true}`; Python swaps the list back in. + +element.addEventListener("click", (ev) => { + if (ev.target.closest("[data-back]")) { + try { history.replaceState(null, "", location.pathname + location.search); } catch (_) { /* sandboxed frame */ } + return trigger("select", { back: true }); + } + const copy = ev.target.closest("[data-copy]"); + if (copy && navigator.clipboard) { + navigator.clipboard.writeText(link() || location.href).then(() => { + copy.innerHTML = `${icon("check", 14)}Copied`; + setTimeout(() => { copy.innerHTML = `${icon("link", 14)}Copy link`; }, 1600); + }); + } +}); + +// This task's own address, from the head itself: a task opened from a run page or the run list +// never went through the task list, which is what sets the hash. +function link() { + const head = element.querySelector(".tp-head[data-spec]"); + return head ? `${location.origin}${location.pathname}${location.search}#task=${encodeURIComponent(head.dataset.spec)}:${head.dataset.index}` : ""; +} + +// A newly opened task starts at its top, wherever the list was scrolled to, and the address bar +// follows it. +watch("value", () => { + const top = element.getBoundingClientRect().top; + if (top < 0) window.scrollBy({ top: top - 12 }); + const to = link(); + if (to && to !== location.href) { try { history.replaceState(null, "", to); } catch (_) { /* sandboxed frame */ } } +}); diff --git a/src/openenv/harbor/ui_assets/task_view.js b/src/openenv/harbor/ui_assets/task_view.js new file mode 100644 index 000000000..2ff6cb545 --- /dev/null +++ b/src/openenv/harbor/ui_assets/task_view.js @@ -0,0 +1,110 @@ +// Task page: the contents list, the file viewer, and links to this task's runs. +// +// The HTML is rendered in Python and replaced each time a task opens, so every handler is delegated +// from `element`, which persists. File contents are fetched on click with `server.hb_file`, never +// shipped up front: a task can hold thousands of files and a person reads three. + +const size = (n) => (n < 1024 ? `${n} B` : n < 1048576 ? `${(n / 1024).toFixed(0)} KB` : `${(n / 1048576).toFixed(1)} MB`); + +const files = {}; // path -> text, for Copy + +function lines(text) { + const rows = text.replace(/\n$/, "").split("\n"); + return rows.map((l) => `${esc(l)}`).join(""); +} + +async function openFile(path) { + const root = element.querySelector(".tp"); + const view = element.querySelector(".tp-code"), head = element.querySelector(".tp-view-h"); + if (!root || !view) return; + element.querySelectorAll("[data-file]").forEach((b) => b.setAttribute("aria-current", String(b.dataset.file === path))); + head.querySelector(".p").textContent = path; + head.querySelector("em").textContent = "loading…"; + view.dataset.path = path; + try { + const f = await server.hb_file([root.dataset.spec, Number(root.dataset.index), path]); + if (view.dataset.path !== path) return; // another file was picked meanwhile + if (f.error) throw new Error(f.error); + head.querySelector("em").textContent = `${size(f.size || 0)}${f.truncated ? " · first 256 KB" : ""}`; + files[path] = f.binary ? "" : f.text; + view.innerHTML = f.binary ? '
    Binary file, not shown.
    ' : lines(f.text); + } catch (e) { + head.querySelector("em").textContent = ""; + view.innerHTML = `
    ${esc(e.message || e)}
    `; + } + view.scrollTop = 0; view.scrollLeft = 0; +} + +function full(on) { + const box = element.querySelector(".tp-files"); + if (!box) return; + const now = on ?? !box.classList.contains("full"); + box.classList.toggle("full", now); + element.querySelectorAll("[data-full]").forEach((b) => { + if (b.closest(".tools")) { b.innerHTML = icon(now ? "minimize" : "maximize", 14); b.title = now ? "Close full view (Esc)" : "Full view (Esc to close)"; } + }); + let scrim = element.querySelector(".hb-scrim"); + if (now && !scrim) { scrim = document.createElement("div"); scrim.className = "hb-scrim"; box.before(scrim); } + if (!now && scrim) scrim.remove(); +} + +function go(sec) { + const target = element.querySelector(`[data-sec="${sec}"]`); + if (target) target.scrollIntoView({ behavior: "smooth", block: "start" }); +} + +// The contents list marks the last section whose top has scrolled past the top of the window. +function spyOn() { + const secs = [...element.querySelectorAll("[data-sec]")]; + const links = [...element.querySelectorAll(".tp-toc a")]; + if (!secs.length || !links.length || !element.offsetParent) return; + let on = secs[0].dataset.sec; + for (const s of secs) if (s.getBoundingClientRect().top < 140) on = s.dataset.sec; + links.forEach((a) => a.classList.toggle("on", a.dataset.to === on)); +} +let ticking = false; +window.addEventListener("scroll", () => { + if (ticking) return; + ticking = true; + requestAnimationFrame(() => { ticking = false; spyOn(); }); +}, { passive: true }); + +function bind() { + const first = element.querySelector(".tp-code[data-path]"); + if (first && first.dataset.path) openFile(first.dataset.path); + spyOn(); +} + +element.addEventListener("click", (ev) => { + const to = ev.target.closest("[data-to]"); + if (to) { ev.preventDefault(); return go(to.dataset.to); } + const file = ev.target.closest("[data-file]"); + if (file) return openFile(file.dataset.file); + const jump = ev.target.closest("[data-file-open]"); + if (jump) { go("files"); return openFile(jump.dataset.fileOpen); } + const run = ev.target.closest("[data-run]"); + if (run) return trigger("select", { run: run.dataset.run }); + if (ev.target.closest("[data-full]")) return full(); + if (ev.target.closest(".hb-scrim")) return full(false); + const wrap = ev.target.closest("[data-wrap]"); + if (wrap) { + const on = wrap.getAttribute("aria-pressed") !== "true"; + wrap.setAttribute("aria-pressed", String(on)); + element.querySelector(".tp-code")?.classList.toggle("wrap", on); + return; + } + const copy = ev.target.closest("[data-copy-file]"); + if (copy && navigator.clipboard) { + const path = element.querySelector(".tp-code")?.dataset.path; + navigator.clipboard.writeText(files[path] || "").then(() => { + copy.innerHTML = icon("check", 14); + setTimeout(() => { copy.innerHTML = icon("copy", 14); }, 1400); + }); + } +}); +document.addEventListener("keydown", (ev) => { + if (ev.key === "Escape" && element.querySelector(".tp-files.full")) full(false); +}); + +watch("value", bind); +bind(); diff --git a/src/openenv/harbor/ui_data.py b/src/openenv/harbor/ui_data.py new file mode 100644 index 000000000..9dab9d2d1 --- /dev/null +++ b/src/openenv/harbor/ui_data.py @@ -0,0 +1,1158 @@ +"""What the Harbor UI shows about a dataset: its tasks, what is in each one, and what a task runs with. + +Discovery (`tasks.py`) lists task directories and deliberately reads nothing inside them, because it +is on the latency path of every Task API call. A person choosing a task needs more than a directory +name: a title to find it by, the files to judge it by, and the settings it will run with. So this +module does read inside task directories, but only for the UI, on demand, and with a per-dataset +cache so a page load never walks a dataset twice. + +Nothing here boots a sandbox or calls a model. +""" + +from __future__ import annotations + +import os +import posixpath +import re +import threading +import tomllib +from concurrent.futures import ThreadPoolExecutor +from pathlib import Path +from typing import Any + +from .tasks import HarborTaskProvider, own_file as _own, resolve_task_dirs, task_root + +# Metadata keys that hold the answer. Harbor puts no rules on `[metadata]`, and some datasets keep +# the gold answer there (`gold_answer = "4"`). The raw `task.toml` stays viewable, the dataset is +# public anyway, but a summary someone skims before running a model should not hand them the answer. +_ANSWER_KEY = re.compile(r"(gold|answer|solution|expected|reference_output)", re.I) + +# The file tree is for reading a task, not for mirroring a repository: a task with a vendored +# codebase can hold tens of thousands of files, and the browser only needs the shape. +MAX_TREE_FILES = 2000 +MAX_FILE_BYTES = 256 * 1024 + +_ROWS: dict[str, list[dict[str, Any]]] = {} +_ROWS_INFLIGHT: dict[str, object] = {} +_ROWS_LOCK = threading.Lock() + + +def _hub_not_found(exc: Exception) -> bool: + """Whether a Hub exception specifically says the requested entry does not exist.""" + if isinstance(exc, FileNotFoundError): + return True + response = getattr(exc, "response", None) + return getattr(response, "status_code", None) == 404 + + +def _first_link(root: Path) -> Path | None: + """A symbolic link under `root`, file or folder, without following any; `None` if there is none.""" + folders = [root] + while folders: + with os.scandir(folders.pop()) as entries: + for entry in entries: + if entry.is_symlink(): + return Path(entry.path) + if entry.is_dir(follow_symlinks=False): + folders.append(Path(entry.path)) + return None + + +def _links_a_task(dataset: Path) -> bool: + """Whether `dataset/tasks`, or a task folder in it, is a link: what discovery would follow. + + One directory listing, so it is cheap enough to ask of every added dataset at startup; links + further down are never followed by the page, the file viewer or `reads_environment`. + """ + tasks = dataset / "tasks" + if tasks.is_symlink(): + return True + with os.scandir(tasks) as entries: + return any(entry.is_symlink() for entry in entries) + + +def _toml(task_dir: Path | None) -> dict[str, Any]: + try: + return tomllib.loads(_own(task_dir, "task.toml").read_text(errors="replace")) + except (OSError, tomllib.TOMLDecodeError): + return {} + + +def _first_line(task_dir: Path | None, limit: int = 160) -> str: + """The first meaningful line of the instruction, used as a title when a task declares none.""" + try: + with _own(task_dir, "instruction.md").open(errors="replace") as fh: + for line in fh: + text = line.strip().lstrip("#").strip() + if text: + return text[:limit] + except OSError: + pass + return "" + + +_MD = re.compile( + r"`+|\*\*|__|^\s*[-*+>]\s+|^\s*\d+[.)]\s+|!?\[([^\]]*)\]\([^)]*\)", re.M +) + + +def _paragraphs(task_dir: Path | None, limit: int = 3, chars: int = 360) -> list[str]: + """The instruction's first few prose paragraphs as plain text, headings and code left out. + + Only the start of the file is read: a brief is two lines on a card. + """ + try: + with _own(task_dir, "instruction.md").open(errors="replace") as fh: + head = fh.read(6000) + except OSError: + return [] + out, block, fenced = [], [], False + for line in head.splitlines() + [""]: + if line.strip().startswith("```"): + fenced = not fenced + continue + if fenced or line.lstrip().startswith("#"): + continue + if line.strip(): + block.append(line.strip()) + continue + if block: + text = _MD.sub(lambda m: m.group(1) or "", " ".join(block)).strip() + if text: + out.append(text[:chars]) + block = [] + if len(out) >= limit: + break + return out + + +def _strings(value: Any) -> list[str]: + if isinstance(value, str): + return [value] + if isinstance(value, (list, tuple)): + return [str(v) for v in value if isinstance(v, (str, int, float))] + return [] + + +def task_row(index: int, task_dir: Path, spec: str | None = None) -> dict[str, Any]: + """One task as the task list shows it. + + Harbor's schema fixes `[task]` (name, description, keywords) but leaves `[metadata]` free-form, + so a title, a category and a difficulty are each looked for under the names datasets actually + use, falling back to the instruction's first line for the title. + + Args: + index (`int`): + The task's position in its dataset, which is its identity everywhere downstream. + task_dir (`Path`): + The task directory. + spec (`str`, *optional*): + Its dataset, which every file read must stay inside (`task_root`). + + Returns: + `dict` with `index`, `name`, `title`, `category`, `difficulty` and `keywords`. + """ + root = task_root(spec, task_dir) + doc = _toml(root) + task, meta = doc.get("task") or {}, doc.get("metadata") or {} + title = ( + meta.get("title") + or task.get("description") + or _first_line(root) + or task_dir.name + ) + keywords: list[str] = [] + for k in _strings(task.get("keywords")) + _strings(meta.get("keywords")): + if k not in keywords and k != meta.get("category"): + keywords.append(k) + return { + "index": index, + "name": task_dir.name, + "title": str(title)[:240], + "category": str(meta.get("category") or meta.get("domain") or ""), + "difficulty": str(meta.get("difficulty") or meta.get("difficulty_tier") or ""), + "keywords": keywords[:8], + "paragraphs": _paragraphs(root), + } + + +def _briefs(rows: list[dict[str, Any]]) -> None: + """Give each row a `brief`: its first paragraph that says something about this task. + + Many datasets open every instruction with the same preamble ("You are an agent, your working + directory is /app"), which on a card reads as the same sentence 2,000 times. A paragraph shared by + more than a fifth of a dataset's tasks is that preamble, so it is skipped, as is one that only + repeats the title. + """ + counts: dict[str, int] = {} + for r in rows: + for p in set(r.get("paragraphs") or []): + counts[p] = counts.get(p, 0) + 1 + common = max(3, len(rows) // 5) + for r in rows: + title = r["title"].strip().rstrip(".").lower() + paras = r.pop("paragraphs", None) or [] + r["brief"] = next( + ( + p[:220] + for p in paras + if counts.get(p, 0) < common + and p.strip().rstrip(".").lower() != title + and not title.startswith(p.strip().rstrip(".").lower()[:60]) + ), + "", + ) + + +def task_rows(spec: str) -> list[dict[str, Any]]: + """Every task in a dataset as the task list shows it, cached per dataset. + + Reading one small `task.toml` per task is milliseconds on local disk but far slower on a + mounted bucket, so the reads run in parallel and the result is kept for the process lifetime. + + Args: + spec (`str`): + Dataset spec: HF dataset repo, local directory, or Harbor registry `name@version`. + + Returns: + `list[dict]`, one row per task, in index order. + """ + with _ROWS_LOCK: + if spec in _ROWS: + return _ROWS[spec] + token = object() + _ROWS_INFLIGHT[spec] = token + try: + dirs = resolve_task_dirs(spec) + workers = int(os.environ.get("OPENENV_UI_INDEX_WORKERS", "16")) + with ThreadPoolExecutor(max_workers=max(1, workers)) as pool: + rows = list( + pool.map(lambda pair: task_row(*pair, spec=spec), enumerate(dirs)) + ) + _briefs(rows) + except BaseException: + with _ROWS_LOCK: + if _ROWS_INFLIGHT.get(spec) is token: + _ROWS_INFLIGHT.pop(spec, None) + raise + with _ROWS_LOCK: + # Removal invalidates both caches while indexing can still be reading files. Do not let that + # in-flight result resurrect rows for a dataset whose backing folder has been deleted. + if _ROWS_INFLIGHT.get(spec) is token: + _ROWS[spec] = rows + _ROWS_INFLIGHT.pop(spec, None) + return rows + + +def common_tags(spec: str) -> set[str]: + """Tags nine in ten of a dataset's tasks carry (`terminal` on a terminal dataset): they tell no two + tasks apart. Empty until the dataset's rows are cached, and for datasets of three tasks or fewer.""" + with _ROWS_LOCK: + rows = _ROWS.get(spec) or [] + if len(rows) <= 3: + return set() + counts: dict[str, int] = {} + for r in rows: + for k in r.get("keywords") or []: + counts[k] = counts.get(k, 0) + 1 + return {k for k, n in counts.items() if n >= 0.9 * len(rows)} + + +def _config_sections(doc: dict[str, Any]) -> list[dict[str, Any]]: + """`task.toml` as the settings a person checks before running: where, with what, for how long.""" + env = doc.get("environment") or {} + agent = doc.get("agent") or {} + verifier = doc.get("verifier") or {} + + def secs(value: Any) -> str: + try: + s = float(value) + except (TypeError, ValueError): + return "" + return f"{s / 60:.0f} min" if s >= 120 else f"{s:.0f} s" + + network = env.get("network_mode") or ( + "internet" + if env.get("allow_internet") is True + else "none" + if env.get("allow_internet") is False + else "" + ) + image = env.get("docker_image") or "built from environment/Dockerfile" + sections = [ + ( + "Environment", + [ + ("Image", image), + ("Working directory", env.get("workdir", "")), + ("CPUs", env.get("cpus", "")), + ("Memory", f"{env['memory_mb']:,} MB" if env.get("memory_mb") else ""), + ( + "Storage", + f"{env['storage_mb']:,} MB" if env.get("storage_mb") else "", + ), + ("GPUs", env.get("gpus", "")), + ("Network", network), + ("Build timeout", secs(env.get("build_timeout_sec"))), + ], + ), + ( + "Agent", + [ + ("Timeout", secs(agent.get("timeout_sec"))), + ("User", agent.get("user", "")), + ("Network", (doc.get("agent") or {}).get("network_mode", "")), + ], + ), + ( + "Verifier", + [ + ("Timeout", secs(verifier.get("timeout_sec"))), + ("User", verifier.get("user", "")), + ("Network", verifier.get("network_mode", "")), + ], + ), + ] + out = [] + for label, rows in sections: + kept = [_row(k, v) for k, v in rows if v not in ("", None)] + if kept: + out.append({"label": label, "rows": kept}) + health = env.get("healthcheck") or {} + if health.get("command"): + out.append( + { + "label": "Setup", + "rows": [ + { + "key": "Before the agent", + "value": f"runs a setup command ({len(str(health['command'])):,} characters)" + + ( + f", up to {secs(health['timeout_sec'])}" + if health.get("timeout_sec") + else "" + ), + } + ], + } + ) + servers = env.get("mcp_servers") or [] + if servers: + out.append( + { + "label": "MCP servers", + "rows": [ + { + "key": str(s.get("name", "server")), + "value": str(s.get("transport") or s.get("url") or ""), + } + for s in servers + if isinstance(s, dict) + ], + } + ) + # Names only. Values are templates such as `${HF_TOKEN}` here, but a dataset could hard-code one, + # and the UI is the wrong place to find that out. + for label, table in ( + ("Environment variables", env.get("env")), + ("Verifier variables", verifier.get("env")), + ): + if isinstance(table, dict) and table: + out.append( + { + "label": label, + "rows": [{"key": k, "value": ""} for k in sorted(table)], + } + ) + return out + + +def _row(key: str, value: Any) -> dict[str, str]: + """A settings row. A digest-pinned image is shortened for reading; `full` keeps it for hover.""" + text = str(value) + if "@sha256:" in text: + name, _, digest = text.partition("@sha256:") + return {"key": key, "value": f"{name}@sha256:{digest[:12]}…", "full": text} + return {"key": key, "value": text} + + +def _metadata(doc: dict[str, Any]) -> tuple[list[dict[str, str]], list[str]]: + """Free-form `[metadata]` fields to show, and the answer-like ones withheld from the summary.""" + shown, hidden = [], [] + for key, value in (doc.get("metadata") or {}).items(): + if _ANSWER_KEY.search(key): + hidden.append(key) + continue + if isinstance(value, (list, tuple)): + value = ", ".join(str(v) for v in value) + if isinstance(value, dict): + continue + text = str(value) + if text: + shown.append({"key": key, "value": text[:300]}) + return shown, hidden + + +def file_tree(task_dir: Path) -> tuple[list[dict[str, Any]], bool]: + """Files under a task directory as `(path, size)`, capped so a vendored repo cannot flood the page. + + Returns: + `tuple` of the file list and whether it was cut at `MAX_TREE_FILES`. + """ + files: list[dict[str, Any]] = [] + root = task_dir.resolve() + # Files at the task's top level first (instruction.md, task.toml), then each folder in turn, + # walked lazily: a vendored tree of a million files is read only as far as the cap. + for folder, dirs, names in os.walk(root): + dirs[:] = sorted(d for d in dirs if not d.startswith(".")) + for name in sorted(n for n in names if not n.startswith(".")): + if len(files) >= MAX_TREE_FILES: + return files, True + path = Path(folder) / name + try: + # a symlink out of the task can't be opened (read_task_file); don't show its size either + if path.is_symlink() and not path.resolve().is_relative_to(root): + continue + size = path.stat().st_size + except OSError: + continue + files.append({"path": path.relative_to(root).as_posix(), "size": size}) + return files, False + + +def task_detail(spec: str, index: int) -> dict[str, Any]: + """Everything the task view shows for one task. Reads files, runs nothing. + + Args: + spec (`str`): + Dataset spec. + index (`int`): + Task index within the dataset. + + Returns: + `dict` with the row fields plus `description`, `instruction`, `config`, `metadata`, + `withheld` (answer-like metadata keys left out of the summary), `files` and + `files_truncated`. + """ + task_dir = HarborTaskProvider([spec]).task_dir(spec, int(index)) + root = task_root(spec, task_dir) + doc = _toml(root) + try: + instruction = _own(root, "instruction.md").read_text(errors="replace") + except OSError: + instruction = "" + shown, hidden = _metadata(doc) + files, truncated = file_tree(root) if root is not None else ([], False) + row = task_row(int(index), task_dir, spec) + row.pop("paragraphs", None) + return { + **row, + "dataset": spec, + "description": str((doc.get("task") or {}).get("description") or ""), + "instruction": instruction, + "config": _config_sections(doc), + "metadata": shown, + "withheld": hidden, + "files": files, + "files_truncated": truncated, + } + + +def reads_environment(spec: str, index: int) -> bool: + """Whether a task would read this server's environment variables or files. + + Harbor resolves `${VAR}` in `task.toml` (`[verifier.env]`, `[environment.env]`) from the process + environment, and Docker Compose interpolates `$VAR` in a compose file the same way; an + `environment:` or build `args:` entry that is a bare name passes the host's value straight + through. Either one hands the task this server's keys. A compose file can also reach past its own + folder: `env_file`, `include` and `extends` read other files, and a host path in a bind mount, a + secret or config `file:`, a build `context`, a build cache, a watch rule or a device puts the + host's files in the container, which a local container backend runs on this machine. So does + asking for more of the host than a folder: `privileged`, added capabilities, the host's + namespaces or Docker socket, another container's volumes, a named volume or network with + settings (the local driver binds any path), or the host's SSH agent during a build. + + Both formats are checked as parsed, since that is what Harbor and Compose act on: an escape such + as `"\\u0024{HF_TOKEN}"` has no `${` on disk. A file that doesn't parse can't be vouched for and + counts as reading. The text patterns stay as a second net. Every YAML file under `environment/` + is checked, since one compose file can include another. + """ + import yaml + + root = task_root(spec, HarborTaskProvider([spec]).task_dir(spec, int(index))) + if root is None: + return True # a link takes the task outside its dataset + env_dir = root / "environment" + toml_path = root / "task.toml" + # A link can put any file on this machine where the task reads its own. + if any(not p.resolve().is_relative_to(root) for p in (toml_path, env_dir)): + return True + if toml_path.is_file(): + try: + text = toml_path.read_text() + data = tomllib.loads(text) + except (OSError, UnicodeDecodeError, tomllib.TOMLDecodeError): + return True + if _HARBOR_VAR.search(text) or any("${" in s for s in _every_string(data)): + return True + for path in sorted(env_dir.rglob("*")) if env_dir.is_dir() else []: + if not path.resolve().is_relative_to(root): + return True + if path.suffix not in (".yml", ".yaml") or not path.is_file(): + continue + try: + if path.stat().st_size > _MAX_COMPOSE_BYTES: + return True + text = path.read_text() + docs = list(yaml.safe_load_all(text)) + except (OSError, UnicodeDecodeError, yaml.YAMLError): + return True + if any(p.search(text) for p in _COMPOSE_READS): + return True + seen: set[int] = set() + if any(_compose_reads(doc, seen) for doc in docs): + return True + return False + + +# Harbor expands only a whole value of `${VAR}` or `${VAR:-default}` (harbor.utils.env), so any `${` +# is flagged; Docker Compose also expands a bare `$VAR`. +_HARBOR_VAR = re.compile(r"\$\{") +_COMPOSE_VAR = re.compile(r"\$(\{|[A-Za-z_])") +_COMPOSE_READS = ( + _COMPOSE_VAR, + re.compile(r"^\s*(env_file|include|extends)\s*:", re.M), + # a bind mount from the host: `- /abs:/x`, `- ~/x:/x`, `- ../x:/x`, or `source: /abs` + re.compile(r"^\s*-\s*[\"']?(/|~|\.\.)[^:\n]*:", re.M), + re.compile(r"^\s*source\s*:\s*[\"']?(/|~|\.\.)", re.M), +) +_MAX_COMPOSE_BYTES = ( + 1_000_000 # compose files are small; a huge one is not worth parsing +) +# long-form mount `source`, secret/config `file`, build `context` and `dockerfile`, `develop.watch` path +_PATH_KEYS = {"source", "file", "context", "dockerfile", "path"} +# Settings that hand a container more than its own folder, whatever their value: extra privileges +# (enough to mount the host's disk on a local backend), other containers' volumes, the host's SSH +# agent during a build, a file of labels read from the host. A task that needs one runs only on a +# dataset the server was started with. +_HOST_ACCESS = { + "privileged", + "cap_add", + "security_opt", + "device_cgroup_rules", + "volumes_from", + "ssh", + "label_file", + # the local volume driver binds any host path with `{type: none, o: bind, device: /etc}` + "driver_opts", + # the Docker API socket (the whole machine), a plugin binary run on the host, a build granted + # the host's network or insecure mode + "use_api_socket", + "provider", + "entitlements", +} +# Namespaces a container can share with the host or another container: `pid: host` alone shows it +# every process's environment on the machine, this server's included. +_NAMESPACES = {"pid", "ipc", "network_mode", "userns_mode", "uts", "cgroup", "network"} + + +def _outside(path: str) -> bool: + """Whether a path in a compose file points outside its own folder: absolute, under `~`, a + Windows drive, or climbing out once normalised (`foo/../../etc` is `../etc`).""" + p = path.strip().replace("\\", "/") + if p.startswith(("/", "~")) or re.match(r"^[A-Za-z]:", p): + return True + n = posixpath.normpath(p) + return n == ".." or n.startswith("../") + + +def _every_string(node: Any): + """Every string in a parsed document, keys included.""" + if isinstance(node, str): + yield node + elif isinstance(node, dict): + for key, value in node.items(): + yield from _every_string(key) + yield from _every_string(value) + elif isinstance(node, (list, tuple)): + for value in node: + yield from _every_string(value) + + +def _passes_host_env(value: Any) -> bool: + """An `environment:` / `args:` entry that takes its value from the host. + + A bare name in the list form (`- HF_TOKEN`) or an empty value in the map form (`HF_TOKEN:`) + copies the host's variable in; a secret's or config's `environment: NAME` is its value. + """ + if isinstance(value, str): + return True + if isinstance(value, list): + return any(not isinstance(e, str) or "=" not in e for e in value) + if isinstance(value, dict): + return any(v is None for v in value.values()) + return False + + +def _compose_reads(node: Any, seen: set[int]) -> bool: + """Whether a parsed compose document reads the host's variables or files. + + YAML aliases share nodes, so each container is walked once: an alias bomb that is tiny on disk + would otherwise take exponential time to walk. + """ + if isinstance(node, str): + return "$" in node + if not isinstance(node, (dict, list)) or id(node) in seen: + return False + seen.add(id(node)) + if isinstance(node, list): + return any(_compose_reads(value, seen) for value in node) + for key, value in node.items(): + key = str(key) + if key in ("env_file", "include", "extends", "devices"): + return True + if key in ("environment", "args") and _passes_host_env(value): + return True + if key in _HOST_ACCESS and value not in (None, False, "", [], {}): + return True + if key == "external" and value: + return True + if ( + key in _NAMESPACES + and isinstance(value, str) + and ( + value.strip().lower() == "host" + or value.strip().startswith("container:") + ) + ): + return True + # A named volume is a map at the top level (a service lists its mounts), and one with any + # setting can bind a host path or reuse a volume that already exists on the machine. + if key == "volumes" and isinstance(value, dict): + if any( + isinstance(c, dict) and set(map(str, c)) - {"labels"} + for c in value.values() + ): + return True + # A network named for one that exists (`host`, another project's), or on a driver that + # attaches to the host's interfaces (`macvlan`, `ipvlan`, `host`). + if key == "networks" and isinstance(value, dict): + for c in value.values(): + if isinstance(c, dict) and ( + "name" in c + or str(c.get("driver", "bridge")) not in ("bridge", "overlay") + ): + return True + # a local build cache reads (`cache_from`) or writes (`cache_to`) a host folder + if key in ("cache_from", "cache_to") and any( + "type=local" in v.replace(" ", "") + for v in ([value] if isinstance(value, str) else value or []) + if isinstance(v, str) + ): + return True + if ( + key == "volumes" + and isinstance(value, list) + and any( + isinstance(e, str) and ":" in e and _outside(e.split(":", 1)[0]) + for e in value + ) + ): + return True + if key in _PATH_KEYS and isinstance(value, str) and _outside(value): + return True + # the map form is `name: path`, the list form `- name=path` + if key == "additional_contexts" and any( + isinstance(v, str) + and _outside(v.split("=", 1)[-1] if isinstance(value, list) else v) + for v in (value.values() if isinstance(value, dict) else value or []) + ): + return True + if _compose_reads(key, seen) or _compose_reads(value, seen): + return True + return False + + +def read_task_file(spec: str, index: int, path: str) -> dict[str, Any]: + """One file from a task directory, for the file viewer. + + The path comes from the browser, so it is resolved and checked to stay inside the task + directory: `..` segments and symlinks pointing elsewhere are refused. + + Args: + spec (`str`): + Dataset spec. + index (`int`): + Task index. + path (`str`): + Path relative to the task directory. + + Returns: + `dict` with `path`, `size`, and either `text` (possibly `truncated`) or `binary: True`, or + `error` when the path is refused or missing. + """ + root = task_root(spec, HarborTaskProvider([spec]).task_dir(spec, int(index))) + target = (root / str(path)).resolve() if root is not None else None + if target is None or not target.is_relative_to(root) or not target.is_file(): + return {"path": path, "error": "not a file in this task"} + size = target.stat().st_size + with target.open("rb") as fh: + raw = fh.read(MAX_FILE_BYTES) + if b"\x00" in raw[:8192]: + return {"path": path, "size": size, "binary": True} + return { + "path": path, + "size": size, + "text": raw.decode("utf-8", errors="replace"), + "truncated": size > MAX_FILE_BYTES, + } + + +def search_hub(query: str = "", limit: int = 40) -> list[dict[str, Any]]: + """Harbor datasets on the Hugging Face Hub, most downloaded first. + + Harbor datasets carry the `harbor` tag (and usually `rl-environment`), which is what makes them + findable without a registry of our own. + + Args: + query (`str`, *optional*, defaults to `""`): + Free-text filter on the dataset id. + limit (`int`, *optional*, defaults to `40`): + Maximum number of results. + + Returns: + `list[dict]` with `id`, `downloads`, `likes` and whether it is tagged `rl-environment`. + """ + from huggingface_hub import HfApi + + # Anonymous: the server's own token would list its private datasets to any visitor. + results = HfApi(token=False).list_datasets( + filter="harbor", + search=(query or "").strip() or None, + sort="downloads", + limit=max(1, min(int(limit), 100)), + ) + return [ + { + "id": d.id, + "downloads": d.downloads or 0, + "likes": d.likes or 0, + "rl_environment": "rl-environment" in (d.tags or []), + } + for d in results + ] + + +HF_ROUTER = "https://router.huggingface.co/v1" +_MODELS: tuple[float, list[dict[str, Any]]] = (0.0, []) +_MODELS_LOCK = threading.Lock() + + +def hf_models(ttl: float = 600.0) -> list[dict[str, Any]]: + """Chat models on Hugging Face Inference Providers, for the model picker. + + The router's model list is public and says, per provider, whether it is live, whether it calls + tools, its context length and its price. An agent is useless without tool calls, so each model + carries `tools` (any live provider supports them) and the picker offers those first. Cached for + `ttl` seconds: it is the same list for every visitor. + + Returns: + `list[dict]` with `id`, `tools`, `context` (the largest across providers) and `providers`, + each `{name, tools, context, price_in, price_out}` in USD per million tokens. + """ + import json + import time + import urllib.request + + global _MODELS + with _MODELS_LOCK: + fetched, cached = _MODELS + if cached and time.time() - fetched < ttl: + return cached + with urllib.request.urlopen(f"{HF_ROUTER}/models", timeout=20) as r: + data = json.loads(r.read()).get("data") or [] + out = [] + for m in data: + providers = [] + for p in m.get("providers") or []: + if p.get("status", "live") != "live": + continue + price = p.get("pricing") or {} + providers.append( + { + "name": p.get("provider", ""), + "tools": bool(p.get("supports_tools")), + "context": p.get("context_length") or 0, + "price_in": price.get("input"), + "price_out": price.get("output"), + } + ) + if not providers or "text" not in ( + (m.get("architecture") or {}).get("output_modalities") or ["text"] + ): + continue + out.append( + { + "id": m.get("id", ""), + "tools": any(p["tools"] for p in providers), + "context": max((p["context"] for p in providers), default=0), + "providers": providers, + } + ) + # Tool-calling models first, then by how widely served, which is a fair proxy for how well used. + out.sort(key=lambda m: (not m["tools"], -len(m["providers"]))) + with _MODELS_LOCK: + _MODELS = (time.time(), out) + return out + + +# ── adding a Hub dataset from the page ────────────────────────────────────────────────────────── +# A download of a task suite is thousands of small files and can take minutes, so an add runs in the +# background and the page polls it: first a check that the dataset is laid out the way Harbor reads +# it (before anything is downloaded), then the download with its file count, then reading the tasks. + +_JOBS: dict[str, dict[str, Any]] = {} +_JOBS_LOCK = threading.Lock() +MAX_ADDS_AT_ONCE = 2 + + +def _max_add_bytes() -> int: + try: + return int(float(os.environ.get("OPENENV_HARBOR_UI_MAX_ADD_GB") or 5) * 1e9) + except ValueError: + return int(5e9) + + +def hub_summary(spec: str) -> dict[str, Any]: + """What adding a Hub dataset would bring: `tasks` held as `tasks//`, the layout Harbor + loads (`None` when it has none), and the repository's size in `bytes`. + + Read from the Hub's listing and metadata, anonymously, so nothing is downloaded to find out. The + size is the whole repository's, an upper bound on what the `tasks/` download takes. + """ + from huggingface_hub import HfApi + from huggingface_hub.hf_api import RepoFolder + + api = HfApi(token=False) + tasks = None + try: + folders = [ + e.path + for e in api.list_repo_tree( + spec, repo_type="dataset", path_in_repo="tasks", recursive=False + ) + if isinstance(e, RepoFolder) + ] + if folders: + first = { + e.path.rsplit("/", 1)[-1] + for e in api.list_repo_tree( + spec, repo_type="dataset", path_in_repo=folders[0], recursive=False + ) + } + if "task.toml" in first: + tasks = len(folders) + else: + # Grouped (`tasks///`), which the loader searches the same way: count + # the task files. Bounded, so a vast repository cannot stall the page. + tasks = 0 + for i, e in enumerate( + api.list_repo_tree( + spec, repo_type="dataset", path_in_repo="tasks", recursive=True + ) + ): + tasks += e.path.endswith("/task.toml") + if i > 50_000: + break + tasks = tasks or None + except Exception as exc: # noqa: BLE001 - Hub exception types vary across supported clients + if not _hub_not_found(exc): + raise + return {"tasks": tasks, "bytes": api.dataset_info(spec).used_storage} + + +def _tasks_bytes(spec: str, cap: int, max_files: int = 500_000) -> int | None: + """The size of a Hub dataset's `tasks/` folder from its file listing, or `None` if unknown. + + Stops counting once the total passes `cap`, since that already settles the question; a listing + longer than `max_files` is not read to the end, and counts as unknown. + """ + from huggingface_hub import HfApi + from huggingface_hub.hf_api import RepoFile + + total = 0 + try: + for n, entry in enumerate( + HfApi(token=False).list_repo_tree( + spec, repo_type="dataset", path_in_repo="tasks", recursive=True + ) + ): + if isinstance(entry, RepoFile): + total += entry.size + if total > cap: + return total + if n >= max_files: + return None + except Exception: # noqa: BLE001 - no listing, no size + return None + return total + + +def _progress(job: dict[str, Any]) -> Any: + """A silent tqdm that records the Hub download's file count in `job`. + + The Hub client uses the same class for each large file's byte counter (`unit="B"`), so only the + bar that counts files is recorded. + """ + import io + + from tqdm import tqdm + + class Progress(tqdm): + def __init__(self, *args: Any, **kwargs: Any) -> None: + kwargs["file"] = io.StringIO() + kwargs.pop("disable", None) + self._files = kwargs.get("unit", "it") != "B" + super().__init__(*args, **kwargs) + if self._files and self.total: + job["total"] = int(self.total) + + def update(self, n: float | None = 1) -> bool | None: + out = super().update(n) + if self._files: + job["done"] = int(self.n) + return out + + return Progress + + +def added_spec(spec: str, settings: Any) -> str: + """What the loader is given for a Hub dataset added from the page. + + With the Space's bucket mounted, the dataset's folder on the mount, the same place `push` puts + the datasets it serves; otherwise the Hub id, which downloads to this machine's cache. + """ + if settings.bucket and settings.bucket_mount: + return str(settings.bucket_mount / spec.replace("/", "__")) + return spec + + +def hub_id(spec: str) -> str: + """The Hub id a dataset spec stands for: `org/name` for a `/data/org__name` mount folder.""" + name = Path(spec).name if spec.startswith("/") else "" + return name.replace("__", "/", 1) if "__" in name else spec + + +def added_in_bucket(settings: Any, served: list[str]) -> list[str]: + """Datasets on the mounted bucket that the server was not started with: the ones added from the + page. The bucket itself is the list, so it survives restarts and needs no file of its own.""" + mount = settings.bucket_mount + if not (settings.bucket and mount): + return [] + out = [] + for p in sorted(mount.iterdir()): + if ( + p.is_dir() + and "__" in p.name + and not p.name.startswith(".") + and (p / "tasks").is_dir() + ): + # Task folders that are links stay out, e.g. a refused add a failed cleanup left behind. + if ( + str(p) not in served + and hub_id(str(p)) not in served + and not _links_a_task(p) + ): + out.append(str(p)) + return out + + +def _forget(spec: str) -> None: + from . import tasks + + with tasks._LOCK: + tasks._CACHE.pop(spec, None) + with _ROWS_LOCK: + _ROWS.pop(spec, None) + _ROWS_INFLIGHT.pop(spec, None) + + +def remove_added(spec: str, settings: Any) -> None: + """Delete a dataset added from the page: its folder in the bucket, or its local download. + + Only ever a folder this module put there: one under the bucket mount named `org__name`, or one in + the dataset cache. The caller checks that `spec` was added from the page, not served. + """ + from . import tasks + + if spec.startswith("/") and settings.bucket and settings.bucket_mount: + folder = Path(spec).resolve() + if folder.parent != settings.bucket_mount.resolve() or "__" not in folder.name: + raise ValueError("not a dataset folder on the bucket") + from huggingface_hub import BucketFile, HfApi + + api = HfApi() + paths = [ + f.path + for f in api.list_bucket_tree( + settings.bucket, prefix=f"{folder.name}/", recursive=True + ) + if isinstance(f, BucketFile) + ] + for i in range(0, len(paths), 1000): + api.batch_bucket_files(settings.bucket, delete=paths[i : i + 1000]) + else: + root = tasks._DATASET_ROOT.resolve() + folder = (tasks._DATASET_ROOT / spec.replace("/", "__")).resolve() + if folder.parent != root: + raise ValueError("not a downloaded dataset") + if folder.is_dir(): + import shutil + + shutil.rmtree(folder) + _forget(spec) + + +def _copy_to_bucket( + spec: str, settings: Any, job: dict[str, Any], expected: int +) -> str: + """Copy a Hub dataset into the Space's bucket, server side, and wait for the mount to show it. + + The same copy `push` makes for the datasets it serves: by content hash, so nothing is downloaded + or uploaded and a suite of thousands of files takes seconds. Only `tasks/` is copied: it is all + the loader reads, the same folder a download fetches, and what the size check measured. The + mount then shows the new folder, which is waited for rather than assumed. + """ + import time + + from huggingface_hub import HfApi + + prefix = spec.replace("/", "__") + job["state"] = "copying" + HfApi().copy_files( + f"hf://datasets/{spec}/tasks/", + f"hf://buckets/{settings.bucket}/{prefix}/tasks/", + ) + job["state"] = "mounting" + folder = settings.bucket_mount / prefix + deadline = time.time() + 180 + from .tasks import _task_dirs_from_directory + + while time.time() < deadline: + try: + if (folder / "tasks").is_dir() and len( + _task_dirs_from_directory(folder) + ) >= expected: + return str(folder) + except OSError: + pass + time.sleep(3) + raise ValueError( + "Copied to the bucket, but the mount has not shown it yet. It appears when the Space restarts." + ) + + +def start_add(spec: str, on_added: Any, settings: Any = None) -> dict[str, Any]: + """Start adding `spec` in the background, or return the add already under way. + + `on_added(loader_spec)` is called once its tasks are read. The caller has checked that `spec` is a + public Hub dataset id; this checks its layout and size, puts it where `added_spec` says (the + bucket, or this machine), and indexes it. + """ + import time + + with _JOBS_LOCK: + job = _JOBS.get(spec) + if job and job["state"] not in ("error", "done"): + return dict(job) + busy = sum(1 for j in _JOBS.values() if j["state"] not in ("error", "done")) + if busy >= MAX_ADDS_AT_ONCE: + return { + "spec": spec, + "state": "error", + "error": "Another dataset is being added. Try again in a moment.", + } + job = { + "spec": spec, + "state": "checking", + "done": 0, + "total": 0, + "tasks": None, + "error": None, + "started": time.time(), + } + _JOBS[spec] = job + + def run() -> None: + try: + summary = hub_summary(spec) + expected = summary["tasks"] + if not expected: + raise ValueError( + "This dataset has no tasks// folder, which is how Harbor datasets are laid out." + ) + if not summary["bytes"]: + # The Hub has no size for some repositories, and reports 0 for one it has not + # measured yet (a fresh upload): measure what the download would take. + summary["bytes"] = _tasks_bytes(spec, _max_add_bytes()) + if summary["bytes"] is None: + raise ValueError( + "Couldn't tell how big this dataset is, so it isn't added from the page." + ) + if summary["bytes"] > _max_add_bytes(): + raise ValueError( + f"This dataset is {summary['bytes'] / 1e9:.1f} GB, over the " + f"{_max_add_bytes() / 1e9:.0f} GB this server adds from the page " + "(OPENENV_HARBOR_UI_MAX_ADD_GB)." + ) + job["bytes"] = summary["bytes"] + job["expected"] = expected + if settings is not None and settings.bucket and settings.bucket_mount: + target = _copy_to_bucket(spec, settings, job, expected) + else: + job["state"] = "downloading" + resolve_task_dirs(spec, tqdm_class=_progress(job)) + target = spec + # A link in a dataset from the Hub can point a task, or the files the page shows by + # itself, at anything on this machine. Refused, and removed so a restart, which lists + # the bucket's folders, does not bring it back. + from . import tasks as _tasks + + root = ( + Path(target) + if target.startswith("/") + else _tasks._DATASET_ROOT / target.replace("/", "__") + ) + link = _first_link(root) + if link is not None: + refusal = ( + f"This dataset contains a symbolic link ({link.relative_to(root)}), which " + "a dataset added from the page may not." + ) + try: + remove_added(target, settings or _NO_BUCKET) + except Exception as exc: # noqa: BLE001 - the refusal stands either way + # What was copied stays. A restart skips it if its task folders are links + # (`added_in_bucket`), and no link further down is ever followed. + refusal += f" Removing the copy failed ({type(exc).__name__})." + raise ValueError(refusal) + job["state"] = "indexing" + job["tasks"] = len(task_rows(target)) + job["target"] = target + on_added(target) + job["state"] = "done" + except Exception as exc: # noqa: BLE001 - reported to the page, never raised + job.update(state="error", error=f"{str(exc)[:240]}") + + threading.Thread(target=run, daemon=True, name=f"harbor-add-{spec}").start() + return dict(job) + + +# `remove_added`'s settings for a server with no bucket: the dataset is a local download +_NO_BUCKET = type("NoBucket", (), {"bucket": None, "bucket_mount": None})() + + +def add_status(spec: str) -> dict[str, Any] | None: + with _JOBS_LOCK: + job = _JOBS.get(spec) + return dict(job) if job else None diff --git a/src/openenv/harbor/ui_icons.py b/src/openenv/harbor/ui_icons.py new file mode 100644 index 000000000..7c3d3ea28 --- /dev/null +++ b/src/openenv/harbor/ui_icons.py @@ -0,0 +1,54 @@ +"""One small icon set for the Harbor UI (24px grid, 1.75 stroke, Lucide-style paths), shared by the +HTML rendered here and the components' JavaScript, so the page never falls back to emoji or glyphs.""" + +from __future__ import annotations + +import html +import json +from functools import lru_cache +from importlib import resources + + +@lru_cache(maxsize=1) +def paths() -> dict[str, str]: + return json.loads( + resources.files("openenv.harbor") + .joinpath("ui_assets", "icons.json") + .read_text() + ) + + +def icon(name: str, size: int = 16, cls: str = "") -> str: + return ( + f'' + ) + + +def json_tag() -> str: + """The icon set, once per page, for every component's `icon()` to read. + + An attribute rather than a `