Repository navigation
Show truthful Standard scan progress and completed review activity #70
Description
Activity
hujm2023 commented
on Jul 29, 2026 AuthorMore actionsUpdate: the scan completed successfully at
44:13.Final summary:
[44:13] Scan complete codex-security: Findings: 37 (15 high, 22 medium). Coverage: partial. codex-security: Elapsed: 2641s. codex-security: Tokens: 135,369,191 input, 129,318,400 cached, 480,582 output. codex-security: Estimated cost: $109.330615 USD.The manifest and turn are both marked
completed; exit code2was expected because coverage waspartial. All documented result artifacts (scan-manifest.json,findings.json,coverage.json,report.md, and SARIF) were generated.The observability concern remains: from
[01:00]until completion at[44:13], the only CLI output was the generic once-per-secondRunning scanheartbeat. There were no phase transitions, worker progress updates, cost updates, or other signals that distinguished this healthy 43-minute in-flight request from a stalled scan.Related observability/cost work has several separate scopes: #63 adds verbose scan diagnostics; #29 asks for ChatGPT allowance visibility; #187 requests rollout-session retention; and #31 / #198 address overlapping cost-polling work. None of those changes alone establishes that a long-running scan is making useful forward progress, so they should be linked without treating this request as a duplicate or already fixed.
- addedarea:cliCommand-line help, arguments, terminal UI, and user interaction.Command-line help, arguments, terminal UI, and user interaction.area:costToken usage, cost estimates, spending limits, and quota visibility.Token usage, cost estimates, spending limits, and quota visibility.bugSomething isn't workingSomething isn't workingpriority:p2Significant defect or product gap with narrower impact or a workaroundSignificant defect or product gap with narrower impact or a workaround
on Aug 15, 2026 mldangelo-oai commented
on Aug 15, 2026 CollaboratorMore actionsThe live activity dashboard landed in #261, but Standard scans can still display an unverified
0 / N reviewedcount. Keeping this open for accurate phase, activity, and completed-review reporting.- changed the title
[-]Standard full-repository scan shows only generic heartbeat for 40+ minutes — expected behavior?[/-][+]Show truthful Standard scan progress and completed review activity[/+]on Aug 15, 2026 Additional reproduction: Standard scan remains in preflight until cost limit
We hit a stronger version of this failure with a very small path-scoped Standard scan: the scan never published a preflight check or review-item completion and stopped only when
--max-costaborted it.Environment
@openai/codex-security@0.1.15- bundled plugin
0.1.22 - bundled Codex runtime
0.148.0-alpha.8 - model
gpt-5.6-sol, reasoning effortxhigh - ChatGPT device authentication
- Node 22 / Python 3
- Linux
arm64under OrbStack's Docker engine on Apple Silicon - non-root uid 1000
- repository scope reduced to a directory containing 2 files
Observed
Across Standard runs:
scans.phaseremainedpreflightfor the whole run;scan_progress.review_items_total = 0andreview_items_completed = 0;preflight_checks_total = 0andpreflight_issues_json = [];- terminal progress nevertheless displayed
reviewing files (...) | Files: 0/2; - usage grew at roughly 1M mostly-cached input tokens/minute while output increased only about 200 tokens/minute;
CODEX_SECURITY_LOG_LEVEL=debug --verboseemitted onlycost.updatedevents, with no shell, helper, sandbox, or source-read activity;- the run ended only at the configured cost limit, whose message became
scans.failure_message.
Sandbox controls
An earlier run had a real Docker seccomp failure preventing Bubblewrap user-namespace creation. That is separate and was eliminated before the reproduction above.
Using the exact post-seccomp container flags and runtime user, a quota-free probe through the bundled Codex sandbox wrapper:
- ran
_bundled_plugin/scripts/config_preflight.pysuccessfully; - returned
status: ready; - read files from the mounted scan workspace.
The same sandbox behavior was reproduced on Linux
x86_64, weakening an ARM64-runtime explanation.--read-onlyandno-new-privilegesdid not affect the helper result.Interpretation
The evidence is consistent with a model/orchestration loop before the Standard-scan agent invokes or publishes capability preflight, rather than a long source-review turn or a Bubblewrap failure. Because the saved session rollout was deleted during credential cleanup, we cannot provide the raw event stream; the workbench counters and verbose-output summary above were captured before cleanup.
Expected behavior
If the agent has not invoked or published preflight after bounded turns/time, stop with a diagnostic before further model usage. At minimum, distinguish an active model request from repeated no-progress turns and avoid displaying
reviewing filesbefore preflight has completed.0.1.16was released after this reproduction; its release notes do not mention preflight or progress-loop changes. We plan one bounded seeded-control retry on that version and can report whether the behavior persists.Follow-up with the bounded
0.1.16seeded-control retry:- Environment: isolated OrbStack Ubuntu machine, Linux arm64, Node 22.14.0. Only a sanitized clone, scan brief, and public knowledge-base files were present.
- CLI
0.1.16; bundled plugin remains0.1.22; Codex remains0.148.0-alpha.8;gpt-5.6-sol/ xhigh. - Exact scope: three files containing three intentional security controls (tenant authorization bypasses and unsafe privileged billing reconciliation).
- Dry run completed immediately and resolved all three files.
- The paid run emitted
worker.preflight delegation="available" configured_slots=8at 28s, then the terminal switched toreviewing filesat 55s. - It hit the
$3limiter after 4m28s: estimated$3.048526, 2,361,440 input tokens, 2,099,712 cached input tokens, and 23,001 output tokens. - Terminal remained
Files: 0/3. - The sealed workbench still recorded
phase=preflight,review_items_total=0,review_items_completed=0,preflight_checks_total=0,finding_occurrences=0, and no registered artifacts. - The only output-dir file was
scoped-source-input.jsonl.
Important correction to the initial interpretation: the Codex rollout logs prove that work was occurring behind the zero counters. The run created one coordinator plus five worker sessions and made 53 sandboxed
execcalls. Workers successfully read the three scoped source files, repository context, migrations, security guidance,security-scan/SKILL.md,scan-prologue.md,config-preflight.md, andcore-scan.md. No bubblewrap, namespace, permission, or read-only-filesystem failure occurred. One initial relative-path read used the output directory and failed with file-not-found; the coordinator immediately retried against$CODEX_SECURITY_REPOSITORYsuccessfully.Worker creation was also highly staggered: coordinator at 18:02:05, workers at 18:04:04, 18:04:27, 18:05:35, 18:05:58, and 18:06:22. The final file-specific worker had only seconds before the cost limiter canceled the run.
So this no longer looks like an ARM64/bubblewrap execution stall. It looks like an observability and cost-bounding problem: active pre-publication worker work is represented as zero review items (and SQLite stays in preflight even after the terminal claims reviewing), while the cost cap can cancel all workers before the first result is committed. A cheap no-progress watchdog still seems valuable, but progress should include worker/tool activity; exposing active/completed worker counts and persisting the actual phase would make the CLI truthful.
- removedarea:costToken usage, cost estimates, spending limits, and quota visibility.Token usage, cost estimates, spending limits, and quota visibility.
on Oct 8, 2026
Summary
A standard full-repository scan has been running for more than 43 minutes while the CLI prints only a generic
Running scanheartbeat once per second. There are no phase changes, completed/planned worker counts, file coverage, estimated cost, or indication of whether the active model request is making progress.The related processes are still alive, so I cannot tell whether this is an expected long-running request or a stalled scan. Is this duration and lack of phase-level progress expected for a standard scan? If so, could the CLI expose enough progress to distinguish a healthy long request from a hang?
The repository identity and contents are intentionally omitted. It is a large private, multi-service Go monorepo (about 9.4 GB on disk).
Environment
0.1.1v26.4.012.0.13.14.6arm64)--auth chatgpt)standardgpt-5.6-solxhighReproduction
The preceding dry-run succeeded:
The actual scan then reported:
No other phase or progress messages appeared between those lines. At 43 minutes, the CLI process, its
codex execchild, andcodex-code-mode-hostwere all still alive.Questions / expected behavior
model request in progress, last progress timestamp, phase, and completed/planned workers?This overlaps with #29 on long-running scan observability, but this report is specifically about the inability to distinguish healthy progress from a stalled standard scan. #53 also shows that scans can run much longer while exposing phase and worker transitions; those transitions are absent here.