Skip to content

Show truthful Standard scan progress and completed review activity #70

Description

@hujm2023

Summary

A standard full-repository scan has been running for more than 43 minutes while the CLI prints only a generic Running scan heartbeat once per second. There are no phase changes, completed/planned worker counts, file coverage, estimated cost, or indication of whether the active model request is making progress.

The related processes are still alive, so I cannot tell whether this is an expected long-running request or a stalled scan. Is this duration and lack of phase-level progress expected for a standard scan? If so, could the CLI expose enough progress to distinguish a healthy long request from a hang?

The repository identity and contents are intentionally omitted. It is a large private, multi-service Go monorepo (about 9.4 GB on disk).

Environment

  • Codex Security CLI: 0.1.1
  • Node.js: v26.4.0
  • npm: 12.0.1
  • Python: 3.14.6
  • OS: macOS 26.5.2 (25F84), Apple Silicon (arm64)
  • Authentication: stored ChatGPT/Codex credentials (--auth chatgpt)
  • Scan mode: standard
  • Default model reported by dry-run: gpt-5.6-sol
  • Default reasoning effort reported by dry-run: xhigh
  • Preflight reported delegation support for up to 6 worker slots

Reproduction

codex-security scan /path/to/private-repository \
  --output-dir /path/outside/repository/results \
  --auth chatgpt

The preceding dry-run succeeded:

[00:00] Validating scan inputs
[00:00] Preflight complete
dryRun: true
mode: standard
authentication:
  method: stored_credentials
  verified: false
model: gpt-5.6-sol
reasoningEffort: xhigh

The actual scan then reported:

[00:05] Authentication: stored Codex credentials.
[00:14] Running scan
[01:00] Preflight: worker delegation supported (up to 6 worker slots).
[43:00] Running scan

No other phase or progress messages appeared between those lines. At 43 minutes, the CLI process, its codex exec child, and codex-code-mode-host were all still alive.

Questions / expected behavior

  1. Is a 40+ minute in-flight period with no phase updates expected for a standard scan of a large repository?
  2. Is there a supported way to inspect current phase, worker completion, files reviewed, or model-request health while the scan runs?
  3. Should the CLI periodically report model request in progress, last progress timestamp, phase, and completed/planned workers?
  4. Should there be a configurable no-progress timeout or warning distinct from a total scan timeout?

This overlaps with #29 on long-running scan observability, but this report is specifically about the inability to distinguish healthy progress from a stalled standard scan. #53 also shows that scans can run much longer while exposing phase and worker transitions; those transitions are absent here.

Activity

  1. hujm2023 commented on Jul 29, 2026

    @hujm2023
    Author

    Update: the scan completed successfully at 44:13.

    Final summary:

    [44:13] Scan complete
    codex-security: Findings: 37 (15 high, 22 medium). Coverage: partial.
    codex-security: Elapsed: 2641s.
    codex-security: Tokens: 135,369,191 input, 129,318,400 cached, 480,582 output.
    codex-security: Estimated cost: $109.330615 USD.
    

    The manifest and turn are both marked completed; exit code 2 was expected because coverage was partial. All documented result artifacts (scan-manifest.json, findings.json, coverage.json, report.md, and SARIF) were generated.

    The observability concern remains: from [01:00] until completion at [44:13], the only CLI output was the generic once-per-second Running scan heartbeat. There were no phase transitions, worker progress updates, cost updates, or other signals that distinguished this healthy 43-minute in-flight request from a stalled scan.

  2. mldangelo-oai commented on Aug 3, 2026

    @mldangelo-oai
    Collaborator

    Related observability/cost work has several separate scopes: #63 adds verbose scan diagnostics; #29 asks for ChatGPT allowance visibility; #187 requests rollout-session retention; and #31 / #198 address overlapping cost-polling work. None of those changes alone establishes that a long-running scan is making useful forward progress, so they should be linked without treating this request as a duplicate or already fixed.

  3. added
    area:cliCommand-line help, arguments, terminal UI, and user interaction.
    area:costToken usage, cost estimates, spending limits, and quota visibility.
    bugSomething isn't working
    priority:p2Significant defect or product gap with narrower impact or a workaround
    on Aug 15, 2026
  4. mldangelo-oai commented on Aug 15, 2026

    @mldangelo-oai
    Collaborator

    The live activity dashboard landed in #261, but Standard scans can still display an unverified 0 / N reviewed count. Keeping this open for accurate phase, activity, and completed-review reporting.

  5. changed the title [-]Standard full-repository scan shows only generic heartbeat for 40+ minutes — expected behavior?[/-] [+]Show truthful Standard scan progress and completed review activity[/+] on Aug 15, 2026
  6. blodebole commented on Aug 20, 2026

    @blodebole

    Additional reproduction: Standard scan remains in preflight until cost limit

    We hit a stronger version of this failure with a very small path-scoped Standard scan: the scan never published a preflight check or review-item completion and stopped only when --max-cost aborted it.

    Environment

    • @openai/codex-security@0.1.15
    • bundled plugin 0.1.22
    • bundled Codex runtime 0.148.0-alpha.8
    • model gpt-5.6-sol, reasoning effort xhigh
    • ChatGPT device authentication
    • Node 22 / Python 3
    • Linux arm64 under OrbStack's Docker engine on Apple Silicon
    • non-root uid 1000
    • repository scope reduced to a directory containing 2 files

    Observed

    Across Standard runs:

    • scans.phase remained preflight for the whole run;
    • scan_progress.review_items_total = 0 and review_items_completed = 0;
    • preflight_checks_total = 0 and preflight_issues_json = [];
    • terminal progress nevertheless displayed reviewing files (...) | Files: 0/2;
    • usage grew at roughly 1M mostly-cached input tokens/minute while output increased only about 200 tokens/minute;
    • CODEX_SECURITY_LOG_LEVEL=debug --verbose emitted only cost.updated events, with no shell, helper, sandbox, or source-read activity;
    • the run ended only at the configured cost limit, whose message became scans.failure_message.

    Sandbox controls

    An earlier run had a real Docker seccomp failure preventing Bubblewrap user-namespace creation. That is separate and was eliminated before the reproduction above.

    Using the exact post-seccomp container flags and runtime user, a quota-free probe through the bundled Codex sandbox wrapper:

    • ran _bundled_plugin/scripts/config_preflight.py successfully;
    • returned status: ready;
    • read files from the mounted scan workspace.

    The same sandbox behavior was reproduced on Linux x86_64, weakening an ARM64-runtime explanation. --read-only and no-new-privileges did not affect the helper result.

    Interpretation

    The evidence is consistent with a model/orchestration loop before the Standard-scan agent invokes or publishes capability preflight, rather than a long source-review turn or a Bubblewrap failure. Because the saved session rollout was deleted during credential cleanup, we cannot provide the raw event stream; the workbench counters and verbose-output summary above were captured before cleanup.

    Expected behavior

    If the agent has not invoked or published preflight after bounded turns/time, stop with a diagnostic before further model usage. At minimum, distinguish an active model request from repeated no-progress turns and avoid displaying reviewing files before preflight has completed.

    0.1.16 was released after this reproduction; its release notes do not mention preflight or progress-loop changes. We plan one bounded seeded-control retry on that version and can report whether the behavior persists.

  7. blodebole commented on Aug 20, 2026

    @blodebole

    Follow-up with the bounded 0.1.16 seeded-control retry:

    • Environment: isolated OrbStack Ubuntu machine, Linux arm64, Node 22.14.0. Only a sanitized clone, scan brief, and public knowledge-base files were present.
    • CLI 0.1.16; bundled plugin remains 0.1.22; Codex remains 0.148.0-alpha.8; gpt-5.6-sol / xhigh.
    • Exact scope: three files containing three intentional security controls (tenant authorization bypasses and unsafe privileged billing reconciliation).
    • Dry run completed immediately and resolved all three files.
    • The paid run emitted worker.preflight delegation="available" configured_slots=8 at 28s, then the terminal switched to reviewing files at 55s.
    • It hit the $3 limiter after 4m28s: estimated $3.048526, 2,361,440 input tokens, 2,099,712 cached input tokens, and 23,001 output tokens.
    • Terminal remained Files: 0/3.
    • The sealed workbench still recorded phase=preflight, review_items_total=0, review_items_completed=0, preflight_checks_total=0, finding_occurrences=0, and no registered artifacts.
    • The only output-dir file was scoped-source-input.jsonl.

    Important correction to the initial interpretation: the Codex rollout logs prove that work was occurring behind the zero counters. The run created one coordinator plus five worker sessions and made 53 sandboxed exec calls. Workers successfully read the three scoped source files, repository context, migrations, security guidance, security-scan/SKILL.md, scan-prologue.md, config-preflight.md, and core-scan.md. No bubblewrap, namespace, permission, or read-only-filesystem failure occurred. One initial relative-path read used the output directory and failed with file-not-found; the coordinator immediately retried against $CODEX_SECURITY_REPOSITORY successfully.

    Worker creation was also highly staggered: coordinator at 18:02:05, workers at 18:04:04, 18:04:27, 18:05:35, 18:05:58, and 18:06:22. The final file-specific worker had only seconds before the cost limiter canceled the run.

    So this no longer looks like an ARM64/bubblewrap execution stall. It looks like an observability and cost-bounding problem: active pre-publication worker work is represented as zero review items (and SQLite stays in preflight even after the terminal claims reviewing), while the cost cap can cancel all workers before the first result is committed. A cheap no-progress watchdog still seems valuable, but progress should include worker/tool activity; exposing active/completed worker counts and persisting the actual phase would make the CLI truthful.

  8. mldangelo-oai commented on Oct 8, 2026

    @mldangelo-oai
    Collaborator

    Merged #882 advances progress when drafts are saved, and #962 exposes observed worker sessions. Open #829 adds completed-file review receipts; #1276 improves usage display and dashboard responsiveness. These cover parts of this request.

  9. removed
    area:costToken usage, cost estimates, spending limits, and quota visibility.
    on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:cliCommand-line help, arguments, terminal UI, and user interaction.bugSomething isn't workingenhancementNew feature or requestpriority:p2Significant defect or product gap with narrower impact or a workaround

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions