Skip to content

Make complexity estimates hold up under gVisor - #76

Merged
naman0r merged 1 commit into
mainfrom
complexity-under-gvisor
Sep 25, 2026
Merged

naman0r merged 1 commit into
mainfrom
complexity-under-gvisor

Conversation

@naman0r

@naman0r naman0r commented Sep 25, 2026

Copy link
Copy Markdown
Owner

What and why

After #75 was deployed, I measured the four reference solutions on the production runner. Two got the wrong class:

  • Valid Parentheses measured "constant", not linear.
  • 3Sum measured linear, not quadratic. The UI would then have said "Matches the expected O(n²)" beside an O(n) label.

Raw timings on the server show why:

  • Every run carries some 150 ms of fixed cost under gVisor, against about 15 ms on a laptop.
  • gVisor reports CPU time in 10 ms steps.
  • Short runs occasionally jump by tens of milliseconds.

The fit gave too much weight to small runs, which were mostly fixed cost and noise. Changes in app/runner/complexity.py:

  • A size counts once the solution's own time is at least the run's measured fixed cost (never below MIN_MEASURABLE_MS), so an error in that estimate stays a fraction of what's measured.
  • Runs shorter than REPEAT_BELOW_SECONDS are timed twice and the faster run is kept, since scheduling only ever adds time. Longer runs aren't repeated, which leaves a slow solution time for one more size.
  • Only the largest FITTED_POINTS sizes are fitted.

Other changes:

  • V14 gives Two Sum one more input size (a million numbers, within its 256 MB), so it clears the new floor at three sizes.
  • ComplexityPanel says "Faster than the expected" instead of "Matches" when the measured class is below the expected one.

Deploying

V14 is a migration. Merge, then run ./infra/deploy.sh on the host.

How this was verified

  • I ran the final complexity.py three times on the production runner through the real sandbox path (runsc, the problems' own memory limits), with the four reference solutions plus nested-loop Two Sum and triple-loop 3Sum. All 18 results were correct:
    • linear 0.81 to 1.34
    • quadratic 1.84 to 2.31
    • cubic 2.88 to 3.18
  • Wall time per analysis on the server is about 5 to 10 s, mostly per-run sandbox start and input generation.
  • Before the change, the same check got 2 of 4 reference solutions wrong. Earlier drafts of this change also got the triple loop wrong in one run of three, which is what led to repeating only short runs.
  • make test: 141 passed, 6 skipped. It includes the check that every analyzable problem's reference solution measures as expected, which now also runs Two Sum at a million numbers. Web lint and type check pass.

In production two of the four reference solutions got the wrong class: runs
there carry some 150 ms of fixed cost and CPU time comes in 10 ms steps, so
the small runs the fit leaned on were mostly noise. A size now counts once the
solution's own time is at least the run's fixed cost, short runs are timed
twice and the faster kept, and only the largest sizes are fitted. Two Sum gets
one more input size so it clears the floor at three sizes. A result faster
than expected no longer reads as a match.
@vercel

vercel Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
tandemcode Ready Ready Preview Sep 25, 2026 6:43pm UTC

@coderabbitai

coderabbitai Bot commented Sep 25, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 21 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 90462321-d472-47be-bab3-13d27443aab6

📥 Commits

Reviewing files that changed from the base of the PR and between 0b43d7d and 5603265.

📒 Files selected for processing (4)
  • apps/backend/app/runner/complexity.py
  • apps/backend/migrations/V14__longer_two_sum_analysis.sql
  • apps/web/src/components/ComplexityPanel.tsx
  • docs/judge-and-sandbox.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@naman0r
naman0r merged commit 14fe8b4 into main Sep 25, 2026
6 checks passed

This branch was successfully deployed

1 active deployment
Preview — 56032651 Deployed Sep 25, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant