Skip to content

Repository files navigation

roborak

roborak

AI code review from the terminal - local diffs, GitLab MRs, and GitHub PRs.

Severity-graded, line-anchored findings with committable fix suggestions.

PyPI Python CI Ruff Providers Forges Docs License

Documentation · Install · Usage · How it works · Configuration · Security · License


What it does

🎯 Anchored, not approximate Every finding points at a line the change actually touched - verified against the files on disk, not the model's word for it.
🧹 Refuses more than it says Low-confidence findings filtered, duplicates collapsed, pre-existing lint debt suppressed, already-posted comments skipped.
🔌 Four sources, one pipeline Local git, GitLab MRs, GitHub PRs and raw paths all normalise into one IR, so output modes can never disagree.
🛠 Static analysis as evidence Runs ruff, mypy, semgrep, eslint and phpstan - plus actionlint, hadolint and checkov for workflows, containers and IaC - with your config, and feeds the results to the model to confirm or explain.
💬 Publishes where you're looking Inline threads for what's worth interrupting for, a summary comment for the rest, incremental so re-runs don't repeat themselves.
✅ Closes what it opened Once later commits fix a finding, roborak says which commit did it and resolves its own thread - but only when the fix is verified against the history, never because the finding stopped appearing.
📦 Reads the lockfile you don't Lockfiles stay out of the model's context - they're generated data - so a parser reads them instead and reports what actually moved: a swapped registry, a lost checksum, a mutable git ref, manifest/lock drift. Changes to CI workflows, Dockerfiles and Terraform get their own review checklists.
🧭 Maps the blast radius Traces changed symbols, routes, events, config keys and env vars out to the unchanged code that depends on them, and says plainly when it could not look.
📋 Issue-aware --issue 42 judges the diff against what was actually asked, and reports the requirements it misses.

Status: feature complete. Local diffs, GitLab MRs, GitHub PRs, issue context, static analysis, custom rules, posting and every output mode work end to end.

Install

uvx roborak review       # try it without installing anything
uv tool install roborak             # or: pipx install roborak
export ANTHROPIC_API_KEY=...        # or OPENAI_API_KEY, GEMINI_API_KEY, …
roborak review           # review everything that differs from the base branch

rk is a shorter alias for the same command. That is the whole quick start. Everything below is optional.

From a checkout instead
uv sync
uv run roborak review
Keys in the config file instead of the shell

Keys can also live in the config file, per provider, which is useful when a checkout needs credentials the shell does not carry:

llm:
  api_keys:
    anthropic: sk-ant-...
    openai: sk-...
  api_base: http://localhost:11434   # optional: proxy, Azure, or a local Ollama

A configured key wins over the provider's environment variable. These are real secrets on disk, so keep them in ~/.config/roborak/.roborak.yaml or in a .roborak.yaml / .roborak.yml your repo ignores - roborak config show redacts them, git does not. roborak config init --global scaffolds that user-wide file and creates it mode 600. Setting api_base alone is enough for endpoints that need no key.

Use

uv run roborak review                       # everything that differs from the base branch
uv run roborak review --base main           # compare against a specific ref
uv run roborak review --uncommitted         # staged and unstaged edits only
uv run roborak review --committed --base main
uv run roborak review --include-untracked
uv run roborak review --no-llm              # static analysis only; no API key needed
uv run roborak review --no-static           # model only, skip the linters
uv run roborak review --no-verify           # don't run the project's own test commands
uv run roborak review --no-walkthrough      # skip the overview; one model call, not two
uv run roborak review --full                # add the agent prompts and the review info
uv run roborak review --panels              # one finding to a panel, not the report
uv run roborak review > review.md           # piped: the raw markdown, chrome on stderr

Directories without git

Point -C at a directory that is not a git repository and roborak reviews every file in it whole, rather than refusing for want of a baseline - useful for extracted archives, generated trees, vendor handoffs and other version-control systems.

uv run roborak review -C /path/to/codebase   # every file, reviewed whole
uv run roborak review -C /path/to/codebase/src   # or just one subtree

The walk never descends into dependency, build-output, cache or VCS directories (node_modules, vendor, dist, build, target, __pycache__, .venv, .git and friends), and it also honours the configured ignore_paths. Binary files and anything over 512 KiB are reported as omissions rather than reviewed. Static analysis still runs; the flags that name a diff (--base, --committed, --uncommitted, --include-untracked) do not apply and are refused.

Forges

uv run roborak review --mr 298              # a GitLab merge request
uv run roborak review --mr https://gitlab.com/acme/web/-/merge_requests/298
uv run roborak review --pr 42               # a GitHub pull request
uv run roborak review --mr 298 --post       # publish inline threads + a summary
uv run roborak review --mr 298 --post --repost   # re-post findings already sent
uv run roborak review --mr 298 --no-discussions # ignore existing MR discussion
uv run roborak review --mr 298 --no-post    # review it, never ask about publishing

Issue context

uv run roborak review --issue 42            # review whatever MR/PR implements issue 42
uv run roborak review --issue https://gitlab.com/acme/web/-/issues/42
uv run roborak review --mr 298 --issue 42   # review MR 298, judged against issue 42
uv run roborak review --issue 42 --base main     # local diff, judged against issue 42

Output and filtering

uv run roborak review --json                # full result as JSON
uv run roborak review --agent               # JSON for another agent to act on
uv run roborak review --prompt-only         # findings as fix instructions
uv run roborak review --markdown report.md  # walkthrough-style markdown report
uv run roborak review -m openai/gpt-5       # any LiteLLM model string
uv run roborak review -s major              # only major and critical
uv run roborak review --fail-on critical    # non-zero exit for CI
uv run roborak review --mr 298 --post --no-check   # comments, but no commit status
uv run roborak review --no-investigate      # skip the evidence-gathering stage
Exit code Meaning
0 Review completed
1 Findings at or above --fail-on
2 Operational error or partial review - failed chunks, unavailable forge patches, or a requested publish that did not complete

Run output and verbosity

Each stage of a review (static analysis, verification, the model passes, the overview, posting) leaves a line on stderr with its outcome and how long it took, and a chunked model review shows its per-pass position while it runs. stdout stays the report alone, so --json, --markdown and a plain pipe are unaffected.

uv run roborak review            # stage lines and warnings on stderr
uv run roborak -v review         # add INFO logs (e.g. how many passes a change needs)
uv run roborak -vv review        # add DEBUG logs
uv run roborak -q review         # errors only: no stage lines

The pre-merge check

Every review ends with a pre-merge check: the verdict, the severity floor it was judged against, and the finding counts that drove it. It is the last section of the report, so it shows in the terminal, in --markdown output, and - because the summary comment is the report - on the merge request too, on every re-run.

The floor is --fail-on when you pass it, and review.block_on (default critical) otherwise. Only --fail-on moves the exit code; without it the block says so rather than implying CI is gated when it is not.

review:
  block_on: major     # the floor the verdict is judged against
output:
  post_check: true    # post it to the forge as a commit status

review.block_on is not review.severity_floor: the floor decides what is reported at all, block_on decides what blocks.

Configurable pre-merge checks reach the same verdict and the same forge status: one set to error blocks a change on its own, with no finding anywhere near the floor. None of them move the exit code, which stays --fail-on's alone.

Blocking takes evidence. A critical or major model finding has to say what makes it true - the trigger and the failure path, a violated contract, or a reproduction - not just how confident it feels. One that cannot is demoted to a minor verification_needed: still reported, still anchored, no longer counted by the verdict. A self-assigned confidence: 0.95 is the model grading its own homework, and it is not grounds to fail a build. Static-analyser findings are exempt, because a tool ran. Set review.require_evidence: false to turn the policy off.

Deployment and runtime risk. The reliability category covers the defects that are locally correct and still fail on rollout: a field required before the version that writes it ships, a migration that locks a hot table or cannot be rolled back, a retry wrapped around a call that is not idempotent, a feature flag defaulting on where it is unset, a connection pool raised past what the database serves, a cache key that stops separating one tenant from the next. The checklist is gated on what the change actually touches - migrations, deploy configuration, public contracts, jobs and queues, retries and timeouts, flags, caches, limits, and logging or metrics that a change removes. A diff crossing none of those surfaces is never asked about rollout, and "add monitoring" is not a finding: an operational label still has to name the deployment or runtime state that breaks.

As a forge status. A review posted with --post also sets a commit status on the change's head commit, named roborak/review, so branch protection and approval rules can gate on it. Re-running a review replaces that status rather than stacking another. Pass --no-check (or set output.post_check: false) to publish comments only.

Token scope Where to require it
GitHub statuses:write, or the classic repo scope Settings → Branches → branch protection rule → Require status checks to pass → add roborak/review
GitLab api Settings → Merge requests → Pipelines must succeed (the status joins the MR's head pipeline)

A token that may comment but not set a status is not an error: the review still publishes and roborak reports the skipped check.

Pre-merge checks

Four checks about the change rather than about the code - the questions a reviewer asks before reading a diff at all. Each takes a level:

Level What it does
off Does not run, render, or reach the verdict.
warning Reported in the report and the summary comment; never blocks.
error Also blocks the pre-merge verdict and the forge commit status.
pre_merge:
  docstring_coverage:
    level: warning
    threshold: 0.8    # fraction of touched symbols that must be documented
  title:
    level: warning
  description:
    level: warning
  linked_issue:
    level: warning

Docstring coverage measures the symbols the diff touched, not whole files, so a one-line fix in a legacy module is not judged by that module. It works for every language a tree-sitter grammar is available for: a leading string in the body, or a comment on the line directly above. A file with no grammar is left out of the count rather than counted as undocumented, and so is one whose content could not be read - reported as unread, which is a different thing. A merge or pull request arrives as a diff, so the check reads the changed files back out of the reviewed commit, reusing the same temporary checkout the blast radius fetches when this repository does not have that commit (impact.forge_checkout).

Title, description and linked issue are deterministic gates - present, long enough, not a placeholder, not an untouched template, naming an issue through Closes #123, an issue URL or --issue. They run under --no-llm too, so a policy you gate merges on does not depend on a provider being reachable. On a local diff, which has no request body, the description and linked-issue checks report not applicable rather than failing.

Raise one of the three text checks to error and roborak also asks the model whether the title and description actually describe the change. That opinion is reported but never blocks on its own: an opinion is not evidence, which is the same bar findings are held to above.

Every check is off with ROBORAK_NO_PRE_MERGE=1, or one at a time with ROBORAK_TITLE_CHECK=error and its siblings.

Other commands

uv run roborak describe                     # title, overview, per-file table, mermaid flow
uv run roborak improve                      # suggestions only, every one committable
uv run roborak fix --dry-run                # preview committable replacements
uv run roborak fix                          # preview, confirm, and apply
uv run roborak fix --yes --json              # apply in scripts, report JSON
uv run roborak ask "why is this locked?"    # a question answered from the diff

uv run roborak rules init                   # scaffold .roborak/rules/ with an example
uv run roborak rules list                   # what roborak will apply here
uv run roborak rules test <rule.md> <file>  # validate a rule and check its scope
uv run roborak setup                        # guided first run: model, key, forge tokens
uv run roborak config init                  # write a commented .roborak.yaml
uv run roborak config init --global         # …or ~/.config/roborak/.roborak.yaml, mode 600
uv run roborak config show                  # the effective config, all layers merged

fix accepts the same source, model, and configuration options as improve. It requires a Git checkout at its repository root. Existing staged and unstaged local edits are supported; fixes remain unstaged and the index is preserved. PR/MR targets require a clean local checkout at the exact reviewed head. Merge conflicts are refused. Files changed during generation or confirmation, ambiguous overlaps, moved anchors, and non-committable suggestions are skipped.

fix --dry-run reports eligible replacements and unified diffs without writing. Interactive runs preview changes and ask once before applying; scripts and pipes must use --yes or --dry-run. Reports separate applied, skipped, failed, and eligible suggestions, with reasons; --json emits a dedicated fix report. Exit 0 means the run completed (including skips, cancellation, or dry-run), and 2 means an operational failure or incomplete generation. Each write temporarily moves the target aside, validates it, and publishes only if the target path is still absent. Concurrent saves at that path are preserved. Failed recovery retains the original in a .roborak-fix-* directory beside the target and reports its location. Filesystems without hard-link support are refused before moving the target. Writers using already-open file descriptors are not locked out; detected changes to the moved original are retained for recovery, but later descriptor writes cannot be guaranteed. Avoid editing files while applying fixes. A failed file does not undo successful changes to other files. No tests, formatting, staging, commits, or publishing run automatically.

Each accepts the same --mr / --pr / --issue / --base targeting as review.

Tokens

--mr needs GITLAB_TOKEN (or ROBORAK_GITLAB_TOKEN, or CI's CI_JOB_TOKEN). --pr needs GITHUB_TOKEN, or an existing gh auth login session, which roborak will use automatically. --issue needs whichever of the two matches the issue's forge, inferred from the URL or the git remote.

Like the LLM keys, these can live in the config file instead of the shell:

forge:
  tokens:
    gitlab: glpat-...
    github: ghp_...

A configured token wins over GITLAB_TOKEN / GITHUB_TOKEN and over the gh session, while ROBORAK_GITLAB_TOKEN / ROBORAK_GITHUB_TOKEN still win over the file. They are secrets on disk, so the same advice applies: keep them in ~/.config/roborak/.roborak.yaml (roborak config init --global creates it mode 600) or in a .roborak.yaml / .roborak.yml your repo ignores. roborak config show redacts them.

Self-hosted instances

A bare --mr 705 has to work out which server it means, which it normally reads off the repository's git remote. Name the instance in the config for the cases where that does not answer it - no remote, or a remote pointing at a mirror:

forge:
  hosts:
    gitlab: gitlab.acme.com
    github: http://gh.local:8080   # https is assumed unless you say otherwise

The git remote still wins: a domain configured user-wide can never hijack a checkout whose remote plainly says otherwise, and a full URL passed to --mr / --pr / --issue beats both. ROBORAK_GITLAB_HOST / ROBORAK_GITHUB_HOST set the same thing from the environment.

A configured host also teaches roborak which forge an unrecognisable domain is, so --issue 24 works in a repo whose remote is something like git.acme.com rather than failing with "could not tell which forge".

How it works

One directional pipeline; each stage only knows the stage before it.

Source → ChangeSet → Compressor → Static pass → Verification → LLM → Validator → Renderer
  • ChangeSet is the universal IR. Local git, GitLab, GitHub and raw paths all normalise into it, so nothing downstream knows where the code came from.
  • Line anchoring is the correctness-critical part. Findings are always in new-file coordinates; Hunk.line_map records each line's position within the diff, and only the publishers translate that into a forge's position payload. tests/test_local_git.py checks the computed numbers against the real files on disk, so an off-by-one cannot agree with itself and pass.
  • The validator drops findings that do not point at a changed line, snapping near misses onto the nearest one, then filters by confidence and severity and collapses duplicates. Most of roborak's usefulness is in what it refuses to say.
  • The verification pass runs the project's own tests, proportionally: the narrowest matching command set that covers the changed files - every configured command whose globs match, deduplicated by argv and capped by max_commands - escalating to the broad check only when the change crosses a shared boundary. It is off until a project configures a command, its commands are read from the base revision so a change cannot define what verifies it, and every result - passed, failed, timed out, could not run, not executed - is recorded on the review rather than left for the reader to run themselves.
  • The static pass runs whichever of ruff, mypy, semgrep, eslint, phpstan, actionlint, hadolint and checkov the repo actually has, using the project's own config - the rules a team already agreed to. Findings on lines the change did not touch are dropped, so a linted file's pre-existing debt never lands on the author. What survives is fed to the model as evidence to confirm or explain, rather than reported raw.
  • The supply-chain pass reads the manifest and lockfile pairs a change touched - npm/yarn/pnpm, Python, Go, Cargo, Composer - straight out of git, and reduces them to a small delta: what was added, removed, upgraded, re-sourced, or is no longer covered by a checksum. The lockfiles themselves never enter the prompt, which is the point: they are excluded by ignore_paths for good reason, and a review that could not see past that exclusion could not notice an unexpected transitive package or a registry swap. When a change touches CI workflows, containers or Terraform, the reviewer additionally gets a checklist written for those trust boundaries rather than for application code. What could not be parsed says so.
  • Every finding is routed, not just printed. A finding that points at a changed line and is worth interrupting for goes inline on the diff, where the author is already looking. A nitpick is folded into the summary instead, so the small stuff cannot drown the review. And one that cannot be anchored is reported in the summary under a warning banner rather than discarded - roborak used to count those as failures and show them only in the terminal, which meant nobody reading the merge request learned they existed. roborak.core.buckets is the one place that decides, so the terminal, the markdown report, the summary comment and the publisher cannot disagree about where a finding belongs.
  • Publishing translates new-file coordinates into each forge's position payload, and only there. GitLab needs all three of base_sha/start_sha/ head_sha from the MR's own diff_refs; GitHub takes one review containing every comment, always as COMMENT - roborak never approves or requests changes on your behalf. A rejected comment never costs you the rest of the review.
  • Incremental review fingerprints each finding independently of its line number, so re-running on a new push posts only what is genuinely new instead of repeating itself. State lives in .roborak/state.json; --repost overrides it.
  • A thread is resolved only after the evidence that closes it has landed. A later --post revisits the actionable threads roborak opened and has not closed, reads the commits between the thread's own anchor and the current head, and asks whether they actually fix what was reported. Where they do, it replies naming those commits and summarising what changed, and only then resolves the thread - a comment closed without the reasoning that closed it is a finding that vanished. Everything short of a verified, attributable fix leaves the thread open: an untrusted checkout, a revision a shallow clone never fetched, a range no commit touched the file in, a model that could not tell. Human comments, the summary, nitpicks and already-resolved threads are never touched, and the markers roborak leaves make a repeated run a no-op rather than a second copy. On GitHub this uses the GraphQL API, which is the only one that can resolve a review thread at all. --repost skips the pass.
  • Existing review discussion is context, not instruction. Forge reviews include bounded unresolved human comments by default, while dropping system notes, bots, stale positions and roborak's own output. --no-discussions disables it.
  • Deciding to publish comes after reading the review. --post has to be chosen before the model has said anything, so an interactive run ends by asking instead - post it, save it as markdown, or neither - showing first how many inline comments are new and how many an earlier run already sent. A local diff has nowhere to post, so only saving is offered. The question is asked only on a terminal: a pipe, a script and a CI job are never prompted, and --no-post or output.confirm_post turns it off for good.
  • The overview is a second pass. review asks for a walkthrough after it has the findings, which is what fills the summary comment and the markdown report's file table. It runs on a copy of the changeset, because compression mutates and shrinking the diff the findings were anchored against would corrupt every line number already reported. A failed overview is logged, never fatal: a review without one is still a review, and must still exit clean. --no-walkthrough skips it.
  • An overview is written once per shape of a change. The published summary carries a digest of its changed files and hunk headers. Re-posting to the same merge request compares that digest against the diff in front of it: unmoved means the story is unmoved, so the model is not asked again and the existing comment stays put. When it has moved, the overview is rewritten and that same comment is edited in place rather than a second one appended. --repost forces a fresh overview, as it does for inline findings.
  • Output modes share one result object, so the terminal report, the markdown file, the JSON payload and the forge comment can never disagree. --json, --agent and --prompt-only write to stdout alone, so they stay pipeable.
  • The report is built for skimming. Findings are grouped into collapsible sections by where they belong, badged with category, severity and the effort a fix will cost, and each one carries a 🤖 prompt a coding agent can act on - plus one collated block for the whole review. A <!-- roborak:v1:… --> marker records each finding's identity in the comment itself, so a published review carries a record of itself that does not depend on local state.
  • There is one document, and you read it before you publish it. One renderer builds it, so what --markdown writes, what a pipe gives back and what --post publishes are byte for byte the same thing - asserted by a test, because it is the invariant a refactor would quietly break. It means the comment repeats the findings that also went out as inline threads, which is the deliberate half of the trade: a comment that omitted them would be a fourth document nobody had read before it was published.
  • How it reaches you depends on who is reading. Redirected or piped it is the raw markdown, so roborak review > review.md gives back exactly the publishable file. At a terminal it is rendered - headings, tables, severity in colour, the flagged lines shown in context from your working tree, syntax-highlighted fixes. Everything roborak says about a run - spinners, errors, the closing question - goes to stderr either way.
  • A terminal cannot fold a section, so it leaves them out instead. A report is built to be skimmed by opening what you want, and every <details> opened at once is the opposite of that: on a twenty-finding review the per-finding agent prompts alone outweigh the findings. The rendered form drops the sections written for a machine - the agent prompts, the review-info tree - and puts what a reader must not lose (an omitted file, a skipped file, an error) in a one-line footer instead. --full restores them. What it never drops is the review: every finding, every badge, every body and every fix is in both forms, which is what tests/test_render.py asserts. --panels is the older view, one finding to a bordered panel.
  • Large diffs are reviewed contract-first in several passes, not truncated. Deterministic roles put public contracts, schemas, migrations, configuration and deployment boundaries ahead of leaf implementation when the eight-pass cap cannot cover everything. Direct consumers and user-authored tests stay with the boundary they exercise when they fit; generated files and low-signal prose go last. Later passes receive bounded contract metadata rather than duplicate diffs, and one final reconciliation pass checks compatibility across chunks. JSON and human-readable coverage show the semantic order and the roles omitted. One failed pass never discards the others, and compression always reports what it skipped.
  • Issue context turns "is this code good?" into "does this code do what was asked?". --issue 42 fetches the issue's title, body, labels and discussion and puts them in the prompt, and - when no other target was named - reviews the merge or pull request linked to it, so --issue 42 alone is enough. Findings of kind requirement_gap name what the issue asked for that the diff does not do. A gap is the one finding with no honest line to point at, so it is exempt from line anchoring and is published in the summary comment rather than inline.
  • AST context (via tree-sitter, installed by default) names the function or class each hunk sits inside. A diff hunk is a window with arbitrary edges; a model that knows it is looking at the middle of run() stops guessing at the surrounding control flow, which is where many false positives come from.

Configuration

Existing users: rename ~/.config/roborak/config.yaml to ~/.config/roborak/.roborak.yaml, preserving its contents and permissions. The old filename is no longer discovered automatically. If the new file already exists, reconcile your settings before replacing either file.

Project config is .roborak.yaml / .roborak.yml in the repo root. .roborak.yaml is the canonical generated default and wins if both exist; the files are not merged. Precedence: CLI flags > ROBORAK_* env vars > project config > ~/.config/roborak/.roborak.yaml > selected profile > defaults. config init writes .roborak.yaml, config init --global writes the user-wide file; both get the same commented template, which ships inside the package rather than being read out of a source checkout.

roborak setup is the guided path to the same files: it asks for a model, a credential for it, and optional forge tokens, then writes only those keys - a sparse file, so every other default stays live across upgrades. It chmods 600 whatever it writes that holds secrets, wherever it wrote it. config init remains the manual path, and the full annotated file to edit.

Review profiles

Start with a preset, then override only the settings you need:

profile: balanced
rk review --profile strict
ROBORAK_PROFILE=fast rk review
rk config show --profile security
Profile When to use it and what changes
balanced Current defaults, unchanged; selected when no profile is given.
fast Local iteration: disables walkthrough, verification, impact mapping, investigation, and temporary forge checkout. Keeps static analysis, supply-chain summaries, and evidence requirements.
strict More thorough review: blocks the verdict on major findings; investigates up to 10 candidates over 3 rounds, 20 files, and 40,000 tokens; maps 24 nodes with 10 consumers each and 3,000 tokens. Verification defaults to broaden_paths: ["**"], 8 commands, and a 600-second timeout.
security Security and infrastructure review: categories [security, reliability], major blocking floor, supply-chain limits of 80 changes / 40 assets / 2,400 tokens, and up to 80 static findings in the prompt. Scanner autodetection stays enabled.

One profile is selected by --profile > ROBORAK_PROFILE > project config > user config > balanced. Profiles never combine. Explicit fields follow the normal precedence chain and always override preset defaults, even when the profile comes from the CLI. This includes values copied from config init: remove or comment out fields you want a profile to control. All unlisted settings retain built-in defaults. No profile grants trusted execution, installs scanners, or supplies verification commands. Only --fail-on changes the exit-code threshold.

Verification selects its profile and explicit settings from the trusted base revision instead of the working tree; CLI/environment/user settings and explicit --config still apply. strict prefers the configured broad fallback when available, otherwise normal targeted selection applies. With no configured commands, nothing runs. config show displays expanded settings, with verification resolved against HEAD, and labels its source and profile. A review against another base revision can resolve verification differently.

In a terminal the closed questions - where the file goes, which model - are arrow-key lists rather than strings to type. Every list ends with Other (type it in)…, because the model list can only ever be a starting point: roborak takes any LiteLLM model string. Keys, tokens and a self-hosted domain stay free text, and keys are never echoed. Run it without a terminal - piped, or in CI - and the same questions come back as plain line prompts reading stdin, so printf '1\n\nsk-ant-…\n\n\n' | rk setup works; with nothing on stdin it writes nothing and exits 0 rather than waiting.

Large diffs are split into bounded model review calls. review.max_chunks limits those calls, and rk review --max-chunks N overrides it for one run. A chunk can hold several related files or part of one oversized file. When more passes are needed, successful ranges are checkpointed in .roborak/state.json and the review is partial. Run the same review again to continue pending or failed ranges automatically. Changed revisions, content, models, or relevant configuration safely start a fresh checkpoint.

version: 1

llm:
  model: anthropic/claude-sonnet-5
  fallback_models: [openai/gpt-5]
  temperature: 0.2

review:
  categories: [security, bug, performance, logic, reliability]
  severity_floor: minor
  max_findings: 25
  max_chunks: 12          # primary model passes allowed per run; later runs resume
  committable_suggestions: true
  min_confidence: 0.5
  require_evidence: true     # a critical/major model finding must show its evidence
  check_requirements: true   # with --issue, report requirements the change misses
  include_discussions: true  # use relevant unresolved MR/PR comments as context
  investigate:               # bounded repository reads that settle a candidate
    enabled: true
    max_candidates: 5        # findings put to the stage, worst first
    max_rounds: 2            # request/result exchanges before it must decide
    max_requests_per_round: 4
    max_files: 10
    max_lines_per_read: 200
    max_search_results: 20
    max_output_chars: 4000
    token_budget: 20000
    timeout_seconds: 30

static:
  enabled: true
  execution: auto     # local direct; CI sandboxed, or skipped if unavailable
  tools: null          # null = autodetect what is on PATH, minus networked tools

supply_chain:
  enabled: true        # dependency delta + CI/container/IaC review checklists
  max_changes: 40      # source, integrity and drift changes sort ahead of bumps
  max_assets: 20
  feed_to_llm: true
  token_budget: 1200   # reserved up front, so it never squeezes out a changed file

verification:
  enabled: true        # no command configured = nothing runs, and no claim is made
  execution: auto      # local direct; CI sandboxed, or skipped if unavailable
  commands:            # the narrowest matching set runs; argv arrays, never shell strings
    - paths: ["src/roborak/context/**"]
      command: ["uv", "run", "pytest", "tests/test_context.py"]
    - paths: ["src/roborak/core/**"]
      command: ["uv", "run", "pytest", "tests/test_pipeline.py", "tests/test_render.py"]
  broaden_paths:       # a change touching one of these runs `fallback` instead
    - "src/roborak/core/models.py"
    - "pyproject.toml"
  fallback: ["uv", "run", "pytest"]
  timeout_seconds: 300
  max_commands: 4      # ceiling on how many commands one review runs
  max_output_lines: 40 # per command; a citation, not a build log
  feed_to_llm: true

pre_merge:            # off | warning | error; only `error` blocks the verdict
  docstring_coverage:
    level: warning
    threshold: 0.8    # fraction of diff-touched symbols that must be documented
  title:
    level: warning
  description:
    level: warning
  linked_issue:
    level: warning

impact:
  enabled: true           # trace changed symbols out to their consumers
  max_nodes: 12           # boundaries traced per review
  max_consumers_per_node: 5
  max_files_scanned: 2000 # ceiling on the fallback walk when git grep is unavailable
  max_snippet_lines: 6
  token_budget: 1500      # prompt tokens the consumer snippets may occupy
  timeout_seconds: 10
  forge_checkout: auto    # fetch a temporary checkout of a PR/MR this repo lacks; off disables
                          # (shared with the docstring-coverage check)
  forge_checkout_timeout_seconds: 60

output:
  walkthrough: true    # spend a second model call on the overview
  confirm_post: true   # offer to publish at the end of an interactive review
  panels: false        # one finding to a panel instead of the report
  full: false          # show the agent prompts and review info the terminal hides

ignore_paths:
  - "**/*.lock"
  - "**/vendor/**"
  - "**/node_modules/**"

language_instructions:
  php: "Laravel 10 with the repository pattern; controllers stay thin."

Custom rules

Standards a linter cannot express go in .roborak/rules/*.md as plain language:

---
id: no-raw-sql
paths: ["app/**/*.php"]
severity: major
category: security
---
Never build SQL by string concatenation. Use the query builder or bound parameters.

Only rules matching the changed files enter the prompt, so token cost stays flat as the rule set grows. Frontmatter is optional - a file containing one sentence is a valid rule, named after the file.

roborak also reads AGENTS.md, CLAUDE.md, .roborak/context.md, or CONTRIBUTING.md (first one found) so reviews match the repo's own conventions. When a base revision is available, conventions are read from that revision so a change cannot rewrite the instructions used to review itself.

Verification

Roborak can run the project's own tests as part of a review, so a finding about changed behaviour has an execution record behind it instead of an invitation to go and check. Selection is proportional: the commands whose paths match the changed files, deduplicated, capped by max_commands. A change touching a broaden_paths entry - a shared contract, a schema, the build configuration - runs fallback instead of the targeted set, since a full suite already contains every subset of itself. A change matching nothing falls back to the broad check, and with no fallback configured it is simply not verified, which the report says in words.

Results reach every surface. The report carries a Verification section - including when nothing ran, because a section that appeared only on green would teach readers that its absence means the tests passed. --json carries summary.verified and a per-command record, and the model is handed failures as evidence with the output framed as data rather than instructions.

Important

Verification runs argv the repository authored, so the commands are read from the base revision - never from the working tree. Reviewing a branch someone else wrote must not run what that branch says to run.

The practical consequence: a change that adds a verification command is not verified by it until it lands. --config <path> is the deliberate override, since choosing that path is itself the trust decision, and your own user config and ROBORAK_* environment are always yours. Beyond that, verification obeys the same execution model as static analysis below: direct locally, sandboxed in CI, skipped in CI when Bubblewrap is unavailable, and --trust-verify (or verification.execution: trusted) as the explicit opt-in. Nothing is installed and nothing is fetched - the suite runs against the checkout as it stands, in a credential-scrubbed environment. Inside the CI sandbox the repository is read-only, so a suite that must write is reported as could not run rather than as a failure.

A failed suite does not move the pre-merge verdict or the exit code: that verdict counts findings. It does print beside it, so "pass" is never read as covering both.

Supply chain and infrastructure

ignore_paths excludes every lockfile from the review by default, and that is the right call: a lockfile is generated data, and sending one to a model spends thousands of tokens asking it to do a diff badly. But excluding it also hides an unexpected transitive package, a swapped registry, a checksum that quietly disappeared, or a manifest edit that never reached the lock.

So the lockfile stays out of the prompt and a deterministic parser reads it instead. What the model gets is a bounded delta - what was added, removed, upgraded, downgraded, re-sourced, or is no longer covered by a hash - plus manifest/lock drift.

Ecosystem Manifests Lockfiles
npm / yarn / pnpm package.json package-lock.json, npm-shrinkwrap.json, yarn.lock (Classic and Berry), pnpm-lock.yaml
Python pyproject.toml, requirements*.txt uv.lock, poetry.lock, pdm.lock
Go go.mod go.sum
Rust Cargo.toml Cargo.lock
PHP composer.json composer.lock

Anything else - Gemfile.lock, Pipfile.lock, gradle.lockfile, mix.lock - is named in the report as unanalysed rather than passed over in silence.

Both sides are read from git, not from the diff: whole-file parses compared against each other, so a reformat that moves every line produces no changes at all. The old side is read at the merge base, so a dependency somebody else bumped on the base branch is not reported as part of this change. A forge diff that is not checked out has no readable sides, and the report says so instead of guessing.

Changes to CI workflows, Dockerfiles, compose files and Terraform additionally turn on review checklists written for those trust boundaries - workflow permissions and secret exposure, container privilege and mutable base images, IAM and public access - gated so that a Terraform-only change never pays for the npm checklist.

Three optional scanners run when the repository already has them: actionlint for workflows, hadolint for Dockerfiles and checkov for IaC. osv-scanner is supported too, but because it queries a vulnerability service it is never autodetected: it runs only when a project names it in static.tools. Everything else stays offline, and nothing is ever installed. A scanner that applied to the change and was not available is named in the report, so a clean section cannot be mistaken for a checked one.

--no-supply-chain (or supply_chain.enabled: false) switches the stage off.

Investigation

require_evidence demotes a critical or major model finding that cannot say what makes it true, which is right - a self-assigned confidence score is not grounds to fail a build - but until now the model had no way to go and find out. A real blocker whose proof sat one function away was demoted anyway, and a false positive that one read would have disproved survived to the report.

So before findings are validated, roborak puts its unproven blockers back to the model with a small set of read-only operations: a bounded line range, a git grep within the repository, the reviewed diff for a file, and where a parser is available, a symbol's declaration. The model asks for what it needs, reads the results, and then confirms, revises or drops each candidate. A confirmed one carries the evidence it found into the report, and keeps its severity; a dropped one never reaches a reader.

The boundary is narrow on purpose. Paths are resolved and proved to be inside the repository before any I/O, so .. and a symlink out of the tree are the same refusal. Searches run as argv with a scrubbed environment, never through a shell. Every result is bounded and says when it was cut, and everything read is fed back as untrusted input. Nothing writes, applies a patch, runs a repository-chosen command, or reaches the network.

Two rules make a failure safe. A candidate the stage cannot settle - a tool error, an exhausted budget, an unreadable reply, a provider outage - is returned exactly as it arrived, so "we could not tell" is never recorded as "we checked". And for a PR or MR, the checkout must be clean and at the reviewed head commit before anything is read from it; otherwise roborak falls back to the file content the forge supplied, or reports the stage as unavailable rather than confirming a finding against an unrelated branch.

It costs one model call on a review that has a critical, major or would-be-demoted finding, and nothing at all on one that does not. --no-investigate (or review.investigate.enabled: false) switches the stage off.

Static-analysis trust

Important

Static analyzers load repository binaries, plugins, and configuration, which can execute code. Outside CI, static.execution: auto treats the checkout as trusted.

In CI it runs through Bubblewrap with a read-only filesystem and no network; when Bubblewrap is unavailable, the static pass is skipped rather than running untrusted code directly. --trust-static (or static.execution: trusted) is the explicit override for a checkout you control. Every static subprocess receives a credential-scrubbed environment in all modes.

CI also ignores .roborak.yaml / .roborak.yml from the working tree entirely, since either could redirect an API key or opt into trusted execution. (The verification section is read from the base revision in every environment, not just CI.) Put CI settings in environment variables, the user config, or pass a trusted base-controlled file with --config.

Design notes

Three decisions carry most of the weight:

  1. ChangeSet is the only thing the pipeline knows. Four sources normalise into it, so adding a fifth touches nothing downstream.
  2. Line numbers are new-file coordinates everywhere, translated to forge position payloads only at the publisher. tests/test_local_git.py checks the computed numbers against files on disk, so an off-by-one cannot agree with itself and pass.
  3. The tool's value is mostly in what it refuses to say. Findings outside the change are dropped, low-confidence ones filtered, duplicates collapsed, pre-existing lint debt suppressed, and already-posted comments skipped.

Development

uv run pytest              # 940 tests
uv run ruff check src tests
uv run ruff format src tests
uv run mypy src/roborak

Live reviewer-quality evaluation is intentionally separate from deterministic PR CI. uv run python evals/run.py exercises 30 labeled defect and clean-control cases, writes token and quality metrics, and enforces the nightly recall, false-positive, anchoring, and parse-success gates.

Conventions, invariants and the PR checklist are in CONTRIBUTING.md.

Documentation

Full documentation lives at roborak.pages.dev - install, every command and flag, configuration reference, custom rules and CI recipes.

License

MIT - see LICENSE.md.


Built with LiteLLM, Typer and Rich.

About

AI code review from the terminal: local diffs, GitLab MRs, and GitHub PRs. Severity-graded, line-anchored findings with committable fixes, backed by linters.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages