Skip to content
newjordanPublic

About

Resources

Contributing

Security policy

Stars

9 stars

Watchers

0 watching

Forks

Repository files navigation

angelX

angelX is a terminal coding agent. It works inside your repository with the models you choose, asks before it runs a command or edits a file, checks its work before it reports done, and keeps what it learns about your project.

angelX starting on GLM-5.3: Excalibur raised, the knight on the summit

curl -fsSL https://raw.githubusercontent.com/newjordan/angelX/main/install.sh | sh
cd your-project && angelX

For Linux x86_64 and arm64 (glibc 2.31 or newer) and macOS on Apple silicon. The first launch finds the models you already have and connects one: Ollama, LM Studio, llama.cpp or vLLM on this machine, an API key, or a ChatGPT or Grok plan login. angelX setup changes it later, and angelX update installs the latest release.

To build from source instead, you need Rust/Cargo, C/C++ tools, Bash, Node.js and Python 3 (on macOS: xcode-select --install, then Rust from https://rustup.rs):

git clone https://github.com/newjordan/angelX.git
cd angelX
./bin/angelX

Model setup · /commands · Feature evidence · Attributions · MIT

How it works

GLM-5.3 reads a small Rust crate to find a unit bug while the knight rides to the Scriptorium on the overworld map

1 · Choose a model. /model lists every connected route with its thinking level; /think changes the level.

Model picker with GLM-5.3 selected

2 · Approve what it does. Commands and edits arrive as action capsules: approve one (y), approve the rest of the turn (a), or deny (n).

Approval capsule for a shell command

3 · Content-checked edits. Each edit is anchored to the file's current content and shows its byte change before it lands. /diff shows the result.

Edit approval: replace once in src/lib.rs, 32 B to 30 B The one-line fix in /diff

4 · Checked before done. The agent runs the checks and reports with receipts: every tool call, its time and the model route.

The turn summary and the final answer with the tests passing

5 · Longer work. /goal sets a durable objective with acceptance criteria and a check. /loop runs it within time, iteration and token limits.

A goal with acceptance criteria and a verify command Loop workshop: time, iterations and token cap

In the cockpit

Formation roster
Formations · model teams, from a lone coordinator to a full roster.
Context budget
Context budget · where the window goes; compaction runs automatically.
Command help and completion
Commands · /help and Tab completion for every command.
The overworld map beside the code
World view · Watch model behavior as an adventure in information.
GLM-5.3 plants a bar chart one graph call at a time while a sprite carries the next point to the field
Graph garden · every graph call grows a crop; a sprite carries each point.
The finished chart in the garden beside its planted values from /world crops
Planted · /world crops reports the exact values it grew.

Features

  • Models and formations — Choose models and thinking levels, configure teams, and run agent graphs.
  • Repository tools — Search files, inspect symbols, follow definitions, and review diffs.
  • Content-checked edits — Hashline editing checks file content before applying anchored changes.
  • Programmable tools — Run tool calls, loops, and filters locally in JavaScript with code_mode.
  • Long-session context — Compact, deduplicate, and age retained context as work continues.
  • Project memory — Review source-linked knowledge in Atlas and reuse it in later tasks.
  • Persistent goals — Resume sessions and autonomous work with configurable run limits.
  • Headless runs — Record task settings, source identity, tool activity, and acceptance evidence.
  • Measured campaigns — Evaluate isolated attempts with verifiers and independent review.
  • Research loops — Use Sloptomizer suggestions, Deli deliberation, and paired experiments.
  • Measured benchmarks — Calculate measured changes from paired benchmark samples.
  • Book of behaviors — The harness steers the model with compact braille stamps; each stamp's English is taught the first time it appears.
  • Adventure world model TUI — Introducing the early stages of Cyberdynamic world tui for reviewing work, presenting data graphs, adventure, and model behavior.
  • Graph garden · 0.1.9 — Farmers and sprites turn actual graph calls into bar, line, and scatter crops on the fields. Visit with /world visit garden, inspect with /world crops, and begin again to regrow the bed. Guide and preview.
  • Calibrated seats · 0.1.9 — Muse Spark runs at minimal effort and Grok 4.7 at low, each picked on its hardest benchmark tasks, and playbook hints no longer send a model after a procedure its task rules out.
  • Live research memory · 0.1.9 line — Sloptomizer connects ordinary checks and measured experiments through compact ⚠ braille advice. Used legends return when the model changes or context is compacted. Ships in 0.1.92; ship notes.
  • A working adventure · 0.1.9 line — The knight follows real loop work through a 3D mine journey, with a School of Magic, study, and underground archive reflecting actual research evidence.
  • The Delve · 0.1.92 — /dungeon walks your knight to the Delve's gate: a co-op knights-vs-monsters dungeon drawn in the realm's own pixel art. Friends join from their own angelX with /dungeon join, and cards, spells and wishes load while you play. Guide.
  • Cartridges · 0.1.93 — Plug a competition into the loop as a folder: a cartridge.toml, and Rust when it needs more. Yours stay outside the tree. Guide and figures.

Benchmarks

136 repository-repair tasks (48 JS, 34 Python, 30 Rust, 24 C++) from the Aider polyglot set. One attempt per task, 600 s limit, graded by each task's tests. Wall: median agent time per attempt. Tokens: totals per cell.

angelX 0.1.8 · polyglot-v1 · 2026-09-30

Agent time against input tokens on DeepSeek V4.1 Flash: angelX solves 135 of 136 twice, in 8.7M tokens and 29.3 agent minutes, then 6.9M and 27.7; OpenCode solves 132 in 67.7M tokens and 58.7 minutes. A failed task is the red stretch of its line.

model solved wall (median) calls / task input tokens cache hit output tokens
DeepSeek V4.1 Flash, thinking off · run 1 135 / 136 9.9 s 6.0 8.7 M 84% 205 k
DeepSeek V4.1 Flash, thinking off · run 2 135 / 136 9.1 s 5.3 6.9 M 80% 195 k

Evaluator: Prime Intellect Verifiers v0.3.1 · angelX 0.1.8 · temperature 0 · 600 s per attempt · fresh environment per attempt · 2026-09-30. The other models' latest results are 0.1.6's, below.

angelX 0.1.6 · polyglot-v1 · 2026-09-23

model solved wall (median) calls / task input tokens cache hit output tokens
DeepSeek V4.1 Flash, thinking off 134 / 136 10.7 s 7.8 12.7 M 89% 310 k
GLM-5.3-Flash, thinking low 134 / 136 46.1 s 7.9 9.4 M 83% 222 k
Grok 4.7, thinking low 136 / 136 17.5 s 5.0 6.7 M 57% 134 k
gpt-6-luna, thinking medium 136 / 136 31.3 s — — — —
Muse Spark, thinking low 136 / 136 25.5 s 4.9 5.9 M 54% 311 k

Evaluator: Prime Intellect Verifiers v0.3.1 · angelX 0.1.6 · temperature 0; Muse Spark 1.0 (Meta's recommended setting); gpt-6-luna on its ChatGPT plan, which reports no token counts · 600 s per attempt · fresh environment per attempt · 2026-09-23

Harness comparison · polyglot-v1 · 2026-09-21

model harness solved wall (median) calls / task input tokens cache hit output tokens
DeepSeek V4.1 Flash, thinking off angelX 133 / 136 9.9 s 7.5 11.5 M 88% 257 k
OpenCode 1.18.31 132 / 136 8.6 s 12.5 67.7 M 97% 349 k
oh-my-pi 18.2.4 86 / 93 * 15.3 s 35.5 200.1 M 99% 866 k
GLM-5.3-Flash, thinking low angelX 135 / 136 50.7 s 8.9 10.9 M 84% 256 k
OpenCode 1.18.31 133 / 136 38.9 s 7.4 10.1 M 85% 179 k
oh-my-pi 18.2.4 133 / 136 37.7 s 8.9 24.1 M 90% 203 k
  • oh-my-pi on DeepSeek reached its 200M-token budget cap after 93 tasks.

The full GLM cohort took 34.5% longer than OpenCode in the September 21 comparison and 17.6% longer in the September 23 Angel run. The later gap is in model time; Angel's measured non-model time was lower. Verification alone does not explain the difference. See the full 136-task slowdown audit for totals, tails, task-level causes and comparison limits.

Tasks solved vs. cumulative agent time Pass/fail per attempt vs. cumulative agent time Output tokens and model calls per task, bars capped at 2.5k Seconds per attempt in run order, with medians Cumulative input tokens, cached and uncached Cache hit rate over the run

Evaluator: Prime Intellect Verifiers v0.3.1 · angelX 98d7340 · OpenCode 1.18.31 · oh-my-pi 18.2.4 · DeepSeek V4.1 Flash, thinking off · GLM-5.3-Flash, thinking low (lowest available) · temperature 0 · 8,192-token output cap · 600 s per attempt · fresh environment per attempt · 2026-09-21

Research and credits

Research and public work that informed Angel:

Code and tooling credits include OpenAI Codex, Grok CLI, oh-my-pi, DeepSeek-Reasonix, Hermes Agent, Prime Agent, Dotmax, ureq, Ratatui, Crossterm, rusty_v8 and V8. Early inspiration: DotAgents and SpeakMCP by aj47. File-tool interface references include Claude Code, aider and OpenHands. Attributions records implementation links, authors and retained licenses; research inspiration and incorporated code are identified separately.

Field results

Field proven: 100+ records on Yukon public research leaderboards, with first-place results in kernel optimization, LLM inference and cryptography research, and 2nd place on the GPU MODE Cholesky leaderboard. https://www.yukon.org/

About

Resources

Contributing

Security policy

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages