Popular repositories Loading
-
stack-eval
stack-eval PublicOfficial implementation for the paper, StackEval: Benchmarking LLMs in Coding Assistance, https://arxiv.org/abs/2412.05288
Repositories
Showing 10 of 21 repositories
- vending-bench Public
An open source vending machine benchmark: a long-horizon agentic business simulation served over MCP, shipped as a Harbor task.
- Murphy Public
- Compass Public
Automatically generates tool-routing prompts for LLM agents, optimizing quality and cost with limited labeled data.
- robot-teleop Public
- prism Public
Give Claude Code persistent memory — captures patterns from your sessions, validates them with AI, and surfaces team knowledge as reusable skills
- prism-registry-demo Public
- sim-starterkit Public
-
- autoresearch-edu Public Forked from fjfok/autoresearch-edu
Educational demo of an AutoResearch ratchet loop using FLAML AutoML on tabular classification benchmarks
Top languages
Loading…
Most used topics
Loading…