Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
-
Updated
Aug 11, 2026 - Python
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
Context & Guide For Reinforcement Learning with Verifiable Rewards with Large Language Models
Score the trustworthiness of outputs from any LLM in real-time
Write loops, not prompts. The loop is easy — the verifier is the whole game. VCN night one at Network School.
Open Arena: SLM verifiers across observability platforms, dataset and model catalogs, and value scenarios for LLM, agentic and harness evals.
An RL Enviorment for AES Inversion
A verifiers RLM environment for testing whether adaptive recursive search outperforms brittle manual RAG choreography on long synthetic corpora.
Reproducible verifier audits, datasheets, agreement metrics, and release gates
A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.
An open reinforcement-learning (RL) environment that trains LLM agents to use the current fact, not the stale one — verifiable reward for temporal fact-currency, built on verifiers / prime-rl (GRPO, LoRA).
Typed asset shapes + visual + headless views for AI agents. One asset definition. Three rendering targets (HTML / Markdown / Text).
A verifiers RL environment that trains models to propose novel, evidence-grounded, falsifiable hypotheses. Rewards novelty with accountability.
Verifiers hello world repo
Research proposal for verifier-gated on-policy distillation with explicit evaluation and claim boundaries
A verifiable RL environment for TRP ion-channel ligand pharmacology, built on Prime Intellect's Verifiers
Running, reproducing, and testing AI agent environments to understand how they work.
Verifiable RL environments for corporate law & governance — deterministic reward, no LLM judge, fully synthetic worlds.
Share the world, resample the consequences: deterministic record/replay of exogenous randomness for branching agent RL. 29x lower advantage variance with no loss of counterfactual fidelity.
Prime-RL / verifiers TSP environment (10-city hard config, lenient parser, eval-ready)
RL environment for AI coding agents: debug a stateful ledger with oracle grading, held-out scenarios, and adversarial soundness checks.
Add a description, image, and links to the verifiers topic page so that developers can more easily learn about it.
To associate your repository with the verifiers topic, visit your repo's landing page and select "manage topics."