A provenance-complete knowledge base architecture for humans and AI agents. A Promptotyping profile for evidence-grounded knowledge work.
This repository is a template. Instantiate it for a project via SETUP.md. The full concept lives in docs/concept.md, and what changed for existing instances is in CHANGELOG.md.
The instance under paper/ applies the architecture to its own internal method manuscript.
Output generated by language models is unauditable in its raw form. A generated report may name its sources, yet its structure offers no way to check an individual statement against the passage that supposedly supports it. Grounded Vault is a repository architecture in which every substantive statement carries a machine-resolvable anchor to the source material that supports it, from the finished output down to the individual passage, dataset row or computation. AI agents produce this structure at scale, and the architecture never asserts that its content is true. A finished vault is a fully prepared audit object in which human expert review can proceed passage by passage.
uv sync # or: pip install pyyaml pytest
python tools/validate.py . # conformance check of the whole vault
python tools/validate.py . --chapter 40_output/01-findings # one chapter and its chain
python -m pytest tests # test suite of all tools
Then follow SETUP.md, which lists the parameters an instance sets and carries a copy-paste prompt for the first agent session that fills them.
The numbered chain runs 00_sources → 10_markdown → 20_distillates → 30_assertions → 40_output, and each layer is checkable on its own. The rules that bind the layers to each other are set out in knowledge/schema.md § Layer model.
The three source types document, publication and data, and the criterion that assigns a source to one of them, are set out in knowledge/schema.md § Source types.
The five numbered folders form the production chain, each one holding the layer that the folder name announces. What each layer and its parts are is defined once in knowledge/index.md § Terminology.
00_sources/holds the sources, the originals as they arrived. Whether an original is committed is set out inknowledge/operations.md§ Acquire.10_markdown/holds the Markdown representations, the full texts indocuments/and the datasets plus their schema description indata/.20_distillates/holds the distillates, one per source, in one subfolder per source type.30_assertions/holds the assertions, one file each, together with one topic map (MOC-*.md) per topic of the controlled topic set.40_output/holds the output, one or more documents such as a report, proposal, thesis or paper, kept as one file per chapter.
The remaining folders lie across the chain rather than inside it.
knowledge/is the governance layer, six project knowledge documents in the Promptotyping sense, holding terminology, parameters, schema, procedures, volatile state and the decision history.references/holds the bibliographic records of citable-only sources as CSL JSON, one array per file, and is needed only while the source typepublicationis active.glossary/holds one file per central technical term of the content, serving as definition, wikilink hub and tag keyword, and is filled as the need arises.tools/holds the validatorvalidate.py, the source inventory generatorinventory.py, the machine review toolreview.py, the migration toolmigrate.pyfor vaults built before the template chain, and the project page generatorbuild_docs.py. An instance with data sources addstools/analysis/, the one folder from which data anchors re-run their deterministic scripts, one script per task.tests/holds the pytest suite of all tools, with the fixture vaults the validator tests run against intests/fixtures/, a minimal conformant vault and a deliberately broken one that carries one specimen per finding class. A coverage test holds every finding code the validator emits against those specimens, so a check added without one fails the suite.docs/holds the concept paper, the generated project pageindex.htmland a dated field report from an instance, all addressed to the template rather than to any instance, so an instance may delete the folder..github/holds the CI workflowchecks.yml, which lints the repository, validates the template and thepaper/instance and runs the test suite on every push and pull request..claude/holds the harness-specific skills that the action layerCLAUDE.mdroutes into, exchangeable together with that file.
docs/index.html is generated from README.md, docs/concept.md, knowledge/index.md, knowledge/schema.md and knowledge/operations.md by python tools/build_docs.py --date <date> and is never edited by hand.
Three instances check the vault, with strictly separated authority. Validation (tools/validate.py) judges resolvability and form, machine review by a language model judges whether a source location supports the statement built on it, and human verification alone establishes evidence. A status is only ever set by a check that actually ran, and a document never stands higher than the anchors it rests on. The terms are defined in knowledge/index.md § Terminology. The contracts, the finding codes, the chapter scope and the status discipline are set out in knowledge/operations.md § Check.
Humans read the vault in Obsidian, following wikilinks from an output footnote down to the supporting passage. Agents enter through CLAUDE.md, an imperative action layer that routes every task onto the declarative rule documents in knowledge/. The Markdown stays portable, and beyond wikilinks and block references no plugin-specific syntax is used.
- Create a repository from this template.
- Follow
SETUP.mdto set the project parameters (purpose, controlled topic set, active source types, output genre, language, verification role, check mechanisms, harness rules), or hand its first-session prompt to an agent. - Run
python tools/validate.py .on every change, andpython -m pytest testswhen you touch the tools. How errors and warnings count is set out inknowledge/operations.md§ Check.
The whole repository is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0), see LICENSE. Third-party research data is excluded from these terms, and rights remain with their respective holders.