Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
d988e5f
chore: add .gitignore, remove archives/backups/pyc from tracking
samfrons Aug 25, 2026
5c5e64e
docs: archive fine-tune/audit prototype to research/, add honest post…
samfrons Aug 25, 2026
abba316
docs: rewrite root README for the agentic rebuild
samfrons Aug 25, 2026
434aac5
feat: scaffold Next.js app with AI SDK deps
samfrons Aug 25, 2026
2ad8162
docs: add ingestion and evals scaffold READMEs
samfrons Aug 25, 2026
be516a7
Keep corpus PDFs out of git, add reproducible fetch script
samfrons Aug 25, 2026
ad76d08
Add standards_chunks schema and hybrid search function
samfrons Aug 25, 2026
79e3009
Add HAI chat backend: system policy, tools, and streaming chat route
samfrons Aug 25, 2026
2c254bc
Add role-based AI enablement playbooks
samfrons Aug 25, 2026
9bd4a39
Add enablement guides and content philosophy README
samfrons Aug 25, 2026
79ed9c5
Run HAI on local Ollama and build the chat UI
samfrons Aug 25, 2026
4501db5
Add corpus ingestion pipeline on a local Ollama + Supabase stack
samfrons Aug 25, 2026
254ecd5
Stop the model inventing citations when retrieval returns nothing
samfrons Aug 25, 2026
10bcf30
Document ingestion usage and the search_standards_hybrid RPC contract
samfrons Aug 25, 2026
f55abdc
Grant service_role its own table privileges; skip the speed probe whe…
samfrons Aug 25, 2026
5286d0a
Let filter_source name a document family so the app's three sources work
samfrons Aug 25, 2026
6f5e9e9
Fit HAI inside Ollama's context window
samfrons Aug 25, 2026
108e5a4
Size chunks to the embedding model's 512-token window
samfrons Aug 25, 2026
c77a5bc
Bound chunks below the embedding window; record the loaded corpus
samfrons Aug 25, 2026
8c020da
Add playbooks/guides UI, try-in-chat, and coach mode
samfrons Aug 25, 2026
9ba2b6f
Replace the PII screen's estimated latency with measured numbers
samfrons Aug 25, 2026
c8d93fb
Add strategy and AI-enablement framework docs
samfrons Aug 25, 2026
9688191
Add the eval harness: live route, independent judge, honest report
samfrons Aug 25, 2026
934c007
Guard streaming turns by stall, not total duration; surface tool use
samfrons Aug 25, 2026
1a16dcb
Add multilingual UI chrome (EN/FR/AR/ES) with RTL support
samfrons Aug 25, 2026
2ac65df
Let an interrupted run reuse the transcripts it already captured
samfrons Aug 25, 2026
fe6d6e1
README accuracy pass, quickstart, demo script, screenshots
samfrons Aug 25, 2026
d7add62
Embed playbooks-index screenshot in README
samfrons Aug 25, 2026
455dcad
Give the judge its own timeout; 6 minutes was measuring our patience
samfrons Aug 25, 2026
bb0ca8b
Fix ingestion drift, write real app README, add retrieval retry
samfrons Aug 25, 2026
69b34d3
Keep tool calls that fail input validation
samfrons Aug 25, 2026
2a9b10a
Publish the first honest smoke report: 0 pass, 3 fail
samfrons Aug 25, 2026
18ee81c
Add corpus candidate research for phase C1 (areas A-D)
samfrons Aug 30, 2026
b9ada72
Add hosted deployment mode alongside local, on free tiers only
samfrons Aug 30, 2026
0ec339b
Print only the host when loading the corpus, not the password
samfrons Aug 30, 2026
c11d0d3
Stop a reasoning model's own output from breaking step two of the too…
samfrons Aug 30, 2026
fb6954b
Vendor content/ into app/ so a deploy from app/ ships the guides
samfrons Aug 30, 2026
0392ce0
Redesign UI in Swiss modern style with custom icon system
samfrons Aug 30, 2026
0ae6750
Speed up hosted chat and add live processing indicators
samfrons Sep 1, 2026
875c13e
Publish full 26-scenario baseline: 1 pass / 2 partial / 23 fail
samfrons Sep 1, 2026
9727379
Require tools before any factual claim, and correct the user's figures
samfrons Sep 1, 2026
ac0777e
Make humanitarian_data say why it has no figure instead of returning []
samfrons Sep 1, 2026
3e8169d
Say when each tool is mandatory, and why an empty search is not permi…
samfrons Sep 1, 2026
178dbe2
Put the grounding rule first, where the local model actually reads it
samfrons Sep 1, 2026
4a706be
Add MIT license (from claude/hai-repo-cleanup-wmrizq)
samfrons Sep 1, 2026
086f8ad
Port fixed auditor + tests from cleanup branch into research archive
samfrons Sep 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file removed .DS_Store
Binary file not shown.
114 changes: 114 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,114 @@
# Continuous integration for HAI.
#
# What this proves: the app lints, typechecks, and passes its unit tests, and
# the two Node side-projects (ingestion, evals) typecheck.
#
# What this deliberately does NOT do: run the eval suite. The evals need a local
# Ollama serving a ~9GB target model and a ~5GB judge model, and a single run
# takes hours on a laptop. A hosted runner has neither, so an eval job here
# could only ever be a stub — and a green "evals passing" badge that never
# graded a transcript is precisely the kind of reassuring-but-meaningless number
# this project already published once and had to retract (see research/README.md).
# Eval numbers come from `pnpm eval` on a machine with the models, and land in
# evals/reports/ with the judge's evidence attached so they can be checked.

name: CI

on:
push:
branches: [main]
pull_request:

concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true

env:
NODE_VERSION: '22'
PNPM_VERSION: '10.13.1'

jobs:
app:
name: app — lint, typecheck, test
runs-on: ubuntu-latest
defaults:
run:
working-directory: app
steps:
- uses: actions/checkout@v4

- uses: pnpm/action-setup@v4
with:
version: ${{ env.PNPM_VERSION }}

- uses: actions/setup-node@v4
with:
node-version: ${{ env.NODE_VERSION }}
cache: pnpm
cache-dependency-path: app/pnpm-lock.yaml

- run: pnpm install --frozen-lockfile

# next-env.d.ts and the generated route types are gitignored, so they have
# to be produced before tsc can see them. `typegen` does that without
# paying for a full build.
- name: Generate Next.js types
run: pnpm exec next typegen

- name: Lint
run: pnpm lint

- name: Typecheck
run: pnpm exec tsc --noEmit

- name: Unit tests
run: pnpm test

evals:
name: evals — typecheck
runs-on: ubuntu-latest
defaults:
run:
working-directory: evals
steps:
- uses: actions/checkout@v4

- uses: pnpm/action-setup@v4
with:
version: ${{ env.PNPM_VERSION }}

- uses: actions/setup-node@v4
with:
node-version: ${{ env.NODE_VERSION }}
cache: pnpm
cache-dependency-path: evals/pnpm-lock.yaml

- run: pnpm install --frozen-lockfile

- name: Typecheck
run: pnpm typecheck

ingestion:
name: ingestion — typecheck
runs-on: ubuntu-latest
defaults:
run:
working-directory: ingestion
steps:
- uses: actions/checkout@v4

- uses: pnpm/action-setup@v4
with:
version: ${{ env.PNPM_VERSION }}

- uses: actions/setup-node@v4
with:
node-version: ${{ env.NODE_VERSION }}
cache: pnpm
cache-dependency-path: ingestion/pnpm-lock.yaml

- name: Install
run: pnpm install --frozen-lockfile

- name: Typecheck
run: pnpm typecheck
83 changes: 83 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# --- Python ---
__pycache__/
*.py[cod]
*$py.class
*.so
.Python
build/
dist/
*.egg-info/
.eggs/
.venv/
venv/
env/
.pytest_cache/
.mypy_cache/
.ruff_cache/
*.ipynb_checkpoints

# --- Node / Next.js ---
node_modules/
.next/
out/
.turbo/
.vercel/
*.tsbuildinfo
next-env.d.ts
coverage/

# --- macOS ---
.DS_Store
.AppleDouble
.LSOverride

# --- Env files (never commit secrets) ---
.env
.env.local
.env.*.local
.env.development
.env.production
!.env.example
!*.env.example

# --- Corpus source documents (licensing) ---
# The Sphere Handbook is all-rights-reserved: local educational/research use is
# permitted, redistribution is not. Other corpus PDFs are looser but are kept out
# of git for consistency. Re-download them with ingestion/fetch-corpus.sh.
# Provenance and per-file licenses stay tracked in ingestion/corpus/SOURCES.md.
ingestion/corpus/*.pdf

# --- Ingestion pipeline artifacts ---
ingestion/.extract-cache/

# --- Deployment seed (licensing) ---
# scripts/deploy/export-corpus.sh writes the extracted chunk text here to seed a
# cloud database. That text *is* the source documents, so committing it would
# redistribute the Sphere Handbook exactly as publishing the PDF would — the
# same restriction that keeps ingestion/corpus/*.pdf out of git. Regenerate it
# from a local ingest instead; it takes seconds.
supabase/seed/*.csv.gz

# --- Archives / backups ---
*.zip
*.tar.gz
*.json.backup_*

# --- Logs ---
*.log
npm-debug.log*
pnpm-debug.log*
yarn-debug.log*
yarn-error.log*

# --- Editor ---
.vscode/
.idea/
*.swp

# --- Vercel ---
# Added by `vercel link`, which also appends a bare `.env*`. That pattern is
# kept narrower here: it sits after the `!.env.example` negation above and would
# otherwise re-ignore the example files the repo deliberately tracks.
.vercel
.env.local
21 changes: 21 additions & 0 deletions LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2025 Sam Frons

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
178 changes: 176 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,2 +1,176 @@
# HAI
Humanitarian AI bot
# HAI — Humanitarian AI Assistant

An agentic, citation-grounded AI assistant for humanitarian work. It answers
questions against primary references — the Sphere Handbook, the Core
Humanitarian Standard, and IASC guidance — and calls live tools for current
crisis data and structured humanitarian datasets, rather than relying on
facts memorized into model weights.

## What it is

- **Retrieval-grounded, not fine-tuned.** Answers cite the source passage
they're grounded in. See [`research/README.md`](research/README.md) for
why an earlier fine-tuning approach was abandoned.
- **Tool-using.** The assistant searches humanitarian standards, pulls
current crisis updates, and queries structured humanitarian data — it
doesn't just generate from a static prompt.
- **Runs entirely on a local model by default.** Chat and embeddings both go
through local Ollama; no key and no per-token cost unless you deliberately
point `LLM_BASE_URL` at a hosted endpoint.
- **Safety layer.** Every message is screened for personally identifiable
data before it reaches the model, refused with the IASC principle it
engages named explicitly, and offered a safe rephrasing — not a bare
"I can't help with that." 114 tests cover it.
- **Evaluated by an independent judge.** A held-out set of 26 domain
scenarios (`petri/seeds/humanitarian_test_scenarios.json`) is graded by a
model from a different family than the one being tested, against explicit
expected facts — not a self-graded, keyword-ratio heuristic. See
[`evals/README.md`](evals/README.md).

![HAI chat, empty state](docs/assets/chat-empty-en.png)

## Architecture

```mermaid
flowchart LR
UI["Next.js UI\n(chat, playbooks, guides)"] --> API["/api/chat\nAI SDK v7"]
API --> Safety["Safety layer\nPII interception"]
Safety --> LLM["Local Ollama\nqwen2.5:14b\n(swappable via env)"]
LLM --> T1["search_standards\nSupabase pgvector\nhybrid search"]
LLM --> T2["crisis_updates\nIFRC GO / ReliefWeb"]
LLM --> T3["humanitarian_data\nHDX HAPI"]
LLM --> I18n["i18n: EN / FR / AR / ES\n(RTL for Arabic)"]
```

- **UI**: Next.js app (`app/`) — chat, six role playbooks, three guides, and
a coach mode that adds a short prompting lesson to each answer.
- **`/api/chat`**: AI SDK v7 route. The model is local `qwen2.5:14b` by
default (`ollama create hai-qwen2.5 -f app/ollama/Modelfile` bakes in the
16k context window HAI needs); `LLM_BASE_URL`/`LLM_MODEL`/`LLM_API_KEY`
point it at any OpenAI-compatible endpoint instead, with no code change.
- **Safety layer**: deterministic regex/heuristic screening
(`app/src/lib/safety/pii.ts`) runs on every message before it reaches the
model — phone numbers, emails, case/registration identifiers, GPS
coordinates, dates of birth, pasted rosters. An optional second-pass LLM
screen (`PII_LLM_SCREEN=true`) catches bare names, off by default because
it costs a full model round-trip per message.
- **Tools**:
- `search_standards` — hybrid (vector + full-text, reciprocal rank fusion)
search over 1,631 ingested chunks of humanitarian standards in a local
Supabase pgvector instance.
- `crisis_updates` — live situation reports. Prefers ReliefWeb, which
since November 2025 requires an OCHA-approved `appname`; without one
(the default) it falls back to **IFRC GO** (no registration required)
and tells the model, and the user, which source answered.
- `humanitarian_data` — structured country indicators from **HDX HAPI**
(population, food security, funding, humanitarian needs). No key
required.
- **i18n**: UI chrome in English, French, Arabic, Spanish, with RTL layout
for Arabic. The model answers in whatever language the user writes in,
independent of the UI locale.
- **Eval harness** (`evals/`): drives the 26-scenario suite against the live
`/api/chat` route and grades transcripts with a judge model from a
different family than the target — run outside the serving path, not on
every request.

## Quickstart

Everything below runs locally with no paid API calls. Expect **10–15
minutes** if the models are already pulled, longer for the first `ollama
pull` (multi-GB) and the first corpus ingestion (embedding ~1,600 chunks
takes about 6 minutes once the corpus is fetched).

**Prerequisites**

- [Ollama](https://ollama.com), running (`ollama serve`)
- [Docker](https://docs.docker.com/get-docker/) (Supabase's local stack runs in it)
- [pnpm](https://pnpm.io) (`corepack enable` gets you `pnpm@10.13.1`, pinned in `app/package.json`)
- [Supabase CLI](https://supabase.com/docs/guides/local-development/cli/getting-started)

**1 — Pull the models and build the HAI variant**

```bash
ollama pull qwen2.5:14b
ollama pull mxbai-embed-large
ollama create hai-qwen2.5 -f app/ollama/Modelfile # bakes in the 16k context HAI needs
```

**2 — Fetch the corpus** (Sphere Handbook, CHS, three IASC guidance docs — not committed; see licensing note in `ingestion/README.md`)

```bash
cd ingestion
./fetch-corpus.sh # downloads and verifies sha256 against known-good copies
```

**3 — Start Supabase and apply migrations**

```bash
cd .. # repo root
supabase start # applies supabase/migrations/ automatically; prints keys
```

Note the `API_URL` (`http://127.0.0.1:54421`) and `SERVICE_ROLE_KEY`/`ANON_KEY`
it prints — the default ports are shifted +100 from Supabase's usual range
because another local project already holds `5432x` on this machine.

**4 — Ingest the corpus**

```bash
cd ingestion
pnpm install
cp .env.example .env # fill in SUPABASE_SERVICE_ROLE_KEY from step 3
pnpm ingest # extract -> chunk -> contextualize -> embed -> load
pnpm search "minimum water supply per person per day" # sanity check
```

**5 — Configure and run the app**

```bash
cd ../app
pnpm install
cp .env.example .env.local # fill in SUPABASE_ANON_KEY from step 3
pnpm dev
```

Open `http://localhost:3000`. `RELIEFWEB_APPNAME` and `HDX_APP_IDENTIFIER`
are optional — both tools work without them (see Architecture above).

Full detail on each ingestion stage, the RPC contract, and known gaps (the
corpus currently loads without contextual-retrieval preambles — retrieval
works, the ranking boost doesn't yet) is in
[`ingestion/README.md`](ingestion/README.md).

| ![PII interception banner](docs/assets/pii-safety-notice.png) | ![Arabic RTL layout](docs/assets/chat-arabic-rtl.png) | ![Playbooks index](docs/assets/playbooks-index.png) |
|---|---|---|
| Safety layer: a case-detail paste caught before it reaches the model, with the IASC principle named | Arabic locale — full RTL layout, not just translated labels | The six role playbooks |

## Docs

| Doc | What's in it |
|---|---|
| [`docs/STRATEGY.md`](docs/STRATEGY.md) | Why this architecture — the case for grounding over fine-tuning |
| [`docs/ENABLEMENT.md`](docs/ENABLEMENT.md) | How an organization adopts AI well; the framework the playbooks/guides implement |
| [`docs/DEMO.md`](docs/DEMO.md) | 5-minute demo script, beat by beat |
| [`research/README.md`](research/README.md) | Honest postmortem: the original fine-tune + audit prototype, and the three bugs that invalidated its results |
| [`content/playbooks/`](content/playbooks/) | Six role-specific playbooks (program, protection, MEAL, comms, grants, logistics) |
| [`content/guides/`](content/guides/) | Effective prompting, responsible use, starting a community of practice |
| [`evals/README.md`](evals/README.md) | How the eval harness works and why the judge is trustworthy |

## Repo layout

| Path | Contents |
|---|---|
| `app/` | Next.js application (UI + `/api/chat` route, safety layer, i18n) |
| `ingestion/` | Corpus → chunking → embeddings → Supabase pgvector pipeline |
| `evals/` | Eval harness driving `petri/seeds/` scenarios against the live app |
| `petri/seeds/` | The 26-scenario eval suite (kept as-is from the prototype) |
| `content/` | Playbooks and guides shown in the app and used by the eval scenarios |
| `research/` | Archived fine-tune/audit prototype + honest postmortem — **start here** for context on why the project pivoted: [`research/README.md`](research/README.md) |

## Prior work

The original prototype attempted a LoRA fine-tune of a local model,
validated by a Petri-style auditor. Three bugs invalidated its results (a
self-judging auditor, a training-data extractor that scraped source code
instead of prose, and a passing threshold that let empty answers through).
Full writeup, with file:line references: [`research/README.md`](research/README.md).
Loading
Loading