-
Notifications
You must be signed in to change notification settings - Fork 0
Grounding hardening, live-hazard tools, agentic deliverables + trace #3
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
1299462
2db9788
37fab57
c9ce425
bcef85f
07ef792
8016563
f27ecc3
7fc827c
5c9f454
75c2363
20164dd
911c0ee
a5c803d
5cb8e94
d4a317e
fbf65ee
42e0563
530dbe7
9654b92
80f2f61
223af1d
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -26,6 +26,27 @@ LLM_BASE_URL=http://localhost:11434/v1 | |
| LLM_MODEL=qwen2.5:14b | ||
| LLM_API_KEY=ollama | ||
|
|
||
| # Optional. The model /api/deliverables uses; defaults to LLM_MODEL. | ||
| # | ||
| # Worth setting on a hosted deployment, because Groq's token buckets are per | ||
| # model — both the per-minute one and the per-day one. Verified against the | ||
| # deployed key: qwen/qwen3.8-27b and openai/gpt-oss-120b each reported their own | ||
| # independent 8,000 tokens/minute and decremented separately. | ||
| # | ||
| # A situation brief costs roughly 25,000 tokens against a 200,000/day ceiling, | ||
| # so sharing one model with chat means a busy chat afternoon quietly consumes | ||
| # the ability to produce a document. Pointing deliverables at a second model | ||
| # gives the two features independent daily budgets on the same key. | ||
| # | ||
| # LLM_DELIVERABLES_MODEL=openai/gpt-oss-120b # hosted demo's value | ||
| LLM_DELIVERABLES_MODEL= | ||
|
|
||
| # Optional. Only if deliverables should use a different provider entirely | ||
| # rather than a second model on the same one. Each falls back to its LLM_* | ||
| # equivalent. | ||
| LLM_DELIVERABLES_BASE_URL= | ||
| LLM_DELIVERABLES_API_KEY= | ||
|
Comment on lines
+47
to
+48
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
rg -n -C 3 'LLM_DELIVERABLES_(BASE_URL|API_KEY)|apiKey:.*base\.apiKey|baseUrl:.*base\.baseUrl' \
app/src/lib/llm/provider.ts app/.env.exampleRepository: samfrons/HAI Length of output: 1773 Sensitive Data Exposure (CWE-200): Exposure of Sensitive Information to an Unauthorized Actor Reachability: Internal · Exploitability: Difficult Require a dedicated key when When 🧰 Tools🪛 dotenv-linter (4.0.0)[warning] 48-48: [UnorderedKey] The LLM_DELIVERABLES_API_KEY key should go before the LLM_DELIVERABLES_BASE_URL key (UnorderedKey) 🤖 Prompt for AI Agents |
||
|
|
||
| # IMPORTANT for local Ollama: the 4096-token default context is too small for | ||
| # HAI, and overflowing it silently drops the system prompt — after which the | ||
| # model loses its grounding and language rules mid-answer. Either build the | ||
|
|
@@ -82,8 +103,30 @@ HDX_APP_IDENTIFIER= | |
| # RATE_LIMIT_RPM paces one client: a per-IP sliding window, default 20/min. It | ||
| # lives in process memory, so on Vercel each serverless instance keeps its own | ||
| # copy and it resets on every deploy — a politeness control, not a spend control. | ||
| # | ||
| # It governs /api/chat only. /api/deliverables keeps its own, much tighter | ||
| # counter (3 runs per 10 minutes, not configurable), because one run is about | ||
| # nineteen model calls rather than one. | ||
| RATE_LIMIT_RPM= | ||
|
|
||
| # LLM_TOKENS_PER_MINUTE no longer needs setting on an ordinary hosted | ||
| # deployment, and its meaning has narrowed. | ||
| # | ||
| # /api/deliverables now reads x-ratelimit-* off every response (see | ||
| # src/lib/llm/rate-limit.ts) and paces against the endpoint's own numbers rather | ||
| # than a configured guess, so Groq's free tier and a paid tier are both paced | ||
| # correctly with no configuration. Pacing switches itself on when an endpoint | ||
| # reports a ceiling and stays off when none is reported — a local Ollama is | ||
| # never paced, and no URL check decides that. | ||
| # | ||
| # What this variable is for now is the endpoint that enforces a ceiling but does | ||
| # not report it in headers: set it to that ceiling in tokens per minute. A | ||
| # reported limit always wins over it. Leave it unset otherwise. | ||
| # | ||
| # Only the deliverables engine reads it; chat is one call at a time and never | ||
| # approaches the limit. | ||
| LLM_TOKENS_PER_MINUTE= | ||
|
|
||
| # MAX_DAILY_REQUESTS is the spend control: one shared counter in Postgres, | ||
| # default 500/day across the whole deployment, enforced atomically so concurrent | ||
| # instances cannot overshoot it. Requires the daily_request_cap migration. | ||
|
|
||
Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Add
/deliverablesto the architecture diagram.The changed UI node lists only chat, playbooks, and guides, but this PR adds
/deliverablesand/api/deliverables. The root architecture description is now incomplete. Add the new workflow surface to the UI node.Proposed documentation update
📝 Committable suggestion
🤖 Prompt for AI Agents