Repository navigation
Keep the userguide fetch honest: production source, and a guard against fetching a shell - #231
Merged
Merged
Conversation
The website session applied my own suggestion back to me -- grep for the hostname rather than reason about which code "should" call production -- and found a second instance on their side in a pre-push hook. Doing the same here turned up twelve references to reactome.org, eleven of which are citation and display links that belong on production: a user-facing citation pointing at beta would be wrong. The twelfth actually fetches. src/data_generation/userguide/urls.py pulls ten pages when the userguide bundle is rebuilt. It stays on production, and the point of this commit is to write down why, because the opposite change is the obvious one to make by analogy. The MCP moved to beta.reactome.org on 2026-09-17 because a gate should not lean on the service it protects, and the two hosts answer identically there -- same release, same species count. That reasoning does not carry to the user guide. The guide documents the site people are actually using, and beta's can describe interface changes that have not shipped, so a bundle built from beta would confidently explain a UI the reader cannot see. Now overridable with REACTOME_USERGUIDE_BASE so a rebuild can be pointed elsewhere for testing, with the default left deliberate rather than incidental. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adam's instruction: nothing here should be making requests to reactome.org. This was the only remaining fetch -- ten pages when the userguide bundle is rebuilt -- and it now defaults to beta.reactome.org, matching the MCP move earlier today. It does not work yet, and that is worth being precise about rather than shipping quietly. Beta serves block-all-automation.conf, which blocks anything self-identifying as automation. This fetcher identifies itself honestly, as ReactomeChatbot/1.0 with a link to the repo, so beta returns 403 for all ten pages. Production carries no such config and returns 200. Checked both. The fix is an allowlist entry on the website side, not a change here -- inventing a browser User-Agent would get through and would be exactly the dishonesty their blocking exists to catch. Asked for on the website session. Until then a rebuild fails, and it fails with a message naming the User-Agent, the reason, and the override, because a bare 403 on an operation this rare would read as a missing page rather than an allowlist gap. The trade is recorded too: beta's guide can describe interface changes that have not reached production, so a bundle built from beta may explain a UI some readers cannot see. That is a content-freshness risk rather than a correctness one, and REACTOME_USERGUIDE_BASE makes either target one variable away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…he guide Moved this to beta on Adam's instruction that nothing here should call reactome.org, then measured what the alternatives return and moved it back. reactome.org/userguide 2,372 words server-rendered Joomla beta.reactome.org/userguide 1,245 words Angular shell; the text is CSS 127.0.0.1:4200/userguide 1,237 words the same Angular app, internally beta and the internal route are the same application and it renders the guide in the browser, so a plain HTTP fetch of either returns font declarations rather than documentation. The installed Release95 bundle was built from the Joomla page and carries the same 2,372 words, which is what makes that comparison meaningful rather than a guess about which number is right. A bundle built from beta would therefore contain stylesheets instead of the user guide, and would pass every structural check on the way: ten files, right embedding dimensions, non-zero document count. It would only be discovered by someone asking the chatbot how to use the pathway browser and getting nothing. The load this avoids is ten requests, a few times a year, only when the bundle is rebuilt. The MCP move earlier today was the one that mattered -- that was per-question traffic on every deploy. The table is in urls.py rather than this message, because the next person to try this will read the module and not the log. It stops being true if the guide is ever server-rendered somewhere else, or the Angular app exposes the content over an API, so it says to re-measure rather than trust it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both sessions looked at /userguide on beta, saw 200, and concluded it was usable. It is a single-page app: the guide is assembled in the browser, so a plain fetch returns inlined CSS. A bundle built from it would have contained stylesheets, passed every check we have -- ten files, right embedding dimensions, non-zero document count -- and answered nothing, surfacing months later as a user asking how to use the pathway browser and getting nothing. So the check is on content, not on the response. The first version of this guard did not work, and its own test caught it. Stripping tags leaves the CSS *inside* <style>, so the shell measured about 1,240 words -- more than the thinnest real page has of prose, and comfortably over the floor meant to exclude it. I had written "a raw count cannot separate them" in the comment and then set a count that could not separate them. Removing script and style contents before counting is what makes it work. Verified against the real pages rather than the synthetic ones: production 746 visible words, the internal Angular app 1. Four tests, including that the shell is large and returns 200 so only the visible text gives it away, and that a rejected page is not cached -- caching one would make the next run succeed against bad content. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Started as "stop hitting reactome.org", became a lesson in what a 200 is worth.
What actually still hit production
One thing: the userguide bundle rebuild — ten pages, a few times a year. The MCP moved to beta earlier today and that was the traffic that mattered, being per-question on every deploy and sweep run. Everything else mentioning
reactome.orgis a citation or display link that users click; we never request them.Why it stays on production
I moved it to beta, then measured what the alternatives return. Same User-Agent, tags stripped:
reactome.org/userguidebeta.reactome.org/userguide127.0.0.1:4200/userguidebeta and the internal route are the same application, and it renders the guide in the browser. Our installed Release95 bundle carries exactly 2,372 words per page — built from Joomla — which is what makes this a comparison rather than a preference.
A bundle built from either non-production source would contain stylesheets instead of documentation, and would pass every check we have on the way.
The guard
Since a 200 clearly isn't evidence, the check is on content. And the first version didn't work — its own test caught it.
Stripping tags leaves the CSS inside
<style>, so the shell measured ~1,240 words: more than the thinnest real page has of prose, and comfortably over the floor meant to exclude it. I'd written "a raw count cannot separate them" in the comment and then set a count that couldn't.Removing script and style contents before counting is what makes it work. Verified against the real pages, not the synthetic ones:
Four tests, including that the shell is large and returns 200 so only visible text gives it away, and that a rejected page isn't cached — caching one would make the next run succeed against bad content.
A better source exists, and it isn't in this PR
The website session pointed at
/site-search-index.json(static, internal, 30,379 words across 11 userguide pages) and at the.mdxfiles those are generated from, inreactome/WebsiteAngular. I checked both:.mdxsourceThe JSON has more text but zero structure — no newlines, no headings — and our loader chunks by section and records
section_title, which is how answers cite a specific part of a page. The.mdxkeeps both and needs no host at all.That's the right destination and it's a real change — a new loader, a cross-repo path, and a before/after retrieval measurement per Principle II. It deserves its own spec rather than being folded in here.
333 tests, ruff and mypy clean.
🤖 Generated with Claude Code