Rebuild the docs site: contract and SDK reference, readable by AI crawlers - #46
Closed
ManulParihar wants to merge 55 commits into
Closed
ManulParihar wants to merge 55 commits into
ManulParihar wants to merge 55 commits into
Conversation
The deployed site was serving six template blog posts, their tag and archive pages, and a placeholder markdown page. That was 11 of the 39 URLs in the sitemap, none of them about Kokio. The blog is now off explicitly rather than commented out, so a future preset default cannot bring it back. Broken markdown links now fail the build instead of printing a warning nobody reads.
intro_deprecated, selfcustodial and smartwallet each repeated sections of intro.md or walletSuite.md word for word. contractSpecs listed ten contracts and had gone stale: the suite has eighteen. Two copies of a claim in one corpus is how a model ends up hedging between them.
Overview and setup both built to live URLs that nothing linked to. A page reachable only by guessing its URL still gets indexed, and it arrives at a reader stripped of the context the sidebar would have given it.
Both linked a .wiki.git clone URL, which git understands and a browser does not. The wikis themselves are fine, so this is just the suffix.
Three things were wrong in the emitted HTML. The home page shipped the template string "Description will go into a meta tag in <head />" as its description and "Hello from Kokio" as its title. The favicon had no file extension, so the page pointed at /images/KokioLogo and got a 404. The name and the descriptions now come from one module instead of being retyped. Repeated strings are what a model builds an entity record out of, so two wordings of the same claim are worse than one.
intro.md opened with a frontmatter block whose only line was a YAML comment, so the sidebar_position it looked like it set was never read. The sidebar is explicit anyway, so the block is gone rather than fixed. Also transperancy, utilies, enbaled, Contacts, and "an mobile".
Every entry carried changefreq weekly and priority 0.5, identical across all 39 URLs. Google ignores both, and a field with the same value everywhere carries no signal, so they are gone. lastmod replaces them, read per file from git. That needs full history, so CI checks out the whole thing now.
There was no robots.txt at all, so the site had no stated position and no sitemap pointer. Retrieval and training crawlers now both get everything, matching kokio.app. Seventeen agents are named individually. The wildcard is what actually grants access; the named groups make the decision auditable and make blocking any one of them a one-line edit later.
A doc page as HTML carries the sidebar, the navbar and a hydration payload, and its code blocks arrive wrapped in syntax-highlighting spans that damage copy-paste. The markdown is a fraction of the bytes and is what a reader actually wanted. /md/<path>.md is one page. /llms-full.txt is all 21 in one fetch, in sidebar order, which is the difference between one request and twenty-one for a question about the contract suite. One script writes both. Two generators over the same corpus drift, and nobody reads either file often enough to notice. Commented-out draft paragraphs are stripped. Two pages keep old text in HTML comments, and it must not reach a surface the website itself does not show.
One file naming every page and what it covers, so a model asking about the contract suite picks two fetches instead of crawling twenty-two. It says plainly that the contract pages are short and points at the source repository for the real reference. Overstating what is here would just produce confident answers built on a paragraph.
Docusaurus derives an anchor from the heading text, so rewording a heading silently breaks every link and citation pointing at it. Written down, the anchor survives the reword.
Koki'o and KOKIO both appeared in prose. Three spellings of one product split it into three weakly supported entities, so the variants now live only in schema, never in a sentence. Also: onchain as one word throughout, per the usual convention; the two commented-out draft paragraphs deleted rather than left to leak into the markdown routes; and the seamless/utilize/empowers filler replaced with what the sentence was actually claiming. All of this reaches llms-full.txt verbatim, so it is not only a style question.
Doc pages compete with StackOverflow and GitHub for the same questions. A typed description with a subject and a modified date gives a retriever grounds to prefer the page it came from. Every doc route now carries one TechArticle beside the breadcrumbs the theme already emitted, and the home page carries a WebSite node. The organisation is referenced by id from kokio.app rather than redefined, or the two hosts would describe two different companies. The subject varies by section. Labelling the eSIM explainer a Solidity document would be a claim a retriever acts on, and a wrong one. No id carries a fragment. A fragment in an @id is a citation target and has to match a real element, and nothing here has one.
Both lockfiles were committed, CI ran yarn and Vercel picked by lockfile, so the two could resolve different trees. One lockfile now, package-lock.json. README also drops the template's GitHub Pages deploy steps, which this repo has never used.
Doc pages referenced the WebSite node by id, but only the home page defined it. A crawler reads one page at a time, so on every page but one that reference pointed at nothing. Each doc page now carries the node in the same graph as its article, and both pages read it from siteCopy so they cannot drift apart. Doc pages also link their own markdown copy, so an agent does not have to guess the /md/ pattern.
The three feature cards were template copy. One offered PayPal, which is not a payment method here, and two sentences ended in a stray second full stop. Rewritten against the description the rest of the site uses.
Ranges meant a fresh install could resolve a different tree from the one that was tested. React moves to 19 and TypeScript to 5.9.3 on the same pass. clsx is gone. One call site passed it a static class string and the other joined two names, neither of which needs a package. React 19 also drops the global JSX namespace, so three return types become ReactNode. engines now says node 20, which is what Docusaurus 3.10 requires.
Seven things break silently: a page missing from the sitemap, a page with no markdown copy, a dead link in llms.txt, a schema id matching nothing, a wrong canonical, a robots group that skipped a disallow, and a missing lastmod, which is how a shallow clone in CI shows up. All of them still render fine. Runs as postbuild, so it fails the build that produced the problem.
Every new contract page lands under docs/contracts with a kebab-case name, so the existing pages move there first rather than sitting beside them in a second naming style for the rest of the phase. Old URLs are not redirected. Ten of them change.
Every address checked with cast against Base Sepolia before it went on the page: code present, recorded codehashes match, proxies point at the listed implementations, both beacons point at the listed wallets.
A page being written has no commit, so no date to read. Only a committed page missing its date means the clone was shallow, which is what the check is actually for.
Seventeen pages under /docs/sdk, ported from kokio-sdk v3.0.1. Mobile and backend split the way the package does, one page per surface. All 243 methods on the two entry points are covered, checked against the interface classes in the SDK source.
It is internal reference, not published SDK surface.
The site had eight one-sentence contract pages and nothing on payments, governance or account abstraction. Now every unit has a page, each with a written intro and the full generated reference under it. scripts/sync-reference.mjs keeps the generated half in step with the repo that produces it. It warns on every build and fails in CI. Contract docs parse as plain markdown now. The generated reference is full of text MDX reads as JSX.
Every page now opens by naming what it is and what Kokio is, carries a title and a description, and runs past the length where a retrieved chunk says nothing. The three shortest pages were the worst of it: resources was a bare link list, the solution page was one run-on paragraph, and the tech stack page stopped mid-sentence in the backend section. Page files are kebab-case now, matching the contract pages, so a file name and its URL always look the same. The eSIM folder, the two Overview pages and the three images move with them.
A page is retrieved as one chunk, without the sidebar around it. Under a hundred words there is nothing in that chunk worth answering from, and a page that never names the product leaves the reader guessing whose docs they landed in. Measured on the markdown output, because that is the surface agents read and frontmatter is stripped from it.
Rules that are cheap to state and expensive to notice: kebab-case file names, a title and a description on every page, one h1 with no skipped levels under it, an explicit id on every heading, and prose without em dashes, filler words, hyphenated onchain or an unescaped dollar sign. Runs as prebuild, so a file breaking one never reaches a build, and CI picks it up through npm run build without a new step.
A diagram is invisible to a screen reader and to every text surface this site generates, so both pages now carry alt text and a table saying what the picture says. Both tables carry more than the picture did: the eSIM page never mentioned the GSMA certificate chain, and the app page never said what the fork on new versus returning user actually changes. The conventions script now fails a build on an image with no alt text, or with alt text short enough to be a caption rather than a description.
The alt text listed a payment adapter that is not in the picture. The diagram predates both the payment and the governance contracts, so it now describes the fiat and crypto paths it does show, and a line underneath names the two contracts it is missing.
Links to kokio.app/manifesto, below Blogs.
Builds the index at build time and searches it in the reader's browser, so no query leaves the site and there is nothing to keep running. The search page holds no text of its own, so it stays out of the sitemap.
Reads the docs the site already publishes as text, picks the few pages closest to the question, and asks Gemini to answer from those alone. When the model is out of daily quota or down, the pages still come back, so the caller can fall back to plain search instead of showing an error.
Replaces the browser-side search index with a box that asks the endpoint. The index was 364 KB every reader downloaded to get word matching; the endpoint ranks the same pages server side and writes an answer as well. The answer check runs in prebuild, so a broken ranking fails the build.
The function used to fetch the site's own text file. A deployment behind Vercel's protection answers that with a login page, which parses into no pages at all, so every question came back as "nothing matches". It now reads the copy shipped beside it, and refuses to start on text that holds no pages rather than answering emptily. When the model declines, the reason goes to the logs instead of to the reader.
Runs the same code the endpoint runs, against the docs on disk or against a deployment, and names what came back instead of throwing a stack.
A preview reading its own address gets the protection login page. The live site is the one copy that always answers, so that is the fallback when no docs shipped in the bundle.
Gemini answers 503 when the model is under load, which happened on the second question anyone asked. It now tries again, then tries the smaller model, and only then falls back to listing pages. A bad key or a used-up quota is final, so those still return at once.
The model spends its own reasoning from the same token budget as the answer, so a 600 token cap ran out before the answer finished. The cap is now large enough, and reasoning is set low, which suits short answers drawn from pages that were already picked out. A truncated answer is still shown, and now says so in the logs.
The thinking option was sent under the wrong name, which the API rejects with a 400, so every question failed and looked like a broken key. The request body is now built in one exported place, with a test pinning each field name against Google's reference, and a rejected config is retried once without the optional tuning. A renamed field costs a slower answer, not a dead box. Retrying is limited to what retrying fixes. A busy model is tried again and then on the smaller model; a bad key or a spent quota returns at once instead of being repeated on every model. Each call has a timeout and the whole request has a budget, so nothing hangs. Ranking changed too. "No KYC" and "no personal information collected" appear on no page, only in the description the file opens with, which was being dropped, so questions about what Kokio collects matched nothing and returned overview pages. That description is now a page of its own, scoring is pivoted so a short page cannot outrank a specific one, and word endings are trimmed so "collected" finds "collects".
Requests from another site are refused, non-JSON bodies are refused, and a per-instance cap sits behind the per-reader one, so one busy instance cannot spend the day's quota on its own. Repeat questions are served from memory for ten minutes, which is most of them on a docs site, and each one saved is a request not spent.
Reading the docs by what people fail to find in them. Sent after the reader has their answer, so a slow webhook costs them nothing, and a failed post is logged rather than turned into a failed search.
A missing webhook returned early in silence, and a mistyped one answers 401 rather than throwing, so both left an empty channel and an empty log.
A function can be frozen the moment it responds, so the post started after that never finished and the channel stayed empty.
Green when the model answered, red when it did not, so a channel of these can be skimmed for the gaps. Pages move to the bottom as links to the live site rather than bare paths.
Quota is counted per model, so a spent limit on the first one says nothing about the second. Costs one extra call on a question that has already failed.
ArpitxGit
approved these changes
Sep 8, 2026
GuyPhy
approved these changes
Sep 8, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
The site was a Docusaurus template with a partial set of hand-written contract pages beside it. This turns it into the full reference for the contracts and the SDK, rewrites the existing pages so each one answers on its own, repaints the site in Kokio's own colours instead of the default template look, and adds the machine-readable surfaces that search engines and AI crawlers actually read. Build-time checks keep it that way.
Content
Contract reference, 16 new pages under
/docs/contracts. One page per contract, plus shared pages for errors, types, interfaces and deployed addresses. The 14 pages underdocs/SmartContracts/that these replace are deleted; several of them documented the same contract twice with different details.SDK reference, 16 new pages under
/docs/sdk, split into the mobile client and the backend client, since the two sign differently and take different arguments.Existing pages rewritten. Every page now names the product and says what it is about in its opening lines, and carries a
titleanddescription. Several were under 100 words and read as fragments of a page that was never there. The two diagrams that only existed as images now have the same information written out as text, including the certificate chain and the new-versus-returning-user payment split that the prose had never mentioned.Template content removed: the sample blog, the tutorial pages, and the Docusaurus artwork.
Machine-readable surfaces
Search and AI crawlers read this site more than people do, so:
/llms.txtand/llms-full.txt, and every page also served as plain markdown at/md/<path>.mdrobots.txtthat answers AI crawlers explicitly instead of by omissionvercel.jsonsets the content types, so/md/*.mdis served as markdown and not downloaded as a filelastmodfrom each page's last commit rather than the build timesrc/siteCopy.mjsholds the name, the description, the JSON-LD ids and the social links in one place, so the site cannot describe itself two waysColour theme and dark mode
The site carried the stock Docusaurus palette, which does not match kokio.app. Colours are now pulled from the marketing site's own design tokens (the beach-sky background, the cashmere orange, the ship-cove and outer-space greys) rather than approximated:
Header and footer links
Checks
Four scripts, wired into
npm run buildso nothing needs a separate CI step.check-conventions.mjsbuild-text.mjs/mdandllms-full.txtverify-site.mjssync-reference.mjsCI runs
sync-reference.mjs --strict, which fails rather than warns, because at that point a person is already looking at a diff.Tooling
yarn.lockdeletednpm overridespin patched versions ofqs,uuidandserialize-javascript. All three arrive through the dev server and the webpack plugins and never reach the built site, but they account for every moderate advisory on this repo.URL changes
All file names are now kebab-case, matching the contract pages.
docs/Overview.mdtodocs/overview.md,mobileApp.mdtomobile-app.md,docs/eSIM/todocs/esim/, and so on. Image files renamed the same way.check-conventions.mjsenforces this from here on.Follow-ups, not in this PR
image-sizestill has 17 open advisories with no fix published. It reads images from this repo at build time, so dismiss them rather than chase them.