Skip to content

Rebuild the docs site: contract and SDK reference, readable by AI crawlers - #46

Closed
ManulParihar wants to merge 55 commits into
mainfrom
upgrade/ai-seo
Closed

ManulParihar wants to merge 55 commits into
mainfrom
upgrade/ai-seo

Conversation

@ManulParihar

@ManulParihar ManulParihar commented Sep 8, 2026 •

Copy link
Copy Markdown
Member

What this does

The site was a Docusaurus template with a partial set of hand-written contract pages beside it. This turns it into the full reference for the contracts and the SDK, rewrites the existing pages so each one answers on its own, repaints the site in Kokio's own colours instead of the default template look, and adds the machine-readable surfaces that search engines and AI crawlers actually read. Build-time checks keep it that way.

Content

Contract reference, 16 new pages under /docs/contracts. One page per contract, plus shared pages for errors, types, interfaces and deployed addresses. The 14 pages under docs/SmartContracts/ that these replace are deleted; several of them documented the same contract twice with different details.

SDK reference, 16 new pages under /docs/sdk, split into the mobile client and the backend client, since the two sign differently and take different arguments.

Existing pages rewritten. Every page now names the product and says what it is about in its opening lines, and carries a title and description. Several were under 100 words and read as fragments of a page that was never there. The two diagrams that only existed as images now have the same information written out as text, including the certificate chain and the new-versus-returning-user payment split that the prose had never mentioned.

Template content removed: the sample blog, the tutorial pages, and the Docusaurus artwork.

Machine-readable surfaces

Search and AI crawlers read this site more than people do, so:

  • /llms.txt and /llms-full.txt, and every page also served as plain markdown at /md/<path>.md
  • JSON-LD on each page, with the site and organisation nodes defined once and referenced rather than repeated
  • robots.txt that answers AI crawlers explicitly instead of by omission
  • vercel.json sets the content types, so /md/*.md is served as markdown and not downloaded as a file
  • sitemap lastmod from each page's last commit rather than the build time
  • src/siteCopy.mjs holds the name, the description, the JSON-LD ids and the social links in one place, so the site cannot describe itself two ways

Colour theme and dark mode

The site carried the stock Docusaurus palette, which does not match kokio.app. Colours are now pulled from the marketing site's own design tokens (the beach-sky background, the cashmere orange, the ship-cove and outer-space greys) rather than approximated:

  • Light and dark themes both defined from these tokens, with contrast checked against WCAG before picking each foreground colour (light-mode links are cashmere-700 at 5.19:1, dark-mode links are cashmere-400 at 6.43:1)
  • Dark mode rebuilt on the same warm dark ground as kokio.app instead of a generic grey, including overriding Infima's own neutral colour ramp, which otherwise paints the sidebar, borders and hover states grey regardless of the theme colours set above it
  • Footer recoloured to match the header instead of a fixed dark bar in both themes
  • Home page rebuilt on the blog's hero layout: the beach illustration from the marketing site, the same heading and body type (Anybody and Lexend, loaded from Google Fonts), and card-style feature panels with hover elevation
  • Favicon changed from the full wordmark logo to the logomark on its own, since a wordmark is unreadable at tab size
  • Fixed the header's "Home" link reading as active on every docs page; it now only highlights on the home page itself and on hover elsewhere

Header and footer links

  • Header: GitHub, X (Twitter) and Telegram now shown as icons instead of text, sized and cropped to match each other
  • Footer Docs column: added Kokio SDK and Smart contracts alongside What is Kokio?
  • Footer Community: Discord removed, Twitter and Telegram kept
  • Footer More: added Blogs, linking to the marketing site's blog

Checks

Four scripts, wired into npm run build so nothing needs a separate CI step.

Script When What it catches
check-conventions.mjs before build file names that are not kebab-case, missing or oversized frontmatter, images with no alt text, filler words, heading problems
build-text.mjs before build generates /md and llms-full.txt
verify-site.mjs after build broken agent-facing links, sitemap gaps, pages too short to answer anything, pages that never name the product
sync-reference.mjs after build contract and SDK pages that have drifted from their source repo

CI runs sync-reference.mjs --strict, which fails rather than warns, because at that point a person is already looking at a diff.

Tooling

  • yarn to npm, and yarn.lock deleted
  • every dependency pinned to an exact version, no carets
  • React 18 to 19
  • npm overrides pin patched versions of qs, uuid and serialize-javascript. All three arrive through the dev server and the webpack plugins and never reach the built site, but they account for every moderate advisory on this repo.

URL changes

All file names are now kebab-case, matching the contract pages. docs/Overview.md to docs/overview.md, mobileApp.md to mobile-app.md, docs/eSIM/ to docs/esim/, and so on. Image files renamed the same way. check-conventions.mjs enforces this from here on.

Follow-ups, not in this PR

  • image-size still has 17 open advisories with no fix published. It reads images from this repo at build time, so dismiss them rather than chase them.

The deployed site was serving six template blog posts, their tag and
archive pages, and a placeholder markdown page. That was 11 of the 39
URLs in the sitemap, none of them about Kokio.

The blog is now off explicitly rather than commented out, so a future
preset default cannot bring it back. Broken markdown links now fail the
build instead of printing a warning nobody reads.
intro_deprecated, selfcustodial and smartwallet each repeated sections
of intro.md or walletSuite.md word for word. contractSpecs listed ten
contracts and had gone stale: the suite has eighteen.

Two copies of a claim in one corpus is how a model ends up hedging
between them.
Overview and setup both built to live URLs that nothing linked to. A
page reachable only by guessing its URL still gets indexed, and it
arrives at a reader stripped of the context the sidebar would have
given it.
Both linked a .wiki.git clone URL, which git understands and a browser
does not. The wikis themselves are fine, so this is just the suffix.
Three things were wrong in the emitted HTML. The home page shipped the
template string "Description will go into a meta tag in <head />" as its
description and "Hello from Kokio" as its title. The favicon had no file
extension, so the page pointed at /images/KokioLogo and got a 404.

The name and the descriptions now come from one module instead of being
retyped. Repeated strings are what a model builds an entity record out
of, so two wordings of the same claim are worse than one.
intro.md opened with a frontmatter block whose only line was a YAML
comment, so the sidebar_position it looked like it set was never read.
The sidebar is explicit anyway, so the block is gone rather than fixed.

Also transperancy, utilies, enbaled, Contacts, and "an mobile".
Every entry carried changefreq weekly and priority 0.5, identical across
all 39 URLs. Google ignores both, and a field with the same value
everywhere carries no signal, so they are gone.

lastmod replaces them, read per file from git. That needs full history,
so CI checks out the whole thing now.
There was no robots.txt at all, so the site had no stated position and
no sitemap pointer. Retrieval and training crawlers now both get
everything, matching kokio.app.

Seventeen agents are named individually. The wildcard is what actually
grants access; the named groups make the decision auditable and make
blocking any one of them a one-line edit later.
A doc page as HTML carries the sidebar, the navbar and a hydration
payload, and its code blocks arrive wrapped in syntax-highlighting spans
that damage copy-paste. The markdown is a fraction of the bytes and is
what a reader actually wanted.

/md/<path>.md is one page. /llms-full.txt is all 21 in one fetch, in
sidebar order, which is the difference between one request and twenty-one
for a question about the contract suite.

One script writes both. Two generators over the same corpus drift, and
nobody reads either file often enough to notice.

Commented-out draft paragraphs are stripped. Two pages keep old text in
HTML comments, and it must not reach a surface the website itself does
not show.
One file naming every page and what it covers, so a model asking about
the contract suite picks two fetches instead of crawling twenty-two.

It says plainly that the contract pages are short and points at the
source repository for the real reference. Overstating what is here would
just produce confident answers built on a paragraph.
Docusaurus derives an anchor from the heading text, so rewording a
heading silently breaks every link and citation pointing at it. Written
down, the anchor survives the reword.
Koki'o and KOKIO both appeared in prose. Three spellings of one product
split it into three weakly supported entities, so the variants now live
only in schema, never in a sentence.

Also: onchain as one word throughout, per the usual convention; the two
commented-out draft paragraphs deleted rather than left to leak into the
markdown routes; and the seamless/utilize/empowers filler replaced with
what the sentence was actually claiming. All of this reaches
llms-full.txt verbatim, so it is not only a style question.
Doc pages compete with StackOverflow and GitHub for the same questions.
A typed description with a subject and a modified date gives a retriever
grounds to prefer the page it came from.

Every doc route now carries one TechArticle beside the breadcrumbs the
theme already emitted, and the home page carries a WebSite node. The
organisation is referenced by id from kokio.app rather than redefined,
or the two hosts would describe two different companies.

The subject varies by section. Labelling the eSIM explainer a Solidity
document would be a claim a retriever acts on, and a wrong one.

No id carries a fragment. A fragment in an @id is a citation target and
has to match a real element, and nothing here has one.
Both lockfiles were committed, CI ran yarn and Vercel picked by lockfile,
so the two could resolve different trees. One lockfile now, package-lock.json.

README also drops the template's GitHub Pages deploy steps, which this repo
has never used.
Doc pages referenced the WebSite node by id, but only the home page defined
it. A crawler reads one page at a time, so on every page but one that
reference pointed at nothing.

Each doc page now carries the node in the same graph as its article, and
both pages read it from siteCopy so they cannot drift apart.

Doc pages also link their own markdown copy, so an agent does not have to
guess the /md/ pattern.
The three feature cards were template copy. One offered PayPal, which is not
a payment method here, and two sentences ended in a stray second full stop.
Rewritten against the description the rest of the site uses.
Ranges meant a fresh install could resolve a different tree from the one
that was tested. React moves to 19 and TypeScript to 5.9.3 on the same pass.

clsx is gone. One call site passed it a static class string and the other
joined two names, neither of which needs a package. React 19 also drops the
global JSX namespace, so three return types become ReactNode.

engines now says node 20, which is what Docusaurus 3.10 requires.
Seven things break silently: a page missing from the sitemap, a page with no
markdown copy, a dead link in llms.txt, a schema id matching nothing, a wrong
canonical, a robots group that skipped a disallow, and a missing lastmod,
which is how a shallow clone in CI shows up. All of them still render fine.

Runs as postbuild, so it fails the build that produced the problem.
Every new contract page lands under docs/contracts with a kebab-case name,
so the existing pages move there first rather than sitting beside them in a
second naming style for the rest of the phase.

Old URLs are not redirected. Ten of them change.
Every address checked with cast against Base Sepolia before it went on
the page: code present, recorded codehashes match, proxies point at the
listed implementations, both beacons point at the listed wallets.
A page being written has no commit, so no date to read. Only a committed
page missing its date means the clone was shallow, which is what the
check is actually for.
Seventeen pages under /docs/sdk, ported from kokio-sdk v3.0.1. Mobile and
backend split the way the package does, one page per surface. All 243
methods on the two entry points are covered, checked against the interface
classes in the SDK source.
It is internal reference, not published SDK surface.
The site had eight one-sentence contract pages and nothing on payments,
governance or account abstraction. Now every unit has a page, each with a
written intro and the full generated reference under it.

scripts/sync-reference.mjs keeps the generated half in step with the repo
that produces it. It warns on every build and fails in CI.

Contract docs parse as plain markdown now. The generated reference is full
of text MDX reads as JSX.
Every page now opens by naming what it is and what Kokio is, carries a
title and a description, and runs past the length where a retrieved chunk
says nothing. The three shortest pages were the worst of it: resources was
a bare link list, the solution page was one run-on paragraph, and the tech
stack page stopped mid-sentence in the backend section.

Page files are kebab-case now, matching the contract pages, so a file name
and its URL always look the same. The eSIM folder, the two Overview pages
and the three images move with them.
A page is retrieved as one chunk, without the sidebar around it. Under a
hundred words there is nothing in that chunk worth answering from, and a
page that never names the product leaves the reader guessing whose docs
they landed in. Measured on the markdown output, because that is the
surface agents read and frontmatter is stripped from it.
Rules that are cheap to state and expensive to notice: kebab-case file
names, a title and a description on every page, one h1 with no skipped
levels under it, an explicit id on every heading, and prose without em
dashes, filler words, hyphenated onchain or an unescaped dollar sign.

Runs as prebuild, so a file breaking one never reaches a build, and CI
picks it up through npm run build without a new step.
A diagram is invisible to a screen reader and to every text surface this
site generates, so both pages now carry alt text and a table saying what
the picture says. Both tables carry more than the picture did: the eSIM
page never mentioned the GSMA certificate chain, and the app page never
said what the fork on new versus returning user actually changes.

The conventions script now fails a build on an image with no alt text, or
with alt text short enough to be a caption rather than a description.
The alt text listed a payment adapter that is not in the picture. The
diagram predates both the payment and the governance contracts, so it now
describes the fiat and crypto paths it does show, and a line underneath
names the two contracts it is missing.
Links to kokio.app/manifesto, below Blogs.
Builds the index at build time and searches it in the reader's browser,
so no query leaves the site and there is nothing to keep running.

The search page holds no text of its own, so it stays out of the sitemap.
Reads the docs the site already publishes as text, picks the few pages
closest to the question, and asks Gemini to answer from those alone.

When the model is out of daily quota or down, the pages still come back,
so the caller can fall back to plain search instead of showing an error.
Replaces the browser-side search index with a box that asks the endpoint.
The index was 364 KB every reader downloaded to get word matching; the
endpoint ranks the same pages server side and writes an answer as well.

The answer check runs in prebuild, so a broken ranking fails the build.
The function used to fetch the site's own text file. A deployment behind
Vercel's protection answers that with a login page, which parses into no
pages at all, so every question came back as "nothing matches".

It now reads the copy shipped beside it, and refuses to start on text that
holds no pages rather than answering emptily. When the model declines, the
reason goes to the logs instead of to the reader.
Runs the same code the endpoint runs, against the docs on disk or against
a deployment, and names what came back instead of throwing a stack.
A preview reading its own address gets the protection login page. The live
site is the one copy that always answers, so that is the fallback when no
docs shipped in the bundle.
Gemini answers 503 when the model is under load, which happened on the
second question anyone asked. It now tries again, then tries the smaller
model, and only then falls back to listing pages.

A bad key or a used-up quota is final, so those still return at once.
The model spends its own reasoning from the same token budget as the answer,
so a 600 token cap ran out before the answer finished. The cap is now large
enough, and reasoning is set low, which suits short answers drawn from pages
that were already picked out.

A truncated answer is still shown, and now says so in the logs.
The thinking option was sent under the wrong name, which the API rejects
with a 400, so every question failed and looked like a broken key. The
request body is now built in one exported place, with a test pinning each
field name against Google's reference, and a rejected config is retried
once without the optional tuning. A renamed field costs a slower answer,
not a dead box.

Retrying is limited to what retrying fixes. A busy model is tried again
and then on the smaller model; a bad key or a spent quota returns at once
instead of being repeated on every model. Each call has a timeout and the
whole request has a budget, so nothing hangs.

Ranking changed too. "No KYC" and "no personal information collected"
appear on no page, only in the description the file opens with, which was
being dropped, so questions about what Kokio collects matched nothing and
returned overview pages. That description is now a page of its own,
scoring is pivoted so a short page cannot outrank a specific one, and word
endings are trimmed so "collected" finds "collects".
Requests from another site are refused, non-JSON bodies are refused, and a
per-instance cap sits behind the per-reader one, so one busy instance
cannot spend the day's quota on its own.

Repeat questions are served from memory for ten minutes, which is most of
them on a docs site, and each one saved is a request not spent.
Reading the docs by what people fail to find in them. Sent after the
reader has their answer, so a slow webhook costs them nothing, and a
failed post is logged rather than turned into a failed search.
A missing webhook returned early in silence, and a mistyped one answers
401 rather than throwing, so both left an empty channel and an empty log.
A function can be frozen the moment it responds, so the post started
after that never finished and the channel stayed empty.
Green when the model answered, red when it did not, so a channel of
these can be skimmed for the gaps. Pages move to the bottom as links to
the live site rather than bare paths.
Quota is counted per model, so a spent limit on the first one says
nothing about the second. Costs one extra call on a question that has
already failed.
@ManulParihar
ManulParihar deleted the upgrade/ai-seo branch September 9, 2026 07:34

This branch was successfully deployed

1 active deployment
Preview — b25235ce Deployed Sep 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants