Skip to content

Make the site legible to search engines and AI assistants- #64 - #65

Merged
ManulParihar merged 43 commits into
mainfrom
dev
Sep 9, 2026
Merged

ManulParihar merged 43 commits into
mainfrom
dev

Conversation

@ManulParihar

Copy link
Copy Markdown
Member

The site was hard for a crawler or an assistant to read: the name was spelled four ways, the home page described itself in three words, there was no structured data, and no plain-text version of anything. This fixes that.

What changed

Name and description. Kokio is now the only spelling in prose and metadata, with the other spellings kept where a machine reads them. One canonical description string is reused across metadata, JSON-LD, OG tags and the text files, so every surface says the same sentence.

Structured data. Organization, WebSite, MobileApplication, FAQPage and DefinedTermSet as JSON-LD. FAQ answers and glossary entries each have an anchor, so a citation can point at one answer instead of the whole page.

Crawler access. robots.ts names the AI crawlers explicitly and allows both retrieval and training. The payment callback is noindex. Sitemap dates come from the content itself, never from build time.

Plain text for agents. /llms.txt for the map, /llms-full.txt for the whole site in one response, and any page as markdown by adding .md to its URL. All of it renders from the same modules the pages do, so it cannot drift.

Glossary. Thirteen definitions of the terms the site uses without explaining them, plus a Kokio entry that ties every spelling to one product.

Roadmap. Milestones now run Q3'26 to Q1'27. The old list ended in Q2'26 and promised a beta that has been overtaken. The launch notice in the banner and the footer is one string that names a month.

Checks. scripts/verify-site.mjs runs after every build and fails on a page missing from the sitemap, a sitemap entry with no text, a dead link in the agent-facing files, a schema id pointing at nothing, a wrong canonical, or a robots group missing a disallow. CI now runs lint too, and checks out full history so the freshness check has something to read.

Cleanup. Deleted page.html (an 85 KB rendered copy of the home page at the repo root) and the two leftover template SVGs.

New URLs

URL What it is
/glossary Definitions, with DefinedTermSet schema
/llms.txt Annotated map of the site
/llms-full.txt Every page as one markdown file
/<any-page>.md That page as markdown, home is /index.md

How to check it

npm run build runs the site checks at the end. npm run verify runs them alone against an existing build.

The site shipped four spellings of the product name and two different
descriptions of what it is. Search and AI systems build an entity record out
of repeated strings, so four surface forms split one product into four weakly
supported ones, and two descriptions give no clear answer to "what is this".

Add lib/site-copy.ts as the single source for the name and description, and
point config and metadata at it. Settle on Kokio: it matches the domain, and
the apostrophe form gets rewritten to a curly quote when MDX renders it, so
the source and the page never agreed anyway.

Also fixes the home page, which had no canonical URL and described itself in
three words, and adds a title template so pages stop hardcoding the suffix.

Blog posts, legal documents and the roadmap copy still use the old spellings.
Those are content changes and are handled separately.
Crawlers had to infer who publishes this site from prose. Emit an Organization
and a WebSite node once from the root layout, so every page carries the same
typed facts: legal name, UEN, registered address, contact, and the social and
source links that let a reader corroborate them elsewhere.

Both nodes list every spelling of the name in alternateName. That is the only
machine-readable place saying the four forms are one company rather than four.

Facts come from the existing legal config rather than new literals, so the
schema and the legal pages cannot drift apart.
The questions lived inside the component that rendered them, which leaves
nowhere for anything else to read them from without copying the text, and
copied text drifts.

Move them to content/faqs.ts alongside the other content modules, and give each
one a stable id for deep links.
Add MobileApplication so the product has a type, a category and a feature list
rather than only marketing prose, and FAQPage so each question and answer is a
retrievable pair instead of text a reader has to find inside the page.

The FAQ schema maps over the content module, so the markup and the rendered
list stay in step.
Each post described the publisher again from scratch, so the same company was
stated in four places with no link between them. Point at the Organization the
layout emits instead, and say which site the post belongs to.

dateModified still copies the publish date. Fixing that needs the file's real
edit time and is handled with the sitemap work.
Nothing tracked the difference between when a post was published and when it
was last revised, so there was no honest source for a last-modified date.

Add an optional updated field to blog frontmatter, defaulting to the publish
date, and a required one on the manifesto. Backfilled from git history.

Deliberately not derived from file mtime or git at build time. Both are wrong
here: a fresh clone gives every file the same checkout mtime, which is already
visible locally, and Vercel clones only the last ten commits, so older files
have no history to read.
Every static page reported the build time as its last-modified date, so an
untouched page looked freshly edited on every deploy. A crawler that re-fetches
on that signal learns it means nothing.

Each entry now reads a real date: posts and the manifesto from frontmatter, the
legal pages from the legal config, the blog index from its newest post. The
home page has no content file, so it carries an explicit constant.

Two consecutive builds with no content change now produce a byte-identical
sitemap.
dateModified copied the publish date, so every post claimed it had never been
touched since it went up. Read the new updated field instead, and add the same
date to the Open Graph tags.
The target is a post every two weeks. The sitemap spec has no value for that,
so the entry stays weekly, which errs towards checking too often rather than
missing a post. Note it so the value can be revisited once the real cadence is
known rather than re-derived from scratch.
alternateName held two different things: ways of writing the word Kokio, and
names like "Kokio App" that are not spellings at all. Prose that lists the
spellings needs only the first group.

Split them, and keep the combined list for schema.
A plain-text summary of the site: what Kokio is, every spelling of the name,
who builds it, and an annotated link to each page worth reading.

Generated from the same modules the pages render from, so it cannot fall behind
them, and it lists only URLs that exist. The glossary and full-text file are
missing on purpose until they ship, since a dead link costs more than an absent
entry.

Worth being honest about the value: measured crawler traffic to these files is
close to nothing, and no major AI vendor has committed to reading them. The
real readers today are coding agents. It is cheap and correct, not a discovery
channel.
The file had a single wildcard rule, which left the stance towards AI crawlers
implicit. Name each one instead, with every token checked against the
operator's own documentation.

Allowing them changes no behaviour, since crawling is permitted by default. It
makes the decision explicit and auditable, and turns a future change of mind
into a one-line edit.

Rules are repeated per agent rather than stated once, because a crawler that
matches a named group reads only that group and never sees the wildcard. A
disallow written once would not have applied to any of them.

Notes on three of the entries are in the file: two are training-control tokens
with no crawler behind them, one agent's operator says it generally ignores
robots.txt, and one publishes no documentation at all.
The page is reached only after a payment and says nothing on its own. Indexing
it produces a search result that looks like a completed transaction belonging
to whoever finds it.

It was already absent from the sitemap and is now disallowed in robots.txt;
this adds the page-level noindex so the three agree.
Adds AI2Bot from the Allen Institute's published crawler notice, and the xAI
names that circulate without one.

Splits the list in two. The first group is agents whose operator documents a
token. The second is names taken from observed traffic and third-party
trackers, so the file no longer implies the same confidence in both.

xAI belongs in the second group and is close to a formality: no documentation
exists, and Grok's fetches are reported to arrive with ordinary browser user
agents from rotating addresses, which no robots.txt rule can match.

Also notes what CCBot really governs. Open-weight models are mostly trained on
corpora derived from Common Crawl rather than on their own crawls, so that one
line reaches further than the rest of the file, and reaches further than it can
be taken back.
Removes the group of agents whose operators publish nothing: Bytespider, the
three xAI names and the reported AI2 variant. All came from third-party
trackers, so none could be verified, and none granted access the wildcard rule
was not already granting.

The list can never be complete anyway. DeepSeek crawls with no user agent, xAI
sends ordinary browser traffic, and the names circulating for Qwen and GLM have
no operator page behind them. A list that is unverifiable and permanently out
of date is worse than a short accurate one next to a wildcard that already says
yes to everybody.

Entries now need one thing to qualify: the operator publishes a page naming the
token. Ten do. Everyone else is covered by the wildcard, which the file now
says plainly.
Each section now has a stable id and an accessible name, so an answer that
cites the page can link to the part it used instead of the whole page.
The page schema already gave each question an @id ending in #faq-<id>.
Nothing in the markup carried that id, so those links went nowhere.
The old paragraph sold distributed ledger technology and seamless
connectivity without naming a single thing the app does. It is the first
prose any reader or crawler meets, so it now carries the same facts as the
canonical description: destinations, payment methods, no KYC, self-custody.
The Terms and the Privacy Policy are structured data, not markdown, so
anything that wants them outside React needs its own serializer.
An agent that has to crawl eight routes usually stops before the end. The
sections are driven by the sitemap, so a page cannot reach it without text
here: the build fails instead of the file quietly going out of date.
Both the full-text file and the per-page markdown routes need the same
rendering, and neither should be the other's dependency.
A blog post is 7 KB of prose this way against 87 KB of markup, so an agent
working to a context budget fits twelve times as much of the site.
Nobody guesses a URL convention. Each page now links its own markdown form,
and llms.txt states the rule once.
The blog index and both legal pages already answered on .md but never said
so. Found by the new checks.
Blog posts reach every surface on their own because they all read from
getAllPosts(). Pages do not: the sitemap is written by hand, and a page
missing from it is missing from the text files and the .md routes too, with
nothing to say so. Seven checks, run after every build.
page.html was an 85 KB copy of the rendered home page sitting at the repo
root. next.svg and vercel.svg came with create-next-app. Nothing links any
of the three.
eslint reported console, process and URL as undefined in scripts/, since the
config only described browser and Next code.
Both are deep links: a citation points at the one item it used, so the id has
to match an element. The check now covers anything a schema set holds.
Thirteen definitions of the terms the site uses without explaining them, plus
a Kokio entry that names every spelling of the product. Definition pages are
precise retrieval targets, and each entry has its own anchor so a citation can
point at one definition.

Reachable from the footer, llms.txt, the sitemap, the full text file and
/glossary.md, all reading the same content file.
Public launch in Q3'26, features and distribution in Q4'26, privacy and
identity partnerships in Q1'27. The old list ended in Q2'26 and had two
quarters that already passed sitting at the front of it.

Each phase now names what it is about, and shipped work comes off the list
rather than staying on it as an apparently missed date.
ManulParihar and others added 13 commits September 7, 2026 15:21
The roadmap answer described a beta launch in Q2 2026 that has been overtaken.
The early access answer pointed at a Sign up for beta button that no longer
exists on the page; the way in is the Telegram group.

Both are written to stand on their own, since an answer is retrieved without
the page around it.
They were held back while their dates were wrong, since plain text is the
format a model quotes back word for word.
Lint was never run in CI, which is why a config gap in scripts/ went unnoticed.
The freshness check reads the commit that last touched a piece of content, and
a shallow clone gives it nothing to read.
The banner said Kokio is launching soon, the footer said Kokio launches soon,
and neither named a date, so both would have stayed plausible long after the
launch. Both now read from one string that names the month, which is wrong
visibly rather than quietly.

The banner also linked X from an icon with no text in it.
The icon and the word X both linked to Twitter, right next to each other. Kept the icon, moved its accessible name into alt text.
Counsel confirmed a three month look-back window and chose SIAC arbitration in
Singapore. Both had been shipping as bracketed placeholders on the live pages,
with the dispute clause still offering the reader both options.

The sentence lost its 'in', since the chosen wording starts with 'by'.
Counsel approved the same spelling used everywhere else. The term defines
itself in the opening paragraph of each document, so the meaning is unchanged.

The registered name KOKIO SG PTE. LTD. is untouched, and the all-caps
liability clause keeps the all-caps rendering it already used.
Contract and policy text can be edited like any other file, and nothing said
that a change needs a lawyer before it ships. The build now warns when either
document has been edited more recently than counselReviewedIso, and both files
carry a header saying what to do about it.

The date moves only after counsel has seen the change, never to quiet the
warning.
Route folder, nav links, sitemap, RSS, llms.txt, and the markdown
mirror all point at /blogs now. Old /blog URLs 308 to their /blogs
equivalent so existing links and RSS subscribers keep working.
The new /blog/:slug* redirect caught the static assets still living
under public/blog/*, sending hero images, author avatars, and inline
images to a /blogs path that doesn't exist. Moved those assets to
public/blog-assets so the asset namespace no longer collides with the
route.
Make the site legible to search engines and AI assistants
@vercel

vercel Bot commented Sep 9, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
kokio-website Ready Ready Preview Sep 9, 2026 6:49am UTC

Request Review

@ManulParihar
ManulParihar merged commit a12cf03 into main Sep 9, 2026
9 checks passed

This branch was successfully deployed

1 active deployment
Preview — 9843f581 Deployed Sep 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant