Make the site legible to search engines and AI assistants - #64
Merged
Merged
Conversation
The site shipped four spellings of the product name and two different descriptions of what it is. Search and AI systems build an entity record out of repeated strings, so four surface forms split one product into four weakly supported ones, and two descriptions give no clear answer to "what is this". Add lib/site-copy.ts as the single source for the name and description, and point config and metadata at it. Settle on Kokio: it matches the domain, and the apostrophe form gets rewritten to a curly quote when MDX renders it, so the source and the page never agreed anyway. Also fixes the home page, which had no canonical URL and described itself in three words, and adds a title template so pages stop hardcoding the suffix. Blog posts, legal documents and the roadmap copy still use the old spellings. Those are content changes and are handled separately.
Crawlers had to infer who publishes this site from prose. Emit an Organization and a WebSite node once from the root layout, so every page carries the same typed facts: legal name, UEN, registered address, contact, and the social and source links that let a reader corroborate them elsewhere. Both nodes list every spelling of the name in alternateName. That is the only machine-readable place saying the four forms are one company rather than four. Facts come from the existing legal config rather than new literals, so the schema and the legal pages cannot drift apart.
The questions lived inside the component that rendered them, which leaves nowhere for anything else to read them from without copying the text, and copied text drifts. Move them to content/faqs.ts alongside the other content modules, and give each one a stable id for deep links.
Add MobileApplication so the product has a type, a category and a feature list rather than only marketing prose, and FAQPage so each question and answer is a retrievable pair instead of text a reader has to find inside the page. The FAQ schema maps over the content module, so the markup and the rendered list stay in step.
Each post described the publisher again from scratch, so the same company was stated in four places with no link between them. Point at the Organization the layout emits instead, and say which site the post belongs to. dateModified still copies the publish date. Fixing that needs the file's real edit time and is handled with the sitemap work.
Nothing tracked the difference between when a post was published and when it was last revised, so there was no honest source for a last-modified date. Add an optional updated field to blog frontmatter, defaulting to the publish date, and a required one on the manifesto. Backfilled from git history. Deliberately not derived from file mtime or git at build time. Both are wrong here: a fresh clone gives every file the same checkout mtime, which is already visible locally, and Vercel clones only the last ten commits, so older files have no history to read.
Every static page reported the build time as its last-modified date, so an untouched page looked freshly edited on every deploy. A crawler that re-fetches on that signal learns it means nothing. Each entry now reads a real date: posts and the manifesto from frontmatter, the legal pages from the legal config, the blog index from its newest post. The home page has no content file, so it carries an explicit constant. Two consecutive builds with no content change now produce a byte-identical sitemap.
dateModified copied the publish date, so every post claimed it had never been touched since it went up. Read the new updated field instead, and add the same date to the Open Graph tags.
The target is a post every two weeks. The sitemap spec has no value for that, so the entry stays weekly, which errs towards checking too often rather than missing a post. Note it so the value can be revisited once the real cadence is known rather than re-derived from scratch.
alternateName held two different things: ways of writing the word Kokio, and names like "Kokio App" that are not spellings at all. Prose that lists the spellings needs only the first group. Split them, and keep the combined list for schema.
A plain-text summary of the site: what Kokio is, every spelling of the name, who builds it, and an annotated link to each page worth reading. Generated from the same modules the pages render from, so it cannot fall behind them, and it lists only URLs that exist. The glossary and full-text file are missing on purpose until they ship, since a dead link costs more than an absent entry. Worth being honest about the value: measured crawler traffic to these files is close to nothing, and no major AI vendor has committed to reading them. The real readers today are coding agents. It is cheap and correct, not a discovery channel.
The file had a single wildcard rule, which left the stance towards AI crawlers implicit. Name each one instead, with every token checked against the operator's own documentation. Allowing them changes no behaviour, since crawling is permitted by default. It makes the decision explicit and auditable, and turns a future change of mind into a one-line edit. Rules are repeated per agent rather than stated once, because a crawler that matches a named group reads only that group and never sees the wildcard. A disallow written once would not have applied to any of them. Notes on three of the entries are in the file: two are training-control tokens with no crawler behind them, one agent's operator says it generally ignores robots.txt, and one publishes no documentation at all.
The page is reached only after a payment and says nothing on its own. Indexing it produces a search result that looks like a completed transaction belonging to whoever finds it. It was already absent from the sitemap and is now disallowed in robots.txt; this adds the page-level noindex so the three agree.
Adds AI2Bot from the Allen Institute's published crawler notice, and the xAI names that circulate without one. Splits the list in two. The first group is agents whose operator documents a token. The second is names taken from observed traffic and third-party trackers, so the file no longer implies the same confidence in both. xAI belongs in the second group and is close to a formality: no documentation exists, and Grok's fetches are reported to arrive with ordinary browser user agents from rotating addresses, which no robots.txt rule can match. Also notes what CCBot really governs. Open-weight models are mostly trained on corpora derived from Common Crawl rather than on their own crawls, so that one line reaches further than the rest of the file, and reaches further than it can be taken back.
Removes the group of agents whose operators publish nothing: Bytespider, the three xAI names and the reported AI2 variant. All came from third-party trackers, so none could be verified, and none granted access the wildcard rule was not already granting. The list can never be complete anyway. DeepSeek crawls with no user agent, xAI sends ordinary browser traffic, and the names circulating for Qwen and GLM have no operator page behind them. A list that is unverifiable and permanently out of date is worse than a short accurate one next to a wildcard that already says yes to everybody. Entries now need one thing to qualify: the operator publishes a page naming the token. Ten do. Everyone else is covered by the wildcard, which the file now says plainly.
Each section now has a stable id and an accessible name, so an answer that cites the page can link to the part it used instead of the whole page.
The page schema already gave each question an @id ending in #faq-<id>. Nothing in the markup carried that id, so those links went nowhere.
The old paragraph sold distributed ledger technology and seamless connectivity without naming a single thing the app does. It is the first prose any reader or crawler meets, so it now carries the same facts as the canonical description: destinations, payment methods, no KYC, self-custody.
The Terms and the Privacy Policy are structured data, not markdown, so anything that wants them outside React needs its own serializer.
An agent that has to crawl eight routes usually stops before the end. The sections are driven by the sitemap, so a page cannot reach it without text here: the build fails instead of the file quietly going out of date.
Both the full-text file and the per-page markdown routes need the same rendering, and neither should be the other's dependency.
A blog post is 7 KB of prose this way against 87 KB of markup, so an agent working to a context budget fits twelve times as much of the site.
Nobody guesses a URL convention. Each page now links its own markdown form, and llms.txt states the rule once.
The blog index and both legal pages already answered on .md but never said so. Found by the new checks.
Blog posts reach every surface on their own because they all read from getAllPosts(). Pages do not: the sitemap is written by hand, and a page missing from it is missing from the text files and the .md routes too, with nothing to say so. Seven checks, run after every build.
page.html was an 85 KB copy of the rendered home page sitting at the repo root. next.svg and vercel.svg came with create-next-app. Nothing links any of the three.
eslint reported console, process and URL as undefined in scripts/, since the config only described browser and Next code.
Both are deep links: a citation points at the one item it used, so the id has to match an element. The check now covers anything a schema set holds.
Thirteen definitions of the terms the site uses without explaining them, plus a Kokio entry that names every spelling of the product. Definition pages are precise retrieval targets, and each entry has its own anchor so a citation can point at one definition. Reachable from the footer, llms.txt, the sitemap, the full text file and /glossary.md, all reading the same content file.
Public launch in Q3'26, features and distribution in Q4'26, privacy and identity partnerships in Q1'27. The old list ended in Q2'26 and had two quarters that already passed sitting at the front of it. Each phase now names what it is about, and shipped work comes off the list rather than staying on it as an apparently missed date.
The roadmap answer described a beta launch in Q2 2026 that has been overtaken. The early access answer pointed at a Sign up for beta button that no longer exists on the page; the way in is the Telegram group. Both are written to stand on their own, since an answer is retrieved without the page around it.
They were held back while their dates were wrong, since plain text is the format a model quotes back word for word.
Lint was never run in CI, which is why a config gap in scripts/ went unnoticed. The freshness check reads the commit that last touched a piece of content, and a shallow clone gives it nothing to read.
The banner said Kokio is launching soon, the footer said Kokio launches soon, and neither named a date, so both would have stayed plausible long after the launch. Both now read from one string that names the month, which is wrong visibly rather than quietly. The banner also linked X from an icon with no text in it.
The icon and the word X both linked to Twitter, right next to each other. Kept the icon, moved its accessible name into alt text.
Counsel confirmed a three month look-back window and chose SIAC arbitration in Singapore. Both had been shipping as bracketed placeholders on the live pages, with the dispute clause still offering the reader both options. The sentence lost its 'in', since the chosen wording starts with 'by'.
Counsel approved the same spelling used everywhere else. The term defines itself in the opening paragraph of each document, so the meaning is unchanged. The registered name KOKIO SG PTE. LTD. is untouched, and the all-caps liability clause keeps the all-caps rendering it already used.
Contract and policy text can be edited like any other file, and nothing said that a change needs a lawyer before it ships. The build now warns when either document has been edited more recently than counselReviewedIso, and both files carry a header saying what to do about it. The date moves only after counsel has seen the change, never to quiet the warning.
Route folder, nav links, sitemap, RSS, llms.txt, and the markdown mirror all point at /blogs now. Old /blog URLs 308 to their /blogs equivalent so existing links and RSS subscribers keep working.
The new /blog/:slug* redirect caught the static assets still living under public/blog/*, sending hero images, author avatars, and inline images to a /blogs path that doesn't exist. Moved those assets to public/blog-assets so the asset namespace no longer collides with the route.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
ManulParihar
added a commit
that referenced
this pull request
Sep 9, 2026
Make the site legible to search engines and AI assistants- #64
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The site was hard for a crawler or an assistant to read: the name was spelled four ways, the home page described itself in three words, there was no structured data, and no plain-text version of anything. This fixes that.
What changed
Name and description.
Kokiois now the only spelling in prose and metadata, with the other spellings kept where a machine reads them. One canonical description string is reused across metadata, JSON-LD, OG tags and the text files, so every surface says the same sentence.Structured data. Organization, WebSite, MobileApplication, FAQPage and DefinedTermSet as JSON-LD. FAQ answers and glossary entries each have an anchor, so a citation can point at one answer instead of the whole page.
Crawler access.
robots.tsnames the AI crawlers explicitly and allows both retrieval and training. The payment callback isnoindex. Sitemap dates come from the content itself, never from build time.Plain text for agents.
/llms.txtfor the map,/llms-full.txtfor the whole site in one response, and any page as markdown by adding.mdto its URL. All of it renders from the same modules the pages do, so it cannot drift.Glossary. Thirteen definitions of the terms the site uses without explaining them, plus a
Kokioentry that ties every spelling to one product.Roadmap. Milestones now run Q3'26 to Q1'27. The old list ended in Q2'26 and promised a beta that has been overtaken. The launch notice in the banner and the footer is one string that names a month.
Checks.
scripts/verify-site.mjsruns after every build and fails on a page missing from the sitemap, a sitemap entry with no text, a dead link in the agent-facing files, a schema id pointing at nothing, a wrong canonical, or a robots group missing a disallow. CI now runs lint too, and checks out full history so the freshness check has something to read.Cleanup. Deleted
page.html(an 85 KB rendered copy of the home page at the repo root) and the two leftover template SVGs.New URLs
/glossaryDefinedTermSetschema/llms.txt/llms-full.txt/<any-page>.md/index.mdHow to check it
npm run buildruns the site checks at the end.npm run verifyruns them alone against an existing build.