Skip to content

ci: make doc-review link check case-sensitive per path segment - #139

Merged
junlin-luo-spacemit merged 1 commit into
mainfrom
ci/harden-doc-review-link-checks
Sep 30, 2026
Merged

junlin-luo-spacemit merged 1 commit into
mainfrom
ci/harden-doc-review-link-checks

Conversation

@junlin-luo-spacemit

Copy link
Copy Markdown
Collaborator

Problem

The repo keeps finding the same bug class: a reference that resolves on case-insensitive Windows/macOS but 404s on Linux CI and the web host. PR #138 fixes the current instances by hand — this PR makes the linter actually catch them, so it stops recurring.

The existing links.py was too permissive in some ways and too strict in others:

  • It only compared the filename, so a wrong-case directory (Static/) slipped through.
  • It ignored HTML <a href>, so anchor-style links were unchecked.
  • It ignored #fragment anchors entirely, so links to headings that no longer exist went unnoticed.
  • It pressured authors to rename the file rather than fix the reference — the opposite of what we want.
  • product_names.py scanned the whole line, including image paths, so a correct lowercase path like static/com260_model.png was reported as a branding violation.

Changes

checks/links.py

  • Verifies every path segment, not just the filename — directories are included.
  • Covers HTML <a href> in addition to markdown links/images and HTML <img src>.
  • Validates #fragment anchors against the target page's real heading slugs (both in-page and page.md#frag).
  • Reports the corrected reference text to copy, instead of implying the file should be renamed.
  • Still skips fenced code blocks and site-absolute (/...) paths.

checks/product_names.py

  • Now prose-only. Strips code spans, link targets (keeping display text), HTML tags, URLs, and path/filename-like tokens before matching, so image paths are no longer flagged as branding violations.

checks/_paths.py (new)

  • Shared helpers for both checks: reference parsing, per-segment case walk, apply_case_fix, and slugify/heading_ids for anchors.
  • Resolves .. segments before the case walk. Without this a relative path escaped the check and produced false "missing file" reports — latent, since the production caller always passes absolute paths.

Cleanup

  • Removed PyGithub from requirements.txt — declared but never imported (the agent uses requests directly).
  • Replaced stale source_of_truth entries in entities/{k1,k3,p1,p1s}.yaml with comments; the paths pointed at key_stone/ and power_stone/ directories that never existed, and nothing reads them.
  • Fixed glossary.yaml header comments that referenced non-existent files.

What this does not do

reviewer.py is intentionally untouched, so the job remains non-blocking. This PR only changes what the checks report — it is a correctness improvement, not a policy change. Turning findings into a merge gate is a separate decision (see below).

Verification

  • All ten check modules import cleanly after the refactor.
  • Reference coverage is unchanged at 671 relative references before and after the rewrite — no detection was lost.
  • Repo-wide run over 94 markdown files: 0 link/heading issues (after PR docs: fix K3 CoM260 image reference casing for case-sensitive hosts #138).
  • Sensitivity tested with deliberately faulted files: a 4-fault file and an 11-fault file were both fully caught, and a decoy inside a fenced code block was correctly ignored — confirming the check can actually fail.
  • get_errors clean on the changed modules.

Follow-up (not in this PR)

If we want this to gate merges, three non-blocking mechanisms have to be removed (continue-on-error: true, the || true, and sys.exit(0) in reviewer.py) and the job marked required in branch protection. Worth noting that a full-repo measurement shows ~48 of 94 files currently trip at least one check, so a hard gate would fail most doc PRs until a baseline or diff-scoped enforcement is in place. Recommend starting with links + headings only, which are clean today.

Risk

Low. Tooling only, no .md files touched — meaning the doc-review.yml workflow (which triggers on paths: "**/*.md") will not even run on this PR. No runtime behavior change for the site build.

links.py now verifies every path segment (directories too, not just filenames) against on-disk casing, covers HTML <a href> in addition to <img src>, validates #fragment anchors against the target page's real heading slugs, and reports the corrected reference text to copy rather than pressuring authors to rename files.

product_names.py now checks prose only, so lowercase image paths are no longer misreported as branding violations. Shared path/slug helpers extracted to _paths.py, which resolves .. segments before the case walk. Removed the unused PyGithub dependency and replaced stale entity source_of_truth/glossary references with comments.

reviewer.py is unchanged, so the job remains non-blocking.
@junlin-luo-spacemit
junlin-luo-spacemit merged commit 2c9e27f into main Sep 30, 2026
1 check passed
@junlin-luo-spacemit
junlin-luo-spacemit deleted the ci/harden-doc-review-link-checks branch September 30, 2026 02:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant