From 61f4b584b52180869e0ac9dd962c2b0e2ab7f5d9 Mon Sep 17 00:00:00 2001 From: nichinichisou Date: Wed, 30 Sep 2026 01:35:51 +0800 Subject: [PATCH 1/2] ci: build, check and publish the music data file on new master data .github/workflows/music-data.yml runs `nnnotes music-data --decoded-master` on moenotes-masterdata-sync's decoded master data, the way the story site workflow reads it, and publishes the file for the chart data page into the story site's bucket under music-data/: music-data.json, jackets/, archive//.json and the build marker build.json. It reuses the story site's triggers (repository_dispatch masterdata-updated, a daily schedule, workflow_dispatch with force and dry_run), its plan/build jobs, secrets, bucket and helpers (story_site.py, apk.sh); no new secret, no master key. plan compares the master data snapshot, the pinned deck commit, the last nnnotes commit and the script's recipe with the published build.json and ends there when nothing changed. build checks the file before anything is uploaded: the JSON Schema, provenance against the snapshot and its manifest, no fewer songs or charts, complete deck statistics and play scenario fields, finite numbers, references and texts, BGM lengths, 0.8 to 2 times the published size, and a smoke test with the chart data page's own modules (MUSIC_DATA_PLAYER_REF, to be set). The jackets and the archive copy go first, music-data.json next, build.json last, each read back by SHA-256. The gate self-test (test_music_data.py) runs in every build and locally in seconds. The workflow needs the fork synced with upstream nnnotes (music-data with the play scenarios and --decoded-master): .github/MUSIC_DATA.md. Co-Authored-By: Claude Opus 5.5 --- .github/MUSIC_DATA.md | 122 ++++ .github/scripts/music_data.py | 847 +++++++++++++++++++++++++++ .github/scripts/music_data_smoke.mjs | 94 +++ .github/scripts/songs_page.sh | 18 + .github/scripts/test_music_data.py | 443 ++++++++++++++ .github/workflows/music-data.yml | 146 +++++ 6 files changed, 1670 insertions(+) create mode 100644 .github/MUSIC_DATA.md create mode 100755 .github/scripts/music_data.py create mode 100644 .github/scripts/music_data_smoke.mjs create mode 100755 .github/scripts/songs_page.sh create mode 100644 .github/scripts/test_music_data.py create mode 100644 .github/workflows/music-data.yml diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md new file mode 100644 index 0000000..2f669c4 --- /dev/null +++ b/.github/MUSIC_DATA.md @@ -0,0 +1,122 @@ +# Music data workflow (StarMoe) + +`.github/workflows/music-data.yml` keeps the music data file of the chart data page (ournotes-player +`examples/songs`) up to date: `nnnotes music-data` of the current master data ([docs/music-data.md](../docs/music-data.md): +every song and chart with the deck model's statistics and the play scenarios), checked by quality gates and +published into the story site's bucket under `music-data/` (`https://storage.bdon.moe/moenotes/music-data/`): + +| Object | Content | Cache-Control | +|---|---|---| +| `music-data.json` | the current file | `no-cache` | +| `jackets/.webp` | every song's jacket (`--jackets`), where the page looks for them | `public, max-age=86400` | +| `archive//.json` | every published file, kept | `public, max-age=31536000, immutable` | +| `build.json` | the build marker: the file's SHA-256, size and counts, what it was made from, the gate results, the run | `no-cache` | + +A run never deletes anything from the bucket. Its helper steps are `.github/scripts/music_data.py` (with the bucket, +HTTP and master data helpers of `story_site.py`), `music_data_smoke.mjs`, `songs_page.sh` and `apk.sh`; the gate +self-test is `test_music_data.py`. Nothing outside `.github/` differs from upstream, so the fork syncs with it as +before. The Cloudflare Pages preview of the page is not part of the workflow. + +## A run + +1. **plan** (seconds): the inputs of a build, from `index.json` of moenotes-masterdata-sync (the snapshot's master + data version, resource version and client version of `MUSIC_DATA_MASTERDATA_REGION`) and the checkout (the + ournotes-deck commit `rust/Cargo.lock` pins, the last nnnotes commit that changed `src/`, `rust/` or + `pyproject.toml`, and `RECIPE` of `music_data.py`), against `inputs` of the published `build.json`. The same + inputs (and a published `music-data.json`): the run ends here. `force` builds anyway. +2. **build**: + - the chart data page's modules (`examples/songs` of `MUSIC_DATA_PLAYER_REF`, not built), nnnotes with its deck + model (the install builds `nnnotes._deck`), the gate self-test, the APK (playfetch with `PLAYFETCH_CREDENTIALS`: + `provenance.client`; nnnotes also reads the bundles the APK carries, as in the story site's builds), the decoded + master data of moenotes-masterdata-sync (every file SHA-256 checked against `index.json`, `MasterManifest.json` + included); + - `nnnotes music-data --decoded-master --jackets jackets -o music-data.json`: the master data as decoded (no + master key), `provenance.master` the manifest's version and SHA-256 of the files as served; the charts, cue + sheets and jackets from the TW catalog, downloaded afresh on every run (never `actions/cache`: nnnotes keeps a + downloaded catalog for good, and the cache holds decrypted game files); + - the gates (below); a failed gate stops the run, the job summary lists why; + - upload: the jackets the bucket lacks (or has at another size; every one with `force`), the archive copy, then + `music-data.json`, `build.json` last. The archive copy, the file and the marker are each read back and checked + against their SHA-256 before the next is written: a failure leaves the previous `build.json`, so the next run + builds again. + +Triggers: `repository_dispatch` `masterdata-updated` (moenotes-masterdata-sync's `dispatch_repositories` already +names this repository for the story site: both workflows run), a daily schedule (03:41 UTC) in case a dispatch was +missed, and `workflow_dispatch`: + +| Input | Meaning | +|---|---| +| `force` | build and publish although the published file was made from the same inputs; upload every jacket again | +| `dry_run` | build and check, then list what would be uploaded instead of uploading | + +Runs do not overlap (`concurrency: music-data`). + +## Gates + +Every one must pass, else nothing is published. Warnings go to the job summary and `build.json` and do not stop it. + +| Gate | Checks | +|---|---| +| (build) | nnnotes' own checks: every table, chart and cue sheet read, every chart measured, the deck statistics cross-checked against the chart facts and the master data (the command writes no file otherwise) | +| `schema` | the file against `docs/schema/music-data.schema.json` of the checkout (JSON Schema 2020-12) | +| `provenance` | `format`; `region` `tw`; `master.source` `api`; `master.version` equal to the snapshot's and its `MasterManifest.json`'s; every table's SHA-256 the manifest's, every decoded table read the one `index.json` lists; the song tables and the deck model's present; `deck.commit` the one `rust/Cargo.lock` pins; `exporter.version` the installed nnnotes; an APK version; a catalog SHA-256 (warning: the APK is another client version than the snapshot's) | +| `counts` | no fewer songs and charts than the published file (warning: ids no longer in it) | +| `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning), `weights[kind][position]` numbers, every check deck within its bound | +| `scenarios` | the play scenario fields: `offSeeds` exactly one entry (seed 0, score, weights, check within its bound), every range's `rankBonusPercents` five ints (the first `rankBonusPercent`), every seed's `scorePerfect`, `rangeWeights` (`[kind][position][range]`) and `rankCheck` (within its bound), every seed range's `rangeScorePerfect` (warnings, none in TW: a null `rangeWeights`, a null kind in it or in `offSeeds`' weights) | +| `finite` | no NaN or infinity (warning: one inside master data rows, `songs[].master`, which the format writes as `1e999`) | +| `references` | texts in every language of `languages` (names and titles not empty); unique ids; songs sorted; the songs' bands, vocal characters and tags in the file; a band or a band name; a jacket, and its file in `jackets/`; a BGM cue; score ranks; charts in difficulty order, score ids unique (warnings: a title without a `zh-Hant` text, a music category on no tab, a character of no band) | +| `bgm` | every song's BGM length: `durationMs = samples * 1000 // sampleRate`, 30 s to 10 min, within 1 s of the cue's `lengthMs`, not ending before a chart's last note (warning: more than a minute after it) | +| `size` | 0.8 to 2 times the published file | +| `page` | `music_data_smoke.mjs`: the page's `catalog.js` and `ranking.js` in Node.js over the file: a row per chart, a plain score-up kind, data for the free, rank and Just scenarios, finite positive figures for every chart the data covers in seven scenarios (Gekisou Live at several ranks, Just rates and a Great share, Free Live), the ranking, frontier and event figures | +| (publish) | read back after upload, SHA-256 checked | + +`counts` and `size` compare with the published `music-data.json` and are skipped while nothing is published. A +legitimate drop (a song the game removed) stops the run: a person checks it, then moves the published +`music-data.json` away (the archive keeps it) or changes the gate in a pull request. + +The self-test runs in each build and locally in seconds, without the network: + +``` +python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py +``` + +with, optionally, `MUSIC_DATA_SCHEMA` (a schema file when the checkout has none), `MUSIC_DATA_PAGE` (an +`examples/songs` directory: the smoke test), `MUSIC_DATA_SAMPLE` (a real file with the play scenario fields: its +content gates pass) and `MUSIC_DATA_OLD_SAMPLE` (one without them: the scenario gate stops it). + +## Settings + +Repository secrets: the story site's (`.github/STORY_SITE.md`), no new one: `NNNOTES_BUNDLE_KEY`, +`NNNOTES_BUNDLE_NONCE_SEED`, `NNNOTES_SERVERS_TW_CDN`, `PLAYFETCH_CREDENTIALS`, `STORY_S3_ACCESS_KEY`, +`STORY_S3_SECRET_KEY`. The master key is not needed: the master data comes decoded. + +Repository variables: + +| Variable | Default | | +|---|---|---| +| `MUSIC_DATA_PLAYER_REF` | none: **required** | the ournotes-player commit whose chart data page reads this file (the page with the play scenarios); a run stops before building without it | +| `MUSIC_DATA_PLAYER_REPOSITORY` | `empty-sekai/ournotes-player` | | +| `MUSIC_DATA_S3_PREFIX` | `music-data` | the key prefix in the bucket | +| `MUSIC_DATA_MASTERDATA_REGION` | `hk-tw-mo` | the region of `index.json`; the build reads the TW catalog (`[catalog] region` `tw`) | +| `STORY_S3_ENDPOINT`, `STORY_S3_BUCKET`, `MASTERDATA_BASE_URL`, `PLAYFETCH_VERSION`, `STORY_APK_PACKAGE` | the story site's | shared with it | + +## Before the first run + +- **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play + scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with + upstream first. Until then `plan` stops naming what is missing. +- **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play + scenario fields, once that page is merged. + +## Notes + +- **Versions.** A new ournotes-deck pin (`rust/Cargo.toml`, `rust/Cargo.lock`) or a new nnnotes commit in `src/`, + `rust/` or `pyproject.toml` reaches this fork with a sync, and the next run builds a new file. The deck + statistics may then differ: `provenance.deck.commit` and `build.json` name the commit. +- **The page and the data.** The page's modules are pinned by `MUSIC_DATA_PLAYER_REF`: after a page release that + reads new fields, move it (a new ref alone does not start a build; run with `force` to check the published data + against the new page). +- **Byte identity.** A build from the same inputs gives the same bytes (the file is canonical); the jackets' + WebP bytes depend on the Pillow version. +- **Logs.** The steps print counts, ids, SHA-256 and field names, not game content; nothing decrypted is cached or + uploaded as an artifact. diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py new file mode 100755 index 0000000..2887336 --- /dev/null +++ b/.github/scripts/music_data.py @@ -0,0 +1,847 @@ +#!/usr/bin/env python3 +"""The music data CI steps (.github/workflows/music-data.yml): `nnnotes music-data` of moenotes-masterdata-sync's +decoded master data, checked by quality gates and published into the story site's bucket under $MUSIC_DATA_S3_PREFIX +(music-data.json, jackets/, archive/, build.json) for the chart data page of ournotes-player (examples/songs). + + plan the inputs of a build (the master data snapshot of $MASTERDATA_REGION in index.json, the + deck commit rust/Cargo.lock pins, the last nnnotes commit of src/, rust/ and pyproject.toml, + RECIPE) against those of the published build.json; GitHub output `build`: true when they + differ, nothing is published or $FORCE is true + master OUT every file of the snapshot of $MASTERDATA_REGION into OUT (SHA-256 checked against + index.json, MasterManifest.json included) and its index entry as OUT.snapshot.json + build OUT `nnnotes music-data --decoded-master --jackets OUT/jackets -o OUT/music-data.json` ([paths] + master: the master step's OUT), its printed summary in OUT/music-data.summary.json + check OUT MASTER PAGE the quality gates (.github/MUSIC_DATA.md) on OUT/music-data.json, with the published file + as the baseline and PAGE the chart data page's modules (examples/songs) for the smoke test; + the report in OUT/check.json, and OUT/build.json (the build marker) when every gate passed; + exits 1 when one failed + publish OUT [--dry-run] + the jackets the bucket lacks or has at another size (every one with $FORCE), the file's + archive copy, then music-data.json, build.json last; each read back and its SHA-256 checked. + Only a file whose check.json passed; never deletes + +Bucket: story_site.Bucket ($STORY_S3_ENDPOINT, $STORY_S3_BUCKET, credentials $STORY_S3_ACCESS_KEY / +$STORY_S3_SECRET_KEY) with the key prefix $MUSIC_DATA_S3_PREFIX (default music-data). The published files are read +over plain HTTP, as the page reads them (the bucket serves public read). +""" +from __future__ import annotations + +import hashlib +import json +import math +import os +import re +import subprocess +import sys +import time +import urllib.error +import urllib.request +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import story_site # noqa: E402 (the bucket, HTTP and master data helpers) +from story_site import env, get, output, summary # noqa: E402 + +FORMAT = "nnnotes.music-data/1" +BUILD_FORMAT = "moenotes.music-data-build/1" +# This script's own version of a build: bump it when what it builds or publishes changes, so that the next run builds +# although the master data, the deck model and nnnotes are the same. +RECIPE = 1 +FILE, MARKER, JACKETS, ARCHIVE = "music-data.json", "build.json", "jackets/", "archive/" +MANIFEST = "MasterManifest.json" +SOURCE_PATHS = ("src", "rust", "pyproject.toml") # nnnotes' code: the commit that last changed one of them +SCHEMA = Path("docs/schema/music-data.schema.json") +SMOKE = Path(__file__).resolve().parent / "music_data_smoke.mjs" +ARCHIVE_CACHE = story_site.ASSET_CACHE # archive//.json: content-addressed +FILE_CACHE = "no-cache" # music-data.json and build.json change in place +JACKET_CACHE = "public, max-age=86400" +SNAPSHOT_KEYS = ("version", "resource_version", "client_version", "verified_at", "manifest_sha256", "table_count") + +# the gates' bounds (MUSIC_DATA.md) +SIZE_RATIO = (0.8, 2.0) # against the published file +BGM_MS = (30_000, 600_000) # a song's BGM length +BGM_CUE_SLACK_MS = 1000 # |durationMs - lengthMs| +BGM_TAIL_MS = 60_000 # BGM after the last note: more is reported +RANKS = 5 +DIFFICULTIES = ("easy", "normal", "hard", "expert") +SONG_TABLES = ("MasterLiveMusic", "MasterLiveMusicScore", "MasterText", "MasterBand", "MasterCharacter", "MasterTag", + "MasterLiveMusicCategory", "MasterSound", "MasterSoundCueSheet", "MasterLiveScoreRank") +SHA256 = re.compile(r"[0-9a-f]{64}") +COMMIT = re.compile(r"[0-9a-f]{40}") +LISTED = 20 # failures and warnings listed per gate + + +def fail(message: str): + sys.exit(f"music_data: {message}") + + +def sha256(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +# ---------------------------------------------------------------- bucket +def s3_prefix() -> str: + p = (os.environ.get("MUSIC_DATA_S3_PREFIX") or "music-data").strip("/") + return f"{p}/" if p else "" + + +def bucket() -> "story_site.Bucket": + b = story_site.Bucket() + b.prefix = s3_prefix() + return b + + +def public_url(key: str) -> str: + return f"{env('STORY_S3_ENDPOINT').rstrip('/')}/{env('STORY_S3_BUCKET')}/{s3_prefix()}{key}" + + +def published(key: str) -> bytes | None: + """A published file (plain HTTP), None when the bucket has none.""" + try: + return get(public_url(key), timeout=300) + except urllib.error.HTTPError as e: + if e.code == 404: + return None + raise + + +def exists(key: str) -> bool: + request = urllib.request.Request(public_url(key), method="HEAD", + headers={"User-Agent": "moenotes-music-data (GitHub Actions)", + "Cache-Control": "no-cache"}) + try: + with urllib.request.urlopen(request, timeout=60): + return True + except urllib.error.HTTPError as e: + if e.code == 404: + return False + raise + + +# ---------------------------------------------------------------- the inputs of a build +def deck_commit(root: Path = Path(".")) -> str: + """The ournotes-deck commit nnnotes builds its deck model with (rust/Cargo.lock).""" + lock = root / "rust" / "Cargo.lock" + if not lock.is_file(): + fail("no rust/Cargo.lock: this nnnotes has no music-data command with the deck model (sync the fork with " + "upstream, .github/MUSIC_DATA.md)") + for block in lock.read_text(encoding="utf-8").split("[[package]]"): + if re.search(r'^name = "ournotes-deck"$', block, re.M): + m = re.search(r'^source = "git\+[^"#]*#([0-9a-f]{40})"$', block, re.M) + if m: + return m.group(1) + fail("rust/Cargo.lock has no ournotes-deck git commit") + + +def nnnotes_commit(root: Path = Path(".")) -> str: + """The commit that last changed nnnotes' code (SOURCE_PATHS); the checkout needs its history.""" + def git(*args): + return subprocess.run(["git", "-C", str(root), *args], capture_output=True, text=True) + if git("rev-parse", "--is-shallow-repository").stdout.strip() != "false": + fail("the checkout is shallow: the nnnotes commit needs the history (actions/checkout fetch-depth: 0)") + commit = git("log", "-1", "--format=%H", "--", *SOURCE_PATHS).stdout.strip() + if not COMMIT.fullmatch(commit): + fail("no commit changed src/, rust/ or pyproject.toml") + return commit + + +def require_decoded_master(root: Path = Path(".")) -> None: + cli = root / "src" / "nnnotes" / "cli.py" + if "--decoded-master" not in cli.read_text(encoding="utf-8"): + fail("this nnnotes has no `music-data --decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with " + "upstream (.github/MUSIC_DATA.md)") + + +def inputs(entry: dict, root: Path = Path(".")) -> dict: + """What a build is made of: the master data snapshot, the deck model, nnnotes and this script.""" + return {"masterRegion": env("MASTERDATA_REGION"), "masterVersion": entry.get("version"), + "resourceVersion": entry.get("resource_version"), "clientVersion": entry.get("client_version"), + "deckCommit": deck_commit(root), "nnnotesCommit": nnnotes_commit(root), "recipe": RECIPE} + + +def short(v) -> str: + return str(v)[:12] if v is not None else "?" + + +# ---------------------------------------------------------------- plan +def cmd_plan() -> None: + _, region = story_site.master_index() + require_decoded_master() + now = inputs(region.get("entry") or {}) + raw = published(MARKER) + marker = json.loads(raw) if raw else None + have = marker.get("inputs") if isinstance(marker, dict) else None + changed = [k for k in now if not isinstance(have, dict) or have.get(k) != now[k]] + if have and not exists(FILE): + changed.append("music-data.json (missing)") + force = os.environ.get("FORCE") == "true" + build = force or bool(changed) + summary(f"### Music data\n\n- master: {now['masterVersion']} of {now['masterRegion']} (resource " + f"{now['resourceVersion']}, client {now['clientVersion']}); deck {short(now['deckCommit'])}; nnnotes " + f"{short(now['nnnotesCommit'])}; recipe {RECIPE}\n- published: " + + (f"master {have.get('masterVersion')}, deck {short(have.get('deckCommit'))}, nnnotes " + f"{short(have.get('nnnotesCommit'))}, built {marker.get('builtAt')}" if isinstance(have, dict) + else "nothing") + + "\n- this run: " + ("builds" + (" (force)" if force else "") + (f", changed: {', '.join(changed)}" + if changed and have else "") + if build else "nothing changed, nothing to build")) + output("build", "true" if build else "false") + + +# ---------------------------------------------------------------- master data +def snapshot_file(master: Path) -> Path: + return master.parent / f"{master.name}.snapshot.json" + + +def cmd_master(out: str) -> None: + url, region = story_site.master_index() + files = region["files"] + entry = region.get("entry") or {} + if MANIFEST not in files: + fail(f"index.json lists no {MANIFEST} for {env('MASTERDATA_REGION')}") + if entry.get("manifest_sha256") and entry["manifest_sha256"] != files[MANIFEST]: + fail(f"index.json: the entry's manifest_sha256 is not the SHA-256 of its {MANIFEST}") + d = Path(out) + d.mkdir(parents=True, exist_ok=True) + story_site.parallel(lambda item: story_site.fetch_table(url, item[0], item[1], d / item[0]), files.items(), 8) + version = json.loads((d / MANIFEST).read_bytes()).get("version") + if entry.get("version") is not None and str(version) != str(entry["version"]): + fail(f"{MANIFEST} has version {version}, index.json {entry['version']} (the service changed snapshots? " + f"run again)") + snapshot = {"region": env("MASTERDATA_REGION"), "entry": {k: entry.get(k) for k in SNAPSHOT_KEYS}, + "files": files} + snapshot_file(d).write_text(json.dumps(snapshot, indent=1, sort_keys=True), encoding="utf-8") + print(f"master data: {version}, {len(files)} files into {d}") + + +# ---------------------------------------------------------------- build +def cmd_build(out: str) -> None: + o = Path(out).resolve() + o.mkdir(parents=True, exist_ok=True) + nnnotes = [sys.executable, "-m", "nnnotes"] + usage = subprocess.run(nnnotes + ["music-data", "--help"], capture_output=True, text=True).stdout + if "--decoded-master" not in usage: + fail("the installed nnnotes has no `music-data --decoded-master` (sync the fork with upstream)") + cmd = nnnotes + ["music-data", "--decoded-master", "--jackets", str(o / "jackets"), "-o", str(o / FILE)] + print("+ " + " ".join(cmd[1:]), flush=True) + with open(o / "music-data.summary.json", "wb") as f: + status = subprocess.run(cmd, stdout=f).returncode + if status: + fail(f"nnnotes music-data exited with {status}") + r = json.loads((o / "music-data.summary.json").read_text(encoding="utf-8")) + summary(f"- built: {r.get('songs')} songs, {r.get('charts')} charts, {r.get('jackets')} jackets, deck " + f"{short(r.get('deck'))}, {r.get('bytes')} bytes, sha256 {short(r.get('sha256'))}") + + +# ---------------------------------------------------------------- the gates +@dataclass +class Context: + """What the gates check the file against; None: the gate (or that part of it) is skipped (the self-test).""" + region: str | None = None # [catalog] region: provenance.region + language: str | None = None # [catalog] language: a song title should have it + master: Path | None = None # the decoded master data read (MasterManifest.json, .json) + snapshot: dict | None = None # its index entry and files (the master step's OUT.snapshot.json) + deck_commit: str | None = None + nnnotes_version: str | None = None + jackets: Path | None = None + schema: Path | None = None + published: bytes | None = None # the published music-data.json (None: nothing is published) + page: Path | None = None # examples/songs of ournotes-player + file: Path | None = None # the file on disk, for the smoke test + + +class Gate: + def __init__(self): + self.failures: list[str] = [] + self.warnings: list[str] = [] + self.note = "" + + def fail(self, message: str): + self.failures.append(message) + + def warn(self, message: str): + self.warnings.append(message) + + +def charts_of(doc: dict): + for song in doc.get("songs") or []: + for chart in song.get("charts") or []: + yield song, chart + + +def where(song: dict, chart: dict) -> str: + return f"chart {chart.get('scoreId')} ({song.get('id')} {chart.get('difficulty')})" + + +def is_int(v) -> bool: + return isinstance(v, int) and not isinstance(v, bool) + + +def is_num(v) -> bool: + return (isinstance(v, (int, float)) and not isinstance(v, bool)) and math.isfinite(v) + + +def within(c) -> bool: + return (isinstance(c, dict) and is_int(c.get("exact")) and is_num(c.get("predicted")) and is_num(c.get("bound")) + and abs(c["exact"] - c["predicted"]) <= c["bound"]) + + +def gate_schema(doc, ctx: Context, g: Gate): + if ctx.schema is None: + g.note = "skipped" + return + if not ctx.schema.is_file(): + g.fail(f"no JSON Schema {ctx.schema}") + return + import jsonschema + schema = json.loads(ctx.schema.read_text(encoding="utf-8")) + validator = jsonschema.Draft202012Validator(schema) + for e in sorted(validator.iter_errors(doc), key=lambda e: list(map(str, e.absolute_path))): + g.fail(f"{'/'.join(map(str, e.absolute_path)) or '(root)'}: {e.message[:160]}") + g.note = ctx.schema.as_posix() + + +def gate_provenance(doc, ctx: Context, g: Gate): + if doc.get("format") != FORMAT: + g.fail(f"format {doc.get('format')!r}, expected {FORMAT}") + p = doc.get("provenance") or {} + if ctx.region is not None and p.get("region") != ctx.region: + g.fail(f"region {p.get('region')!r}, expected {ctx.region!r}") + m = p.get("master") or {} + if m.get("source") != "api": + g.fail(f"master.source {m.get('source')!r}, expected 'api'") + tables = m.get("tables") or {} + missing = [t for t in SONG_TABLES if t not in tables] + if missing: + g.fail(f"master.tables lacks {', '.join(missing)}") + if doc.get("deck") is not None and len(tables) <= len(SONG_TABLES): + g.fail("master.tables has only the song tables, but the file has deck statistics") + if ctx.snapshot is not None: + v = ctx.snapshot["entry"].get("version") + if m.get("version") != v: + g.fail(f"master.version {m.get('version')!r}, the snapshot's {v!r}") + if ctx.master is not None: + manifest = json.loads((ctx.master / MANIFEST).read_bytes()) + if m.get("version") != manifest.get("version"): + g.fail(f"master.version {m.get('version')!r}, {MANIFEST}'s {manifest.get('version')!r}") + listed = {f.get("name"): str(f.get("hash") or "").lower() for f in manifest.get("files") or []} + files = (ctx.snapshot or {}).get("files") or {} + for t, v in sorted(tables.items()): + if (v or {}).get("sha256") != listed.get(f"{t}.bin"): + g.fail(f"master.tables.{t}.sha256 is not {MANIFEST}'s {t}.bin") + decoded = ctx.master / f"{t}.json" + if not decoded.is_file(): + g.fail(f"no decoded table {t}.json") + elif files and sha256(decoded.read_bytes()) != files.get(f"{t}.json"): + g.fail(f"{t}.json read is not the one index.json lists") + deck = p.get("deck") + if doc.get("deck") is not None and not isinstance(deck, dict): + g.fail("provenance.deck is null, but the file has deck statistics") + if isinstance(deck, dict): + if deck.get("name") != "ournotes-deck" or not COMMIT.fullmatch(str(deck.get("commit"))): + g.fail("provenance.deck names no ournotes-deck commit") + elif ctx.deck_commit is not None and deck["commit"] != ctx.deck_commit: + g.fail(f"deck commit {deck['commit'][:12]}, rust/Cargo.lock pins {ctx.deck_commit[:12]}") + ex = p.get("exporter") or {} + if ex.get("name") != "nnnotes": + g.fail(f"exporter {ex.get('name')!r}") + elif ctx.nnnotes_version is not None and ex.get("version") != ctx.nnnotes_version: + g.fail(f"exporter version {ex.get('version')!r}, installed nnnotes {ctx.nnnotes_version!r}") + client = p.get("client") or {} + if not isinstance(client.get("versionName"), str) or not is_int(client.get("versionCode")): + g.fail("provenance.client has no APK version (was the APK read?)") + elif ctx.snapshot is not None and ctx.snapshot["entry"].get("client_version") not in (None, + client["versionName"]): + g.warn(f"the APK is {client['versionName']}, the snapshot names client " + f"{ctx.snapshot['entry']['client_version']}") + if not SHA256.fullmatch(str((p.get("catalog") or {}).get("sha256"))): + g.fail("provenance.catalog has no SHA-256") + g.note = f"master {m.get('version')}, deck {short((deck or {}).get('commit'))}, nnnotes {ex.get('version')}" + + +def baseline(ctx: Context, g: Gate): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return None + try: + return json.loads(ctx.published) + except ValueError: + g.warn("the published music-data.json is not JSON: skipped") + return None + + +def gate_counts(doc, ctx: Context, g: Gate): + old = baseline(ctx, g) + if old is None: + return + songs = {s.get("id") for s in doc.get("songs") or []} + charts = {c.get("scoreId") for _, c in charts_of(doc)} + old_songs = {s.get("id") for s in old.get("songs") or []} + old_charts = {c.get("scoreId") for _, c in charts_of(old)} + if len(songs) < len(old_songs): + g.fail(f"{len(songs)} songs, the published file has {len(old_songs)}") + if len(charts) < len(old_charts): + g.fail(f"{len(charts)} charts, the published file has {len(old_charts)}") + if old_songs - songs: + g.warn(f"songs no longer in the file: {', '.join(map(str, sorted(old_songs - songs)))}") + if old_charts - charts: + g.warn(f"charts no longer in the file: {', '.join(map(str, sorted(old_charts - charts)))}") + g.note = f"songs {len(old_songs)} -> {len(songs)}, charts {len(old_charts)} -> {len(charts)}" + + +def gate_size(doc, ctx: Context, g: Gate, raw: bytes): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return + ratio = len(raw) / max(1, len(ctx.published)) + if not SIZE_RATIO[0] <= ratio <= SIZE_RATIO[1]: + g.fail(f"{len(raw)} bytes, {ratio:.2f} times the published {len(ctx.published)} (bounds {SIZE_RATIO[0]} to " + f"{SIZE_RATIO[1]})") + g.note = f"{len(ctx.published)} -> {len(raw)} bytes ({ratio:.2f})" + + +def weights_shape(w, kinds: int, positions: int, nullable: bool) -> bool: + return isinstance(w, list) and len(w) == kinds and all( + (nullable and k is None) or (isinstance(k, list) and len(k) == positions and all(map(is_num, k))) for k in w) + + +def gate_deck(doc, ctx: Context, g: Gate): + deck = doc.get("deck") + if not isinstance(deck, dict): + g.fail("deck is null: no deck statistics (made with --no-deck?)") + return + kinds = len(deck.get("kinds") or []) + if not kinds: + g.fail("deck.kinds is empty") + power = (deck.get("model") or {}).get("power") + if not (is_num(power) and power > 0): + g.fail(f"deck.model.power {power!r}") + n = unplayable = seeds = 0 + for song, chart in charts_of(doc): + n += 1 + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, events = d.get("positions"), d.get("events") or [] + if len(events) != len(chart.get("skillEventsMs") or []): + g.fail(f"{w}: {len(events)} skill events, the chart has {len(chart.get('skillEventsMs') or [])}") + if not is_int(positions) or (events and positions != max(e[0] for e in events) + 1): + g.fail(f"{w}: positions {positions!r} do not match the events") + continue + if d.get("unplayable"): + unplayable += 1 + g.warn(f"{w}: unplayable with Gekisou on ({d['unplayable']})") + if d.get("seeds"): + g.fail(f"{w}: unplayable, but has Gekisou on seeds") + elif not d.get("seeds"): + g.fail(f"{w}: no seeds") + for seed in d.get("seeds") or []: + seeds += 1 + s = f"{w} seed {seed.get('seed')}" + if not is_int(seed.get("score")): + g.fail(f"{s}: score {seed.get('score')!r}") + if not weights_shape(seed.get("weights"), kinds, positions, nullable=False): + g.fail(f"{s}: weights are not [kind][position] numbers") + if len(seed.get("ranges") or []) != len(d.get("ranges") or []): + g.fail(f"{s}: {len(seed.get('ranges') or [])} range results for {len(d.get('ranges') or [])} ranges") + if not within(seed.get("check")): + g.fail(f"{s}: the check deck is not within its bound") + g.note = f"{n} charts, {kinds} kinds, {seeds} seeds, {unplayable} unplayable" + + +def gate_scenarios(doc, ctx: Context, g: Gate): + """The play scenario fields (Gekisou off, every rank, the Perfect play): offSeeds exactly one, every range's + rankBonusPercents five ints, every seed scorePerfect, rangeWeights and rankCheck, every seed range + rangeScorePerfect. A null rangeWeights, a null kind in it or in offSeeds' weights is a warning (none in TW).""" + kinds = len(((doc.get("deck") or {}).get("kinds")) or []) + null_rw = null_kind = null_off = 0 + for song, chart in charts_of(doc): + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, ranges = d.get("positions"), d.get("ranges") or [] + off = d.get("offSeeds") + if not isinstance(off, list) or len(off) != 1: + g.fail(f"{w}: offSeeds {'missing' if off is None else f'has {len(off)} entries'}, expected exactly one") + else: + o = off[0] + if o.get("seed") != 0 or not is_int(o.get("score")): + g.fail(f"{w}: Gekisou off seed {o.get('seed')!r} score {o.get('score')!r}") + if not weights_shape(o.get("weights"), kinds, positions, nullable=True): + g.fail(f"{w}: Gekisou off weights are not [kind][position] numbers") + else: + null_off += sum(k is None for k in o["weights"]) + if not within(o.get("check")): + g.fail(f"{w}: the Gekisou off check deck is not within its bound") + for i, r in enumerate(ranges): + p = r.get("rankBonusPercents") + if not (isinstance(p, list) and len(p) == RANKS and all(map(is_int, p))): + g.fail(f"{w} range {i}: rankBonusPercents {'missing' if p is None else 'not five ints'}") + elif p[0] != r.get("rankBonusPercent"): + g.fail(f"{w} range {i}: rankBonusPercents[0] {p[0]} is not rankBonusPercent " + f"{r.get('rankBonusPercent')}") + for seed in d.get("seeds") or []: + s = f"{w} seed {seed.get('seed')}" + absent = [k for k in ("scorePerfect", "rangeWeights", "rankCheck") if k not in seed] + if absent: + g.fail(f"{s}: no {', '.join(absent)}") + continue + if not is_int(seed["scorePerfect"]): + g.fail(f"{s}: scorePerfect {seed['scorePerfect']!r}") + for i, r in enumerate(seed.get("ranges") or []): + if not is_int(r.get("rangeScorePerfect")): + why = "missing" if "rangeScorePerfect" not in r else "not an int" + g.fail(f"{s} range {i}: rangeScorePerfect {why}") + rw = seed["rangeWeights"] + if rw is None: + null_rw += 1 + elif not (isinstance(rw, list) and len(rw) == kinds and all( + k is None or (isinstance(k, list) and len(k) == positions and all( + isinstance(x, list) and len(x) == len(ranges) and all(map(is_num, x)) for x in k)) + for k in rw)): + g.fail(f"{s}: rangeWeights are not [kind][position][range] numbers") + else: + null_kind += sum(k is None for k in rw) + rc = seed["rankCheck"] + if rc is not None: + if len(rc.get("ranks") or []) != len(ranges) or not all( + is_int(x) and 1 <= x <= RANKS for x in rc.get("ranks") or []): + g.fail(f"{s}: rankCheck ranks are not one rank per range") + elif not within(rc): + g.fail(f"{s}: the rank check deck is not within its bound") + if null_rw: + g.warn(f"{null_rw} seeds have rangeWeights null (overlapping ranges: no rank scenarios on those charts)") + if null_kind: + g.warn(f"{null_kind} rangeWeights kinds are null (conditions on the confirmed rank)") + if null_off: + g.warn(f"{null_off} Gekisou off weight kinds are null (conditions on the Gekisou state)") + g.note = "offSeeds, rankBonusPercents, scorePerfect, rangeWeights, rankCheck, rangeScorePerfect" + + +def nonfinite(v, path: str, out: list): + if isinstance(v, float): + if not math.isfinite(v): + out.append(path) + elif isinstance(v, list): + for i, x in enumerate(v): + nonfinite(x, f"{path}[{i}]", out) + elif isinstance(v, dict): + for k, x in v.items(): + nonfinite(x, f"{path}.{k}" if path else k, out) + + +def gate_finite(doc, ctx: Context, g: Gate): + """No NaN or infinity; one inside master data rows as served (songs[].master, master: `1e999` is how the format + writes a binary32 infinity) is reported, not failed.""" + found: list[str] = [] + nonfinite(doc, "", found) + for path in found: + if re.match(r"(songs\[\d+\]\.master|master)\b", path): + g.warn(f"{path}: not finite (master data as served)") + else: + g.fail(f"{path}: not finite") + g.note = f"{len(found)} non-finite numbers" + + +def gate_references(doc, ctx: Context, g: Gate): + langs = doc.get("languages") or [] + if not langs: + g.fail("no languages") + + def text(t, what: str, required: bool) -> bool: + if t is None: + if required: + g.fail(f"{what}: no text") + return False + if not isinstance(t, dict) or set(t) != set(langs) or not all(isinstance(v, str) for v in t.values()): + g.fail(f"{what}: not a text in {', '.join(langs)}") + return False + if required and not any(t.values()): + g.fail(f"{what}: empty in every language") + return False + return True + + def ids(rows, what: str) -> set: + seen = [r.get("id") for r in rows] + if len(set(seen)) != len(seen) or not all(map(is_int, seen)): + g.fail(f"{what}: ids are not unique ints") + return set(seen) + + bands = ids(doc.get("bands") or [], "bands") + characters = ids(doc.get("characters") or [], "characters") + tags = ids(doc.get("tags") or [], "tags") + ids(doc.get("categories") or [], "categories") + for b in doc.get("bands") or []: + text(b.get("name"), f"band {b.get('id')} name", True) + for c in doc.get("characters") or []: + text(c.get("name"), f"character {c.get('id')} name", True) + text(c.get("shortName"), f"character {c.get('id')} short name", False) + if c.get("bandId") not in bands: + g.warn(f"character {c.get('id')}: band {c.get('bandId')} is not in bands") + for t in doc.get("tags") or []: + text(t.get("name"), f"tag {t.get('id')} name", True) + categories = set() + for c in doc.get("categories") or []: + text(c.get("name"), f"category {c.get('id')} name", True) + categories.update(c.get("musicCategories") or []) + songs = doc.get("songs") or [] + if not songs: + g.fail("no songs") + ids(songs, "songs") + if [s.get("id") for s in songs] != sorted(s.get("id") for s in songs if is_int(s.get("id"))): + g.fail("songs are not sorted by id") + score_ids: set = set() + untitled = [] + for s in songs: + sid = f"song {s.get('id')}" + if text(s.get("title"), f"{sid} title", True) and ctx.language and not s["title"].get(ctx.language): + untitled.append(s.get("id")) + for k in ("ruby", "phonetic", "bandName", "lyricist", "composer", "arranger"): + text(s.get(k), f"{sid} {k}", False) + for key, known, what in (("bandIds", bands, "band"), ("vocalCharacterIds", characters, "character"), + ("bestMusicTagIds", tags, "tag")): + unknown = [x for x in s.get(key) or [] if x not in known] + if unknown: + g.fail(f"{sid}: {what} {', '.join(map(str, unknown))} not in the file's {what}s") + if not s.get("bandIds") and s.get("bandName") is None: + g.fail(f"{sid}: neither a band nor a band name") + other = [x for x in s.get("musicCategories") or [] if x not in categories] + if other: + g.warn(f"{sid}: music categories {', '.join(map(str, other))} are on no category tab") + jacket = s.get("jacket") + if not isinstance(jacket, str) or not jacket: + g.fail(f"{sid}: no jacket") + elif ctx.jackets is not None: + f = ctx.jackets / f"{jacket}.webp" + if not f.is_file() or not f.stat().st_size: + g.fail(f"{sid}: no jacket file jackets/{jacket}.webp") + bgm = s.get("bgm") or {} + if not (is_int(bgm.get("soundId")) and bgm.get("cueSheet") and bgm.get("cue")): + g.fail(f"{sid}: no BGM cue") + ranks = [r.get("rank") for r in s.get("scoreRanks") or []] + if not ranks or len(set(ranks)) != len(ranks): + g.fail(f"{sid}: score ranks {ranks!r}") + charts = s.get("charts") or [] + order = [c.get("difficulty") for c in charts] + if not charts or order != [d for d in DIFFICULTIES if d in order] or len(set(order)) != len(order): + g.fail(f"{sid}: charts {order!r}") + for c in charts: + if c.get("scoreId") in score_ids: + g.fail(f"{where(s, c)}: the score id occurs twice") + score_ids.add(c.get("scoreId")) + if untitled: + g.warn(f"songs without a {ctx.language} title: {', '.join(map(str, untitled))}") + g.note = (f"{len(songs)} songs, {len(bands)} bands, {len(characters)} characters, {len(tags)} tags" + + ("" if ctx.jackets is None else ", jackets present")) + + +def gate_bgm(doc, ctx: Context, g: Gate): + lengths = [] + for s in doc.get("songs") or []: + sid = f"song {s.get('id')}" + L = (s.get("bgm") or {}).get("length") + if not isinstance(L, dict): + g.fail(f"{sid}: no BGM length") + continue + dur, samples, rate = L.get("durationMs"), L.get("samples"), L.get("sampleRate") + if not (is_int(dur) and is_int(samples) and is_int(rate) and rate > 0 and samples > 0): + g.fail(f"{sid}: BGM length {dur!r} ms, {samples!r} samples at {rate!r} Hz") + continue + if dur != samples * 1000 // rate: + g.fail(f"{sid}: BGM durationMs {dur} is not samples * 1000 // sampleRate") + if not BGM_MS[0] <= dur <= BGM_MS[1]: + g.fail(f"{sid}: BGM of {dur} ms (bounds {BGM_MS[0]} to {BGM_MS[1]})") + if is_int(L.get("lengthMs")) and abs(dur - L["lengthMs"]) > BGM_CUE_SLACK_MS: + g.fail(f"{sid}: BGM stream {dur} ms, cue length {L['lengthMs']} ms") + last = max((c.get("lastNoteMs") or 0 for c in s.get("charts") or []), default=0) + if dur < last: + g.fail(f"{sid}: BGM of {dur} ms ends before the last note at {last} ms") + elif dur - last > BGM_TAIL_MS: + g.warn(f"{sid}: BGM plays {dur - last} ms after the last note") + lengths.append(dur) + if lengths: + g.note = f"{min(lengths) / 1000:.1f} s to {max(lengths) / 1000:.1f} s" + + +def gate_page(doc, ctx: Context, g: Gate): + """The chart data page's own modules (catalog.js, ranking.js) over the file in Node.js: music_data_smoke.mjs.""" + if ctx.page is None: + g.note = "skipped" + return + if not (ctx.page / "catalog.js").is_file() or not (ctx.page / "ranking.js").is_file(): + g.fail(f"{ctx.page}: no catalog.js / ranking.js (the chart data page, examples/songs)") + return + r = subprocess.run(["node", str(SMOKE), str(ctx.page), str(ctx.file)], capture_output=True, text=True, + timeout=600) + lines = [x for x in (r.stdout + r.stderr).splitlines() if x.strip()] + if r.returncode: + for x in lines or [f"node exited with {r.returncode}"]: + g.fail(x[:300]) + else: + g.note = lines[-1][:200] if lines else "passed" + + +GATES = (("schema", gate_schema), ("provenance", gate_provenance), ("counts", gate_counts), ("deck", gate_deck), + ("scenarios", gate_scenarios), ("finite", gate_finite), ("references", gate_references), ("bgm", gate_bgm), + ("size", gate_size), ("page", gate_page)) + + +def gates(raw: bytes, ctx: Context, only=None) -> dict: + """The report of the gates (`only`: their names) on the file's bytes.""" + try: + doc = json.loads(raw.decode("utf-8")) + if not isinstance(doc, dict): + raise ValueError("not an object") + except ValueError as e: + doc, results = None, [{"gate": "json", "passed": False, "failures": [f"not JSON: {str(e)[:160]}"], + "failureCount": 1, "warnings": [], "warningCount": 0, "note": ""}] + if doc is not None: + results = [] + for name, fn in GATES: + if only is not None and name not in only: + continue + g = Gate() + try: + fn(doc, ctx, g, raw) if name == "size" else fn(doc, ctx, g) + except Exception as e: # a malformed file the gate did not foresee fails the gate + g.fail(f"the gate stopped: {type(e).__name__}: {str(e)[:160]}") + results.append({"gate": name, "passed": not g.failures, "failures": g.failures[:LISTED], + "failureCount": len(g.failures), "warnings": g.warnings[:LISTED], + "warningCount": len(g.warnings), "note": g.note}) + return {"passed": all(r["passed"] for r in results), "sha256": sha256(raw), "bytes": len(raw), + "gates": results} + + +def report_markdown(report: dict) -> str: + rows = ["| Gate | Result | |", "|---|---|---|"] + for r in report["gates"]: + result = "passed" if r["passed"] else f"**failed** ({r['failureCount']})" + if r["warningCount"]: + result += f", {r['warningCount']} warnings" + rows.append(f"| {r['gate']} | {result} | {r['note']} |") + details = [] + for r in report["gates"]: + for kind, key, n in (("failure", "failures", "failureCount"), ("warning", "warnings", "warningCount")): + for x in r[key]: + details.append(f"- {r['gate']} {kind}: {x}") + if r[n] > len(r[key]): + details.append(f"- {r['gate']}: {r[n] - len(r[key])} more {kind}s") + return "\n".join(["", "#### Gates", ""] + rows + ([""] + details if details else [])) + + +def cmd_check(out: str, master: str, page: str) -> None: + import nnnotes + o, m = Path(out), Path(master) + raw = (o / FILE).read_bytes() + snapshot = json.loads(snapshot_file(m).read_text(encoding="utf-8")) + ctx = Context(region=env("NNNOTES_CATALOG_REGION"), language=os.environ.get("NNNOTES_CATALOG_LANGUAGE"), + master=m, snapshot=snapshot, deck_commit=deck_commit(), nnnotes_version=nnnotes.__version__, + jackets=o / "jackets", schema=SCHEMA, published=published(FILE), page=Path(page), file=o / FILE) + report = gates(raw, ctx) + (o / "check.json").write_text(json.dumps(report, indent=1), encoding="utf-8") + summary(report_markdown(report)) + if not report["passed"]: + fail("a gate failed: nothing is published (the job summary lists why)") + doc = json.loads(raw) + p = doc["provenance"] + version = re.sub(r"[^A-Za-z0-9._-]", "_", str(p["master"]["version"])) + marker = { + "format": BUILD_FORMAT, + "file": FILE, "sha256": report["sha256"], "bytes": report["bytes"], + "archive": f"{ARCHIVE}{version}/{report['sha256']}.json", + "songs": len(doc["songs"]), "charts": sum(len(s["charts"]) for s in doc["songs"]), + "jackets": len(list((o / "jackets").glob("*.webp"))), + "inputs": inputs(snapshot["entry"]), + "provenance": {"region": p["region"], "client": p["client"], "masterVersion": p["master"]["version"], + "deckCommit": p["deck"]["commit"], "exporterVersion": p["exporter"]["version"]}, + "player": {"repository": os.environ.get("PLAYER_REPOSITORY"), "ref": os.environ.get("PLAYER_REF")}, + "gates": {r["gate"]: {"passed": r["passed"], "warnings": r["warningCount"]} for r in report["gates"]}, + "builtAt": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"), + "run": (f"{os.environ['GITHUB_SERVER_URL']}/{os.environ['GITHUB_REPOSITORY']}/actions/runs/" + f"{os.environ['GITHUB_RUN_ID']}" if os.environ.get("GITHUB_RUN_ID") else None), + } + (o / MARKER).write_text(json.dumps(marker, indent=1) + "\n", encoding="utf-8") + + +# ---------------------------------------------------------------- publish +def upload(b, key: str, src: Path, cache: str) -> None: + extra = {"ContentType": story_site.TYPES.get(src.suffix.lower(), "application/octet-stream"), + "CacheControl": cache} + b.s3.upload_file(str(src), b.name, b.prefix + key, ExtraArgs=extra) + + +def read_back(b, key: str, digest: str) -> None: + """The object as stored (S3 GET), else as served (plain HTTP), must have the SHA-256 uploaded.""" + for attempt in range(4): + try: + if sha256(b.s3.get_object(Bucket=b.name, Key=b.prefix + key)["Body"].read()) == digest: + return + except Exception as e: # Cloudflare in front of the store rejects a signed GET at times + print(f"read back {key}: {type(e).__name__}", flush=True) + time.sleep(3 * (attempt + 1)) + try: + if sha256(get(public_url(key), timeout=300)) == digest: + return + except urllib.error.URLError as e: + print(f"read back {key} over HTTP: {e}", flush=True) + fail(f"{key}: the bucket does not serve what was uploaded (SHA-256 {digest[:12]})") + + +def cmd_publish(out: str, dry_run: bool = False) -> None: + o = Path(out) + raw = (o / FILE).read_bytes() + report = json.loads((o / "check.json").read_text(encoding="utf-8")) + if not report.get("passed") or report.get("sha256") != sha256(raw) or not (o / MARKER).is_file(): + fail("the file has not passed the gates (check.json): nothing is published") + marker = json.loads((o / MARKER).read_text(encoding="utf-8")) + b = bucket() + if not b.writable and not dry_run: + fail("publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") + force = os.environ.get("FORCE") == "true" + have = b.keys(JACKETS) + jackets = sorted((o / "jackets").glob("*.webp")) + new = [j for j in jackets if force or have.get(JACKETS + j.name) != j.stat().st_size] + archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) + steps = [(JACKETS + j.name, j, JACKET_CACHE, False) for j in new] + if not archived: + steps.append((marker["archive"], o / FILE, ARCHIVE_CACHE, True)) + # the file before the marker that names it, the jackets and the archive copy before the file + steps += [(FILE, o / FILE, FILE_CACHE, True), (MARKER, o / MARKER, FILE_CACHE, True)] + if dry_run: + for key, _, _, _ in steps: + print(f"would upload {b.prefix}{key}") + else: + story_site.parallel(lambda s: upload(b, s[0], s[1], s[2]), [s for s in steps if not s[3]]) + for key, src, cache, check in steps: + if check: + upload(b, key, src, cache) + read_back(b, key, sha256(src.read_bytes())) + summary(f"- {'would publish' if dry_run else 'published'} {len(new)} jackets (of {len(jackets)}), " + f"{'the archive copy, ' if not archived else ''}{FILE} ({len(raw)} bytes, sha256 " + f"{short(marker['sha256'])}), {MARKER}: {public_url(FILE)}") + + +def main(argv: list[str]) -> None: + if not argv: + sys.exit(__doc__) + cmd, args = argv[0], argv[1:] + if cmd == "plan" and not args: + cmd_plan() + elif cmd == "master" and len(args) == 1: + cmd_master(args[0]) + elif cmd == "build" and len(args) == 1: + cmd_build(args[0]) + elif cmd == "check" and len(args) == 3: + cmd_check(*args) + elif cmd == "publish" and len(args) in (1, 2) and args[1:] in ([], ["--dry-run"]): + cmd_publish(args[0], dry_run=bool(args[1:])) + else: + sys.exit(__doc__) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/.github/scripts/music_data_smoke.mjs b/.github/scripts/music_data_smoke.mjs new file mode 100644 index 0000000..a711edc --- /dev/null +++ b/.github/scripts/music_data_smoke.mjs @@ -0,0 +1,94 @@ +// The chart data page's smoke test (music_data.py check, gate "page"): the page's own pure modules (ournotes-player +// examples/songs: catalog.js, ranking.js) over a music-data.json in Node.js, the way the page reads it. Every chart +// must get a row, every playable chart its figures in every play scenario (Gekisou Live at several ranks and Just +// rates, Free Live, a Great share), finite and positive, and the rankings must work on them. +// +// node music_data_smoke.mjs +// +// Prints one line per problem (at most 40) and exits 1, or a one-line summary. +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { pathToFileURL } from "node:url"; + +const [dir, file] = process.argv.slice(2); +if (!dir || !file) { + console.error("usage: node music_data_smoke.mjs "); + process.exit(2); +} +const catalog = await import(pathToFileURL(path.join(dir, "catalog.js")).href); +const ranking = await import(pathToFileURL(path.join(dir, "ranking.js")).href); +const data = JSON.parse(readFileSync(file, "utf8")); + +const problems = []; +const problem = (m) => problems.push(m); +const finite = (v) => typeof v === "number" && Number.isFinite(v); + +const charts = (data.songs || []).flatMap((s) => (s.charts || []).map((c) => ({ song: s, chart: c }))); +const rows = catalog.chartRows(data); +if (rows.length !== charts.length) problem(`chartRows: ${rows.length} rows for ${charts.length} charts`); +const kind = ranking.plainKind(data); +if (kind === null) problem("plainKind: the file has no plain score-up kind (effect 2000, 5 s, no targets)"); + +// Whether the data has what a scenario needs on a chart (a null rangeWeights or kind leaves a chart without figures in +// the rank and Just scenarios: music_data.py reports those as warnings). +const computable = (deck, scenario) => { + if (!deck) return false; + const has = (s) => s && Array.isArray((s.weights || [])[kind]); + if (scenario && scenario.mode === "free") return (deck.offSeeds || []).length > 0 && deck.offSeeds.every(has); + if (deck.unplayable || !(deck.seeds || []).length) return false; + const rank1 = !scenario || ((scenario.ranks || []).every((r) => r === 1) && (scenario.just ?? 1) >= 1); + return deck.seeds.every((s) => has(s) && (rank1 || Array.isArray((s.rangeWeights || [])[kind]))); +}; +const has = ranking.scenarioData(data); +for (const k of ["free", "ranks", "just"]) if (!has[k]) problem(`scenarioData: no data for the ${k} scenario`); + +const SCENARIOS = [ + ["Gekisou Live, rank 1", null], + ["Gekisou Live, rank 5", { mode: "battle", ranks: [5, 5, 5], just: 1, great: 0 }], + ["Gekisou Live, ranks 2 3 4", { mode: "battle", ranks: [2, 3, 4], just: 1, great: 0 }], + ["Gekisou Live, Just 0", { mode: "battle", ranks: [1, 1, 1], just: 0, great: 0 }], + ["Gekisou Live, rank 3, Just 0.5, Great 0.2", { mode: "battle", ranks: [3, 3, 3], just: 0.5, great: 0.2 }], + ["Free Live", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0 }], + ["Free Live, Great 0.5", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0.5 }], +]; +const SKILLS = [1, 1, 1, 1, 1]; +let figures = 0; +for (const [name, scenario] of SCENARIOS) { + const joined = ranking.joinCharts(data, scenario); + const expected = charts.filter(({ chart }) => computable(chart.deck, scenario)).length; + if (joined.length !== expected) problem(`${name}: ${joined.length} charts with figures, expected ${expected}`); + for (const r of joined) { + figures++; + if (!finite(r.base) || r.base <= 0) problem(`${name}: chart ${r.scoreId}: base ${r.base}`); + if (!Array.isArray(r.weights) || !r.weights.length || !r.weights.every(finite)) { + problem(`${name}: chart ${r.scoreId}: weights are not finite numbers`); + } + } + if (!joined.length) continue; + const ranked = ranking.rank(joined, { skills: SKILLS, source: "bgm", overheadMs: 30000 }); + if (!ranked.some((r) => r.frontier)) problem(`${name}: no chart on the frontier`); + for (const r of ranked) { + if (r.lengthMs === null) problem(`${name}: chart ${r.scoreId}: no play length`); + else if (!finite(r.perMinute) || !finite(r.rate)) problem(`${name}: chart ${r.scoreId}: rate ${r.rate}`); + } + for (const room of [0, 5]) { + ranking.eventDominance(ranked, "bgm", ranking.X_MAX, room); + for (const r of ranked) { + const p = ranking.requiredPower(r, SKILLS, "S", 1, room); + if (p !== null && !(finite(p) && p >= 0)) problem(`${name}: chart ${r.scoreId}: required power ${p}`); + } + } + const one = ranked[0]; + const chance = ranking.reachChance(one, SKILLS, 300000, "S", 1, 0); + if (chance !== null && !(chance >= 0 && chance <= 1)) problem(`${name}: reach chance ${chance}`); + catalog.refigure(rows, data, scenario); +} +catalog.histogram(rows, (r) => r.level); + +if (problems.length) { + for (const m of problems.slice(0, 40)) console.log(m); + if (problems.length > 40) console.log(`${problems.length - 40} more problems`); + process.exit(1); +} +console.log(`page smoke test: ${rows.length} charts, ${SCENARIOS.length} scenarios, ${figures} chart figures; ` + + `plain kind ${kind}, scenarios free/ranks/just`); diff --git a/.github/scripts/songs_page.sh b/.github/scripts/songs_page.sh new file mode 100755 index 0000000..71849aa --- /dev/null +++ b/.github/scripts/songs_page.sh @@ -0,0 +1,18 @@ +#!/usr/bin/env bash +# The chart data page of ournotes-player ($PLAYER_REPOSITORY at $PLAYER_REF: examples/songs, the page that reads +# music-data.json) into $1, for the smoke test of music_data.py check. Only that directory is checked out and nothing +# is built: the test imports the page's pure modules (catalog.js, ranking.js) in Node.js. +set -euo pipefail +dir="$1" +if [ -z "${PLAYER_REF:-}" ]; then + echo "::error::set the repository variable MUSIC_DATA_PLAYER_REF: the ournotes-player commit whose chart data page reads this music data (.github/MUSIC_DATA.md)" + exit 1 +fi +rm -rf "$dir" +git clone --quiet --filter=blob:none --no-checkout "https://github.com/$PLAYER_REPOSITORY.git" "$dir" +git -C "$dir" sparse-checkout set examples/songs +git -C "$dir" checkout --quiet "$PLAYER_REF" +for f in catalog.js ranking.js; do + [ -f "$dir/examples/songs/$f" ] || { echo "::error::$PLAYER_REPOSITORY $PLAYER_REF has no examples/songs/$f"; exit 1; } +done +echo "ournotes-player $(git -C "$dir" rev-parse --short HEAD): examples/songs" diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py new file mode 100644 index 0000000..e601026 --- /dev/null +++ b/.github/scripts/test_music_data.py @@ -0,0 +1,443 @@ +"""Self-test of the music data gates (music_data.py): a synthetic music-data.json passes every gate, and each gate +fails on the defect it is there for. Runs in seconds, without the network: + + python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + +The JSON Schema gate uses docs/schema/music-data.schema.json (or $MUSIC_DATA_SCHEMA) when the checkout has it; the +page smoke test runs when $MUSIC_DATA_PAGE names ournotes-player's examples/songs (and Node.js is installed). Real +files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass) and +$MUSIC_DATA_OLD_SAMPLE (one without them: the scenario gate stops it). +""" +import copy +import hashlib +import importlib.util +import io +import json +import os +import shutil +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import music_data # noqa: E402 +from music_data import Context, gates # noqa: E402 + +ROOT = Path(__file__).resolve().parents[2] +LANGS = ["ja", "en", "zh-Hant", "zh-Hans", "ko"] +TABLES = music_data.SONG_TABLES + ("MasterLiveSkillEffect", "MasterLiveGekisouRankingScoreBonus") +BIN = {t: hashlib.sha256(t.encode()).hexdigest() for t in TABLES} # the manifest's hashes of the served files +DECK = "d4" * 20 +CONTENT = ("deck", "scenarios", "finite", "references", "bgm") + + +def schema_path(): + p = Path(os.environ.get("MUSIC_DATA_SCHEMA") or ROOT / music_data.SCHEMA) + if not p.is_file(): + return None + return p if importlib.util.find_spec("jsonschema") else None + + +def page_path(): + p = os.environ.get("MUSIC_DATA_PAGE") + return Path(p) if p and Path(p, "ranking.js").is_file() and shutil.which("node") else None + + +# ---------------------------------------------------------------- a synthetic file +def text(stem): + return {lang: f"{stem}-{lang}" for lang in LANGS} + + +def check_deck(exact): + return {"deck": [[0, 5000], None], "exact": exact, "predicted": exact + 0.25, "bound": 7.0} + + +def deck_chart(): + seed = {"seed": 0, "score": 120000, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, + "justCount": 0, "lotResults": [0, 0, 0, 0], + "rangeScorePerfect": 4000}], + "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000, + "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} + return {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, + "ranges": [{"index": 0, "mission": 1, "startMs": 1000, "endMs": 5000, "rankBonusPercent": 250, + "rankBonusPercents": [250, 190, 160, 100, 100]}], + "justNotes": 0, "seeds": [seed], + "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], + "unplayable": None} + + +def chart(difficulty, score_id, last=60000): + return {"difficulty": difficulty, "scoreId": score_id, "level": 10, "displayLevel": 10.5, "fullComboCount": 20, + "asset": {"key": f"Live/MusicScore/c/c_{score_id}", "sha256": "ab" * 32}, + "notes": {"judged": 20, "total": 22, "byOperateType": {"1": 20, "120": 2}}, + "bpm": {"main": 120.0, "min": 120.0, "max": 120.0, "changes": [{"timeMs": 0, "bpm": 120.0}]}, + "firstNoteMs": 1000, "lastJudgedNoteMs": last, "lastNoteMs": last, "musicLengthMs": last + 1000, + "skillEventsMs": [1000, 3000], "fevers": [[1000, 5000]], "deck": deck_chart()} + + +def song(i, charts): + ranks = ["D", "C", "B", "A", "S", "SS"] + return {"id": i, "sortOrder": i, "startAt": "2026/01/01 0:00:00", "defaultUnlock": True, "title": text(f"t{i}"), + "ruby": None, "phonetic": text(f"p{i}"), "bandIds": [1], "bandName": None, "vocalCharacterIds": [1], + "lyricist": text("l"), "composer": text("c"), "arranger": None, "musicType": 1, "musicCategories": [1], + "bestMusicTagIds": [1], "jacket": f"jkt_{i}", "gekisouMissions": [1, 2, 3], + "bgm": {"soundId": i, "cueSheet": f"Bgm{i}", "cue": f"song{i}", + "length": {"lengthMs": 90000, "samples": 48000 * 90, "sampleRate": 48000, "durationMs": 90000}}, + "scoreRanks": [{"rank": r, "requiredScore": n * 1000, "battleRequiredScore": n * 2000} + for n, r in enumerate(ranks)], + "charts": charts, "master": {"MasterLiveMusic": {"_id": i, "_rate": 1.5}, "MasterLiveScoreRank": []}} + + +def sample() -> dict: + kind = {"id": 0, "effectType": 2000, "activationTimeSecond": 5.0, "durationMs": 5000, "skillTargetIds": [], + "skillConditionGroup": 0, "skillReleaseConditionGroup": 0, "effectLimitCount": 0, + "effectExecuteLimitCount": 0, "effectExecuteLimitResetConditionGroup": 0, "rows": [1], "values": [10000]} + return { + "format": "nnnotes.music-data/1", + "provenance": { + "region": "tw", "client": {"versionName": "1.0.1", "versionCode": 25}, + "catalog": {"resourceVersion": None, "sha256": "cd" * 32}, + "master": {"source": "api", "version": "v-test", "tables": {t: {"sha256": BIN[t]} for t in TABLES}}, + "exporter": {"name": "nnnotes", "version": "0.1.2", "chartFormat": "nnnotes.live-score/1"}, + "deck": {"name": "ournotes-deck", "version": "0.0.1", + "source": "https://github.com/empty-sekai/ournotes-deck", "commit": DECK, + "format": "ournotes-deck.chart-stats/2"}}, + "languages": LANGS, + "bands": [{"id": 1, "name": text("band"), "mainColor": "#3388BB", "subColor": "#FFFFFF"}], + "characters": [{"id": 1, "bandId": 1, "name": text("ch"), "shortName": text("c"), "mainColor": "#77BBDD"}], + "tags": [{"id": 1, "name": text("tag")}], + "categories": [{"id": 1, "musicCategories": [1], "name": text("cat")}], + "deck": {"model": {"power": 300000, "checkPower": 1000003}, "kinds": [kind]}, + "songs": [song(100001, [chart("easy", 10), chart("expert", 30)]), song(100002, [chart("expert", 40)])], + } + + +def context(tmp_path: Path, doc: dict, **kw) -> tuple[bytes, Context]: + """The file's bytes and what check reads next to it: the decoded master data and its snapshot, the jackets (the + page smoke test only with page=page_path(): it starts Node.js).""" + master = tmp_path / "master" + master.mkdir(exist_ok=True) + files = {} + for t in TABLES: + data = json.dumps({"_allData": []}).encode() + (master / f"{t}.json").write_bytes(data) + files[f"{t}.json"] = hashlib.sha256(data).hexdigest() + listed = [{"name": f"{t}.bin", "hash": BIN[t], "size": 1} for t in TABLES] + (master / music_data.MANIFEST).write_text(json.dumps({"version": "v-test", "files": listed}), encoding="utf-8") + jackets = tmp_path / "jackets" + jackets.mkdir(exist_ok=True) + for s in doc.get("songs") or []: + if s.get("jacket"): + (jackets / f"{s['jacket']}.webp").write_bytes(b"RIFF....WEBP") + # an infinity as nnnotes writes it (1e999: JSON that JavaScript reads too); a NaN stays NaN (nnnotes writes none) + raw = json.dumps(doc, ensure_ascii=False).replace("Infinity", "1e999").encode("utf-8") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(region="tw", language="zh-Hant", master=master, + snapshot={"entry": {"version": "v-test", "client_version": "1.0.1"}, "files": files}, + deck_commit=DECK, nnnotes_version="0.1.2", jackets=jackets, schema=schema_path(), published=None, + page=None, file=tmp_path / "music-data.json") + for k, v in kw.items(): + setattr(ctx, k, v) + return raw, ctx + + +def run(tmp_path, doc, **kw) -> dict: + return gates(*context(tmp_path, doc, **kw)) + + +def gate(report: dict, name: str) -> dict: + return next(g for g in report["gates"] if g["gate"] == name) + + +def failures(report: dict) -> list: + return [(g["gate"], g["failures"]) for g in report["gates"] if not g["passed"]] + + +def seed(doc, song=0, chart=0): + return doc["songs"][song]["charts"][chart]["deck"]["seeds"][0] + + +# ---------------------------------------------------------------- the gates +def test_the_sample_passes_every_gate(tmp_path): + r = run(tmp_path, sample(), page=page_path()) + assert r["passed"], failures(r) + assert [g["gate"] for g in r["gates"]] == [name for name, _ in music_data.GATES] + assert not [g for g in r["gates"] if g["warningCount"]] + assert r["sha256"] == hashlib.sha256((tmp_path / "music-data.json").read_bytes()).hexdigest() + + +def test_schema(tmp_path): + if schema_path() is None: + pytest.skip("no docs/schema/music-data.schema.json (a fork not synced with upstream) or no jsonschema") + doc = sample() + doc["songs"][0]["charts"][0]["level"] = "10" + g = gate(run(tmp_path, doc), "schema") + assert not g["passed"] and any("songs/0/charts/0/level" in f for f in g["failures"]) + + +def test_page_smoke(tmp_path): + if page_path() is None: + pytest.skip("set MUSIC_DATA_PAGE to ournotes-player's examples/songs (and install Node.js)") + assert gate(run(tmp_path, sample(), page=page_path()), "page")["passed"] + doc = sample() + for s in doc["songs"]: + for c in s["charts"]: + c["deck"].pop("offSeeds") + g = gate(run(tmp_path, doc, page=page_path()), "page") + assert not g["passed"] and any("free scenario" in f for f in g["failures"]) + + +def unplayable(doc): + doc["songs"][1]["charts"][0]["deck"]["unplayable"] = "more than three fevers" # its seeds kept + + +@pytest.mark.parametrize("change, name, match", [ + # the play scenario fields + (lambda d: d["songs"][0]["charts"][0]["deck"].pop("offSeeds"), "scenarios", "offSeeds missing"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["offSeeds"].append({}), "scenarios", "offSeeds has 2 entries"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].pop("rankBonusPercents"), "scenarios", + "rankBonusPercents missing"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[250, 190, 160, 100]), + "scenarios", "rankBonusPercents not five ints"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[251, 190, 160, 100, 100]), + "scenarios", "is not rankBonusPercent"), + (lambda d: seed(d).pop("scorePerfect"), "scenarios", "no scorePerfect"), + (lambda d: seed(d).pop("rangeWeights"), "scenarios", "no rangeWeights"), + (lambda d: seed(d).pop("rankCheck"), "scenarios", "no rankCheck"), + (lambda d: seed(d)["ranges"][0].pop("rangeScorePerfect"), "scenarios", "rangeScorePerfect missing"), + (lambda d: seed(d).update(rangeWeights=[[[0.1]]]), "scenarios", "rangeWeights are not"), + (lambda d: seed(d)["rankCheck"].update(exact=99999), "scenarios", "rank check deck is not within"), + # the deck statistics + (lambda d: d.update(deck=None), "deck", "deck is null"), + (lambda d: d["songs"][0]["charts"][0].update(deck=None), "deck", "no deck statistics"), + (lambda d: seed(d)["check"].update(exact=99999), "deck", "check deck is not within"), + (lambda d: seed(d).update(weights=[[0.5]]), "deck", "weights are not"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(seeds=[]), "deck", "no seeds"), + (unplayable, "deck", "unplayable, but has Gekisou on seeds"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(events=[[0, 1000]]), "deck", "1 skill events"), + # numbers + (lambda d: seed(d)["weights"][0].__setitem__(1, float("nan")), "finite", "weights[0][1]: not finite"), + (lambda d: d["songs"][1]["charts"][0]["bpm"].update(main=float("inf")), "finite", "bpm.main: not finite"), + # references and texts + (lambda d: d["songs"][1]["bandIds"].append(9), "references", "band 9 not in"), + (lambda d: d["songs"][1]["vocalCharacterIds"].append(7), "references", "character 7 not in"), + (lambda d: d["songs"][0].update(title=None), "references", "title: no text"), + (lambda d: d["songs"][0].update(title={lang: "" for lang in LANGS}), "references", "empty in every language"), + (lambda d: d["songs"][0]["composer"].pop("ko"), "references", "composer: not a text"), + (lambda d: d["songs"][0].update(jacket=None), "references", "no jacket"), + (lambda d: d["songs"].reverse(), "references", "not sorted"), + (lambda d: d["songs"][1]["charts"][0].update(scoreId=10), "references", "occurs twice"), + (lambda d: d["songs"][0]["charts"].reverse(), "references", "charts ['expert', 'easy']"), + (lambda d: d["songs"][0].update(scoreRanks=[]), "references", "score ranks"), + (lambda d: d["songs"][0].update(bandIds=[]), "references", "neither a band nor a band name"), + # BGM + (lambda d: d["songs"][0]["bgm"].update(length=None), "bgm", "no BGM length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=50000, samples=48000 * 50), "bgm", + "ends before the last note"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=90500), "bgm", "is not samples"), + (lambda d: d["songs"][0]["bgm"]["length"].update(lengthMs=95000), "bgm", "cue length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=1200000, samples=48000 * 1200), "bgm", + "bounds"), + # provenance + (lambda d: d.update(format="nnnotes.music-data/2"), "provenance", "format"), + (lambda d: d["provenance"].update(region="en"), "provenance", "region 'en'"), + (lambda d: d["provenance"]["master"].update(version="other"), "provenance", "the snapshot's 'v-test'"), + (lambda d: d["provenance"]["master"].update(source="embedded"), "provenance", "master.source"), + (lambda d: d["provenance"]["master"]["tables"]["MasterText"].update(sha256="00" * 32), "provenance", + "MasterText.sha256 is not"), + (lambda d: d["provenance"]["master"]["tables"].pop("MasterBand"), "provenance", "lacks MasterBand"), + (lambda d: d["provenance"]["deck"].update(commit="ee" * 20), "provenance", "rust/Cargo.lock pins"), + (lambda d: d["provenance"]["exporter"].update(version="0.0.9"), "provenance", "installed nnnotes"), + (lambda d: d["provenance"]["client"].update(versionName=None), "provenance", "no APK version"), +]) +def test_a_gate_fails(tmp_path, change, name, match): + doc = sample() + change(doc) + r = run(tmp_path, doc) + g = gate(r, name) + assert not r["passed"] and not g["passed"] and any(match in f for f in g["failures"]), g + + +def test_a_table_read_is_not_the_listed_one(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.master / "MasterText.json").write_text('{"_allData": [{}]}', encoding="utf-8") + g = gate(gates(raw, ctx), "provenance") + assert not g["passed"] and g["failures"] == ["MasterText.json read is not the one index.json lists"] + + +def test_a_jacket_file_is_missing(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.jackets / "jkt_100002.webp").unlink() + g = gate(gates(raw, ctx), "references") + assert not g["passed"] and g["failures"] == ["song 100002: no jacket file jackets/jkt_100002.webp"] + + +def test_warnings_do_not_fail(tmp_path): + doc = sample() + seed(doc).update(rangeWeights=None, rankCheck=None) # overlapping ranges + seed(doc, 0, 1)["rangeWeights"][0] = None # a kind reading the confirmed rank + doc["songs"][1]["charts"][0]["deck"]["offSeeds"][0]["weights"][0] = None + doc["songs"][0]["master"]["MasterLiveMusic"]["_rate"] = float("inf") # 1e999 in master data as served + doc["songs"][1]["title"]["zh-Hant"] = "" + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 3 and gate(r, "finite")["warningCount"] == 1 + assert gate(r, "references")["warnings"] == ["songs without a zh-Hant title: 100002"] + + +def test_against_the_published_file(tmp_path): + doc = sample() + raw, ctx = context(tmp_path, doc) + more = copy.deepcopy(doc) + more["songs"].append(song(100003, [chart("expert", 50)])) + ctx.published = json.dumps(more).encode() + r = gates(raw, ctx) + assert not gate(r, "counts")["passed"] and gate(r, "counts")["failures"] == [ + "2 songs, the published file has 3", "3 charts, the published file has 4"] + assert gate(r, "counts")["warnings"] == ["songs no longer in the file: 100003", "charts no longer in the file: 50"] + ctx.published = raw + b" " * (len(raw) * 2) # the same songs in a file three times as big + r = gates(raw, ctx) + assert gate(r, "counts")["passed"] and not gate(r, "size")["passed"] + ctx.published = raw[:len(raw) // 3] # a third: not JSON, and twice exceeded + r = gates(raw, ctx) + assert gate(r, "counts")["warnings"] == ["the published music-data.json is not JSON: skipped"] + assert not gate(r, "size")["passed"] + ctx.published = raw + assert gates(raw, ctx)["passed"] + + +def test_not_json(tmp_path): + r = gates(b"{", Context()) + assert not r["passed"] and r["gates"][0]["gate"] == "json" + + +def test_report_markdown(tmp_path): + doc = sample() + doc["songs"][0]["charts"][0]["deck"].pop("offSeeds") + text_ = music_data.report_markdown(run(tmp_path, doc)) + assert "| scenarios | **failed** (1) |" in text_ + assert "- scenarios failure: chart 10 (100001 easy): offSeeds missing, expected exactly one" in text_ + + +def test_deck_commit_from_cargo_lock(tmp_path): + lock = tmp_path / "rust" / "Cargo.lock" + lock.parent.mkdir() + lock.write_text('[[package]]\nname = "pyo3"\nversion = "0.29.0"\n\n[[package]]\nname = "ournotes-deck"\n' + 'version = "0.0.1"\nsource = "git+https://github.com/empty-sekai/ournotes-deck?rev=' + "a" * 40 + + "#" + "b" * 40 + '"\n', encoding="utf-8") + assert music_data.deck_commit(tmp_path) == "b" * 40 + with pytest.raises(SystemExit, match="sync the fork"): + music_data.deck_commit(tmp_path / "nowhere") + + +# ---------------------------------------------------------------- real files (when named) +def real(name): + p = os.environ.get(name) + if not p or not Path(p).is_file(): + pytest.skip(f"set {name} to a music-data.json") + return Path(p).read_bytes() + + +def test_a_real_file_without_the_scenario_fields(): + raw = real("MUSIC_DATA_OLD_SAMPLE") + r = gates(raw, Context(language="zh-Hant"), only=CONTENT) + assert failures(r) == [("scenarios", gate(r, "scenarios")["failures"])] + g = gate(r, "scenarios") + charts = sum(len(s["charts"]) for s in json.loads(raw)["songs"]) + assert g["failureCount"] >= charts and all("offSeeds missing" in f or "rankBonusPercents missing" in f + or "no scorePerfect" in f for f in g["failures"]) + + +def test_a_real_file_with_the_scenario_fields(tmp_path): + raw = real("MUSIC_DATA_SAMPLE") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(language="zh-Hant", page=page_path(), file=tmp_path / "music-data.json", published=raw) + r = gates(raw, ctx, only=CONTENT + ("counts", "size", "page")) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 0 + + +# ---------------------------------------------------------------- publish (a stand-in bucket) +class FakeS3: + def __init__(self, corrupt=None): + self.store, self.log, self.corrupt = {}, [], corrupt + + def upload_file(self, src, bucket, key, ExtraArgs): + self.store[key] = b"not it" if key == self.corrupt else Path(src).read_bytes() + self.log.append((key, ExtraArgs["CacheControl"], ExtraArgs["ContentType"])) + + def get_object(self, Bucket, Key): + return {"Body": io.BytesIO(self.store[Key])} + + +class FakeBucket: + name, prefix, writable = "moenotes", "music-data/", True + + def __init__(self, s3): + self.s3 = s3 + + def keys(self, sub=""): + return {k[len(self.prefix):]: len(v) for k, v in self.s3.store.items() if k.startswith(self.prefix + sub)} + + +def published_out(tmp_path, monkeypatch, s3): + out = tmp_path / "out" + out.mkdir() + raw, ctx = context(out, sample()) + report = gates(raw, ctx) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "build.json").write_text(json.dumps({"sha256": report["sha256"], + "archive": f"archive/v-test/{report['sha256']}.json"}), + encoding="utf-8") + monkeypatch.setattr(music_data, "bucket", lambda: FakeBucket(s3)) + monkeypatch.setattr(music_data.time, "sleep", lambda s: None) + monkeypatch.setenv("STORY_S3_ENDPOINT", "https://storage.example") + monkeypatch.setenv("STORY_S3_BUCKET", "moenotes") + monkeypatch.delenv("FORCE", raising=False) + return out, report + + +def test_publish_order_and_read_back(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + music_data.cmd_publish(str(out)) + keys = [k for k, _, _ in s3.log] + archive = f"music-data/archive/v-test/{report['sha256']}.json" + assert sorted(keys[:2]) == ["music-data/jackets/jkt_100001.webp", "music-data/jackets/jkt_100002.webp"] + assert keys[2:] == [archive, "music-data/music-data.json", "music-data/build.json"] + caches = {k: c for k, c, _ in s3.log} + assert caches[archive].endswith("immutable") and caches["music-data/music-data.json"] == "no-cache" + assert s3.store["music-data/music-data.json"] == (out / "music-data.json").read_bytes() + s3.log.clear() + music_data.cmd_publish(str(out)) # again: the jackets and the archive copy are there + assert [k for k, _, _ in s3.log] == ["music-data/music-data.json", "music-data/build.json"] + monkeypatch.setenv("FORCE", "true") + s3.log.clear() + music_data.cmd_publish(str(out), dry_run=True) + assert s3.log == [] + + +def test_publish_stops_before_the_marker_when_the_file_does_not_read_back(tmp_path, monkeypatch): + s3 = FakeS3(corrupt="music-data/music-data.json") + out, _ = published_out(tmp_path, monkeypatch, s3) + + def offline(*a, **kw): + raise music_data.urllib.error.URLError("offline") + monkeypatch.setattr(music_data, "get", offline) + with pytest.raises(SystemExit, match="music-data.json: the bucket does not serve what was uploaded"): + music_data.cmd_publish(str(out)) + assert "music-data/build.json" not in s3.store + + +def test_publish_needs_passed_gates(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + (out / "check.json").write_text(json.dumps(dict(report, passed=False)), encoding="utf-8") + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "music-data.json").write_bytes(b"{}") # not the file that was checked + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + assert s3.store == {} diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml new file mode 100644 index 0000000..a5dffd9 --- /dev/null +++ b/.github/workflows/music-data.yml @@ -0,0 +1,146 @@ +name: Music data + +# Publishes the music data file of the current master data for the chart data page (ournotes-player examples/songs): +# `nnnotes music-data` (every song and chart with the deck model's statistics, docs/music-data.md) of +# moenotes-masterdata-sync's decoded master data, into the story site's bucket under music-data/: music-data.json, +# jackets/, archive//.json and the build marker build.json. A run first compares what it would +# build (the master data snapshot, the deck commit, the nnnotes commit) with the published build.json and ends there +# when nothing changed; else it builds, checks the file (quality gates and a smoke test with the page's own modules) +# and uploads it: the jackets and the archive copy first, then music-data.json, the build marker last. A file that +# fails a gate is not published. Nothing is deleted from the bucket. +# +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new +# master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds although nothing changed +# (`force`) or does a dry run. docs: .github/MUSIC_DATA.md + +on: + repository_dispatch: + types: [masterdata-updated] + workflow_dispatch: + inputs: + force: + description: "build and publish although the published file was made from the same inputs" + type: boolean + default: false + dry_run: + description: "build and check, but upload nothing" + type: boolean + default: false + schedule: + - cron: "41 3 * * *" + +permissions: + contents: read + +# One run at a time: a run replaces music-data.json and build.json. +concurrency: + group: music-data + cancel-in-progress: false + +env: + # The story site's bucket (its repository variables) and the key prefix of the music data. + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || 'music-data' }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} + # The ournotes-player whose chart data page reads the file: its modules smoke-test every build. No default yet: the + # page with the play scenarios is not merged; a run stops before building until MUSIC_DATA_PLAYER_REF is set. + PLAYER_REPOSITORY: ${{ vars.MUSIC_DATA_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.MUSIC_DATA_PLAYER_REF || '' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the bundle key and the CDN come from secrets in the build step. + NNNOTES_CATALOG_REGION: tw + NNNOTES_CATALOG_LANGUAGE: zh-Hant + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 10 + outputs: + build: ${{ steps.plan.outputs.build }} + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 # the nnnotes commit: the last one that changed src/, rust/ or pyproject.toml + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - name: Inputs against the published build + id: plan + env: + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py plan + + build: + needs: plan + if: needs.plan.outputs.build == 'true' + runs-on: ubuntu-latest + timeout-minutes: 90 + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: The chart data page + run: .github/scripts/songs_page.sh "$WORK/page" + + - uses: Swatinem/rust-cache@v2 + with: + workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - name: Install nnnotes + # builds the extension module nnnotes._deck: the deck model at the commit rust/Cargo.toml pins + run: uv pip install --system -e ".[test]" boto3 + + - name: Gate self-test + env: + MUSIC_DATA_PAGE: ${{ env.WORK }}/page/examples/songs + run: python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/music_data.py master "$WORK/master" + + - name: Build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + # A new cache on every run, never actions/cache: nnnotes keeps a downloaded catalog file for good (a cached + # one would never see a new resource version), and the cache holds decrypted game files. + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + run: python .github/scripts/music_data.py build "$WORK/out" + + - name: Gates + run: python .github/scripts/music_data.py check "$WORK/out" "$WORK/master" "$WORK/page/examples/songs" + + - name: Publish + # only after every gate passed (a failed step skips it); a dry run lists what it would upload + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py publish "$WORK/out" ${{ github.event.inputs.dry_run == 'true' && '--dry-run' || '' }} From 8877ca8c69e0e89c82e5f94e90e7b7bea808db66 Mon Sep 17 00:00:00 2001 From: nichinichisou Date: Wed, 30 Sep 2026 01:50:36 +0800 Subject: [PATCH 2/2] ci: publish the music data only with MUSIC_DATA_PUBLISH true A publishing switch, the repository variable MUSIC_DATA_PUBLISH (unset: off): unless it is `true` every run of music-data.yml, whatever its trigger, is a dry run. It builds, runs every gate and the page smoke test and lists what it would upload; the publish step gets no bucket key, and `music_data.py publish` itself refuses to upload. MUSIC_DATA.md says how to turn it on. Co-Authored-By: Claude Opus 5.5 --- .github/MUSIC_DATA.md | 13 +++++++++++-- .github/scripts/music_data.py | 24 ++++++++++++++++++++---- .github/scripts/test_music_data.py | 13 +++++++++++++ .github/workflows/music-data.yml | 14 +++++++++----- 4 files changed, 53 insertions(+), 11 deletions(-) diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md index 2f669c4..f8b22ab 100644 --- a/.github/MUSIC_DATA.md +++ b/.github/MUSIC_DATA.md @@ -51,6 +51,13 @@ missed, and `workflow_dispatch`: Runs do not overlap (`concurrency: music-data`). +**Publishing switch.** Nothing is uploaded unless the repository variable `MUSIC_DATA_PUBLISH` is `true` (unset: +off). Off, every run, whatever its trigger (the schedule, `masterdata-updated`, `workflow_dispatch` with or without +`dry_run`), is a dry run: it builds, runs every gate and the smoke test, and lists what it would upload; the publish +step then gets no bucket key (anonymous, it can only read) and `music_data.py publish` itself refuses to upload. +Turn it on (`gh variable set MUSIC_DATA_PUBLISH --body true`) once the published format is final; while it is off, +nothing being published, `plan` finds no `build.json` and every run builds. + ## Gates Every one must pass, else nothing is published. Warnings go to the job summary and `build.json` and do not stop it. @@ -95,6 +102,7 @@ Repository variables: | Variable | Default | | |---|---|---| | `MUSIC_DATA_PLAYER_REF` | none: **required** | the ournotes-player commit whose chart data page reads this file (the page with the play scenarios); a run stops before building without it | +| `MUSIC_DATA_PUBLISH` | none: off | `true`: upload; anything else: every run is a dry run | | `MUSIC_DATA_PLAYER_REPOSITORY` | `empty-sekai/ournotes-player` | | | `MUSIC_DATA_S3_PREFIX` | `music-data` | the key prefix in the bucket | | `MUSIC_DATA_MASTERDATA_REGION` | `hk-tw-mo` | the region of `index.json`; the build reads the TW catalog (`[catalog] region` `tw`) | @@ -103,10 +111,11 @@ Repository variables: ## Before the first run - **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play - scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with - upstream first. Until then `plan` stops naming what is missing. + scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6, `12df2a6`): sync the + fork with upstream first. Until then `plan` stops naming what is missing. - **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play scenario fields, once that page is merged. +- **Publishing.** Set `MUSIC_DATA_PUBLISH` to `true` last, when dry runs pass and the published format is final. ## Notes diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py index 2887336..0c336b5 100755 --- a/.github/scripts/music_data.py +++ b/.github/scripts/music_data.py @@ -18,7 +18,8 @@ publish OUT [--dry-run] the jackets the bucket lacks or has at another size (every one with $FORCE), the file's archive copy, then music-data.json, build.json last; each read back and its SHA-256 checked. - Only a file whose check.json passed; never deletes + Only a file whose check.json passed; never deletes. A dry run unless $MUSIC_DATA_PUBLISH is + `true` (the publishing switch, off by default): it lists what it would upload Bucket: story_site.Bucket ($STORY_S3_ENDPOINT, $STORY_S3_BUCKET, credentials $STORY_S3_ACCESS_KEY / $STORY_S3_SECRET_KEY) with the key prefix $MUSIC_DATA_S3_PREFIX (default music-data). The published files are read @@ -187,6 +188,7 @@ def cmd_plan() -> None: + "\n- this run: " + ("builds" + (" (force)" if force else "") + (f", changed: {', '.join(changed)}" if changed and have else "") if build else "nothing changed, nothing to build")) + summary(f"- publishing: {'on' if publishing() else 'off (MUSIC_DATA_PUBLISH is not `true`: a dry run)'}") output("build", "true" if build else "false") @@ -791,7 +793,15 @@ def read_back(b, key: str, digest: str) -> None: fail(f"{key}: the bucket does not serve what was uploaded (SHA-256 {digest[:12]})") +def publishing() -> bool: + """The publishing switch: uploads only with MUSIC_DATA_PUBLISH `true` (a repository variable, unset: off).""" + return os.environ.get("MUSIC_DATA_PUBLISH") == "true" + + def cmd_publish(out: str, dry_run: bool = False) -> None: + if not publishing() and not dry_run: + summary("- publishing is off (the repository variable MUSIC_DATA_PUBLISH is not `true`): a dry run") + dry_run = True o = Path(out) raw = (o / FILE).read_bytes() report = json.loads((o / "check.json").read_text(encoding="utf-8")) @@ -802,16 +812,22 @@ def cmd_publish(out: str, dry_run: bool = False) -> None: if not b.writable and not dry_run: fail("publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") force = os.environ.get("FORCE") == "true" - have = b.keys(JACKETS) + try: + have = b.keys(JACKETS) + archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) + except Exception as e: # a dry run without a key where the bucket lists to none + if not dry_run: + raise + print(f"cannot list the bucket ({type(e).__name__}): the dry run lists every object", flush=True) + have, archived = {}, False jackets = sorted((o / "jackets").glob("*.webp")) new = [j for j in jackets if force or have.get(JACKETS + j.name) != j.stat().st_size] - archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) steps = [(JACKETS + j.name, j, JACKET_CACHE, False) for j in new] if not archived: steps.append((marker["archive"], o / FILE, ARCHIVE_CACHE, True)) # the file before the marker that names it, the jackets and the archive copy before the file steps += [(FILE, o / FILE, FILE_CACHE, True), (MARKER, o / MARKER, FILE_CACHE, True)] - if dry_run: + if dry_run or not publishing(): for key, _, _, _ in steps: print(f"would upload {b.prefix}{key}") else: diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py index e601026..bc86b88 100644 --- a/.github/scripts/test_music_data.py +++ b/.github/scripts/test_music_data.py @@ -395,6 +395,7 @@ def published_out(tmp_path, monkeypatch, s3): monkeypatch.setenv("STORY_S3_ENDPOINT", "https://storage.example") monkeypatch.setenv("STORY_S3_BUCKET", "moenotes") monkeypatch.delenv("FORCE", raising=False) + monkeypatch.setenv("MUSIC_DATA_PUBLISH", "true") return out, report @@ -430,6 +431,18 @@ def offline(*a, **kw): assert "music-data/build.json" not in s3.store +@pytest.mark.parametrize("switch", [None, "", "false", "True", "1"]) +def test_publishing_is_off_unless_the_switch_is_true(tmp_path, monkeypatch, switch): + s3 = FakeS3() + out, _ = published_out(tmp_path, monkeypatch, s3) + if switch is None: + monkeypatch.delenv("MUSIC_DATA_PUBLISH") + else: + monkeypatch.setenv("MUSIC_DATA_PUBLISH", switch) + music_data.cmd_publish(str(out)) # a dry run, whatever the command line says + assert s3.log == [] and s3.store == {} + + def test_publish_needs_passed_gates(tmp_path, monkeypatch): s3 = FakeS3() out, report = published_out(tmp_path, monkeypatch, s3) diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml index a5dffd9..d0c9478 100644 --- a/.github/workflows/music-data.yml +++ b/.github/workflows/music-data.yml @@ -11,7 +11,8 @@ name: Music data # # Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new # master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds although nothing changed -# (`force`) or does a dry run. docs: .github/MUSIC_DATA.md +# (`force`) or does a dry run. Publishing switch: unless the repository variable MUSIC_DATA_PUBLISH is `true`, every +# run, whatever its trigger, is a dry run (it builds and checks, and uploads nothing). docs: .github/MUSIC_DATA.md on: repository_dispatch: @@ -42,6 +43,8 @@ env: STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || 'music-data' }} + # The publishing switch: anything but 'true' makes every run a dry run. + MUSIC_DATA_PUBLISH: ${{ vars.MUSIC_DATA_PUBLISH || 'false' }} # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} MASTERDATA_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} @@ -138,9 +141,10 @@ jobs: run: python .github/scripts/music_data.py check "$WORK/out" "$WORK/master" "$WORK/page/examples/songs" - name: Publish - # only after every gate passed (a failed step skips it); a dry run lists what it would upload + # Only after every gate passed (a failed step skips it). A dry run (the dry_run input, or the publishing switch + # MUSIC_DATA_PUBLISH off) lists what it would upload; it gets no bucket key: anonymous, it can only read. env: - STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} - STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + STORY_S3_ACCESS_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_ACCESS_KEY || '' }} + STORY_S3_SECRET_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_SECRET_KEY || '' }} FORCE: ${{ github.event.inputs.force }} - run: python .github/scripts/music_data.py publish "$WORK/out" ${{ github.event.inputs.dry_run == 'true' && '--dry-run' || '' }} + run: python .github/scripts/music_data.py publish "$WORK/out" ${{ (github.event.inputs.dry_run == 'true' || vars.MUSIC_DATA_PUBLISH != 'true') && '--dry-run' || '' }}