diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md new file mode 100644 index 0000000..f8b22ab --- /dev/null +++ b/.github/MUSIC_DATA.md @@ -0,0 +1,131 @@ +# Music data workflow (StarMoe) + +`.github/workflows/music-data.yml` keeps the music data file of the chart data page (ournotes-player +`examples/songs`) up to date: `nnnotes music-data` of the current master data ([docs/music-data.md](../docs/music-data.md): +every song and chart with the deck model's statistics and the play scenarios), checked by quality gates and +published into the story site's bucket under `music-data/` (`https://storage.bdon.moe/moenotes/music-data/`): + +| Object | Content | Cache-Control | +|---|---|---| +| `music-data.json` | the current file | `no-cache` | +| `jackets/.webp` | every song's jacket (`--jackets`), where the page looks for them | `public, max-age=86400` | +| `archive//.json` | every published file, kept | `public, max-age=31536000, immutable` | +| `build.json` | the build marker: the file's SHA-256, size and counts, what it was made from, the gate results, the run | `no-cache` | + +A run never deletes anything from the bucket. Its helper steps are `.github/scripts/music_data.py` (with the bucket, +HTTP and master data helpers of `story_site.py`), `music_data_smoke.mjs`, `songs_page.sh` and `apk.sh`; the gate +self-test is `test_music_data.py`. Nothing outside `.github/` differs from upstream, so the fork syncs with it as +before. The Cloudflare Pages preview of the page is not part of the workflow. + +## A run + +1. **plan** (seconds): the inputs of a build, from `index.json` of moenotes-masterdata-sync (the snapshot's master + data version, resource version and client version of `MUSIC_DATA_MASTERDATA_REGION`) and the checkout (the + ournotes-deck commit `rust/Cargo.lock` pins, the last nnnotes commit that changed `src/`, `rust/` or + `pyproject.toml`, and `RECIPE` of `music_data.py`), against `inputs` of the published `build.json`. The same + inputs (and a published `music-data.json`): the run ends here. `force` builds anyway. +2. **build**: + - the chart data page's modules (`examples/songs` of `MUSIC_DATA_PLAYER_REF`, not built), nnnotes with its deck + model (the install builds `nnnotes._deck`), the gate self-test, the APK (playfetch with `PLAYFETCH_CREDENTIALS`: + `provenance.client`; nnnotes also reads the bundles the APK carries, as in the story site's builds), the decoded + master data of moenotes-masterdata-sync (every file SHA-256 checked against `index.json`, `MasterManifest.json` + included); + - `nnnotes music-data --decoded-master --jackets jackets -o music-data.json`: the master data as decoded (no + master key), `provenance.master` the manifest's version and SHA-256 of the files as served; the charts, cue + sheets and jackets from the TW catalog, downloaded afresh on every run (never `actions/cache`: nnnotes keeps a + downloaded catalog for good, and the cache holds decrypted game files); + - the gates (below); a failed gate stops the run, the job summary lists why; + - upload: the jackets the bucket lacks (or has at another size; every one with `force`), the archive copy, then + `music-data.json`, `build.json` last. The archive copy, the file and the marker are each read back and checked + against their SHA-256 before the next is written: a failure leaves the previous `build.json`, so the next run + builds again. + +Triggers: `repository_dispatch` `masterdata-updated` (moenotes-masterdata-sync's `dispatch_repositories` already +names this repository for the story site: both workflows run), a daily schedule (03:41 UTC) in case a dispatch was +missed, and `workflow_dispatch`: + +| Input | Meaning | +|---|---| +| `force` | build and publish although the published file was made from the same inputs; upload every jacket again | +| `dry_run` | build and check, then list what would be uploaded instead of uploading | + +Runs do not overlap (`concurrency: music-data`). + +**Publishing switch.** Nothing is uploaded unless the repository variable `MUSIC_DATA_PUBLISH` is `true` (unset: +off). Off, every run, whatever its trigger (the schedule, `masterdata-updated`, `workflow_dispatch` with or without +`dry_run`), is a dry run: it builds, runs every gate and the smoke test, and lists what it would upload; the publish +step then gets no bucket key (anonymous, it can only read) and `music_data.py publish` itself refuses to upload. +Turn it on (`gh variable set MUSIC_DATA_PUBLISH --body true`) once the published format is final; while it is off, +nothing being published, `plan` finds no `build.json` and every run builds. + +## Gates + +Every one must pass, else nothing is published. Warnings go to the job summary and `build.json` and do not stop it. + +| Gate | Checks | +|---|---| +| (build) | nnnotes' own checks: every table, chart and cue sheet read, every chart measured, the deck statistics cross-checked against the chart facts and the master data (the command writes no file otherwise) | +| `schema` | the file against `docs/schema/music-data.schema.json` of the checkout (JSON Schema 2020-12) | +| `provenance` | `format`; `region` `tw`; `master.source` `api`; `master.version` equal to the snapshot's and its `MasterManifest.json`'s; every table's SHA-256 the manifest's, every decoded table read the one `index.json` lists; the song tables and the deck model's present; `deck.commit` the one `rust/Cargo.lock` pins; `exporter.version` the installed nnnotes; an APK version; a catalog SHA-256 (warning: the APK is another client version than the snapshot's) | +| `counts` | no fewer songs and charts than the published file (warning: ids no longer in it) | +| `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning), `weights[kind][position]` numbers, every check deck within its bound | +| `scenarios` | the play scenario fields: `offSeeds` exactly one entry (seed 0, score, weights, check within its bound), every range's `rankBonusPercents` five ints (the first `rankBonusPercent`), every seed's `scorePerfect`, `rangeWeights` (`[kind][position][range]`) and `rankCheck` (within its bound), every seed range's `rangeScorePerfect` (warnings, none in TW: a null `rangeWeights`, a null kind in it or in `offSeeds`' weights) | +| `finite` | no NaN or infinity (warning: one inside master data rows, `songs[].master`, which the format writes as `1e999`) | +| `references` | texts in every language of `languages` (names and titles not empty); unique ids; songs sorted; the songs' bands, vocal characters and tags in the file; a band or a band name; a jacket, and its file in `jackets/`; a BGM cue; score ranks; charts in difficulty order, score ids unique (warnings: a title without a `zh-Hant` text, a music category on no tab, a character of no band) | +| `bgm` | every song's BGM length: `durationMs = samples * 1000 // sampleRate`, 30 s to 10 min, within 1 s of the cue's `lengthMs`, not ending before a chart's last note (warning: more than a minute after it) | +| `size` | 0.8 to 2 times the published file | +| `page` | `music_data_smoke.mjs`: the page's `catalog.js` and `ranking.js` in Node.js over the file: a row per chart, a plain score-up kind, data for the free, rank and Just scenarios, finite positive figures for every chart the data covers in seven scenarios (Gekisou Live at several ranks, Just rates and a Great share, Free Live), the ranking, frontier and event figures | +| (publish) | read back after upload, SHA-256 checked | + +`counts` and `size` compare with the published `music-data.json` and are skipped while nothing is published. A +legitimate drop (a song the game removed) stops the run: a person checks it, then moves the published +`music-data.json` away (the archive keeps it) or changes the gate in a pull request. + +The self-test runs in each build and locally in seconds, without the network: + +``` +python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py +``` + +with, optionally, `MUSIC_DATA_SCHEMA` (a schema file when the checkout has none), `MUSIC_DATA_PAGE` (an +`examples/songs` directory: the smoke test), `MUSIC_DATA_SAMPLE` (a real file with the play scenario fields: its +content gates pass) and `MUSIC_DATA_OLD_SAMPLE` (one without them: the scenario gate stops it). + +## Settings + +Repository secrets: the story site's (`.github/STORY_SITE.md`), no new one: `NNNOTES_BUNDLE_KEY`, +`NNNOTES_BUNDLE_NONCE_SEED`, `NNNOTES_SERVERS_TW_CDN`, `PLAYFETCH_CREDENTIALS`, `STORY_S3_ACCESS_KEY`, +`STORY_S3_SECRET_KEY`. The master key is not needed: the master data comes decoded. + +Repository variables: + +| Variable | Default | | +|---|---|---| +| `MUSIC_DATA_PLAYER_REF` | none: **required** | the ournotes-player commit whose chart data page reads this file (the page with the play scenarios); a run stops before building without it | +| `MUSIC_DATA_PUBLISH` | none: off | `true`: upload; anything else: every run is a dry run | +| `MUSIC_DATA_PLAYER_REPOSITORY` | `empty-sekai/ournotes-player` | | +| `MUSIC_DATA_S3_PREFIX` | `music-data` | the key prefix in the bucket | +| `MUSIC_DATA_MASTERDATA_REGION` | `hk-tw-mo` | the region of `index.json`; the build reads the TW catalog (`[catalog] region` `tw`) | +| `STORY_S3_ENDPOINT`, `STORY_S3_BUCKET`, `MASTERDATA_BASE_URL`, `PLAYFETCH_VERSION`, `STORY_APK_PACKAGE` | the story site's | shared with it | + +## Before the first run + +- **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play + scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6, `12df2a6`): sync the + fork with upstream first. Until then `plan` stops naming what is missing. +- **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play + scenario fields, once that page is merged. +- **Publishing.** Set `MUSIC_DATA_PUBLISH` to `true` last, when dry runs pass and the published format is final. + +## Notes + +- **Versions.** A new ournotes-deck pin (`rust/Cargo.toml`, `rust/Cargo.lock`) or a new nnnotes commit in `src/`, + `rust/` or `pyproject.toml` reaches this fork with a sync, and the next run builds a new file. The deck + statistics may then differ: `provenance.deck.commit` and `build.json` name the commit. +- **The page and the data.** The page's modules are pinned by `MUSIC_DATA_PLAYER_REF`: after a page release that + reads new fields, move it (a new ref alone does not start a build; run with `force` to check the published data + against the new page). +- **Byte identity.** A build from the same inputs gives the same bytes (the file is canonical); the jackets' + WebP bytes depend on the Pillow version. +- **Logs.** The steps print counts, ids, SHA-256 and field names, not game content; nothing decrypted is cached or + uploaded as an artifact. diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py new file mode 100755 index 0000000..0c336b5 --- /dev/null +++ b/.github/scripts/music_data.py @@ -0,0 +1,863 @@ +#!/usr/bin/env python3 +"""The music data CI steps (.github/workflows/music-data.yml): `nnnotes music-data` of moenotes-masterdata-sync's +decoded master data, checked by quality gates and published into the story site's bucket under $MUSIC_DATA_S3_PREFIX +(music-data.json, jackets/, archive/, build.json) for the chart data page of ournotes-player (examples/songs). + + plan the inputs of a build (the master data snapshot of $MASTERDATA_REGION in index.json, the + deck commit rust/Cargo.lock pins, the last nnnotes commit of src/, rust/ and pyproject.toml, + RECIPE) against those of the published build.json; GitHub output `build`: true when they + differ, nothing is published or $FORCE is true + master OUT every file of the snapshot of $MASTERDATA_REGION into OUT (SHA-256 checked against + index.json, MasterManifest.json included) and its index entry as OUT.snapshot.json + build OUT `nnnotes music-data --decoded-master --jackets OUT/jackets -o OUT/music-data.json` ([paths] + master: the master step's OUT), its printed summary in OUT/music-data.summary.json + check OUT MASTER PAGE the quality gates (.github/MUSIC_DATA.md) on OUT/music-data.json, with the published file + as the baseline and PAGE the chart data page's modules (examples/songs) for the smoke test; + the report in OUT/check.json, and OUT/build.json (the build marker) when every gate passed; + exits 1 when one failed + publish OUT [--dry-run] + the jackets the bucket lacks or has at another size (every one with $FORCE), the file's + archive copy, then music-data.json, build.json last; each read back and its SHA-256 checked. + Only a file whose check.json passed; never deletes. A dry run unless $MUSIC_DATA_PUBLISH is + `true` (the publishing switch, off by default): it lists what it would upload + +Bucket: story_site.Bucket ($STORY_S3_ENDPOINT, $STORY_S3_BUCKET, credentials $STORY_S3_ACCESS_KEY / +$STORY_S3_SECRET_KEY) with the key prefix $MUSIC_DATA_S3_PREFIX (default music-data). The published files are read +over plain HTTP, as the page reads them (the bucket serves public read). +""" +from __future__ import annotations + +import hashlib +import json +import math +import os +import re +import subprocess +import sys +import time +import urllib.error +import urllib.request +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import story_site # noqa: E402 (the bucket, HTTP and master data helpers) +from story_site import env, get, output, summary # noqa: E402 + +FORMAT = "nnnotes.music-data/1" +BUILD_FORMAT = "moenotes.music-data-build/1" +# This script's own version of a build: bump it when what it builds or publishes changes, so that the next run builds +# although the master data, the deck model and nnnotes are the same. +RECIPE = 1 +FILE, MARKER, JACKETS, ARCHIVE = "music-data.json", "build.json", "jackets/", "archive/" +MANIFEST = "MasterManifest.json" +SOURCE_PATHS = ("src", "rust", "pyproject.toml") # nnnotes' code: the commit that last changed one of them +SCHEMA = Path("docs/schema/music-data.schema.json") +SMOKE = Path(__file__).resolve().parent / "music_data_smoke.mjs" +ARCHIVE_CACHE = story_site.ASSET_CACHE # archive//.json: content-addressed +FILE_CACHE = "no-cache" # music-data.json and build.json change in place +JACKET_CACHE = "public, max-age=86400" +SNAPSHOT_KEYS = ("version", "resource_version", "client_version", "verified_at", "manifest_sha256", "table_count") + +# the gates' bounds (MUSIC_DATA.md) +SIZE_RATIO = (0.8, 2.0) # against the published file +BGM_MS = (30_000, 600_000) # a song's BGM length +BGM_CUE_SLACK_MS = 1000 # |durationMs - lengthMs| +BGM_TAIL_MS = 60_000 # BGM after the last note: more is reported +RANKS = 5 +DIFFICULTIES = ("easy", "normal", "hard", "expert") +SONG_TABLES = ("MasterLiveMusic", "MasterLiveMusicScore", "MasterText", "MasterBand", "MasterCharacter", "MasterTag", + "MasterLiveMusicCategory", "MasterSound", "MasterSoundCueSheet", "MasterLiveScoreRank") +SHA256 = re.compile(r"[0-9a-f]{64}") +COMMIT = re.compile(r"[0-9a-f]{40}") +LISTED = 20 # failures and warnings listed per gate + + +def fail(message: str): + sys.exit(f"music_data: {message}") + + +def sha256(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +# ---------------------------------------------------------------- bucket +def s3_prefix() -> str: + p = (os.environ.get("MUSIC_DATA_S3_PREFIX") or "music-data").strip("/") + return f"{p}/" if p else "" + + +def bucket() -> "story_site.Bucket": + b = story_site.Bucket() + b.prefix = s3_prefix() + return b + + +def public_url(key: str) -> str: + return f"{env('STORY_S3_ENDPOINT').rstrip('/')}/{env('STORY_S3_BUCKET')}/{s3_prefix()}{key}" + + +def published(key: str) -> bytes | None: + """A published file (plain HTTP), None when the bucket has none.""" + try: + return get(public_url(key), timeout=300) + except urllib.error.HTTPError as e: + if e.code == 404: + return None + raise + + +def exists(key: str) -> bool: + request = urllib.request.Request(public_url(key), method="HEAD", + headers={"User-Agent": "moenotes-music-data (GitHub Actions)", + "Cache-Control": "no-cache"}) + try: + with urllib.request.urlopen(request, timeout=60): + return True + except urllib.error.HTTPError as e: + if e.code == 404: + return False + raise + + +# ---------------------------------------------------------------- the inputs of a build +def deck_commit(root: Path = Path(".")) -> str: + """The ournotes-deck commit nnnotes builds its deck model with (rust/Cargo.lock).""" + lock = root / "rust" / "Cargo.lock" + if not lock.is_file(): + fail("no rust/Cargo.lock: this nnnotes has no music-data command with the deck model (sync the fork with " + "upstream, .github/MUSIC_DATA.md)") + for block in lock.read_text(encoding="utf-8").split("[[package]]"): + if re.search(r'^name = "ournotes-deck"$', block, re.M): + m = re.search(r'^source = "git\+[^"#]*#([0-9a-f]{40})"$', block, re.M) + if m: + return m.group(1) + fail("rust/Cargo.lock has no ournotes-deck git commit") + + +def nnnotes_commit(root: Path = Path(".")) -> str: + """The commit that last changed nnnotes' code (SOURCE_PATHS); the checkout needs its history.""" + def git(*args): + return subprocess.run(["git", "-C", str(root), *args], capture_output=True, text=True) + if git("rev-parse", "--is-shallow-repository").stdout.strip() != "false": + fail("the checkout is shallow: the nnnotes commit needs the history (actions/checkout fetch-depth: 0)") + commit = git("log", "-1", "--format=%H", "--", *SOURCE_PATHS).stdout.strip() + if not COMMIT.fullmatch(commit): + fail("no commit changed src/, rust/ or pyproject.toml") + return commit + + +def require_decoded_master(root: Path = Path(".")) -> None: + cli = root / "src" / "nnnotes" / "cli.py" + if "--decoded-master" not in cli.read_text(encoding="utf-8"): + fail("this nnnotes has no `music-data --decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with " + "upstream (.github/MUSIC_DATA.md)") + + +def inputs(entry: dict, root: Path = Path(".")) -> dict: + """What a build is made of: the master data snapshot, the deck model, nnnotes and this script.""" + return {"masterRegion": env("MASTERDATA_REGION"), "masterVersion": entry.get("version"), + "resourceVersion": entry.get("resource_version"), "clientVersion": entry.get("client_version"), + "deckCommit": deck_commit(root), "nnnotesCommit": nnnotes_commit(root), "recipe": RECIPE} + + +def short(v) -> str: + return str(v)[:12] if v is not None else "?" + + +# ---------------------------------------------------------------- plan +def cmd_plan() -> None: + _, region = story_site.master_index() + require_decoded_master() + now = inputs(region.get("entry") or {}) + raw = published(MARKER) + marker = json.loads(raw) if raw else None + have = marker.get("inputs") if isinstance(marker, dict) else None + changed = [k for k in now if not isinstance(have, dict) or have.get(k) != now[k]] + if have and not exists(FILE): + changed.append("music-data.json (missing)") + force = os.environ.get("FORCE") == "true" + build = force or bool(changed) + summary(f"### Music data\n\n- master: {now['masterVersion']} of {now['masterRegion']} (resource " + f"{now['resourceVersion']}, client {now['clientVersion']}); deck {short(now['deckCommit'])}; nnnotes " + f"{short(now['nnnotesCommit'])}; recipe {RECIPE}\n- published: " + + (f"master {have.get('masterVersion')}, deck {short(have.get('deckCommit'))}, nnnotes " + f"{short(have.get('nnnotesCommit'))}, built {marker.get('builtAt')}" if isinstance(have, dict) + else "nothing") + + "\n- this run: " + ("builds" + (" (force)" if force else "") + (f", changed: {', '.join(changed)}" + if changed and have else "") + if build else "nothing changed, nothing to build")) + summary(f"- publishing: {'on' if publishing() else 'off (MUSIC_DATA_PUBLISH is not `true`: a dry run)'}") + output("build", "true" if build else "false") + + +# ---------------------------------------------------------------- master data +def snapshot_file(master: Path) -> Path: + return master.parent / f"{master.name}.snapshot.json" + + +def cmd_master(out: str) -> None: + url, region = story_site.master_index() + files = region["files"] + entry = region.get("entry") or {} + if MANIFEST not in files: + fail(f"index.json lists no {MANIFEST} for {env('MASTERDATA_REGION')}") + if entry.get("manifest_sha256") and entry["manifest_sha256"] != files[MANIFEST]: + fail(f"index.json: the entry's manifest_sha256 is not the SHA-256 of its {MANIFEST}") + d = Path(out) + d.mkdir(parents=True, exist_ok=True) + story_site.parallel(lambda item: story_site.fetch_table(url, item[0], item[1], d / item[0]), files.items(), 8) + version = json.loads((d / MANIFEST).read_bytes()).get("version") + if entry.get("version") is not None and str(version) != str(entry["version"]): + fail(f"{MANIFEST} has version {version}, index.json {entry['version']} (the service changed snapshots? " + f"run again)") + snapshot = {"region": env("MASTERDATA_REGION"), "entry": {k: entry.get(k) for k in SNAPSHOT_KEYS}, + "files": files} + snapshot_file(d).write_text(json.dumps(snapshot, indent=1, sort_keys=True), encoding="utf-8") + print(f"master data: {version}, {len(files)} files into {d}") + + +# ---------------------------------------------------------------- build +def cmd_build(out: str) -> None: + o = Path(out).resolve() + o.mkdir(parents=True, exist_ok=True) + nnnotes = [sys.executable, "-m", "nnnotes"] + usage = subprocess.run(nnnotes + ["music-data", "--help"], capture_output=True, text=True).stdout + if "--decoded-master" not in usage: + fail("the installed nnnotes has no `music-data --decoded-master` (sync the fork with upstream)") + cmd = nnnotes + ["music-data", "--decoded-master", "--jackets", str(o / "jackets"), "-o", str(o / FILE)] + print("+ " + " ".join(cmd[1:]), flush=True) + with open(o / "music-data.summary.json", "wb") as f: + status = subprocess.run(cmd, stdout=f).returncode + if status: + fail(f"nnnotes music-data exited with {status}") + r = json.loads((o / "music-data.summary.json").read_text(encoding="utf-8")) + summary(f"- built: {r.get('songs')} songs, {r.get('charts')} charts, {r.get('jackets')} jackets, deck " + f"{short(r.get('deck'))}, {r.get('bytes')} bytes, sha256 {short(r.get('sha256'))}") + + +# ---------------------------------------------------------------- the gates +@dataclass +class Context: + """What the gates check the file against; None: the gate (or that part of it) is skipped (the self-test).""" + region: str | None = None # [catalog] region: provenance.region + language: str | None = None # [catalog] language: a song title should have it + master: Path | None = None # the decoded master data read (MasterManifest.json, .json) + snapshot: dict | None = None # its index entry and files (the master step's OUT.snapshot.json) + deck_commit: str | None = None + nnnotes_version: str | None = None + jackets: Path | None = None + schema: Path | None = None + published: bytes | None = None # the published music-data.json (None: nothing is published) + page: Path | None = None # examples/songs of ournotes-player + file: Path | None = None # the file on disk, for the smoke test + + +class Gate: + def __init__(self): + self.failures: list[str] = [] + self.warnings: list[str] = [] + self.note = "" + + def fail(self, message: str): + self.failures.append(message) + + def warn(self, message: str): + self.warnings.append(message) + + +def charts_of(doc: dict): + for song in doc.get("songs") or []: + for chart in song.get("charts") or []: + yield song, chart + + +def where(song: dict, chart: dict) -> str: + return f"chart {chart.get('scoreId')} ({song.get('id')} {chart.get('difficulty')})" + + +def is_int(v) -> bool: + return isinstance(v, int) and not isinstance(v, bool) + + +def is_num(v) -> bool: + return (isinstance(v, (int, float)) and not isinstance(v, bool)) and math.isfinite(v) + + +def within(c) -> bool: + return (isinstance(c, dict) and is_int(c.get("exact")) and is_num(c.get("predicted")) and is_num(c.get("bound")) + and abs(c["exact"] - c["predicted"]) <= c["bound"]) + + +def gate_schema(doc, ctx: Context, g: Gate): + if ctx.schema is None: + g.note = "skipped" + return + if not ctx.schema.is_file(): + g.fail(f"no JSON Schema {ctx.schema}") + return + import jsonschema + schema = json.loads(ctx.schema.read_text(encoding="utf-8")) + validator = jsonschema.Draft202012Validator(schema) + for e in sorted(validator.iter_errors(doc), key=lambda e: list(map(str, e.absolute_path))): + g.fail(f"{'/'.join(map(str, e.absolute_path)) or '(root)'}: {e.message[:160]}") + g.note = ctx.schema.as_posix() + + +def gate_provenance(doc, ctx: Context, g: Gate): + if doc.get("format") != FORMAT: + g.fail(f"format {doc.get('format')!r}, expected {FORMAT}") + p = doc.get("provenance") or {} + if ctx.region is not None and p.get("region") != ctx.region: + g.fail(f"region {p.get('region')!r}, expected {ctx.region!r}") + m = p.get("master") or {} + if m.get("source") != "api": + g.fail(f"master.source {m.get('source')!r}, expected 'api'") + tables = m.get("tables") or {} + missing = [t for t in SONG_TABLES if t not in tables] + if missing: + g.fail(f"master.tables lacks {', '.join(missing)}") + if doc.get("deck") is not None and len(tables) <= len(SONG_TABLES): + g.fail("master.tables has only the song tables, but the file has deck statistics") + if ctx.snapshot is not None: + v = ctx.snapshot["entry"].get("version") + if m.get("version") != v: + g.fail(f"master.version {m.get('version')!r}, the snapshot's {v!r}") + if ctx.master is not None: + manifest = json.loads((ctx.master / MANIFEST).read_bytes()) + if m.get("version") != manifest.get("version"): + g.fail(f"master.version {m.get('version')!r}, {MANIFEST}'s {manifest.get('version')!r}") + listed = {f.get("name"): str(f.get("hash") or "").lower() for f in manifest.get("files") or []} + files = (ctx.snapshot or {}).get("files") or {} + for t, v in sorted(tables.items()): + if (v or {}).get("sha256") != listed.get(f"{t}.bin"): + g.fail(f"master.tables.{t}.sha256 is not {MANIFEST}'s {t}.bin") + decoded = ctx.master / f"{t}.json" + if not decoded.is_file(): + g.fail(f"no decoded table {t}.json") + elif files and sha256(decoded.read_bytes()) != files.get(f"{t}.json"): + g.fail(f"{t}.json read is not the one index.json lists") + deck = p.get("deck") + if doc.get("deck") is not None and not isinstance(deck, dict): + g.fail("provenance.deck is null, but the file has deck statistics") + if isinstance(deck, dict): + if deck.get("name") != "ournotes-deck" or not COMMIT.fullmatch(str(deck.get("commit"))): + g.fail("provenance.deck names no ournotes-deck commit") + elif ctx.deck_commit is not None and deck["commit"] != ctx.deck_commit: + g.fail(f"deck commit {deck['commit'][:12]}, rust/Cargo.lock pins {ctx.deck_commit[:12]}") + ex = p.get("exporter") or {} + if ex.get("name") != "nnnotes": + g.fail(f"exporter {ex.get('name')!r}") + elif ctx.nnnotes_version is not None and ex.get("version") != ctx.nnnotes_version: + g.fail(f"exporter version {ex.get('version')!r}, installed nnnotes {ctx.nnnotes_version!r}") + client = p.get("client") or {} + if not isinstance(client.get("versionName"), str) or not is_int(client.get("versionCode")): + g.fail("provenance.client has no APK version (was the APK read?)") + elif ctx.snapshot is not None and ctx.snapshot["entry"].get("client_version") not in (None, + client["versionName"]): + g.warn(f"the APK is {client['versionName']}, the snapshot names client " + f"{ctx.snapshot['entry']['client_version']}") + if not SHA256.fullmatch(str((p.get("catalog") or {}).get("sha256"))): + g.fail("provenance.catalog has no SHA-256") + g.note = f"master {m.get('version')}, deck {short((deck or {}).get('commit'))}, nnnotes {ex.get('version')}" + + +def baseline(ctx: Context, g: Gate): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return None + try: + return json.loads(ctx.published) + except ValueError: + g.warn("the published music-data.json is not JSON: skipped") + return None + + +def gate_counts(doc, ctx: Context, g: Gate): + old = baseline(ctx, g) + if old is None: + return + songs = {s.get("id") for s in doc.get("songs") or []} + charts = {c.get("scoreId") for _, c in charts_of(doc)} + old_songs = {s.get("id") for s in old.get("songs") or []} + old_charts = {c.get("scoreId") for _, c in charts_of(old)} + if len(songs) < len(old_songs): + g.fail(f"{len(songs)} songs, the published file has {len(old_songs)}") + if len(charts) < len(old_charts): + g.fail(f"{len(charts)} charts, the published file has {len(old_charts)}") + if old_songs - songs: + g.warn(f"songs no longer in the file: {', '.join(map(str, sorted(old_songs - songs)))}") + if old_charts - charts: + g.warn(f"charts no longer in the file: {', '.join(map(str, sorted(old_charts - charts)))}") + g.note = f"songs {len(old_songs)} -> {len(songs)}, charts {len(old_charts)} -> {len(charts)}" + + +def gate_size(doc, ctx: Context, g: Gate, raw: bytes): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return + ratio = len(raw) / max(1, len(ctx.published)) + if not SIZE_RATIO[0] <= ratio <= SIZE_RATIO[1]: + g.fail(f"{len(raw)} bytes, {ratio:.2f} times the published {len(ctx.published)} (bounds {SIZE_RATIO[0]} to " + f"{SIZE_RATIO[1]})") + g.note = f"{len(ctx.published)} -> {len(raw)} bytes ({ratio:.2f})" + + +def weights_shape(w, kinds: int, positions: int, nullable: bool) -> bool: + return isinstance(w, list) and len(w) == kinds and all( + (nullable and k is None) or (isinstance(k, list) and len(k) == positions and all(map(is_num, k))) for k in w) + + +def gate_deck(doc, ctx: Context, g: Gate): + deck = doc.get("deck") + if not isinstance(deck, dict): + g.fail("deck is null: no deck statistics (made with --no-deck?)") + return + kinds = len(deck.get("kinds") or []) + if not kinds: + g.fail("deck.kinds is empty") + power = (deck.get("model") or {}).get("power") + if not (is_num(power) and power > 0): + g.fail(f"deck.model.power {power!r}") + n = unplayable = seeds = 0 + for song, chart in charts_of(doc): + n += 1 + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, events = d.get("positions"), d.get("events") or [] + if len(events) != len(chart.get("skillEventsMs") or []): + g.fail(f"{w}: {len(events)} skill events, the chart has {len(chart.get('skillEventsMs') or [])}") + if not is_int(positions) or (events and positions != max(e[0] for e in events) + 1): + g.fail(f"{w}: positions {positions!r} do not match the events") + continue + if d.get("unplayable"): + unplayable += 1 + g.warn(f"{w}: unplayable with Gekisou on ({d['unplayable']})") + if d.get("seeds"): + g.fail(f"{w}: unplayable, but has Gekisou on seeds") + elif not d.get("seeds"): + g.fail(f"{w}: no seeds") + for seed in d.get("seeds") or []: + seeds += 1 + s = f"{w} seed {seed.get('seed')}" + if not is_int(seed.get("score")): + g.fail(f"{s}: score {seed.get('score')!r}") + if not weights_shape(seed.get("weights"), kinds, positions, nullable=False): + g.fail(f"{s}: weights are not [kind][position] numbers") + if len(seed.get("ranges") or []) != len(d.get("ranges") or []): + g.fail(f"{s}: {len(seed.get('ranges') or [])} range results for {len(d.get('ranges') or [])} ranges") + if not within(seed.get("check")): + g.fail(f"{s}: the check deck is not within its bound") + g.note = f"{n} charts, {kinds} kinds, {seeds} seeds, {unplayable} unplayable" + + +def gate_scenarios(doc, ctx: Context, g: Gate): + """The play scenario fields (Gekisou off, every rank, the Perfect play): offSeeds exactly one, every range's + rankBonusPercents five ints, every seed scorePerfect, rangeWeights and rankCheck, every seed range + rangeScorePerfect. A null rangeWeights, a null kind in it or in offSeeds' weights is a warning (none in TW).""" + kinds = len(((doc.get("deck") or {}).get("kinds")) or []) + null_rw = null_kind = null_off = 0 + for song, chart in charts_of(doc): + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, ranges = d.get("positions"), d.get("ranges") or [] + off = d.get("offSeeds") + if not isinstance(off, list) or len(off) != 1: + g.fail(f"{w}: offSeeds {'missing' if off is None else f'has {len(off)} entries'}, expected exactly one") + else: + o = off[0] + if o.get("seed") != 0 or not is_int(o.get("score")): + g.fail(f"{w}: Gekisou off seed {o.get('seed')!r} score {o.get('score')!r}") + if not weights_shape(o.get("weights"), kinds, positions, nullable=True): + g.fail(f"{w}: Gekisou off weights are not [kind][position] numbers") + else: + null_off += sum(k is None for k in o["weights"]) + if not within(o.get("check")): + g.fail(f"{w}: the Gekisou off check deck is not within its bound") + for i, r in enumerate(ranges): + p = r.get("rankBonusPercents") + if not (isinstance(p, list) and len(p) == RANKS and all(map(is_int, p))): + g.fail(f"{w} range {i}: rankBonusPercents {'missing' if p is None else 'not five ints'}") + elif p[0] != r.get("rankBonusPercent"): + g.fail(f"{w} range {i}: rankBonusPercents[0] {p[0]} is not rankBonusPercent " + f"{r.get('rankBonusPercent')}") + for seed in d.get("seeds") or []: + s = f"{w} seed {seed.get('seed')}" + absent = [k for k in ("scorePerfect", "rangeWeights", "rankCheck") if k not in seed] + if absent: + g.fail(f"{s}: no {', '.join(absent)}") + continue + if not is_int(seed["scorePerfect"]): + g.fail(f"{s}: scorePerfect {seed['scorePerfect']!r}") + for i, r in enumerate(seed.get("ranges") or []): + if not is_int(r.get("rangeScorePerfect")): + why = "missing" if "rangeScorePerfect" not in r else "not an int" + g.fail(f"{s} range {i}: rangeScorePerfect {why}") + rw = seed["rangeWeights"] + if rw is None: + null_rw += 1 + elif not (isinstance(rw, list) and len(rw) == kinds and all( + k is None or (isinstance(k, list) and len(k) == positions and all( + isinstance(x, list) and len(x) == len(ranges) and all(map(is_num, x)) for x in k)) + for k in rw)): + g.fail(f"{s}: rangeWeights are not [kind][position][range] numbers") + else: + null_kind += sum(k is None for k in rw) + rc = seed["rankCheck"] + if rc is not None: + if len(rc.get("ranks") or []) != len(ranges) or not all( + is_int(x) and 1 <= x <= RANKS for x in rc.get("ranks") or []): + g.fail(f"{s}: rankCheck ranks are not one rank per range") + elif not within(rc): + g.fail(f"{s}: the rank check deck is not within its bound") + if null_rw: + g.warn(f"{null_rw} seeds have rangeWeights null (overlapping ranges: no rank scenarios on those charts)") + if null_kind: + g.warn(f"{null_kind} rangeWeights kinds are null (conditions on the confirmed rank)") + if null_off: + g.warn(f"{null_off} Gekisou off weight kinds are null (conditions on the Gekisou state)") + g.note = "offSeeds, rankBonusPercents, scorePerfect, rangeWeights, rankCheck, rangeScorePerfect" + + +def nonfinite(v, path: str, out: list): + if isinstance(v, float): + if not math.isfinite(v): + out.append(path) + elif isinstance(v, list): + for i, x in enumerate(v): + nonfinite(x, f"{path}[{i}]", out) + elif isinstance(v, dict): + for k, x in v.items(): + nonfinite(x, f"{path}.{k}" if path else k, out) + + +def gate_finite(doc, ctx: Context, g: Gate): + """No NaN or infinity; one inside master data rows as served (songs[].master, master: `1e999` is how the format + writes a binary32 infinity) is reported, not failed.""" + found: list[str] = [] + nonfinite(doc, "", found) + for path in found: + if re.match(r"(songs\[\d+\]\.master|master)\b", path): + g.warn(f"{path}: not finite (master data as served)") + else: + g.fail(f"{path}: not finite") + g.note = f"{len(found)} non-finite numbers" + + +def gate_references(doc, ctx: Context, g: Gate): + langs = doc.get("languages") or [] + if not langs: + g.fail("no languages") + + def text(t, what: str, required: bool) -> bool: + if t is None: + if required: + g.fail(f"{what}: no text") + return False + if not isinstance(t, dict) or set(t) != set(langs) or not all(isinstance(v, str) for v in t.values()): + g.fail(f"{what}: not a text in {', '.join(langs)}") + return False + if required and not any(t.values()): + g.fail(f"{what}: empty in every language") + return False + return True + + def ids(rows, what: str) -> set: + seen = [r.get("id") for r in rows] + if len(set(seen)) != len(seen) or not all(map(is_int, seen)): + g.fail(f"{what}: ids are not unique ints") + return set(seen) + + bands = ids(doc.get("bands") or [], "bands") + characters = ids(doc.get("characters") or [], "characters") + tags = ids(doc.get("tags") or [], "tags") + ids(doc.get("categories") or [], "categories") + for b in doc.get("bands") or []: + text(b.get("name"), f"band {b.get('id')} name", True) + for c in doc.get("characters") or []: + text(c.get("name"), f"character {c.get('id')} name", True) + text(c.get("shortName"), f"character {c.get('id')} short name", False) + if c.get("bandId") not in bands: + g.warn(f"character {c.get('id')}: band {c.get('bandId')} is not in bands") + for t in doc.get("tags") or []: + text(t.get("name"), f"tag {t.get('id')} name", True) + categories = set() + for c in doc.get("categories") or []: + text(c.get("name"), f"category {c.get('id')} name", True) + categories.update(c.get("musicCategories") or []) + songs = doc.get("songs") or [] + if not songs: + g.fail("no songs") + ids(songs, "songs") + if [s.get("id") for s in songs] != sorted(s.get("id") for s in songs if is_int(s.get("id"))): + g.fail("songs are not sorted by id") + score_ids: set = set() + untitled = [] + for s in songs: + sid = f"song {s.get('id')}" + if text(s.get("title"), f"{sid} title", True) and ctx.language and not s["title"].get(ctx.language): + untitled.append(s.get("id")) + for k in ("ruby", "phonetic", "bandName", "lyricist", "composer", "arranger"): + text(s.get(k), f"{sid} {k}", False) + for key, known, what in (("bandIds", bands, "band"), ("vocalCharacterIds", characters, "character"), + ("bestMusicTagIds", tags, "tag")): + unknown = [x for x in s.get(key) or [] if x not in known] + if unknown: + g.fail(f"{sid}: {what} {', '.join(map(str, unknown))} not in the file's {what}s") + if not s.get("bandIds") and s.get("bandName") is None: + g.fail(f"{sid}: neither a band nor a band name") + other = [x for x in s.get("musicCategories") or [] if x not in categories] + if other: + g.warn(f"{sid}: music categories {', '.join(map(str, other))} are on no category tab") + jacket = s.get("jacket") + if not isinstance(jacket, str) or not jacket: + g.fail(f"{sid}: no jacket") + elif ctx.jackets is not None: + f = ctx.jackets / f"{jacket}.webp" + if not f.is_file() or not f.stat().st_size: + g.fail(f"{sid}: no jacket file jackets/{jacket}.webp") + bgm = s.get("bgm") or {} + if not (is_int(bgm.get("soundId")) and bgm.get("cueSheet") and bgm.get("cue")): + g.fail(f"{sid}: no BGM cue") + ranks = [r.get("rank") for r in s.get("scoreRanks") or []] + if not ranks or len(set(ranks)) != len(ranks): + g.fail(f"{sid}: score ranks {ranks!r}") + charts = s.get("charts") or [] + order = [c.get("difficulty") for c in charts] + if not charts or order != [d for d in DIFFICULTIES if d in order] or len(set(order)) != len(order): + g.fail(f"{sid}: charts {order!r}") + for c in charts: + if c.get("scoreId") in score_ids: + g.fail(f"{where(s, c)}: the score id occurs twice") + score_ids.add(c.get("scoreId")) + if untitled: + g.warn(f"songs without a {ctx.language} title: {', '.join(map(str, untitled))}") + g.note = (f"{len(songs)} songs, {len(bands)} bands, {len(characters)} characters, {len(tags)} tags" + + ("" if ctx.jackets is None else ", jackets present")) + + +def gate_bgm(doc, ctx: Context, g: Gate): + lengths = [] + for s in doc.get("songs") or []: + sid = f"song {s.get('id')}" + L = (s.get("bgm") or {}).get("length") + if not isinstance(L, dict): + g.fail(f"{sid}: no BGM length") + continue + dur, samples, rate = L.get("durationMs"), L.get("samples"), L.get("sampleRate") + if not (is_int(dur) and is_int(samples) and is_int(rate) and rate > 0 and samples > 0): + g.fail(f"{sid}: BGM length {dur!r} ms, {samples!r} samples at {rate!r} Hz") + continue + if dur != samples * 1000 // rate: + g.fail(f"{sid}: BGM durationMs {dur} is not samples * 1000 // sampleRate") + if not BGM_MS[0] <= dur <= BGM_MS[1]: + g.fail(f"{sid}: BGM of {dur} ms (bounds {BGM_MS[0]} to {BGM_MS[1]})") + if is_int(L.get("lengthMs")) and abs(dur - L["lengthMs"]) > BGM_CUE_SLACK_MS: + g.fail(f"{sid}: BGM stream {dur} ms, cue length {L['lengthMs']} ms") + last = max((c.get("lastNoteMs") or 0 for c in s.get("charts") or []), default=0) + if dur < last: + g.fail(f"{sid}: BGM of {dur} ms ends before the last note at {last} ms") + elif dur - last > BGM_TAIL_MS: + g.warn(f"{sid}: BGM plays {dur - last} ms after the last note") + lengths.append(dur) + if lengths: + g.note = f"{min(lengths) / 1000:.1f} s to {max(lengths) / 1000:.1f} s" + + +def gate_page(doc, ctx: Context, g: Gate): + """The chart data page's own modules (catalog.js, ranking.js) over the file in Node.js: music_data_smoke.mjs.""" + if ctx.page is None: + g.note = "skipped" + return + if not (ctx.page / "catalog.js").is_file() or not (ctx.page / "ranking.js").is_file(): + g.fail(f"{ctx.page}: no catalog.js / ranking.js (the chart data page, examples/songs)") + return + r = subprocess.run(["node", str(SMOKE), str(ctx.page), str(ctx.file)], capture_output=True, text=True, + timeout=600) + lines = [x for x in (r.stdout + r.stderr).splitlines() if x.strip()] + if r.returncode: + for x in lines or [f"node exited with {r.returncode}"]: + g.fail(x[:300]) + else: + g.note = lines[-1][:200] if lines else "passed" + + +GATES = (("schema", gate_schema), ("provenance", gate_provenance), ("counts", gate_counts), ("deck", gate_deck), + ("scenarios", gate_scenarios), ("finite", gate_finite), ("references", gate_references), ("bgm", gate_bgm), + ("size", gate_size), ("page", gate_page)) + + +def gates(raw: bytes, ctx: Context, only=None) -> dict: + """The report of the gates (`only`: their names) on the file's bytes.""" + try: + doc = json.loads(raw.decode("utf-8")) + if not isinstance(doc, dict): + raise ValueError("not an object") + except ValueError as e: + doc, results = None, [{"gate": "json", "passed": False, "failures": [f"not JSON: {str(e)[:160]}"], + "failureCount": 1, "warnings": [], "warningCount": 0, "note": ""}] + if doc is not None: + results = [] + for name, fn in GATES: + if only is not None and name not in only: + continue + g = Gate() + try: + fn(doc, ctx, g, raw) if name == "size" else fn(doc, ctx, g) + except Exception as e: # a malformed file the gate did not foresee fails the gate + g.fail(f"the gate stopped: {type(e).__name__}: {str(e)[:160]}") + results.append({"gate": name, "passed": not g.failures, "failures": g.failures[:LISTED], + "failureCount": len(g.failures), "warnings": g.warnings[:LISTED], + "warningCount": len(g.warnings), "note": g.note}) + return {"passed": all(r["passed"] for r in results), "sha256": sha256(raw), "bytes": len(raw), + "gates": results} + + +def report_markdown(report: dict) -> str: + rows = ["| Gate | Result | |", "|---|---|---|"] + for r in report["gates"]: + result = "passed" if r["passed"] else f"**failed** ({r['failureCount']})" + if r["warningCount"]: + result += f", {r['warningCount']} warnings" + rows.append(f"| {r['gate']} | {result} | {r['note']} |") + details = [] + for r in report["gates"]: + for kind, key, n in (("failure", "failures", "failureCount"), ("warning", "warnings", "warningCount")): + for x in r[key]: + details.append(f"- {r['gate']} {kind}: {x}") + if r[n] > len(r[key]): + details.append(f"- {r['gate']}: {r[n] - len(r[key])} more {kind}s") + return "\n".join(["", "#### Gates", ""] + rows + ([""] + details if details else [])) + + +def cmd_check(out: str, master: str, page: str) -> None: + import nnnotes + o, m = Path(out), Path(master) + raw = (o / FILE).read_bytes() + snapshot = json.loads(snapshot_file(m).read_text(encoding="utf-8")) + ctx = Context(region=env("NNNOTES_CATALOG_REGION"), language=os.environ.get("NNNOTES_CATALOG_LANGUAGE"), + master=m, snapshot=snapshot, deck_commit=deck_commit(), nnnotes_version=nnnotes.__version__, + jackets=o / "jackets", schema=SCHEMA, published=published(FILE), page=Path(page), file=o / FILE) + report = gates(raw, ctx) + (o / "check.json").write_text(json.dumps(report, indent=1), encoding="utf-8") + summary(report_markdown(report)) + if not report["passed"]: + fail("a gate failed: nothing is published (the job summary lists why)") + doc = json.loads(raw) + p = doc["provenance"] + version = re.sub(r"[^A-Za-z0-9._-]", "_", str(p["master"]["version"])) + marker = { + "format": BUILD_FORMAT, + "file": FILE, "sha256": report["sha256"], "bytes": report["bytes"], + "archive": f"{ARCHIVE}{version}/{report['sha256']}.json", + "songs": len(doc["songs"]), "charts": sum(len(s["charts"]) for s in doc["songs"]), + "jackets": len(list((o / "jackets").glob("*.webp"))), + "inputs": inputs(snapshot["entry"]), + "provenance": {"region": p["region"], "client": p["client"], "masterVersion": p["master"]["version"], + "deckCommit": p["deck"]["commit"], "exporterVersion": p["exporter"]["version"]}, + "player": {"repository": os.environ.get("PLAYER_REPOSITORY"), "ref": os.environ.get("PLAYER_REF")}, + "gates": {r["gate"]: {"passed": r["passed"], "warnings": r["warningCount"]} for r in report["gates"]}, + "builtAt": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"), + "run": (f"{os.environ['GITHUB_SERVER_URL']}/{os.environ['GITHUB_REPOSITORY']}/actions/runs/" + f"{os.environ['GITHUB_RUN_ID']}" if os.environ.get("GITHUB_RUN_ID") else None), + } + (o / MARKER).write_text(json.dumps(marker, indent=1) + "\n", encoding="utf-8") + + +# ---------------------------------------------------------------- publish +def upload(b, key: str, src: Path, cache: str) -> None: + extra = {"ContentType": story_site.TYPES.get(src.suffix.lower(), "application/octet-stream"), + "CacheControl": cache} + b.s3.upload_file(str(src), b.name, b.prefix + key, ExtraArgs=extra) + + +def read_back(b, key: str, digest: str) -> None: + """The object as stored (S3 GET), else as served (plain HTTP), must have the SHA-256 uploaded.""" + for attempt in range(4): + try: + if sha256(b.s3.get_object(Bucket=b.name, Key=b.prefix + key)["Body"].read()) == digest: + return + except Exception as e: # Cloudflare in front of the store rejects a signed GET at times + print(f"read back {key}: {type(e).__name__}", flush=True) + time.sleep(3 * (attempt + 1)) + try: + if sha256(get(public_url(key), timeout=300)) == digest: + return + except urllib.error.URLError as e: + print(f"read back {key} over HTTP: {e}", flush=True) + fail(f"{key}: the bucket does not serve what was uploaded (SHA-256 {digest[:12]})") + + +def publishing() -> bool: + """The publishing switch: uploads only with MUSIC_DATA_PUBLISH `true` (a repository variable, unset: off).""" + return os.environ.get("MUSIC_DATA_PUBLISH") == "true" + + +def cmd_publish(out: str, dry_run: bool = False) -> None: + if not publishing() and not dry_run: + summary("- publishing is off (the repository variable MUSIC_DATA_PUBLISH is not `true`): a dry run") + dry_run = True + o = Path(out) + raw = (o / FILE).read_bytes() + report = json.loads((o / "check.json").read_text(encoding="utf-8")) + if not report.get("passed") or report.get("sha256") != sha256(raw) or not (o / MARKER).is_file(): + fail("the file has not passed the gates (check.json): nothing is published") + marker = json.loads((o / MARKER).read_text(encoding="utf-8")) + b = bucket() + if not b.writable and not dry_run: + fail("publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") + force = os.environ.get("FORCE") == "true" + try: + have = b.keys(JACKETS) + archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) + except Exception as e: # a dry run without a key where the bucket lists to none + if not dry_run: + raise + print(f"cannot list the bucket ({type(e).__name__}): the dry run lists every object", flush=True) + have, archived = {}, False + jackets = sorted((o / "jackets").glob("*.webp")) + new = [j for j in jackets if force or have.get(JACKETS + j.name) != j.stat().st_size] + steps = [(JACKETS + j.name, j, JACKET_CACHE, False) for j in new] + if not archived: + steps.append((marker["archive"], o / FILE, ARCHIVE_CACHE, True)) + # the file before the marker that names it, the jackets and the archive copy before the file + steps += [(FILE, o / FILE, FILE_CACHE, True), (MARKER, o / MARKER, FILE_CACHE, True)] + if dry_run or not publishing(): + for key, _, _, _ in steps: + print(f"would upload {b.prefix}{key}") + else: + story_site.parallel(lambda s: upload(b, s[0], s[1], s[2]), [s for s in steps if not s[3]]) + for key, src, cache, check in steps: + if check: + upload(b, key, src, cache) + read_back(b, key, sha256(src.read_bytes())) + summary(f"- {'would publish' if dry_run else 'published'} {len(new)} jackets (of {len(jackets)}), " + f"{'the archive copy, ' if not archived else ''}{FILE} ({len(raw)} bytes, sha256 " + f"{short(marker['sha256'])}), {MARKER}: {public_url(FILE)}") + + +def main(argv: list[str]) -> None: + if not argv: + sys.exit(__doc__) + cmd, args = argv[0], argv[1:] + if cmd == "plan" and not args: + cmd_plan() + elif cmd == "master" and len(args) == 1: + cmd_master(args[0]) + elif cmd == "build" and len(args) == 1: + cmd_build(args[0]) + elif cmd == "check" and len(args) == 3: + cmd_check(*args) + elif cmd == "publish" and len(args) in (1, 2) and args[1:] in ([], ["--dry-run"]): + cmd_publish(args[0], dry_run=bool(args[1:])) + else: + sys.exit(__doc__) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/.github/scripts/music_data_smoke.mjs b/.github/scripts/music_data_smoke.mjs new file mode 100644 index 0000000..a711edc --- /dev/null +++ b/.github/scripts/music_data_smoke.mjs @@ -0,0 +1,94 @@ +// The chart data page's smoke test (music_data.py check, gate "page"): the page's own pure modules (ournotes-player +// examples/songs: catalog.js, ranking.js) over a music-data.json in Node.js, the way the page reads it. Every chart +// must get a row, every playable chart its figures in every play scenario (Gekisou Live at several ranks and Just +// rates, Free Live, a Great share), finite and positive, and the rankings must work on them. +// +// node music_data_smoke.mjs +// +// Prints one line per problem (at most 40) and exits 1, or a one-line summary. +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { pathToFileURL } from "node:url"; + +const [dir, file] = process.argv.slice(2); +if (!dir || !file) { + console.error("usage: node music_data_smoke.mjs "); + process.exit(2); +} +const catalog = await import(pathToFileURL(path.join(dir, "catalog.js")).href); +const ranking = await import(pathToFileURL(path.join(dir, "ranking.js")).href); +const data = JSON.parse(readFileSync(file, "utf8")); + +const problems = []; +const problem = (m) => problems.push(m); +const finite = (v) => typeof v === "number" && Number.isFinite(v); + +const charts = (data.songs || []).flatMap((s) => (s.charts || []).map((c) => ({ song: s, chart: c }))); +const rows = catalog.chartRows(data); +if (rows.length !== charts.length) problem(`chartRows: ${rows.length} rows for ${charts.length} charts`); +const kind = ranking.plainKind(data); +if (kind === null) problem("plainKind: the file has no plain score-up kind (effect 2000, 5 s, no targets)"); + +// Whether the data has what a scenario needs on a chart (a null rangeWeights or kind leaves a chart without figures in +// the rank and Just scenarios: music_data.py reports those as warnings). +const computable = (deck, scenario) => { + if (!deck) return false; + const has = (s) => s && Array.isArray((s.weights || [])[kind]); + if (scenario && scenario.mode === "free") return (deck.offSeeds || []).length > 0 && deck.offSeeds.every(has); + if (deck.unplayable || !(deck.seeds || []).length) return false; + const rank1 = !scenario || ((scenario.ranks || []).every((r) => r === 1) && (scenario.just ?? 1) >= 1); + return deck.seeds.every((s) => has(s) && (rank1 || Array.isArray((s.rangeWeights || [])[kind]))); +}; +const has = ranking.scenarioData(data); +for (const k of ["free", "ranks", "just"]) if (!has[k]) problem(`scenarioData: no data for the ${k} scenario`); + +const SCENARIOS = [ + ["Gekisou Live, rank 1", null], + ["Gekisou Live, rank 5", { mode: "battle", ranks: [5, 5, 5], just: 1, great: 0 }], + ["Gekisou Live, ranks 2 3 4", { mode: "battle", ranks: [2, 3, 4], just: 1, great: 0 }], + ["Gekisou Live, Just 0", { mode: "battle", ranks: [1, 1, 1], just: 0, great: 0 }], + ["Gekisou Live, rank 3, Just 0.5, Great 0.2", { mode: "battle", ranks: [3, 3, 3], just: 0.5, great: 0.2 }], + ["Free Live", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0 }], + ["Free Live, Great 0.5", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0.5 }], +]; +const SKILLS = [1, 1, 1, 1, 1]; +let figures = 0; +for (const [name, scenario] of SCENARIOS) { + const joined = ranking.joinCharts(data, scenario); + const expected = charts.filter(({ chart }) => computable(chart.deck, scenario)).length; + if (joined.length !== expected) problem(`${name}: ${joined.length} charts with figures, expected ${expected}`); + for (const r of joined) { + figures++; + if (!finite(r.base) || r.base <= 0) problem(`${name}: chart ${r.scoreId}: base ${r.base}`); + if (!Array.isArray(r.weights) || !r.weights.length || !r.weights.every(finite)) { + problem(`${name}: chart ${r.scoreId}: weights are not finite numbers`); + } + } + if (!joined.length) continue; + const ranked = ranking.rank(joined, { skills: SKILLS, source: "bgm", overheadMs: 30000 }); + if (!ranked.some((r) => r.frontier)) problem(`${name}: no chart on the frontier`); + for (const r of ranked) { + if (r.lengthMs === null) problem(`${name}: chart ${r.scoreId}: no play length`); + else if (!finite(r.perMinute) || !finite(r.rate)) problem(`${name}: chart ${r.scoreId}: rate ${r.rate}`); + } + for (const room of [0, 5]) { + ranking.eventDominance(ranked, "bgm", ranking.X_MAX, room); + for (const r of ranked) { + const p = ranking.requiredPower(r, SKILLS, "S", 1, room); + if (p !== null && !(finite(p) && p >= 0)) problem(`${name}: chart ${r.scoreId}: required power ${p}`); + } + } + const one = ranked[0]; + const chance = ranking.reachChance(one, SKILLS, 300000, "S", 1, 0); + if (chance !== null && !(chance >= 0 && chance <= 1)) problem(`${name}: reach chance ${chance}`); + catalog.refigure(rows, data, scenario); +} +catalog.histogram(rows, (r) => r.level); + +if (problems.length) { + for (const m of problems.slice(0, 40)) console.log(m); + if (problems.length > 40) console.log(`${problems.length - 40} more problems`); + process.exit(1); +} +console.log(`page smoke test: ${rows.length} charts, ${SCENARIOS.length} scenarios, ${figures} chart figures; ` + + `plain kind ${kind}, scenarios free/ranks/just`); diff --git a/.github/scripts/songs_page.sh b/.github/scripts/songs_page.sh new file mode 100755 index 0000000..71849aa --- /dev/null +++ b/.github/scripts/songs_page.sh @@ -0,0 +1,18 @@ +#!/usr/bin/env bash +# The chart data page of ournotes-player ($PLAYER_REPOSITORY at $PLAYER_REF: examples/songs, the page that reads +# music-data.json) into $1, for the smoke test of music_data.py check. Only that directory is checked out and nothing +# is built: the test imports the page's pure modules (catalog.js, ranking.js) in Node.js. +set -euo pipefail +dir="$1" +if [ -z "${PLAYER_REF:-}" ]; then + echo "::error::set the repository variable MUSIC_DATA_PLAYER_REF: the ournotes-player commit whose chart data page reads this music data (.github/MUSIC_DATA.md)" + exit 1 +fi +rm -rf "$dir" +git clone --quiet --filter=blob:none --no-checkout "https://github.com/$PLAYER_REPOSITORY.git" "$dir" +git -C "$dir" sparse-checkout set examples/songs +git -C "$dir" checkout --quiet "$PLAYER_REF" +for f in catalog.js ranking.js; do + [ -f "$dir/examples/songs/$f" ] || { echo "::error::$PLAYER_REPOSITORY $PLAYER_REF has no examples/songs/$f"; exit 1; } +done +echo "ournotes-player $(git -C "$dir" rev-parse --short HEAD): examples/songs" diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py new file mode 100644 index 0000000..bc86b88 --- /dev/null +++ b/.github/scripts/test_music_data.py @@ -0,0 +1,456 @@ +"""Self-test of the music data gates (music_data.py): a synthetic music-data.json passes every gate, and each gate +fails on the defect it is there for. Runs in seconds, without the network: + + python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + +The JSON Schema gate uses docs/schema/music-data.schema.json (or $MUSIC_DATA_SCHEMA) when the checkout has it; the +page smoke test runs when $MUSIC_DATA_PAGE names ournotes-player's examples/songs (and Node.js is installed). Real +files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass) and +$MUSIC_DATA_OLD_SAMPLE (one without them: the scenario gate stops it). +""" +import copy +import hashlib +import importlib.util +import io +import json +import os +import shutil +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import music_data # noqa: E402 +from music_data import Context, gates # noqa: E402 + +ROOT = Path(__file__).resolve().parents[2] +LANGS = ["ja", "en", "zh-Hant", "zh-Hans", "ko"] +TABLES = music_data.SONG_TABLES + ("MasterLiveSkillEffect", "MasterLiveGekisouRankingScoreBonus") +BIN = {t: hashlib.sha256(t.encode()).hexdigest() for t in TABLES} # the manifest's hashes of the served files +DECK = "d4" * 20 +CONTENT = ("deck", "scenarios", "finite", "references", "bgm") + + +def schema_path(): + p = Path(os.environ.get("MUSIC_DATA_SCHEMA") or ROOT / music_data.SCHEMA) + if not p.is_file(): + return None + return p if importlib.util.find_spec("jsonschema") else None + + +def page_path(): + p = os.environ.get("MUSIC_DATA_PAGE") + return Path(p) if p and Path(p, "ranking.js").is_file() and shutil.which("node") else None + + +# ---------------------------------------------------------------- a synthetic file +def text(stem): + return {lang: f"{stem}-{lang}" for lang in LANGS} + + +def check_deck(exact): + return {"deck": [[0, 5000], None], "exact": exact, "predicted": exact + 0.25, "bound": 7.0} + + +def deck_chart(): + seed = {"seed": 0, "score": 120000, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, + "justCount": 0, "lotResults": [0, 0, 0, 0], + "rangeScorePerfect": 4000}], + "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000, + "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} + return {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, + "ranges": [{"index": 0, "mission": 1, "startMs": 1000, "endMs": 5000, "rankBonusPercent": 250, + "rankBonusPercents": [250, 190, 160, 100, 100]}], + "justNotes": 0, "seeds": [seed], + "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], + "unplayable": None} + + +def chart(difficulty, score_id, last=60000): + return {"difficulty": difficulty, "scoreId": score_id, "level": 10, "displayLevel": 10.5, "fullComboCount": 20, + "asset": {"key": f"Live/MusicScore/c/c_{score_id}", "sha256": "ab" * 32}, + "notes": {"judged": 20, "total": 22, "byOperateType": {"1": 20, "120": 2}}, + "bpm": {"main": 120.0, "min": 120.0, "max": 120.0, "changes": [{"timeMs": 0, "bpm": 120.0}]}, + "firstNoteMs": 1000, "lastJudgedNoteMs": last, "lastNoteMs": last, "musicLengthMs": last + 1000, + "skillEventsMs": [1000, 3000], "fevers": [[1000, 5000]], "deck": deck_chart()} + + +def song(i, charts): + ranks = ["D", "C", "B", "A", "S", "SS"] + return {"id": i, "sortOrder": i, "startAt": "2026/01/01 0:00:00", "defaultUnlock": True, "title": text(f"t{i}"), + "ruby": None, "phonetic": text(f"p{i}"), "bandIds": [1], "bandName": None, "vocalCharacterIds": [1], + "lyricist": text("l"), "composer": text("c"), "arranger": None, "musicType": 1, "musicCategories": [1], + "bestMusicTagIds": [1], "jacket": f"jkt_{i}", "gekisouMissions": [1, 2, 3], + "bgm": {"soundId": i, "cueSheet": f"Bgm{i}", "cue": f"song{i}", + "length": {"lengthMs": 90000, "samples": 48000 * 90, "sampleRate": 48000, "durationMs": 90000}}, + "scoreRanks": [{"rank": r, "requiredScore": n * 1000, "battleRequiredScore": n * 2000} + for n, r in enumerate(ranks)], + "charts": charts, "master": {"MasterLiveMusic": {"_id": i, "_rate": 1.5}, "MasterLiveScoreRank": []}} + + +def sample() -> dict: + kind = {"id": 0, "effectType": 2000, "activationTimeSecond": 5.0, "durationMs": 5000, "skillTargetIds": [], + "skillConditionGroup": 0, "skillReleaseConditionGroup": 0, "effectLimitCount": 0, + "effectExecuteLimitCount": 0, "effectExecuteLimitResetConditionGroup": 0, "rows": [1], "values": [10000]} + return { + "format": "nnnotes.music-data/1", + "provenance": { + "region": "tw", "client": {"versionName": "1.0.1", "versionCode": 25}, + "catalog": {"resourceVersion": None, "sha256": "cd" * 32}, + "master": {"source": "api", "version": "v-test", "tables": {t: {"sha256": BIN[t]} for t in TABLES}}, + "exporter": {"name": "nnnotes", "version": "0.1.2", "chartFormat": "nnnotes.live-score/1"}, + "deck": {"name": "ournotes-deck", "version": "0.0.1", + "source": "https://github.com/empty-sekai/ournotes-deck", "commit": DECK, + "format": "ournotes-deck.chart-stats/2"}}, + "languages": LANGS, + "bands": [{"id": 1, "name": text("band"), "mainColor": "#3388BB", "subColor": "#FFFFFF"}], + "characters": [{"id": 1, "bandId": 1, "name": text("ch"), "shortName": text("c"), "mainColor": "#77BBDD"}], + "tags": [{"id": 1, "name": text("tag")}], + "categories": [{"id": 1, "musicCategories": [1], "name": text("cat")}], + "deck": {"model": {"power": 300000, "checkPower": 1000003}, "kinds": [kind]}, + "songs": [song(100001, [chart("easy", 10), chart("expert", 30)]), song(100002, [chart("expert", 40)])], + } + + +def context(tmp_path: Path, doc: dict, **kw) -> tuple[bytes, Context]: + """The file's bytes and what check reads next to it: the decoded master data and its snapshot, the jackets (the + page smoke test only with page=page_path(): it starts Node.js).""" + master = tmp_path / "master" + master.mkdir(exist_ok=True) + files = {} + for t in TABLES: + data = json.dumps({"_allData": []}).encode() + (master / f"{t}.json").write_bytes(data) + files[f"{t}.json"] = hashlib.sha256(data).hexdigest() + listed = [{"name": f"{t}.bin", "hash": BIN[t], "size": 1} for t in TABLES] + (master / music_data.MANIFEST).write_text(json.dumps({"version": "v-test", "files": listed}), encoding="utf-8") + jackets = tmp_path / "jackets" + jackets.mkdir(exist_ok=True) + for s in doc.get("songs") or []: + if s.get("jacket"): + (jackets / f"{s['jacket']}.webp").write_bytes(b"RIFF....WEBP") + # an infinity as nnnotes writes it (1e999: JSON that JavaScript reads too); a NaN stays NaN (nnnotes writes none) + raw = json.dumps(doc, ensure_ascii=False).replace("Infinity", "1e999").encode("utf-8") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(region="tw", language="zh-Hant", master=master, + snapshot={"entry": {"version": "v-test", "client_version": "1.0.1"}, "files": files}, + deck_commit=DECK, nnnotes_version="0.1.2", jackets=jackets, schema=schema_path(), published=None, + page=None, file=tmp_path / "music-data.json") + for k, v in kw.items(): + setattr(ctx, k, v) + return raw, ctx + + +def run(tmp_path, doc, **kw) -> dict: + return gates(*context(tmp_path, doc, **kw)) + + +def gate(report: dict, name: str) -> dict: + return next(g for g in report["gates"] if g["gate"] == name) + + +def failures(report: dict) -> list: + return [(g["gate"], g["failures"]) for g in report["gates"] if not g["passed"]] + + +def seed(doc, song=0, chart=0): + return doc["songs"][song]["charts"][chart]["deck"]["seeds"][0] + + +# ---------------------------------------------------------------- the gates +def test_the_sample_passes_every_gate(tmp_path): + r = run(tmp_path, sample(), page=page_path()) + assert r["passed"], failures(r) + assert [g["gate"] for g in r["gates"]] == [name for name, _ in music_data.GATES] + assert not [g for g in r["gates"] if g["warningCount"]] + assert r["sha256"] == hashlib.sha256((tmp_path / "music-data.json").read_bytes()).hexdigest() + + +def test_schema(tmp_path): + if schema_path() is None: + pytest.skip("no docs/schema/music-data.schema.json (a fork not synced with upstream) or no jsonschema") + doc = sample() + doc["songs"][0]["charts"][0]["level"] = "10" + g = gate(run(tmp_path, doc), "schema") + assert not g["passed"] and any("songs/0/charts/0/level" in f for f in g["failures"]) + + +def test_page_smoke(tmp_path): + if page_path() is None: + pytest.skip("set MUSIC_DATA_PAGE to ournotes-player's examples/songs (and install Node.js)") + assert gate(run(tmp_path, sample(), page=page_path()), "page")["passed"] + doc = sample() + for s in doc["songs"]: + for c in s["charts"]: + c["deck"].pop("offSeeds") + g = gate(run(tmp_path, doc, page=page_path()), "page") + assert not g["passed"] and any("free scenario" in f for f in g["failures"]) + + +def unplayable(doc): + doc["songs"][1]["charts"][0]["deck"]["unplayable"] = "more than three fevers" # its seeds kept + + +@pytest.mark.parametrize("change, name, match", [ + # the play scenario fields + (lambda d: d["songs"][0]["charts"][0]["deck"].pop("offSeeds"), "scenarios", "offSeeds missing"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["offSeeds"].append({}), "scenarios", "offSeeds has 2 entries"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].pop("rankBonusPercents"), "scenarios", + "rankBonusPercents missing"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[250, 190, 160, 100]), + "scenarios", "rankBonusPercents not five ints"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[251, 190, 160, 100, 100]), + "scenarios", "is not rankBonusPercent"), + (lambda d: seed(d).pop("scorePerfect"), "scenarios", "no scorePerfect"), + (lambda d: seed(d).pop("rangeWeights"), "scenarios", "no rangeWeights"), + (lambda d: seed(d).pop("rankCheck"), "scenarios", "no rankCheck"), + (lambda d: seed(d)["ranges"][0].pop("rangeScorePerfect"), "scenarios", "rangeScorePerfect missing"), + (lambda d: seed(d).update(rangeWeights=[[[0.1]]]), "scenarios", "rangeWeights are not"), + (lambda d: seed(d)["rankCheck"].update(exact=99999), "scenarios", "rank check deck is not within"), + # the deck statistics + (lambda d: d.update(deck=None), "deck", "deck is null"), + (lambda d: d["songs"][0]["charts"][0].update(deck=None), "deck", "no deck statistics"), + (lambda d: seed(d)["check"].update(exact=99999), "deck", "check deck is not within"), + (lambda d: seed(d).update(weights=[[0.5]]), "deck", "weights are not"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(seeds=[]), "deck", "no seeds"), + (unplayable, "deck", "unplayable, but has Gekisou on seeds"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(events=[[0, 1000]]), "deck", "1 skill events"), + # numbers + (lambda d: seed(d)["weights"][0].__setitem__(1, float("nan")), "finite", "weights[0][1]: not finite"), + (lambda d: d["songs"][1]["charts"][0]["bpm"].update(main=float("inf")), "finite", "bpm.main: not finite"), + # references and texts + (lambda d: d["songs"][1]["bandIds"].append(9), "references", "band 9 not in"), + (lambda d: d["songs"][1]["vocalCharacterIds"].append(7), "references", "character 7 not in"), + (lambda d: d["songs"][0].update(title=None), "references", "title: no text"), + (lambda d: d["songs"][0].update(title={lang: "" for lang in LANGS}), "references", "empty in every language"), + (lambda d: d["songs"][0]["composer"].pop("ko"), "references", "composer: not a text"), + (lambda d: d["songs"][0].update(jacket=None), "references", "no jacket"), + (lambda d: d["songs"].reverse(), "references", "not sorted"), + (lambda d: d["songs"][1]["charts"][0].update(scoreId=10), "references", "occurs twice"), + (lambda d: d["songs"][0]["charts"].reverse(), "references", "charts ['expert', 'easy']"), + (lambda d: d["songs"][0].update(scoreRanks=[]), "references", "score ranks"), + (lambda d: d["songs"][0].update(bandIds=[]), "references", "neither a band nor a band name"), + # BGM + (lambda d: d["songs"][0]["bgm"].update(length=None), "bgm", "no BGM length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=50000, samples=48000 * 50), "bgm", + "ends before the last note"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=90500), "bgm", "is not samples"), + (lambda d: d["songs"][0]["bgm"]["length"].update(lengthMs=95000), "bgm", "cue length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=1200000, samples=48000 * 1200), "bgm", + "bounds"), + # provenance + (lambda d: d.update(format="nnnotes.music-data/2"), "provenance", "format"), + (lambda d: d["provenance"].update(region="en"), "provenance", "region 'en'"), + (lambda d: d["provenance"]["master"].update(version="other"), "provenance", "the snapshot's 'v-test'"), + (lambda d: d["provenance"]["master"].update(source="embedded"), "provenance", "master.source"), + (lambda d: d["provenance"]["master"]["tables"]["MasterText"].update(sha256="00" * 32), "provenance", + "MasterText.sha256 is not"), + (lambda d: d["provenance"]["master"]["tables"].pop("MasterBand"), "provenance", "lacks MasterBand"), + (lambda d: d["provenance"]["deck"].update(commit="ee" * 20), "provenance", "rust/Cargo.lock pins"), + (lambda d: d["provenance"]["exporter"].update(version="0.0.9"), "provenance", "installed nnnotes"), + (lambda d: d["provenance"]["client"].update(versionName=None), "provenance", "no APK version"), +]) +def test_a_gate_fails(tmp_path, change, name, match): + doc = sample() + change(doc) + r = run(tmp_path, doc) + g = gate(r, name) + assert not r["passed"] and not g["passed"] and any(match in f for f in g["failures"]), g + + +def test_a_table_read_is_not_the_listed_one(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.master / "MasterText.json").write_text('{"_allData": [{}]}', encoding="utf-8") + g = gate(gates(raw, ctx), "provenance") + assert not g["passed"] and g["failures"] == ["MasterText.json read is not the one index.json lists"] + + +def test_a_jacket_file_is_missing(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.jackets / "jkt_100002.webp").unlink() + g = gate(gates(raw, ctx), "references") + assert not g["passed"] and g["failures"] == ["song 100002: no jacket file jackets/jkt_100002.webp"] + + +def test_warnings_do_not_fail(tmp_path): + doc = sample() + seed(doc).update(rangeWeights=None, rankCheck=None) # overlapping ranges + seed(doc, 0, 1)["rangeWeights"][0] = None # a kind reading the confirmed rank + doc["songs"][1]["charts"][0]["deck"]["offSeeds"][0]["weights"][0] = None + doc["songs"][0]["master"]["MasterLiveMusic"]["_rate"] = float("inf") # 1e999 in master data as served + doc["songs"][1]["title"]["zh-Hant"] = "" + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 3 and gate(r, "finite")["warningCount"] == 1 + assert gate(r, "references")["warnings"] == ["songs without a zh-Hant title: 100002"] + + +def test_against_the_published_file(tmp_path): + doc = sample() + raw, ctx = context(tmp_path, doc) + more = copy.deepcopy(doc) + more["songs"].append(song(100003, [chart("expert", 50)])) + ctx.published = json.dumps(more).encode() + r = gates(raw, ctx) + assert not gate(r, "counts")["passed"] and gate(r, "counts")["failures"] == [ + "2 songs, the published file has 3", "3 charts, the published file has 4"] + assert gate(r, "counts")["warnings"] == ["songs no longer in the file: 100003", "charts no longer in the file: 50"] + ctx.published = raw + b" " * (len(raw) * 2) # the same songs in a file three times as big + r = gates(raw, ctx) + assert gate(r, "counts")["passed"] and not gate(r, "size")["passed"] + ctx.published = raw[:len(raw) // 3] # a third: not JSON, and twice exceeded + r = gates(raw, ctx) + assert gate(r, "counts")["warnings"] == ["the published music-data.json is not JSON: skipped"] + assert not gate(r, "size")["passed"] + ctx.published = raw + assert gates(raw, ctx)["passed"] + + +def test_not_json(tmp_path): + r = gates(b"{", Context()) + assert not r["passed"] and r["gates"][0]["gate"] == "json" + + +def test_report_markdown(tmp_path): + doc = sample() + doc["songs"][0]["charts"][0]["deck"].pop("offSeeds") + text_ = music_data.report_markdown(run(tmp_path, doc)) + assert "| scenarios | **failed** (1) |" in text_ + assert "- scenarios failure: chart 10 (100001 easy): offSeeds missing, expected exactly one" in text_ + + +def test_deck_commit_from_cargo_lock(tmp_path): + lock = tmp_path / "rust" / "Cargo.lock" + lock.parent.mkdir() + lock.write_text('[[package]]\nname = "pyo3"\nversion = "0.29.0"\n\n[[package]]\nname = "ournotes-deck"\n' + 'version = "0.0.1"\nsource = "git+https://github.com/empty-sekai/ournotes-deck?rev=' + "a" * 40 + + "#" + "b" * 40 + '"\n', encoding="utf-8") + assert music_data.deck_commit(tmp_path) == "b" * 40 + with pytest.raises(SystemExit, match="sync the fork"): + music_data.deck_commit(tmp_path / "nowhere") + + +# ---------------------------------------------------------------- real files (when named) +def real(name): + p = os.environ.get(name) + if not p or not Path(p).is_file(): + pytest.skip(f"set {name} to a music-data.json") + return Path(p).read_bytes() + + +def test_a_real_file_without_the_scenario_fields(): + raw = real("MUSIC_DATA_OLD_SAMPLE") + r = gates(raw, Context(language="zh-Hant"), only=CONTENT) + assert failures(r) == [("scenarios", gate(r, "scenarios")["failures"])] + g = gate(r, "scenarios") + charts = sum(len(s["charts"]) for s in json.loads(raw)["songs"]) + assert g["failureCount"] >= charts and all("offSeeds missing" in f or "rankBonusPercents missing" in f + or "no scorePerfect" in f for f in g["failures"]) + + +def test_a_real_file_with_the_scenario_fields(tmp_path): + raw = real("MUSIC_DATA_SAMPLE") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(language="zh-Hant", page=page_path(), file=tmp_path / "music-data.json", published=raw) + r = gates(raw, ctx, only=CONTENT + ("counts", "size", "page")) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 0 + + +# ---------------------------------------------------------------- publish (a stand-in bucket) +class FakeS3: + def __init__(self, corrupt=None): + self.store, self.log, self.corrupt = {}, [], corrupt + + def upload_file(self, src, bucket, key, ExtraArgs): + self.store[key] = b"not it" if key == self.corrupt else Path(src).read_bytes() + self.log.append((key, ExtraArgs["CacheControl"], ExtraArgs["ContentType"])) + + def get_object(self, Bucket, Key): + return {"Body": io.BytesIO(self.store[Key])} + + +class FakeBucket: + name, prefix, writable = "moenotes", "music-data/", True + + def __init__(self, s3): + self.s3 = s3 + + def keys(self, sub=""): + return {k[len(self.prefix):]: len(v) for k, v in self.s3.store.items() if k.startswith(self.prefix + sub)} + + +def published_out(tmp_path, monkeypatch, s3): + out = tmp_path / "out" + out.mkdir() + raw, ctx = context(out, sample()) + report = gates(raw, ctx) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "build.json").write_text(json.dumps({"sha256": report["sha256"], + "archive": f"archive/v-test/{report['sha256']}.json"}), + encoding="utf-8") + monkeypatch.setattr(music_data, "bucket", lambda: FakeBucket(s3)) + monkeypatch.setattr(music_data.time, "sleep", lambda s: None) + monkeypatch.setenv("STORY_S3_ENDPOINT", "https://storage.example") + monkeypatch.setenv("STORY_S3_BUCKET", "moenotes") + monkeypatch.delenv("FORCE", raising=False) + monkeypatch.setenv("MUSIC_DATA_PUBLISH", "true") + return out, report + + +def test_publish_order_and_read_back(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + music_data.cmd_publish(str(out)) + keys = [k for k, _, _ in s3.log] + archive = f"music-data/archive/v-test/{report['sha256']}.json" + assert sorted(keys[:2]) == ["music-data/jackets/jkt_100001.webp", "music-data/jackets/jkt_100002.webp"] + assert keys[2:] == [archive, "music-data/music-data.json", "music-data/build.json"] + caches = {k: c for k, c, _ in s3.log} + assert caches[archive].endswith("immutable") and caches["music-data/music-data.json"] == "no-cache" + assert s3.store["music-data/music-data.json"] == (out / "music-data.json").read_bytes() + s3.log.clear() + music_data.cmd_publish(str(out)) # again: the jackets and the archive copy are there + assert [k for k, _, _ in s3.log] == ["music-data/music-data.json", "music-data/build.json"] + monkeypatch.setenv("FORCE", "true") + s3.log.clear() + music_data.cmd_publish(str(out), dry_run=True) + assert s3.log == [] + + +def test_publish_stops_before_the_marker_when_the_file_does_not_read_back(tmp_path, monkeypatch): + s3 = FakeS3(corrupt="music-data/music-data.json") + out, _ = published_out(tmp_path, monkeypatch, s3) + + def offline(*a, **kw): + raise music_data.urllib.error.URLError("offline") + monkeypatch.setattr(music_data, "get", offline) + with pytest.raises(SystemExit, match="music-data.json: the bucket does not serve what was uploaded"): + music_data.cmd_publish(str(out)) + assert "music-data/build.json" not in s3.store + + +@pytest.mark.parametrize("switch", [None, "", "false", "True", "1"]) +def test_publishing_is_off_unless_the_switch_is_true(tmp_path, monkeypatch, switch): + s3 = FakeS3() + out, _ = published_out(tmp_path, monkeypatch, s3) + if switch is None: + monkeypatch.delenv("MUSIC_DATA_PUBLISH") + else: + monkeypatch.setenv("MUSIC_DATA_PUBLISH", switch) + music_data.cmd_publish(str(out)) # a dry run, whatever the command line says + assert s3.log == [] and s3.store == {} + + +def test_publish_needs_passed_gates(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + (out / "check.json").write_text(json.dumps(dict(report, passed=False)), encoding="utf-8") + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "music-data.json").write_bytes(b"{}") # not the file that was checked + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + assert s3.store == {} diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml new file mode 100644 index 0000000..d0c9478 --- /dev/null +++ b/.github/workflows/music-data.yml @@ -0,0 +1,150 @@ +name: Music data + +# Publishes the music data file of the current master data for the chart data page (ournotes-player examples/songs): +# `nnnotes music-data` (every song and chart with the deck model's statistics, docs/music-data.md) of +# moenotes-masterdata-sync's decoded master data, into the story site's bucket under music-data/: music-data.json, +# jackets/, archive//.json and the build marker build.json. A run first compares what it would +# build (the master data snapshot, the deck commit, the nnnotes commit) with the published build.json and ends there +# when nothing changed; else it builds, checks the file (quality gates and a smoke test with the page's own modules) +# and uploads it: the jackets and the archive copy first, then music-data.json, the build marker last. A file that +# fails a gate is not published. Nothing is deleted from the bucket. +# +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new +# master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds although nothing changed +# (`force`) or does a dry run. Publishing switch: unless the repository variable MUSIC_DATA_PUBLISH is `true`, every +# run, whatever its trigger, is a dry run (it builds and checks, and uploads nothing). docs: .github/MUSIC_DATA.md + +on: + repository_dispatch: + types: [masterdata-updated] + workflow_dispatch: + inputs: + force: + description: "build and publish although the published file was made from the same inputs" + type: boolean + default: false + dry_run: + description: "build and check, but upload nothing" + type: boolean + default: false + schedule: + - cron: "41 3 * * *" + +permissions: + contents: read + +# One run at a time: a run replaces music-data.json and build.json. +concurrency: + group: music-data + cancel-in-progress: false + +env: + # The story site's bucket (its repository variables) and the key prefix of the music data. + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || 'music-data' }} + # The publishing switch: anything but 'true' makes every run a dry run. + MUSIC_DATA_PUBLISH: ${{ vars.MUSIC_DATA_PUBLISH || 'false' }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} + # The ournotes-player whose chart data page reads the file: its modules smoke-test every build. No default yet: the + # page with the play scenarios is not merged; a run stops before building until MUSIC_DATA_PLAYER_REF is set. + PLAYER_REPOSITORY: ${{ vars.MUSIC_DATA_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.MUSIC_DATA_PLAYER_REF || '' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the bundle key and the CDN come from secrets in the build step. + NNNOTES_CATALOG_REGION: tw + NNNOTES_CATALOG_LANGUAGE: zh-Hant + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 10 + outputs: + build: ${{ steps.plan.outputs.build }} + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 # the nnnotes commit: the last one that changed src/, rust/ or pyproject.toml + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - name: Inputs against the published build + id: plan + env: + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py plan + + build: + needs: plan + if: needs.plan.outputs.build == 'true' + runs-on: ubuntu-latest + timeout-minutes: 90 + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: The chart data page + run: .github/scripts/songs_page.sh "$WORK/page" + + - uses: Swatinem/rust-cache@v2 + with: + workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - name: Install nnnotes + # builds the extension module nnnotes._deck: the deck model at the commit rust/Cargo.toml pins + run: uv pip install --system -e ".[test]" boto3 + + - name: Gate self-test + env: + MUSIC_DATA_PAGE: ${{ env.WORK }}/page/examples/songs + run: python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/music_data.py master "$WORK/master" + + - name: Build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + # A new cache on every run, never actions/cache: nnnotes keeps a downloaded catalog file for good (a cached + # one would never see a new resource version), and the cache holds decrypted game files. + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + run: python .github/scripts/music_data.py build "$WORK/out" + + - name: Gates + run: python .github/scripts/music_data.py check "$WORK/out" "$WORK/master" "$WORK/page/examples/songs" + + - name: Publish + # Only after every gate passed (a failed step skips it). A dry run (the dry_run input, or the publishing switch + # MUSIC_DATA_PUBLISH off) lists what it would upload; it gets no bucket key: anonymous, it can only read. + env: + STORY_S3_ACCESS_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_ACCESS_KEY || '' }} + STORY_S3_SECRET_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_SECRET_KEY || '' }} + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py publish "$WORK/out" ${{ (github.event.inputs.dry_run == 'true' || vars.MUSIC_DATA_PUBLISH != 'true') && '--dry-run' || '' }}