From f8b09ea3dec8dfd95eb593d915cf67f59aad3ade Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 02:28:28 +0800 Subject: [PATCH 01/14] ci: build the StarMoe story site incrementally and publish it to its bucket The story-site workflow adds the stories the published site lacks: it plans from the master data of moenotes-masterdata-sync against the bucket's story manifests, builds the missing stories with `nnnotes web --story` on top of the fetched manifests (fonts, tools, player and APK pinned), rewrites the indexes with `--player-only`, and uploads new assets, then manifests, then the indexes; it never deletes. Triggered by the masterdata-updated dispatch, daily, or by hand (story ids, force, dry run). Everything is under .github/, so the fork keeps syncing with upstream. Co-Authored-By: Claude Opus 5.5 (1M context) --- .github/STORY_SITE.md | 67 ++++++++ .github/scripts/apk.sh | 26 +++ .github/scripts/fonts.sh | 27 +++ .github/scripts/player.sh | 11 ++ .github/scripts/story_site.py | 278 +++++++++++++++++++++++++++++++ .github/scripts/tools.sh | 16 ++ .github/workflows/story-site.yml | 175 +++++++++++++++++++ 7 files changed, 600 insertions(+) create mode 100644 .github/STORY_SITE.md create mode 100755 .github/scripts/apk.sh create mode 100755 .github/scripts/fonts.sh create mode 100755 .github/scripts/player.sh create mode 100755 .github/scripts/story_site.py create mode 100755 .github/scripts/tools.sh create mode 100644 .github/workflows/story-site.yml diff --git a/.github/STORY_SITE.md b/.github/STORY_SITE.md new file mode 100644 index 0000000..c28ee84 --- /dev/null +++ b/.github/STORY_SITE.md @@ -0,0 +1,67 @@ +# Story site workflow (StarMoe) + +`.github/workflows/story-site.yml` keeps the StarMoe story site up to date: it builds the stories the published site +lacks with this repository's `nnnotes web --story` and uploads them to the bucket that serves the site +(`https://storage.bdon.moe/moenotes/`, the layout `nnnotes web` writes: `stories.json`, `stories/`, `models/`, +`assets/`, `story/`). It only adds: a run never deletes anything from the bucket. Its helper steps are in +`.github/scripts/`; nothing outside `.github/` differs from upstream, so the fork syncs with it as before. + +## A run + +1. **plan** (a few seconds): the MasterAdv ids of the decoded master data of moenotes-masterdata-sync + (`MasterAdv.json`, SHA-256 checked against its `index.json`) against the `stories/.json` objects of the bucket. + The ids without a manifest, at most `STORY_LIMIT` (40) in id order, are the run's stories; none: the run ends here. +2. **build** (only when there is something to build): + - fonts (pinned by SHA-256: the files the published stories record in `ui/fonts.json`), vgmstream, ffmpeg, the + built ournotes-player (`STORY_PLAYER_REF`), the APK (playfetch with the account in `PLAYFETCH_CREDENTIALS`), the + decoded master data; + - every object of the site except `assets/` (the manifests and indexes, about 150 MB); + - `nnnotes web site --story ...` (with the Live2D models these stories load that the site lacks), then + `nnnotes web site --player-only`, which rewrites `stories.json`, `models.json`, `charts.json` and the player + pages from every manifest present, also after a failed build; + - upload: the assets the bucket lacks, then the new or changed manifests and player files, the indexes last. + Stories that failed have no manifest, so the next run builds them again; their errors are in the job summary. + +Triggers: `repository_dispatch` `masterdata-updated` (moenotes-masterdata-sync sends it when a region serves a new +snapshot: `dispatch_repositories`), a daily schedule (03:23 UTC) in case a dispatch was missed, and +`workflow_dispatch`: + +| Input | Meaning | +|---|---| +| `stories` | MasterAdv ids to build (spaces or commas); empty: every story the site lacks | +| `force` | rebuild the given stories and their Live2D models although their manifests exist | +| `dry_run` | build, then list what would be uploaded instead of uploading | + +Runs do not overlap (`concurrency: story-site`). + +## Settings + +Repository secrets (Settings → Secrets and variables → Actions → Secrets): + +| Secret | Value | +|---|---| +| `NNNOTES_BUNDLE_KEY` | `[bundle] key` (32 hex digits) | +| `NNNOTES_BUNDLE_NONCE_SEED` | `[bundle] nonce_seed` | +| `NNNOTES_SERVERS_TW_CDN` | `[servers.tw] cdn`: the TW CDN base URL | +| `PLAYFETCH_CREDENTIALS` | the whole `credentials.json` of `playfetch login` (the account that can pull `com.bilibili.sirius`) | +| `STORY_S3_ACCESS_KEY`, `STORY_S3_SECRET_KEY` | an S3 key that can list, read and write the bucket | + +Repository variables (optional; the defaults are the StarMoe site): `STORY_S3_ENDPOINT` (`https://storage.bdon.moe`), +`STORY_S3_BUCKET` (`moenotes`), `STORY_S3_PREFIX` (empty: the bucket root), `MASTERDATA_BASE_URL` +(`https://metadata.bdon.moe`), `STORY_MASTERDATA_REGION` (`hk-tw-mo`), `STORY_LIMIT` (`40`), `STORY_PLAYER_REPOSITORY` +(`empty-sekai/ournotes-player`), `STORY_PLAYER_REF` (`3774d8ac3987`), `PLAYFETCH_VERSION` (`v0.92`), +`STORY_APK_PACKAGE` (`com.bilibili.sirius`). + +## Notes + +- **The player version.** The site's player pages come from `STORY_PLAYER_REF`, and its data must be what that player + reads. After a player release that changes the data format, sync this fork with upstream nnnotes and set + `STORY_PLAYER_REF` to the matching player; new stories are then written for it. Stories already published are + not rebuilt (run with `stories` and `force` for that). +- **The APK.** Google Play serves the current game version. When nnnotes' type trees do not match it, the build stops + naming the class and the Unity version: sync the fork with upstream once nnnotes supports that version. +- **Bytes.** A rebuild on the runner is not byte-identical to a build elsewhere: the AAC encoder (ffmpeg) and the PNG + encoder (zlib) differ between machines. The pixels and every JSON value are the same, so new stories and the + published ones fit together; only a forced rebuild re-uploads such files. +- **Master data.** The build reads moenotes-masterdata-sync's current snapshot; stories published from another master + source keep their texts until they are rebuilt. diff --git a/.github/scripts/apk.sh b/.github/scripts/apk.sh new file mode 100755 index 0000000..9aa5f52 --- /dev/null +++ b/.github/scripts/apk.sh @@ -0,0 +1,26 @@ +#!/usr/bin/env bash +# The game's split APKs (base.apk; split_config.arm64_v8a.apk holds the CRI Lips library) into $1, with playfetch +# ($PLAYFETCH_VERSION) and the account store in the secret PLAYFETCH_CREDENTIALS (the credentials.json of +# `playfetch login`). Google Play serves the current version; nnnotes stops, naming the class, when its type trees do not +# match the APK's Unity version. +set -euo pipefail +dir="$1" +if [ -z "${PLAYFETCH_CREDENTIALS_JSON:-}" ]; then + echo "::error::set the repository secret PLAYFETCH_CREDENTIALS (the credentials.json of playfetch login)" + exit 1 +fi +bin="$RUNNER_TEMP/playfetch" +curl -fsSL -o "$bin" "https://github.com/Exmeaning/playfetch/releases/download/$PLAYFETCH_VERSION/playfetch-$PLAYFETCH_VERSION-linux-amd64" +curl -fsSL -o "$RUNNER_TEMP/SHA256SUMS" "https://github.com/Exmeaning/playfetch/releases/download/$PLAYFETCH_VERSION/SHA256SUMS" +(cd "$RUNNER_TEMP" && grep " playfetch-$PLAYFETCH_VERSION-linux-amd64\$" SHA256SUMS | sed "s| playfetch-$PLAYFETCH_VERSION-linux-amd64| playfetch|" | sha256sum -c -) +chmod 755 "$bin" +export PLAYFETCH_CREDENTIALS="$RUNNER_TEMP/playfetch-credentials.json" +umask 077 +printf '%s' "$PLAYFETCH_CREDENTIALS_JSON" > "$PLAYFETCH_CREDENTIALS" +"$bin" pull "$APK_PACKAGE" -out-root "$RUNNER_TEMP/apk-downloads" -mode split +got="$(find "$RUNNER_TEMP/apk-downloads/$APK_PACKAGE" -name base.apk | sort | tail -1)" +[ -n "$got" ] || { echo "::error::playfetch pulled no base.apk"; exit 1; } +mkdir -p "$dir" +cp "$(dirname "$got")"/*.apk "$dir/" +rm -f "$PLAYFETCH_CREDENTIALS" +ls -l "$dir" diff --git a/.github/scripts/fonts.sh b/.github/scripts/fonts.sh new file mode 100755 index 0000000..40dd4f8 --- /dev/null +++ b/.github/scripts/fonts.sh @@ -0,0 +1,27 @@ +#!/usr/bin/env bash +# The open fonts the site's story text is drawn with (the files the published stories record in ui/fonts.json), into +# $1, each checked against its SHA-256: Noto Sans CJK 2.004 (ja and en: JP, zh-Hant: TC, zh-Hans: SC), Pretendard 1.3.9 +# SemiBold (ko), Noto Color Emoji 2.051 (the emoji sprites). All under the SIL Open Font License 1.1. +set -euo pipefail +dir="$1" +mkdir -p "$dir" +cd "$dir" +cjk=https://github.com/notofonts/noto-cjk/raw/165c01b46ea533872e002e0785ff17e44f6d97d8/Sans/OTF +fetch() { # fetch FILE URL SHA256 + if [ ! -f "$1" ] || ! echo "$3 $1" | sha256sum -c --quiet - 2>/dev/null; then + curl -fsSL -o "$1" "$2" + echo "$3 $1" | sha256sum -c - + fi +} +fetch NotoSansCJKjp-Regular.otf "$cjk/Japanese/NotoSansCJKjp-Regular.otf" 68a3fc98800b2a27b371f2fb79991daf3633bd89309d4ffaa6946fd587f375b5 +fetch NotoSansCJKtc-Regular.otf "$cjk/TraditionalChinese/NotoSansCJKtc-Regular.otf" dce08bd4fd91aa8aa76ed8fea4b694c2dfb8550f67871e326843212ddbeb88b4 +fetch NotoSansCJKsc-Regular.otf "$cjk/SimplifiedChinese/NotoSansCJKsc-Regular.otf" 2c76254f6fc379fddfce0a7e84fb5385bb135d3e399294f6eeb6680d0365b74b +fetch NotoColorEmoji.ttf https://github.com/googlefonts/noto-emoji/raw/v2.051/fonts/NotoColorEmoji.ttf 72a635cb3d2f3524c51620cdde406b217204e8a6a06c6a096ff8ed4b5fd6e27b +if [ ! -f Pretendard-SemiBold.otf ] || ! echo "c89bc43027dc7cde5726e96223376f8eec09302b2fc1f8147fd5b57cfc376118 Pretendard-SemiBold.otf" | sha256sum -c --quiet - 2>/dev/null; then + curl -fsSL -o pretendard.zip https://github.com/orioncactus/pretendard/releases/download/v1.3.9/Pretendard-1.3.9.zip + echo "04be351a74d6bf7d60c480a3087e51d185485d35a52023142af1df19eb8c428a pretendard.zip" | sha256sum -c - + unzip -q -o -j pretendard.zip public/static/Pretendard-SemiBold.otf + rm pretendard.zip + echo "c89bc43027dc7cde5726e96223376f8eec09302b2fc1f8147fd5b57cfc376118 Pretendard-SemiBold.otf" | sha256sum -c - +fi +ls -l diff --git a/.github/scripts/player.sh b/.github/scripts/player.sh new file mode 100755 index 0000000..fb9d3b9 --- /dev/null +++ b/.github/scripts/player.sh @@ -0,0 +1,11 @@ +#!/usr/bin/env bash +# A built ournotes-player checkout ($PLAYER_REPOSITORY at $PLAYER_REF) in $1: nnnotes writes its page files into the site. +set -euo pipefail +dir="$1" +if [ ! -f "$dir/dist/ournotes-player.element.min.js" ] || [ "$(git -C "$dir" rev-parse HEAD 2>/dev/null)" != "$(git -C "$dir" rev-parse "$PLAYER_REF^{commit}" 2>/dev/null)" ]; then + rm -rf "$dir" + git clone --quiet --filter=blob:none "https://github.com/$PLAYER_REPOSITORY.git" "$dir" + git -C "$dir" checkout --quiet "$PLAYER_REF" + (cd "$dir" && npm ci --ignore-scripts --no-audit --no-fund && npm run build) +fi +echo "ournotes-player $(git -C "$dir" rev-parse --short HEAD) ($(node -p "require('$dir/package.json').version"))" diff --git a/.github/scripts/story_site.py b/.github/scripts/story_site.py new file mode 100755 index 0000000..2690bf2 --- /dev/null +++ b/.github/scripts/story_site.py @@ -0,0 +1,278 @@ +#!/usr/bin/env python3 +"""The story site's CI steps (.github/workflows/story-site.yml): the published site lives in an S3 bucket, a run adds +the stories it lacks. + + plan the stories to build: $REQUESTED, else every MasterAdv id without a manifest in the bucket + (at most $STORY_LIMIT, in id order); GitHub outputs `stories` (space separated) and `count` + master OUT the decoded master data of $MASTERDATA_REGION from moenotes-masterdata-sync (index.json: + every table with its SHA-256), for `nnnotes web` + fetch SITE every object of the site except assets/ into SITE (the manifests `nnnotes web` skips and the + indexes it rewrites from), with their SHA-256 in SITE.fetched.json + build SITE ID... `nnnotes web SITE --story ID ...` ($FORCE: --force), then `--player-only`, which rewrites the + player files and the indexes from every manifest present even after a failed build + publish SITE [--dry-run] + the new assets (those the bucket lacks), then every other file that is new or changed since + `fetch`, the indexes (stories.json, models.json, charts.json) last; never deletes + +Bucket: $STORY_S3_ENDPOINT, $STORY_S3_BUCKET, $STORY_S3_PREFIX (a key prefix, may be empty), credentials +$STORY_S3_ACCESS_KEY / $STORY_S3_SECRET_KEY (unset: anonymous, reads only). Objects are stored as nnnotes writes them: +a compressed asset (.gz / .br) as it is, without Content-Encoding (the player decodes it). +""" +from __future__ import annotations + +import hashlib +import json +import os +import re +import subprocess +import sys +import urllib.request +from concurrent.futures import ThreadPoolExecutor +from pathlib import Path + +INDEXES = ("stories.json", "models.json", "charts.json") +ASSETS = "assets/" +STORY_MANIFEST = re.compile(r"stories/(\d+)\.json") +ASSET_CACHE = "public, max-age=31536000, immutable" # content-addressed: never changes +OTHER_CACHE = "public, max-age=300" +TYPES = {".json": "application/json", ".gz": "application/gzip", ".br": "application/octet-stream", + ".m4a": "audio/mp4", ".mp4": "video/mp4", ".webm": "video/webm", ".png": "image/png", ".jpg": "image/jpeg", + ".glsl": "text/plain; charset=utf-8", ".html": "text/html; charset=utf-8", + ".js": "text/javascript; charset=utf-8", ".mjs": "text/javascript; charset=utf-8", + ".css": "text/css; charset=utf-8", ".map": "application/json", ".txt": "text/plain; charset=utf-8", + ".svg": "image/svg+xml", ".wasm": "application/wasm", ".bin": "application/octet-stream"} + + +def get(url: str, timeout: int = 120) -> bytes: + # an explicit User-Agent: Cloudflare in front of the services refuses Python-urllib's + request = urllib.request.Request(url, headers={"User-Agent": "moenotes-story-site (GitHub Actions)", + "Cache-Control": "no-cache"}) + with urllib.request.urlopen(request, timeout=timeout) as r: + return r.read() + + +def env(name: str, default: str | None = None) -> str: + value = os.environ.get(name, "") + if value: + return value + if default is None: + sys.exit(f"story_site: set {name}") + return default + + +def output(name: str, value: str) -> None: + path = os.environ.get("GITHUB_OUTPUT") + if path: + with open(path, "a", encoding="utf-8") as f: + f.write(f"{name}={value}\n") + print(f"{name}={value}") + + +def summary(text: str) -> None: + path = os.environ.get("GITHUB_STEP_SUMMARY") + if path: + with open(path, "a", encoding="utf-8") as f: + f.write(text + "\n") + print(text) + + +# ---------------------------------------------------------------- bucket +class Bucket: + def __init__(self): + import boto3 + from botocore import UNSIGNED + from botocore.config import Config + key, secret = os.environ.get("STORY_S3_ACCESS_KEY"), os.environ.get("STORY_S3_SECRET_KEY") + # S3-compatible stores (SeaweedFS here) take path-style requests and no default CRC checksums + config = Config(s3={"addressing_style": "path"}, retries={"max_attempts": 8, "mode": "adaptive"}, + request_checksum_calculation="when_required", response_checksum_validation="when_required", + max_pool_connections=32, **({} if key and secret else {"signature_version": UNSIGNED})) + self.writable = bool(key and secret) + self.s3 = boto3.client("s3", endpoint_url=env("STORY_S3_ENDPOINT"), aws_access_key_id=key or None, + aws_secret_access_key=secret or None, region_name="us-east-1", config=config) + self.name = env("STORY_S3_BUCKET") + prefix = os.environ.get("STORY_S3_PREFIX", "").strip("/") + self.prefix = f"{prefix}/" if prefix else "" + + def keys(self, sub: str = "") -> dict[str, int]: + """{site path: size} of the objects under `sub` (a site path prefix).""" + out = {} + for page in self.s3.get_paginator("list_objects_v2").paginate(Bucket=self.name, Prefix=self.prefix + sub): + for o in page.get("Contents", []): + out[o["Key"][len(self.prefix):]] = o["Size"] + return out + + def download(self, path: str, dest: Path) -> None: + dest.parent.mkdir(parents=True, exist_ok=True) + self.s3.download_file(self.name, self.prefix + path, str(dest)) + + def upload(self, path: str, src: Path) -> None: + suffix = Path(path).suffix.lower() + extra = {"ContentType": TYPES.get(suffix, "application/octet-stream"), + "CacheControl": ASSET_CACHE if path.startswith(ASSETS) else OTHER_CACHE} + self.s3.upload_file(str(src), self.name, self.prefix + path, ExtraArgs=extra) + + +def parallel(fn, items, workers: int = 16) -> None: + with ThreadPoolExecutor(workers) as ex: + for _ in ex.map(fn, items): + pass + + +def sha256(path: Path) -> str: + h = hashlib.sha256() + with open(path, "rb") as f: + for block in iter(lambda: f.read(1 << 20), b""): + h.update(block) + return h.hexdigest() + + +# ---------------------------------------------------------------- master data +def master_index() -> tuple[str, dict]: + base = env("MASTERDATA_BASE_URL").rstrip("/") + index = json.loads(get(f"{base}/index.json", 60)) + region = index["regions"][env("MASTERDATA_REGION")] + if not re.fullmatch(r"/([a-z]+/)?master/", region["path"]): + sys.exit(f"story_site: unexpected master path {region['path']!r}") + return base + region["path"], region + + +def fetch_table(url: str, name: str, digest: str, dest: Path) -> None: + if not re.fullmatch(r"[A-Za-z0-9_-]+\.json", name): + sys.exit(f"story_site: unexpected master file name {name!r}") + for _ in range(4): + data = get(url + name) + if hashlib.sha256(data).hexdigest() == digest: + dest.write_bytes(data) + return + sys.exit(f"story_site: {name}: SHA-256 differs from index.json (the service changed snapshots? run again)") + + +def cmd_master(out: str) -> None: + url, region = master_index() + out_dir = Path(out) + out_dir.mkdir(parents=True, exist_ok=True) + parallel(lambda item: fetch_table(url, item[0], item[1], out_dir / item[0]), region["files"].items(), 8) + print(f"master data: {len(region['files'])} tables into {out_dir}") + + +# ---------------------------------------------------------------- plan +def parse_ids(text: str) -> list[int]: + ids = [] + for token in re.split(r"[\s,]+", text.strip()): + if token: + if not token.isdigit(): + sys.exit(f"story_site: not a story id: {token!r}") + ids.append(int(token)) + return list(dict.fromkeys(ids)) + + +def cmd_plan() -> None: + url, region = master_index() + data = get(url + "MasterAdv.json") + if hashlib.sha256(data).hexdigest() != region["files"]["MasterAdv.json"]: + sys.exit("story_site: MasterAdv.json: SHA-256 differs from index.json (run again)") + stories = sorted(row["_id"] for row in json.loads(data)["_allData"]) # nnnotes storysite.all_stories + have = {int(m.group(1)) for k in Bucket().keys("stories/") if (m := STORY_MANIFEST.fullmatch(k))} + requested = parse_ids(os.environ.get("REQUESTED", "")) + if requested: + unknown = [i for i in requested if i not in set(stories)] + if unknown: + sys.exit(f"story_site: no MasterAdv row for {', '.join(map(str, unknown))}") + todo, missing = requested, [i for i in requested if i not in have] + else: + missing = [i for i in stories if i not in have] + todo = missing[:int(env("STORY_LIMIT", "40"))] + summary(f"### Story site\n\n- master: {region.get('entry', {}).get('version', '?')} " + f"({len(stories)} stories), the site has {len(have)}\n- missing: {len(missing)}; this run builds " + f"{len(todo)}" + (f": {' '.join(map(str, todo))}" if todo else "")) + output("stories", " ".join(map(str, todo))) + output("count", str(len(todo))) + + +# ---------------------------------------------------------------- fetch / build / publish +def fetched_file(site: Path) -> Path: + return site.parent / f"{site.name}.fetched.json" + + +def cmd_fetch(site_dir: str) -> None: + site, bucket = Path(site_dir), Bucket() + keys = [k for k in bucket.keys() if not k.startswith(ASSETS) and not k.endswith("/")] + parallel(lambda k: bucket.download(k, site / k), keys) + (site / "assets").mkdir(parents=True, exist_ok=True) + digests = {k: sha256(site / k) for k in keys} + fetched_file(site).write_text(json.dumps(digests, indent=1, sort_keys=True), encoding="utf-8") + print(f"fetched {len(keys)} files ({sum((site / k).stat().st_size for k in keys) / 1e6:.1f} MB) into {site}") + + +def cmd_build(site_dir: str, ids: list[str]) -> None: + site = Path(site_dir).resolve() + stories = parse_ids(" ".join(ids)) + tmp = site.parent / f"{site.name}.tmp" + nnnotes = [sys.executable, "-m", "nnnotes", "web", str(site), "--tmp", str(tmp)] + status = 0 + if stories: + cmd = nnnotes + [a for i in stories for a in ("--story", str(i))] + cmd += ["--workers", env("STORY_WORKERS", "2")] + if os.environ.get("FORCE") == "true": + cmd.append("--force") + print("+ " + " ".join(cmd[1:]), flush=True) + status = subprocess.run(cmd, stdout=open(site.parent / f"{site.name}.build.json", "wb")).returncode + # The indexes and player files from every manifest present, whatever the build did. + index = subprocess.run(nnnotes + ["--player-only"], stdout=subprocess.PIPE).returncode + report = site.parent / f"{site.name}.build.json" + if report.is_file(): + try: + r = json.loads(report.read_text(encoding="utf-8")) + summary(f"- built {len(r.get('storiesBuilt', []))}, failed {len(r.get('storiesFailed', []))}, " + f"skipped {len(r.get('storiesSkipped', []))}; models built " + f"{len((r.get('storyModels') or {}).get('modelsBuilt') or [])} in {r.get('storySeconds', '?')} s") + except ValueError: + pass + if status or index: + sys.exit(f"story_site: nnnotes web exited with {status or index}") + + +def cmd_publish(site_dir: str, dry_run: bool = False) -> None: + site, bucket = Path(site_dir), Bucket() + if not bucket.writable and not dry_run: + sys.exit("story_site: publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") + before = json.loads(fetched_file(site).read_text(encoding="utf-8")) + local = sorted(p.relative_to(site).as_posix() for p in site.rglob("*") if p.is_file()) + assets = [k for k in local if k.startswith(ASSETS)] + have = bucket.keys(ASSETS) + new_assets = [k for k in assets if have.get(k) != (site / k).stat().st_size] + changed = [k for k in local if not k.startswith(ASSETS) and before.get(k) != sha256(site / k)] + last = [k for k in changed if k in INDEXES] + first = [k for k in changed if k not in INDEXES] + if dry_run: + for k in first + last: + print(f"would upload {k}") + else: + # manifests before the indexes that list them, assets before the manifests that name them + for group in (new_assets, first, last): + parallel(lambda k: bucket.upload(k, site / k), group) + summary(f"- {'would publish' if dry_run else 'published'} {len(new_assets)} assets ({sum((site / k).stat().st_size for k in new_assets) / 1e6:.1f} MB)" + f", {len(first)} files, indexes: {', '.join(last) or 'unchanged'}") + + +def main(argv: list[str]) -> None: + if not argv: + sys.exit(__doc__) + cmd, args = argv[0], argv[1:] + if cmd == "plan" and not args: + cmd_plan() + elif cmd == "master" and len(args) == 1: + cmd_master(args[0]) + elif cmd == "fetch" and len(args) == 1: + cmd_fetch(args[0]) + elif cmd == "build" and args: + cmd_build(args[0], args[1:]) + elif cmd == "publish" and len(args) in (1, 2) and args[1:] in ([], ["--dry-run"]): + cmd_publish(args[0], dry_run=bool(args[1:])) + else: + sys.exit(__doc__) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/.github/scripts/tools.sh b/.github/scripts/tools.sh new file mode 100755 index 0000000..5d925d4 --- /dev/null +++ b/.github/scripts/tools.sh @@ -0,0 +1,16 @@ +#!/usr/bin/env bash +# vgmstream-cli (the CRI HCA voices and music), pinned by version and SHA-256, into $1. ffmpeg comes from apt. +set -euo pipefail +dir="$1" +version=r2117 +digest=2f98c77f756079f63fbd119939067f1ed461d77e70993bc4cc372736d859c84a +mkdir -p "$dir" +if [ ! -x "$dir/vgmstream-cli" ]; then + curl -fsSL -o "$dir/vgmstream.zip" "https://github.com/vgmstream/vgmstream/releases/download/$version/vgmstream-linux.zip" + echo "$digest $dir/vgmstream.zip" | sha256sum -c - + unzip -q -o "$dir/vgmstream.zip" vgmstream-cli -d "$dir" + rm "$dir/vgmstream.zip" + chmod 755 "$dir/vgmstream-cli" +fi +"$dir/vgmstream-cli" -V 2>/dev/null | head -1 || true +ffmpeg -version | head -1 diff --git a/.github/workflows/story-site.yml b/.github/workflows/story-site.yml new file mode 100644 index 0000000..470464d --- /dev/null +++ b/.github/workflows/story-site.yml @@ -0,0 +1,175 @@ +name: Story site + +# Adds the stories that the published story site does not have yet to it. The site (stories.json, stories/, models/, +# assets/, story/: the layout `nnnotes web` writes) lives in an S3 bucket; a run fetches its manifests (not its assets), +# builds the missing stories with `nnnotes web --story` on top of them, and uploads what the build added or changed: +# new assets first, then the story and model manifests, the indexes last. Nothing is deleted from the bucket. +# +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new +# master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds given stories (with `force`: +# again) or does a dry run. docs: .github/STORY_SITE.md + +on: + repository_dispatch: + types: [masterdata-updated] + workflow_dispatch: + inputs: + stories: + description: "MasterAdv ids to build, separated by spaces or commas (empty: every story the site lacks)" + required: false + default: "" + force: + description: "rebuild the given stories (and their Live2D models) even though their manifests exist" + type: boolean + default: false + dry_run: + description: "build, but upload nothing" + type: boolean + default: false + schedule: + - cron: "23 3 * * *" + +permissions: + contents: read + +# One run at a time: every run rewrites the site's indexes from the manifests it fetched. +concurrency: + group: story-site + cancel-in-progress: false + +env: + # The bucket of the site (repository variables; the defaults are the StarMoe site). + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + STORY_S3_PREFIX: ${{ vars.STORY_S3_PREFIX || '' }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every table with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ vars.STORY_MASTERDATA_REGION || 'hk-tw-mo' }} + # At most this many stories per run (a run must end within the job limit); the next run builds the rest. + STORY_LIMIT: ${{ vars.STORY_LIMIT || '40' }} + # The ournotes-player the site is written for (its page files are part of the site). + PLAYER_REPOSITORY: ${{ vars.STORY_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.STORY_PLAYER_REF || '3774d8ac3987' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the keys and the CDN come from secrets in the build step. + NNNOTES_CATALOG_REGION: tw + NNNOTES_CATALOG_LANGUAGE: zh-Hant + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 15 + outputs: + stories: ${{ steps.plan.outputs.stories }} + count: ${{ steps.plan.outputs.count }} + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - run: python -m pip install --quiet boto3 + - name: Stories to build + id: plan + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + REQUESTED: ${{ github.event.inputs.stories }} + run: python .github/scripts/story_site.py plan + + build: + needs: plan + if: needs.plan.outputs.count != '0' + runs-on: ubuntu-latest + timeout-minutes: 340 + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: Free disk space + # The runner's preinstalled SDKs take most of its disk; a story build keeps its bundles and work files. + run: sudo rm -rf /usr/local/lib/android /usr/share/dotnet /opt/ghc /usr/local/.ghcup /opt/hostedtoolcache/CodeQL && df -h "$GITHUB_WORKSPACE" + + - name: Install nnnotes + run: | + python -m pip install --upgrade pip + python -m pip install -e ".[fonts]" boto3 + + - name: Tools + run: | + sudo apt-get update -qq && sudo apt-get install -y -qq ffmpeg > /dev/null + .github/scripts/tools.sh "$WORK/tools" + + - name: Fonts + uses: actions/cache@v6 + with: + path: ${{ env.WORK }}/fonts + key: story-fonts-${{ hashFiles('.github/scripts/fonts.sh') }} + - run: .github/scripts/fonts.sh "$WORK/fonts" + + - name: ournotes-player + uses: actions/cache@v6 + with: + path: ${{ env.WORK }}/player + key: story-player-${{ env.PLAYER_REPOSITORY }}-${{ env.PLAYER_REF }} + - run: .github/scripts/player.sh "$WORK/player" + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/story_site.py master "$WORK/master" + + - name: The published site + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + run: python .github/scripts/story_site.py fetch "$WORK/site" + + - name: Build + id: build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + NNNOTES_PATHS_PLAYER: ${{ env.WORK }}/player + NNNOTES_PATHS_VGMSTREAM: ${{ env.WORK }}/tools/vgmstream-cli + NNNOTES_PATHS_FONTS_JA: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf + NNNOTES_PATHS_FONTS_EN: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf + NNNOTES_PATHS_FONTS_ZH_HANT: ${{ env.WORK }}/fonts/NotoSansCJKtc-Regular.otf + NNNOTES_PATHS_FONTS_ZH_HANS: ${{ env.WORK }}/fonts/NotoSansCJKsc-Regular.otf + NNNOTES_PATHS_FONTS_KO: ${{ env.WORK }}/fonts/Pretendard-SemiBold.otf + NNNOTES_PATHS_FONTS_EMOJI: ${{ env.WORK }}/fonts/NotoColorEmoji.ttf + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/story_site.py build "$WORK/site" ${{ needs.plan.outputs.stories }} + + - name: Publish + # Also after a partial build: the stories that were built are published, the failed ones are built again by + # the next run (they have no manifest). A dry run lists what it would upload. + if: always() && steps.build.outcome != 'skipped' && steps.build.outcome != 'cancelled' + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + run: python .github/scripts/story_site.py publish "$WORK/site" ${{ github.event.inputs.dry_run == 'true' && '--dry-run' || '' }} + + - name: Failures + if: always() + run: | + f="$WORK/site.story-failures.json" + if [ -f "$f" ]; then echo "::error::stories failed, see the job summary"; { echo '### Failed stories'; echo '```json'; cat "$f"; echo '```'; } >> "$GITHUB_STEP_SUMMARY"; fi From a04c32fe984acee6404835596a1eded9fe50e826 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 17:54:47 +0800 Subject: [PATCH 02/14] story-site: fetch the site over plain HTTP (the bucket is public; signed ranged downloads were rejected by Cloudflare occasionally) The fetch step downloaded every manifest with S3's signed download_file (ranged), which Cloudflare's edge intermittently answered with SignatureDoesNotMatch. The bucket serves public read anyway and plan already lists it anonymously, so fetch now GETs each file with the same client plan uses; publish keeps its signed uploads. Co-Authored-By: Claude Code --- .github/scripts/story_site.py | 11 ++++++++++- 1 file changed, 10 insertions(+), 1 deletion(-) diff --git a/.github/scripts/story_site.py b/.github/scripts/story_site.py index 2690bf2..0bdc827 100755 --- a/.github/scripts/story_site.py +++ b/.github/scripts/story_site.py @@ -198,7 +198,16 @@ def fetched_file(site: Path) -> Path: def cmd_fetch(site_dir: str) -> None: site, bucket = Path(site_dir), Bucket() keys = [k for k in bucket.keys() if not k.startswith(ASSETS) and not k.endswith("/")] - parallel(lambda k: bucket.download(k, site / k), keys) + endpoint = env("STORY_S3_ENDPOINT").rstrip("/") + + # The bucket serves public read + list, so the fetch is plain HTTP like the player's, not a signed S3 + # call: Cloudflare's edge intermittently returned SignatureDoesNotMatch on signed ranged downloads. + def fetch(key: str) -> None: + dest = site / key + dest.parent.mkdir(parents=True, exist_ok=True) + dest.write_bytes(get(f"{endpoint}/{env('STORY_S3_BUCKET')}/{bucket.prefix}{key}", timeout=300)) + + parallel(fetch, keys) (site / "assets").mkdir(parents=True, exist_ok=True) digests = {k: sha256(site / k) for k in keys} fetched_file(site).write_text(json.dumps(digests, indent=1, sort_keys=True), encoding="utf-8") From bdcfaf2aa0fd256fd0c766103033fa6a690939ed Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 17:56:17 +0800 Subject: [PATCH 03/14] story-site: docs note the plain-HTTP fetch Co-Authored-By: Claude Code --- .github/STORY_SITE.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/.github/STORY_SITE.md b/.github/STORY_SITE.md index c28ee84..466dbbd 100644 --- a/.github/STORY_SITE.md +++ b/.github/STORY_SITE.md @@ -15,7 +15,9 @@ lacks with this repository's `nnnotes web --story` and uploads them to the bucke - fonts (pinned by SHA-256: the files the published stories record in `ui/fonts.json`), vgmstream, ffmpeg, the built ournotes-player (`STORY_PLAYER_REF`), the APK (playfetch with the account in `PLAYFETCH_CREDENTIALS`), the decoded master data; - - every object of the site except `assets/` (the manifests and indexes, about 150 MB); + - every object of the site except `assets/` (the manifests and indexes, about 150 MB), over plain HTTP like the + player's browser: the bucket serves public read, and Cloudflare's S3-signed ranged downloads were rejected + intermittently with `SignatureDoesNotMatch`; - `nnnotes web site --story ...` (with the Live2D models these stories load that the site lacks), then `nnnotes web site --player-only`, which rewrites `stories.json`, `models.json`, `charts.json` and the player pages from every manifest present, also after a failed build; From 12929724b0c102c75c282ed517e25654666a0774 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 18:19:08 +0800 Subject: [PATCH 04/14] story-site: pin UnityPy to the 1.24 line (1.25.x requires etcpak, which has no manylinux x86_64 wheel) UnityPy 1.25.0 declares Requires-Dist: etcpak, but etcpak's PyPI project ships neither manylinux x86_64 wheels nor a source build usable on a fresh Ubuntu runner with Python 3.13, so pip fails with 'no matching distribution' while 1.25 is the resolved version. The runner's UnityPy build has always been from sdist anyway; 1.24.2 publishes one and does not need etcpak. Co-Authored-By: Claude Code --- .github/workflows/story-site.yml | 3 +++ pyproject.toml | 4 +++- 2 files changed, 6 insertions(+), 1 deletion(-) diff --git a/.github/workflows/story-site.yml b/.github/workflows/story-site.yml index 470464d..1c7892d 100644 --- a/.github/workflows/story-site.yml +++ b/.github/workflows/story-site.yml @@ -104,6 +104,9 @@ jobs: - name: Install nnnotes run: | python -m pip install --upgrade pip + # UnityPy 1.25.x requires etcpak, which has no manylinux x86_64 wheel on PyPI (only PyPy / cp37), so the + # runner's cp313 environment cannot resolve it. Pin to the 1.24 line, which does not require etcpak. + python -m pip install "UnityPy>=1.24,<1.25" python -m pip install -e ".[fonts]" boto3 - name: Tools diff --git a/pyproject.toml b/pyproject.toml index 001a9d2..d108f5a 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -11,7 +11,9 @@ license-files = ["LICENSE"] requires-python = ">=3.11" authors = [{ name = "MetaMiku" }, { name = "emptysekai" }] dependencies = [ - "UnityPy>=1.25", + # 1.25.x requires etcpak, whose PyPI wheels lack manylinux x86_64 for new CPython, so the runner cannot install + # it. Stay on the 1.24 line (sdist on PyPI, built from source like before) until UnityPy fixes its etcpak dep. + "UnityPy>=1.24,<1.25", "numpy>=2.0", "Pillow>=10.0", "pycryptodome>=3.20", From 17a379e6545abc1e867c8e0f71463f8f0103c9d1 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 18:26:38 +0800 Subject: [PATCH 05/14] story-site: install with uv instead of pip pip 26's new resolver pins a cached unitypy==1.25.3 into the editable install's constraint set, which conflicts with the required etcpak (which has no manylinux x86_64 wheel on the runner's Python 3.13). uv's resolver does not carry pip's installed set into a fresh venv and picks unitypy==1.24.2, which is what pyproject.toml already allows. The two pip install steps (plan's boto3, build's editable install) now go through astral-sh/setup-uv with --system. Co-Authored-By: Claude Code --- .github/workflows/story-site.yml | 17 ++++++++++------- pyproject.toml | 2 +- 2 files changed, 11 insertions(+), 8 deletions(-) diff --git a/.github/workflows/story-site.yml b/.github/workflows/story-site.yml index 1c7892d..3c4e950 100644 --- a/.github/workflows/story-site.yml +++ b/.github/workflows/story-site.yml @@ -72,7 +72,10 @@ jobs: - uses: actions/setup-python@v7 with: python-version: ${{ env.PYTHON_VERSION }} - - run: python -m pip install --quiet boto3 + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - run: uv pip install --system --quiet boto3 - name: Stories to build id: plan env: @@ -101,13 +104,13 @@ jobs: # The runner's preinstalled SDKs take most of its disk; a story build keeps its bundles and work files. run: sudo rm -rf /usr/local/lib/android /usr/share/dotnet /opt/ghc /usr/local/.ghcup /opt/hostedtoolcache/CodeQL && df -h "$GITHUB_WORKSPACE" + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true - name: Install nnnotes - run: | - python -m pip install --upgrade pip - # UnityPy 1.25.x requires etcpak, which has no manylinux x86_64 wheel on PyPI (only PyPy / cp37), so the - # runner's cp313 environment cannot resolve it. Pin to the 1.24 line, which does not require etcpak. - python -m pip install "UnityPy>=1.24,<1.25" - python -m pip install -e ".[fonts]" boto3 + # UnityPy 1.25.x requires etcpak, which has no manylinux x86_64 wheel on PyPI (only PyPy / cp37), so the + # runner's cp313 environment cannot resolve it. pyproject.toml already pins it to the 1.24 line. + run: uv pip install --system -e ".[fonts]" boto3 - name: Tools run: | diff --git a/pyproject.toml b/pyproject.toml index d108f5a..4ed803d 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -12,7 +12,7 @@ requires-python = ">=3.11" authors = [{ name = "MetaMiku" }, { name = "emptysekai" }] dependencies = [ # 1.25.x requires etcpak, whose PyPI wheels lack manylinux x86_64 for new CPython, so the runner cannot install - # it. Stay on the 1.24 line (sdist on PyPI, built from source like before) until UnityPy fixes its etcpak dep. + # it (uv and pip both refuse). Stay on the 1.24 line until UnityPy fixes its etcpak dependency. "UnityPy>=1.24,<1.25", "numpy>=2.0", "Pillow>=10.0", From 7dbce8858055c9b5718edd74c624b4c62af9e206 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 19:14:48 +0800 Subject: [PATCH 06/14] ci: install with uv too, and pin goldens to UnityPy 1.24.2 The upstream CI workflow (push to main) was still using pip and its goldens job pinned UnityPy==1.25.3, which cannot resolve on the runner because etcpak has no manylinux x86_64 wheel. Install through uv like the story-site workflow and pin goldens to the 1.24 line. Co-Authored-By: Claude Code --- .github/goldens-constraints.txt | 2 +- .github/workflows/ci.yml | 14 ++++++++------ 2 files changed, 9 insertions(+), 7 deletions(-) diff --git a/.github/goldens-constraints.txt b/.github/goldens-constraints.txt index 9d81f28..08398e4 100644 --- a/.github/goldens-constraints.txt +++ b/.github/goldens-constraints.txt @@ -1,4 +1,4 @@ -UnityPy==1.25.3 +UnityPy==1.24.2 Pillow==12.3.0 numpy==2.4.6 texture2ddecoder==1.0.6 diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ad56323..89a9a6b 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -26,10 +26,11 @@ jobs: python-version: ${{ matrix.python-version }} cache: pip cache-dependency-path: pyproject.toml + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true - name: Install - run: | - python -m pip install --upgrade pip - python -m pip install -e ".[test]" + run: uv pip install --system -e ".[test]" - name: Lint run: python -m pyflakes src tests - name: Command line @@ -50,10 +51,11 @@ jobs: cache-dependency-path: | pyproject.toml .github/goldens-constraints.txt + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true - name: Install (pinned) - run: | - python -m pip install --upgrade pip - python -m pip install -c .github/goldens-constraints.txt -e ".[test]" + run: uv pip install --system -c .github/goldens-constraints.txt -e ".[test]" - name: Golden records env: GOLDENS_STRICT: "1" From a87313720dfe1bec826c4d1966235e5334eb8d2a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Mon, 28 Sep 2026 19:40:33 +0800 Subject: [PATCH 07/14] revert: UnityPy back to >=1.25 now that uv resolves it UnityPy 1.25.3's PyPI metadata originally required etcpak (which has no manylinux x86_64 wheel), but the current metadata for the same version declares tpk-ar instead. pip 26's resolver keeps the stale etcpak requirement in its cache and fails, while uv fetches the fresh metadata and resolves UnityPy 1.25.3 cleanly. Stay on >=1.25 and let uv handle it. Co-Authored-By: Claude Code --- .github/goldens-constraints.txt | 2 +- pyproject.toml | 6 +++--- 2 files changed, 4 insertions(+), 4 deletions(-) diff --git a/.github/goldens-constraints.txt b/.github/goldens-constraints.txt index 08398e4..9d81f28 100644 --- a/.github/goldens-constraints.txt +++ b/.github/goldens-constraints.txt @@ -1,4 +1,4 @@ -UnityPy==1.24.2 +UnityPy==1.25.3 Pillow==12.3.0 numpy==2.4.6 texture2ddecoder==1.0.6 diff --git a/pyproject.toml b/pyproject.toml index 4ed803d..3baadeb 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -11,9 +11,9 @@ license-files = ["LICENSE"] requires-python = ">=3.11" authors = [{ name = "MetaMiku" }, { name = "emptysekai" }] dependencies = [ - # 1.25.x requires etcpak, whose PyPI wheels lack manylinux x86_64 for new CPython, so the runner cannot install - # it (uv and pip both refuse). Stay on the 1.24 line until UnityPy fixes its etcpak dependency. - "UnityPy>=1.24,<1.25", + # 1.25.3's PyPI metadata briefly required etcpak (no manylinux x86_64 wheel); the same version now declares + # tpk-ar instead. pip 26's resolver keeps the cached old metadata and fails, so we install with uv. + "UnityPy>=1.25", "numpy>=2.0", "Pillow>=10.0", "pycryptodome>=3.20", From 61f4b584b52180869e0ac9dd962c2b0e2ab7f5d9 Mon Sep 17 00:00:00 2001 From: nichinichisou Date: Wed, 30 Sep 2026 01:35:51 +0800 Subject: [PATCH 08/14] ci: build, check and publish the music data file on new master data .github/workflows/music-data.yml runs `nnnotes music-data --decoded-master` on moenotes-masterdata-sync's decoded master data, the way the story site workflow reads it, and publishes the file for the chart data page into the story site's bucket under music-data/: music-data.json, jackets/, archive//.json and the build marker build.json. It reuses the story site's triggers (repository_dispatch masterdata-updated, a daily schedule, workflow_dispatch with force and dry_run), its plan/build jobs, secrets, bucket and helpers (story_site.py, apk.sh); no new secret, no master key. plan compares the master data snapshot, the pinned deck commit, the last nnnotes commit and the script's recipe with the published build.json and ends there when nothing changed. build checks the file before anything is uploaded: the JSON Schema, provenance against the snapshot and its manifest, no fewer songs or charts, complete deck statistics and play scenario fields, finite numbers, references and texts, BGM lengths, 0.8 to 2 times the published size, and a smoke test with the chart data page's own modules (MUSIC_DATA_PLAYER_REF, to be set). The jackets and the archive copy go first, music-data.json next, build.json last, each read back by SHA-256. The gate self-test (test_music_data.py) runs in every build and locally in seconds. The workflow needs the fork synced with upstream nnnotes (music-data with the play scenarios and --decoded-master): .github/MUSIC_DATA.md. Co-Authored-By: Claude Opus 5.5 --- .github/MUSIC_DATA.md | 122 ++++ .github/scripts/music_data.py | 847 +++++++++++++++++++++++++++ .github/scripts/music_data_smoke.mjs | 94 +++ .github/scripts/songs_page.sh | 18 + .github/scripts/test_music_data.py | 443 ++++++++++++++ .github/workflows/music-data.yml | 146 +++++ 6 files changed, 1670 insertions(+) create mode 100644 .github/MUSIC_DATA.md create mode 100755 .github/scripts/music_data.py create mode 100644 .github/scripts/music_data_smoke.mjs create mode 100755 .github/scripts/songs_page.sh create mode 100644 .github/scripts/test_music_data.py create mode 100644 .github/workflows/music-data.yml diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md new file mode 100644 index 0000000..2f669c4 --- /dev/null +++ b/.github/MUSIC_DATA.md @@ -0,0 +1,122 @@ +# Music data workflow (StarMoe) + +`.github/workflows/music-data.yml` keeps the music data file of the chart data page (ournotes-player +`examples/songs`) up to date: `nnnotes music-data` of the current master data ([docs/music-data.md](../docs/music-data.md): +every song and chart with the deck model's statistics and the play scenarios), checked by quality gates and +published into the story site's bucket under `music-data/` (`https://storage.bdon.moe/moenotes/music-data/`): + +| Object | Content | Cache-Control | +|---|---|---| +| `music-data.json` | the current file | `no-cache` | +| `jackets/.webp` | every song's jacket (`--jackets`), where the page looks for them | `public, max-age=86400` | +| `archive//.json` | every published file, kept | `public, max-age=31536000, immutable` | +| `build.json` | the build marker: the file's SHA-256, size and counts, what it was made from, the gate results, the run | `no-cache` | + +A run never deletes anything from the bucket. Its helper steps are `.github/scripts/music_data.py` (with the bucket, +HTTP and master data helpers of `story_site.py`), `music_data_smoke.mjs`, `songs_page.sh` and `apk.sh`; the gate +self-test is `test_music_data.py`. Nothing outside `.github/` differs from upstream, so the fork syncs with it as +before. The Cloudflare Pages preview of the page is not part of the workflow. + +## A run + +1. **plan** (seconds): the inputs of a build, from `index.json` of moenotes-masterdata-sync (the snapshot's master + data version, resource version and client version of `MUSIC_DATA_MASTERDATA_REGION`) and the checkout (the + ournotes-deck commit `rust/Cargo.lock` pins, the last nnnotes commit that changed `src/`, `rust/` or + `pyproject.toml`, and `RECIPE` of `music_data.py`), against `inputs` of the published `build.json`. The same + inputs (and a published `music-data.json`): the run ends here. `force` builds anyway. +2. **build**: + - the chart data page's modules (`examples/songs` of `MUSIC_DATA_PLAYER_REF`, not built), nnnotes with its deck + model (the install builds `nnnotes._deck`), the gate self-test, the APK (playfetch with `PLAYFETCH_CREDENTIALS`: + `provenance.client`; nnnotes also reads the bundles the APK carries, as in the story site's builds), the decoded + master data of moenotes-masterdata-sync (every file SHA-256 checked against `index.json`, `MasterManifest.json` + included); + - `nnnotes music-data --decoded-master --jackets jackets -o music-data.json`: the master data as decoded (no + master key), `provenance.master` the manifest's version and SHA-256 of the files as served; the charts, cue + sheets and jackets from the TW catalog, downloaded afresh on every run (never `actions/cache`: nnnotes keeps a + downloaded catalog for good, and the cache holds decrypted game files); + - the gates (below); a failed gate stops the run, the job summary lists why; + - upload: the jackets the bucket lacks (or has at another size; every one with `force`), the archive copy, then + `music-data.json`, `build.json` last. The archive copy, the file and the marker are each read back and checked + against their SHA-256 before the next is written: a failure leaves the previous `build.json`, so the next run + builds again. + +Triggers: `repository_dispatch` `masterdata-updated` (moenotes-masterdata-sync's `dispatch_repositories` already +names this repository for the story site: both workflows run), a daily schedule (03:41 UTC) in case a dispatch was +missed, and `workflow_dispatch`: + +| Input | Meaning | +|---|---| +| `force` | build and publish although the published file was made from the same inputs; upload every jacket again | +| `dry_run` | build and check, then list what would be uploaded instead of uploading | + +Runs do not overlap (`concurrency: music-data`). + +## Gates + +Every one must pass, else nothing is published. Warnings go to the job summary and `build.json` and do not stop it. + +| Gate | Checks | +|---|---| +| (build) | nnnotes' own checks: every table, chart and cue sheet read, every chart measured, the deck statistics cross-checked against the chart facts and the master data (the command writes no file otherwise) | +| `schema` | the file against `docs/schema/music-data.schema.json` of the checkout (JSON Schema 2020-12) | +| `provenance` | `format`; `region` `tw`; `master.source` `api`; `master.version` equal to the snapshot's and its `MasterManifest.json`'s; every table's SHA-256 the manifest's, every decoded table read the one `index.json` lists; the song tables and the deck model's present; `deck.commit` the one `rust/Cargo.lock` pins; `exporter.version` the installed nnnotes; an APK version; a catalog SHA-256 (warning: the APK is another client version than the snapshot's) | +| `counts` | no fewer songs and charts than the published file (warning: ids no longer in it) | +| `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning), `weights[kind][position]` numbers, every check deck within its bound | +| `scenarios` | the play scenario fields: `offSeeds` exactly one entry (seed 0, score, weights, check within its bound), every range's `rankBonusPercents` five ints (the first `rankBonusPercent`), every seed's `scorePerfect`, `rangeWeights` (`[kind][position][range]`) and `rankCheck` (within its bound), every seed range's `rangeScorePerfect` (warnings, none in TW: a null `rangeWeights`, a null kind in it or in `offSeeds`' weights) | +| `finite` | no NaN or infinity (warning: one inside master data rows, `songs[].master`, which the format writes as `1e999`) | +| `references` | texts in every language of `languages` (names and titles not empty); unique ids; songs sorted; the songs' bands, vocal characters and tags in the file; a band or a band name; a jacket, and its file in `jackets/`; a BGM cue; score ranks; charts in difficulty order, score ids unique (warnings: a title without a `zh-Hant` text, a music category on no tab, a character of no band) | +| `bgm` | every song's BGM length: `durationMs = samples * 1000 // sampleRate`, 30 s to 10 min, within 1 s of the cue's `lengthMs`, not ending before a chart's last note (warning: more than a minute after it) | +| `size` | 0.8 to 2 times the published file | +| `page` | `music_data_smoke.mjs`: the page's `catalog.js` and `ranking.js` in Node.js over the file: a row per chart, a plain score-up kind, data for the free, rank and Just scenarios, finite positive figures for every chart the data covers in seven scenarios (Gekisou Live at several ranks, Just rates and a Great share, Free Live), the ranking, frontier and event figures | +| (publish) | read back after upload, SHA-256 checked | + +`counts` and `size` compare with the published `music-data.json` and are skipped while nothing is published. A +legitimate drop (a song the game removed) stops the run: a person checks it, then moves the published +`music-data.json` away (the archive keeps it) or changes the gate in a pull request. + +The self-test runs in each build and locally in seconds, without the network: + +``` +python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py +``` + +with, optionally, `MUSIC_DATA_SCHEMA` (a schema file when the checkout has none), `MUSIC_DATA_PAGE` (an +`examples/songs` directory: the smoke test), `MUSIC_DATA_SAMPLE` (a real file with the play scenario fields: its +content gates pass) and `MUSIC_DATA_OLD_SAMPLE` (one without them: the scenario gate stops it). + +## Settings + +Repository secrets: the story site's (`.github/STORY_SITE.md`), no new one: `NNNOTES_BUNDLE_KEY`, +`NNNOTES_BUNDLE_NONCE_SEED`, `NNNOTES_SERVERS_TW_CDN`, `PLAYFETCH_CREDENTIALS`, `STORY_S3_ACCESS_KEY`, +`STORY_S3_SECRET_KEY`. The master key is not needed: the master data comes decoded. + +Repository variables: + +| Variable | Default | | +|---|---|---| +| `MUSIC_DATA_PLAYER_REF` | none: **required** | the ournotes-player commit whose chart data page reads this file (the page with the play scenarios); a run stops before building without it | +| `MUSIC_DATA_PLAYER_REPOSITORY` | `empty-sekai/ournotes-player` | | +| `MUSIC_DATA_S3_PREFIX` | `music-data` | the key prefix in the bucket | +| `MUSIC_DATA_MASTERDATA_REGION` | `hk-tw-mo` | the region of `index.json`; the build reads the TW catalog (`[catalog] region` `tw`) | +| `STORY_S3_ENDPOINT`, `STORY_S3_BUCKET`, `MASTERDATA_BASE_URL`, `PLAYFETCH_VERSION`, `STORY_APK_PACKAGE` | the story site's | shared with it | + +## Before the first run + +- **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play + scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with + upstream first. Until then `plan` stops naming what is missing. +- **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play + scenario fields, once that page is merged. + +## Notes + +- **Versions.** A new ournotes-deck pin (`rust/Cargo.toml`, `rust/Cargo.lock`) or a new nnnotes commit in `src/`, + `rust/` or `pyproject.toml` reaches this fork with a sync, and the next run builds a new file. The deck + statistics may then differ: `provenance.deck.commit` and `build.json` name the commit. +- **The page and the data.** The page's modules are pinned by `MUSIC_DATA_PLAYER_REF`: after a page release that + reads new fields, move it (a new ref alone does not start a build; run with `force` to check the published data + against the new page). +- **Byte identity.** A build from the same inputs gives the same bytes (the file is canonical); the jackets' + WebP bytes depend on the Pillow version. +- **Logs.** The steps print counts, ids, SHA-256 and field names, not game content; nothing decrypted is cached or + uploaded as an artifact. diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py new file mode 100755 index 0000000..2887336 --- /dev/null +++ b/.github/scripts/music_data.py @@ -0,0 +1,847 @@ +#!/usr/bin/env python3 +"""The music data CI steps (.github/workflows/music-data.yml): `nnnotes music-data` of moenotes-masterdata-sync's +decoded master data, checked by quality gates and published into the story site's bucket under $MUSIC_DATA_S3_PREFIX +(music-data.json, jackets/, archive/, build.json) for the chart data page of ournotes-player (examples/songs). + + plan the inputs of a build (the master data snapshot of $MASTERDATA_REGION in index.json, the + deck commit rust/Cargo.lock pins, the last nnnotes commit of src/, rust/ and pyproject.toml, + RECIPE) against those of the published build.json; GitHub output `build`: true when they + differ, nothing is published or $FORCE is true + master OUT every file of the snapshot of $MASTERDATA_REGION into OUT (SHA-256 checked against + index.json, MasterManifest.json included) and its index entry as OUT.snapshot.json + build OUT `nnnotes music-data --decoded-master --jackets OUT/jackets -o OUT/music-data.json` ([paths] + master: the master step's OUT), its printed summary in OUT/music-data.summary.json + check OUT MASTER PAGE the quality gates (.github/MUSIC_DATA.md) on OUT/music-data.json, with the published file + as the baseline and PAGE the chart data page's modules (examples/songs) for the smoke test; + the report in OUT/check.json, and OUT/build.json (the build marker) when every gate passed; + exits 1 when one failed + publish OUT [--dry-run] + the jackets the bucket lacks or has at another size (every one with $FORCE), the file's + archive copy, then music-data.json, build.json last; each read back and its SHA-256 checked. + Only a file whose check.json passed; never deletes + +Bucket: story_site.Bucket ($STORY_S3_ENDPOINT, $STORY_S3_BUCKET, credentials $STORY_S3_ACCESS_KEY / +$STORY_S3_SECRET_KEY) with the key prefix $MUSIC_DATA_S3_PREFIX (default music-data). The published files are read +over plain HTTP, as the page reads them (the bucket serves public read). +""" +from __future__ import annotations + +import hashlib +import json +import math +import os +import re +import subprocess +import sys +import time +import urllib.error +import urllib.request +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import story_site # noqa: E402 (the bucket, HTTP and master data helpers) +from story_site import env, get, output, summary # noqa: E402 + +FORMAT = "nnnotes.music-data/1" +BUILD_FORMAT = "moenotes.music-data-build/1" +# This script's own version of a build: bump it when what it builds or publishes changes, so that the next run builds +# although the master data, the deck model and nnnotes are the same. +RECIPE = 1 +FILE, MARKER, JACKETS, ARCHIVE = "music-data.json", "build.json", "jackets/", "archive/" +MANIFEST = "MasterManifest.json" +SOURCE_PATHS = ("src", "rust", "pyproject.toml") # nnnotes' code: the commit that last changed one of them +SCHEMA = Path("docs/schema/music-data.schema.json") +SMOKE = Path(__file__).resolve().parent / "music_data_smoke.mjs" +ARCHIVE_CACHE = story_site.ASSET_CACHE # archive//.json: content-addressed +FILE_CACHE = "no-cache" # music-data.json and build.json change in place +JACKET_CACHE = "public, max-age=86400" +SNAPSHOT_KEYS = ("version", "resource_version", "client_version", "verified_at", "manifest_sha256", "table_count") + +# the gates' bounds (MUSIC_DATA.md) +SIZE_RATIO = (0.8, 2.0) # against the published file +BGM_MS = (30_000, 600_000) # a song's BGM length +BGM_CUE_SLACK_MS = 1000 # |durationMs - lengthMs| +BGM_TAIL_MS = 60_000 # BGM after the last note: more is reported +RANKS = 5 +DIFFICULTIES = ("easy", "normal", "hard", "expert") +SONG_TABLES = ("MasterLiveMusic", "MasterLiveMusicScore", "MasterText", "MasterBand", "MasterCharacter", "MasterTag", + "MasterLiveMusicCategory", "MasterSound", "MasterSoundCueSheet", "MasterLiveScoreRank") +SHA256 = re.compile(r"[0-9a-f]{64}") +COMMIT = re.compile(r"[0-9a-f]{40}") +LISTED = 20 # failures and warnings listed per gate + + +def fail(message: str): + sys.exit(f"music_data: {message}") + + +def sha256(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +# ---------------------------------------------------------------- bucket +def s3_prefix() -> str: + p = (os.environ.get("MUSIC_DATA_S3_PREFIX") or "music-data").strip("/") + return f"{p}/" if p else "" + + +def bucket() -> "story_site.Bucket": + b = story_site.Bucket() + b.prefix = s3_prefix() + return b + + +def public_url(key: str) -> str: + return f"{env('STORY_S3_ENDPOINT').rstrip('/')}/{env('STORY_S3_BUCKET')}/{s3_prefix()}{key}" + + +def published(key: str) -> bytes | None: + """A published file (plain HTTP), None when the bucket has none.""" + try: + return get(public_url(key), timeout=300) + except urllib.error.HTTPError as e: + if e.code == 404: + return None + raise + + +def exists(key: str) -> bool: + request = urllib.request.Request(public_url(key), method="HEAD", + headers={"User-Agent": "moenotes-music-data (GitHub Actions)", + "Cache-Control": "no-cache"}) + try: + with urllib.request.urlopen(request, timeout=60): + return True + except urllib.error.HTTPError as e: + if e.code == 404: + return False + raise + + +# ---------------------------------------------------------------- the inputs of a build +def deck_commit(root: Path = Path(".")) -> str: + """The ournotes-deck commit nnnotes builds its deck model with (rust/Cargo.lock).""" + lock = root / "rust" / "Cargo.lock" + if not lock.is_file(): + fail("no rust/Cargo.lock: this nnnotes has no music-data command with the deck model (sync the fork with " + "upstream, .github/MUSIC_DATA.md)") + for block in lock.read_text(encoding="utf-8").split("[[package]]"): + if re.search(r'^name = "ournotes-deck"$', block, re.M): + m = re.search(r'^source = "git\+[^"#]*#([0-9a-f]{40})"$', block, re.M) + if m: + return m.group(1) + fail("rust/Cargo.lock has no ournotes-deck git commit") + + +def nnnotes_commit(root: Path = Path(".")) -> str: + """The commit that last changed nnnotes' code (SOURCE_PATHS); the checkout needs its history.""" + def git(*args): + return subprocess.run(["git", "-C", str(root), *args], capture_output=True, text=True) + if git("rev-parse", "--is-shallow-repository").stdout.strip() != "false": + fail("the checkout is shallow: the nnnotes commit needs the history (actions/checkout fetch-depth: 0)") + commit = git("log", "-1", "--format=%H", "--", *SOURCE_PATHS).stdout.strip() + if not COMMIT.fullmatch(commit): + fail("no commit changed src/, rust/ or pyproject.toml") + return commit + + +def require_decoded_master(root: Path = Path(".")) -> None: + cli = root / "src" / "nnnotes" / "cli.py" + if "--decoded-master" not in cli.read_text(encoding="utf-8"): + fail("this nnnotes has no `music-data --decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with " + "upstream (.github/MUSIC_DATA.md)") + + +def inputs(entry: dict, root: Path = Path(".")) -> dict: + """What a build is made of: the master data snapshot, the deck model, nnnotes and this script.""" + return {"masterRegion": env("MASTERDATA_REGION"), "masterVersion": entry.get("version"), + "resourceVersion": entry.get("resource_version"), "clientVersion": entry.get("client_version"), + "deckCommit": deck_commit(root), "nnnotesCommit": nnnotes_commit(root), "recipe": RECIPE} + + +def short(v) -> str: + return str(v)[:12] if v is not None else "?" + + +# ---------------------------------------------------------------- plan +def cmd_plan() -> None: + _, region = story_site.master_index() + require_decoded_master() + now = inputs(region.get("entry") or {}) + raw = published(MARKER) + marker = json.loads(raw) if raw else None + have = marker.get("inputs") if isinstance(marker, dict) else None + changed = [k for k in now if not isinstance(have, dict) or have.get(k) != now[k]] + if have and not exists(FILE): + changed.append("music-data.json (missing)") + force = os.environ.get("FORCE") == "true" + build = force or bool(changed) + summary(f"### Music data\n\n- master: {now['masterVersion']} of {now['masterRegion']} (resource " + f"{now['resourceVersion']}, client {now['clientVersion']}); deck {short(now['deckCommit'])}; nnnotes " + f"{short(now['nnnotesCommit'])}; recipe {RECIPE}\n- published: " + + (f"master {have.get('masterVersion')}, deck {short(have.get('deckCommit'))}, nnnotes " + f"{short(have.get('nnnotesCommit'))}, built {marker.get('builtAt')}" if isinstance(have, dict) + else "nothing") + + "\n- this run: " + ("builds" + (" (force)" if force else "") + (f", changed: {', '.join(changed)}" + if changed and have else "") + if build else "nothing changed, nothing to build")) + output("build", "true" if build else "false") + + +# ---------------------------------------------------------------- master data +def snapshot_file(master: Path) -> Path: + return master.parent / f"{master.name}.snapshot.json" + + +def cmd_master(out: str) -> None: + url, region = story_site.master_index() + files = region["files"] + entry = region.get("entry") or {} + if MANIFEST not in files: + fail(f"index.json lists no {MANIFEST} for {env('MASTERDATA_REGION')}") + if entry.get("manifest_sha256") and entry["manifest_sha256"] != files[MANIFEST]: + fail(f"index.json: the entry's manifest_sha256 is not the SHA-256 of its {MANIFEST}") + d = Path(out) + d.mkdir(parents=True, exist_ok=True) + story_site.parallel(lambda item: story_site.fetch_table(url, item[0], item[1], d / item[0]), files.items(), 8) + version = json.loads((d / MANIFEST).read_bytes()).get("version") + if entry.get("version") is not None and str(version) != str(entry["version"]): + fail(f"{MANIFEST} has version {version}, index.json {entry['version']} (the service changed snapshots? " + f"run again)") + snapshot = {"region": env("MASTERDATA_REGION"), "entry": {k: entry.get(k) for k in SNAPSHOT_KEYS}, + "files": files} + snapshot_file(d).write_text(json.dumps(snapshot, indent=1, sort_keys=True), encoding="utf-8") + print(f"master data: {version}, {len(files)} files into {d}") + + +# ---------------------------------------------------------------- build +def cmd_build(out: str) -> None: + o = Path(out).resolve() + o.mkdir(parents=True, exist_ok=True) + nnnotes = [sys.executable, "-m", "nnnotes"] + usage = subprocess.run(nnnotes + ["music-data", "--help"], capture_output=True, text=True).stdout + if "--decoded-master" not in usage: + fail("the installed nnnotes has no `music-data --decoded-master` (sync the fork with upstream)") + cmd = nnnotes + ["music-data", "--decoded-master", "--jackets", str(o / "jackets"), "-o", str(o / FILE)] + print("+ " + " ".join(cmd[1:]), flush=True) + with open(o / "music-data.summary.json", "wb") as f: + status = subprocess.run(cmd, stdout=f).returncode + if status: + fail(f"nnnotes music-data exited with {status}") + r = json.loads((o / "music-data.summary.json").read_text(encoding="utf-8")) + summary(f"- built: {r.get('songs')} songs, {r.get('charts')} charts, {r.get('jackets')} jackets, deck " + f"{short(r.get('deck'))}, {r.get('bytes')} bytes, sha256 {short(r.get('sha256'))}") + + +# ---------------------------------------------------------------- the gates +@dataclass +class Context: + """What the gates check the file against; None: the gate (or that part of it) is skipped (the self-test).""" + region: str | None = None # [catalog] region: provenance.region + language: str | None = None # [catalog] language: a song title should have it + master: Path | None = None # the decoded master data read (MasterManifest.json, .json) + snapshot: dict | None = None # its index entry and files (the master step's OUT.snapshot.json) + deck_commit: str | None = None + nnnotes_version: str | None = None + jackets: Path | None = None + schema: Path | None = None + published: bytes | None = None # the published music-data.json (None: nothing is published) + page: Path | None = None # examples/songs of ournotes-player + file: Path | None = None # the file on disk, for the smoke test + + +class Gate: + def __init__(self): + self.failures: list[str] = [] + self.warnings: list[str] = [] + self.note = "" + + def fail(self, message: str): + self.failures.append(message) + + def warn(self, message: str): + self.warnings.append(message) + + +def charts_of(doc: dict): + for song in doc.get("songs") or []: + for chart in song.get("charts") or []: + yield song, chart + + +def where(song: dict, chart: dict) -> str: + return f"chart {chart.get('scoreId')} ({song.get('id')} {chart.get('difficulty')})" + + +def is_int(v) -> bool: + return isinstance(v, int) and not isinstance(v, bool) + + +def is_num(v) -> bool: + return (isinstance(v, (int, float)) and not isinstance(v, bool)) and math.isfinite(v) + + +def within(c) -> bool: + return (isinstance(c, dict) and is_int(c.get("exact")) and is_num(c.get("predicted")) and is_num(c.get("bound")) + and abs(c["exact"] - c["predicted"]) <= c["bound"]) + + +def gate_schema(doc, ctx: Context, g: Gate): + if ctx.schema is None: + g.note = "skipped" + return + if not ctx.schema.is_file(): + g.fail(f"no JSON Schema {ctx.schema}") + return + import jsonschema + schema = json.loads(ctx.schema.read_text(encoding="utf-8")) + validator = jsonschema.Draft202012Validator(schema) + for e in sorted(validator.iter_errors(doc), key=lambda e: list(map(str, e.absolute_path))): + g.fail(f"{'/'.join(map(str, e.absolute_path)) or '(root)'}: {e.message[:160]}") + g.note = ctx.schema.as_posix() + + +def gate_provenance(doc, ctx: Context, g: Gate): + if doc.get("format") != FORMAT: + g.fail(f"format {doc.get('format')!r}, expected {FORMAT}") + p = doc.get("provenance") or {} + if ctx.region is not None and p.get("region") != ctx.region: + g.fail(f"region {p.get('region')!r}, expected {ctx.region!r}") + m = p.get("master") or {} + if m.get("source") != "api": + g.fail(f"master.source {m.get('source')!r}, expected 'api'") + tables = m.get("tables") or {} + missing = [t for t in SONG_TABLES if t not in tables] + if missing: + g.fail(f"master.tables lacks {', '.join(missing)}") + if doc.get("deck") is not None and len(tables) <= len(SONG_TABLES): + g.fail("master.tables has only the song tables, but the file has deck statistics") + if ctx.snapshot is not None: + v = ctx.snapshot["entry"].get("version") + if m.get("version") != v: + g.fail(f"master.version {m.get('version')!r}, the snapshot's {v!r}") + if ctx.master is not None: + manifest = json.loads((ctx.master / MANIFEST).read_bytes()) + if m.get("version") != manifest.get("version"): + g.fail(f"master.version {m.get('version')!r}, {MANIFEST}'s {manifest.get('version')!r}") + listed = {f.get("name"): str(f.get("hash") or "").lower() for f in manifest.get("files") or []} + files = (ctx.snapshot or {}).get("files") or {} + for t, v in sorted(tables.items()): + if (v or {}).get("sha256") != listed.get(f"{t}.bin"): + g.fail(f"master.tables.{t}.sha256 is not {MANIFEST}'s {t}.bin") + decoded = ctx.master / f"{t}.json" + if not decoded.is_file(): + g.fail(f"no decoded table {t}.json") + elif files and sha256(decoded.read_bytes()) != files.get(f"{t}.json"): + g.fail(f"{t}.json read is not the one index.json lists") + deck = p.get("deck") + if doc.get("deck") is not None and not isinstance(deck, dict): + g.fail("provenance.deck is null, but the file has deck statistics") + if isinstance(deck, dict): + if deck.get("name") != "ournotes-deck" or not COMMIT.fullmatch(str(deck.get("commit"))): + g.fail("provenance.deck names no ournotes-deck commit") + elif ctx.deck_commit is not None and deck["commit"] != ctx.deck_commit: + g.fail(f"deck commit {deck['commit'][:12]}, rust/Cargo.lock pins {ctx.deck_commit[:12]}") + ex = p.get("exporter") or {} + if ex.get("name") != "nnnotes": + g.fail(f"exporter {ex.get('name')!r}") + elif ctx.nnnotes_version is not None and ex.get("version") != ctx.nnnotes_version: + g.fail(f"exporter version {ex.get('version')!r}, installed nnnotes {ctx.nnnotes_version!r}") + client = p.get("client") or {} + if not isinstance(client.get("versionName"), str) or not is_int(client.get("versionCode")): + g.fail("provenance.client has no APK version (was the APK read?)") + elif ctx.snapshot is not None and ctx.snapshot["entry"].get("client_version") not in (None, + client["versionName"]): + g.warn(f"the APK is {client['versionName']}, the snapshot names client " + f"{ctx.snapshot['entry']['client_version']}") + if not SHA256.fullmatch(str((p.get("catalog") or {}).get("sha256"))): + g.fail("provenance.catalog has no SHA-256") + g.note = f"master {m.get('version')}, deck {short((deck or {}).get('commit'))}, nnnotes {ex.get('version')}" + + +def baseline(ctx: Context, g: Gate): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return None + try: + return json.loads(ctx.published) + except ValueError: + g.warn("the published music-data.json is not JSON: skipped") + return None + + +def gate_counts(doc, ctx: Context, g: Gate): + old = baseline(ctx, g) + if old is None: + return + songs = {s.get("id") for s in doc.get("songs") or []} + charts = {c.get("scoreId") for _, c in charts_of(doc)} + old_songs = {s.get("id") for s in old.get("songs") or []} + old_charts = {c.get("scoreId") for _, c in charts_of(old)} + if len(songs) < len(old_songs): + g.fail(f"{len(songs)} songs, the published file has {len(old_songs)}") + if len(charts) < len(old_charts): + g.fail(f"{len(charts)} charts, the published file has {len(old_charts)}") + if old_songs - songs: + g.warn(f"songs no longer in the file: {', '.join(map(str, sorted(old_songs - songs)))}") + if old_charts - charts: + g.warn(f"charts no longer in the file: {', '.join(map(str, sorted(old_charts - charts)))}") + g.note = f"songs {len(old_songs)} -> {len(songs)}, charts {len(old_charts)} -> {len(charts)}" + + +def gate_size(doc, ctx: Context, g: Gate, raw: bytes): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return + ratio = len(raw) / max(1, len(ctx.published)) + if not SIZE_RATIO[0] <= ratio <= SIZE_RATIO[1]: + g.fail(f"{len(raw)} bytes, {ratio:.2f} times the published {len(ctx.published)} (bounds {SIZE_RATIO[0]} to " + f"{SIZE_RATIO[1]})") + g.note = f"{len(ctx.published)} -> {len(raw)} bytes ({ratio:.2f})" + + +def weights_shape(w, kinds: int, positions: int, nullable: bool) -> bool: + return isinstance(w, list) and len(w) == kinds and all( + (nullable and k is None) or (isinstance(k, list) and len(k) == positions and all(map(is_num, k))) for k in w) + + +def gate_deck(doc, ctx: Context, g: Gate): + deck = doc.get("deck") + if not isinstance(deck, dict): + g.fail("deck is null: no deck statistics (made with --no-deck?)") + return + kinds = len(deck.get("kinds") or []) + if not kinds: + g.fail("deck.kinds is empty") + power = (deck.get("model") or {}).get("power") + if not (is_num(power) and power > 0): + g.fail(f"deck.model.power {power!r}") + n = unplayable = seeds = 0 + for song, chart in charts_of(doc): + n += 1 + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, events = d.get("positions"), d.get("events") or [] + if len(events) != len(chart.get("skillEventsMs") or []): + g.fail(f"{w}: {len(events)} skill events, the chart has {len(chart.get('skillEventsMs') or [])}") + if not is_int(positions) or (events and positions != max(e[0] for e in events) + 1): + g.fail(f"{w}: positions {positions!r} do not match the events") + continue + if d.get("unplayable"): + unplayable += 1 + g.warn(f"{w}: unplayable with Gekisou on ({d['unplayable']})") + if d.get("seeds"): + g.fail(f"{w}: unplayable, but has Gekisou on seeds") + elif not d.get("seeds"): + g.fail(f"{w}: no seeds") + for seed in d.get("seeds") or []: + seeds += 1 + s = f"{w} seed {seed.get('seed')}" + if not is_int(seed.get("score")): + g.fail(f"{s}: score {seed.get('score')!r}") + if not weights_shape(seed.get("weights"), kinds, positions, nullable=False): + g.fail(f"{s}: weights are not [kind][position] numbers") + if len(seed.get("ranges") or []) != len(d.get("ranges") or []): + g.fail(f"{s}: {len(seed.get('ranges') or [])} range results for {len(d.get('ranges') or [])} ranges") + if not within(seed.get("check")): + g.fail(f"{s}: the check deck is not within its bound") + g.note = f"{n} charts, {kinds} kinds, {seeds} seeds, {unplayable} unplayable" + + +def gate_scenarios(doc, ctx: Context, g: Gate): + """The play scenario fields (Gekisou off, every rank, the Perfect play): offSeeds exactly one, every range's + rankBonusPercents five ints, every seed scorePerfect, rangeWeights and rankCheck, every seed range + rangeScorePerfect. A null rangeWeights, a null kind in it or in offSeeds' weights is a warning (none in TW).""" + kinds = len(((doc.get("deck") or {}).get("kinds")) or []) + null_rw = null_kind = null_off = 0 + for song, chart in charts_of(doc): + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, ranges = d.get("positions"), d.get("ranges") or [] + off = d.get("offSeeds") + if not isinstance(off, list) or len(off) != 1: + g.fail(f"{w}: offSeeds {'missing' if off is None else f'has {len(off)} entries'}, expected exactly one") + else: + o = off[0] + if o.get("seed") != 0 or not is_int(o.get("score")): + g.fail(f"{w}: Gekisou off seed {o.get('seed')!r} score {o.get('score')!r}") + if not weights_shape(o.get("weights"), kinds, positions, nullable=True): + g.fail(f"{w}: Gekisou off weights are not [kind][position] numbers") + else: + null_off += sum(k is None for k in o["weights"]) + if not within(o.get("check")): + g.fail(f"{w}: the Gekisou off check deck is not within its bound") + for i, r in enumerate(ranges): + p = r.get("rankBonusPercents") + if not (isinstance(p, list) and len(p) == RANKS and all(map(is_int, p))): + g.fail(f"{w} range {i}: rankBonusPercents {'missing' if p is None else 'not five ints'}") + elif p[0] != r.get("rankBonusPercent"): + g.fail(f"{w} range {i}: rankBonusPercents[0] {p[0]} is not rankBonusPercent " + f"{r.get('rankBonusPercent')}") + for seed in d.get("seeds") or []: + s = f"{w} seed {seed.get('seed')}" + absent = [k for k in ("scorePerfect", "rangeWeights", "rankCheck") if k not in seed] + if absent: + g.fail(f"{s}: no {', '.join(absent)}") + continue + if not is_int(seed["scorePerfect"]): + g.fail(f"{s}: scorePerfect {seed['scorePerfect']!r}") + for i, r in enumerate(seed.get("ranges") or []): + if not is_int(r.get("rangeScorePerfect")): + why = "missing" if "rangeScorePerfect" not in r else "not an int" + g.fail(f"{s} range {i}: rangeScorePerfect {why}") + rw = seed["rangeWeights"] + if rw is None: + null_rw += 1 + elif not (isinstance(rw, list) and len(rw) == kinds and all( + k is None or (isinstance(k, list) and len(k) == positions and all( + isinstance(x, list) and len(x) == len(ranges) and all(map(is_num, x)) for x in k)) + for k in rw)): + g.fail(f"{s}: rangeWeights are not [kind][position][range] numbers") + else: + null_kind += sum(k is None for k in rw) + rc = seed["rankCheck"] + if rc is not None: + if len(rc.get("ranks") or []) != len(ranges) or not all( + is_int(x) and 1 <= x <= RANKS for x in rc.get("ranks") or []): + g.fail(f"{s}: rankCheck ranks are not one rank per range") + elif not within(rc): + g.fail(f"{s}: the rank check deck is not within its bound") + if null_rw: + g.warn(f"{null_rw} seeds have rangeWeights null (overlapping ranges: no rank scenarios on those charts)") + if null_kind: + g.warn(f"{null_kind} rangeWeights kinds are null (conditions on the confirmed rank)") + if null_off: + g.warn(f"{null_off} Gekisou off weight kinds are null (conditions on the Gekisou state)") + g.note = "offSeeds, rankBonusPercents, scorePerfect, rangeWeights, rankCheck, rangeScorePerfect" + + +def nonfinite(v, path: str, out: list): + if isinstance(v, float): + if not math.isfinite(v): + out.append(path) + elif isinstance(v, list): + for i, x in enumerate(v): + nonfinite(x, f"{path}[{i}]", out) + elif isinstance(v, dict): + for k, x in v.items(): + nonfinite(x, f"{path}.{k}" if path else k, out) + + +def gate_finite(doc, ctx: Context, g: Gate): + """No NaN or infinity; one inside master data rows as served (songs[].master, master: `1e999` is how the format + writes a binary32 infinity) is reported, not failed.""" + found: list[str] = [] + nonfinite(doc, "", found) + for path in found: + if re.match(r"(songs\[\d+\]\.master|master)\b", path): + g.warn(f"{path}: not finite (master data as served)") + else: + g.fail(f"{path}: not finite") + g.note = f"{len(found)} non-finite numbers" + + +def gate_references(doc, ctx: Context, g: Gate): + langs = doc.get("languages") or [] + if not langs: + g.fail("no languages") + + def text(t, what: str, required: bool) -> bool: + if t is None: + if required: + g.fail(f"{what}: no text") + return False + if not isinstance(t, dict) or set(t) != set(langs) or not all(isinstance(v, str) for v in t.values()): + g.fail(f"{what}: not a text in {', '.join(langs)}") + return False + if required and not any(t.values()): + g.fail(f"{what}: empty in every language") + return False + return True + + def ids(rows, what: str) -> set: + seen = [r.get("id") for r in rows] + if len(set(seen)) != len(seen) or not all(map(is_int, seen)): + g.fail(f"{what}: ids are not unique ints") + return set(seen) + + bands = ids(doc.get("bands") or [], "bands") + characters = ids(doc.get("characters") or [], "characters") + tags = ids(doc.get("tags") or [], "tags") + ids(doc.get("categories") or [], "categories") + for b in doc.get("bands") or []: + text(b.get("name"), f"band {b.get('id')} name", True) + for c in doc.get("characters") or []: + text(c.get("name"), f"character {c.get('id')} name", True) + text(c.get("shortName"), f"character {c.get('id')} short name", False) + if c.get("bandId") not in bands: + g.warn(f"character {c.get('id')}: band {c.get('bandId')} is not in bands") + for t in doc.get("tags") or []: + text(t.get("name"), f"tag {t.get('id')} name", True) + categories = set() + for c in doc.get("categories") or []: + text(c.get("name"), f"category {c.get('id')} name", True) + categories.update(c.get("musicCategories") or []) + songs = doc.get("songs") or [] + if not songs: + g.fail("no songs") + ids(songs, "songs") + if [s.get("id") for s in songs] != sorted(s.get("id") for s in songs if is_int(s.get("id"))): + g.fail("songs are not sorted by id") + score_ids: set = set() + untitled = [] + for s in songs: + sid = f"song {s.get('id')}" + if text(s.get("title"), f"{sid} title", True) and ctx.language and not s["title"].get(ctx.language): + untitled.append(s.get("id")) + for k in ("ruby", "phonetic", "bandName", "lyricist", "composer", "arranger"): + text(s.get(k), f"{sid} {k}", False) + for key, known, what in (("bandIds", bands, "band"), ("vocalCharacterIds", characters, "character"), + ("bestMusicTagIds", tags, "tag")): + unknown = [x for x in s.get(key) or [] if x not in known] + if unknown: + g.fail(f"{sid}: {what} {', '.join(map(str, unknown))} not in the file's {what}s") + if not s.get("bandIds") and s.get("bandName") is None: + g.fail(f"{sid}: neither a band nor a band name") + other = [x for x in s.get("musicCategories") or [] if x not in categories] + if other: + g.warn(f"{sid}: music categories {', '.join(map(str, other))} are on no category tab") + jacket = s.get("jacket") + if not isinstance(jacket, str) or not jacket: + g.fail(f"{sid}: no jacket") + elif ctx.jackets is not None: + f = ctx.jackets / f"{jacket}.webp" + if not f.is_file() or not f.stat().st_size: + g.fail(f"{sid}: no jacket file jackets/{jacket}.webp") + bgm = s.get("bgm") or {} + if not (is_int(bgm.get("soundId")) and bgm.get("cueSheet") and bgm.get("cue")): + g.fail(f"{sid}: no BGM cue") + ranks = [r.get("rank") for r in s.get("scoreRanks") or []] + if not ranks or len(set(ranks)) != len(ranks): + g.fail(f"{sid}: score ranks {ranks!r}") + charts = s.get("charts") or [] + order = [c.get("difficulty") for c in charts] + if not charts or order != [d for d in DIFFICULTIES if d in order] or len(set(order)) != len(order): + g.fail(f"{sid}: charts {order!r}") + for c in charts: + if c.get("scoreId") in score_ids: + g.fail(f"{where(s, c)}: the score id occurs twice") + score_ids.add(c.get("scoreId")) + if untitled: + g.warn(f"songs without a {ctx.language} title: {', '.join(map(str, untitled))}") + g.note = (f"{len(songs)} songs, {len(bands)} bands, {len(characters)} characters, {len(tags)} tags" + + ("" if ctx.jackets is None else ", jackets present")) + + +def gate_bgm(doc, ctx: Context, g: Gate): + lengths = [] + for s in doc.get("songs") or []: + sid = f"song {s.get('id')}" + L = (s.get("bgm") or {}).get("length") + if not isinstance(L, dict): + g.fail(f"{sid}: no BGM length") + continue + dur, samples, rate = L.get("durationMs"), L.get("samples"), L.get("sampleRate") + if not (is_int(dur) and is_int(samples) and is_int(rate) and rate > 0 and samples > 0): + g.fail(f"{sid}: BGM length {dur!r} ms, {samples!r} samples at {rate!r} Hz") + continue + if dur != samples * 1000 // rate: + g.fail(f"{sid}: BGM durationMs {dur} is not samples * 1000 // sampleRate") + if not BGM_MS[0] <= dur <= BGM_MS[1]: + g.fail(f"{sid}: BGM of {dur} ms (bounds {BGM_MS[0]} to {BGM_MS[1]})") + if is_int(L.get("lengthMs")) and abs(dur - L["lengthMs"]) > BGM_CUE_SLACK_MS: + g.fail(f"{sid}: BGM stream {dur} ms, cue length {L['lengthMs']} ms") + last = max((c.get("lastNoteMs") or 0 for c in s.get("charts") or []), default=0) + if dur < last: + g.fail(f"{sid}: BGM of {dur} ms ends before the last note at {last} ms") + elif dur - last > BGM_TAIL_MS: + g.warn(f"{sid}: BGM plays {dur - last} ms after the last note") + lengths.append(dur) + if lengths: + g.note = f"{min(lengths) / 1000:.1f} s to {max(lengths) / 1000:.1f} s" + + +def gate_page(doc, ctx: Context, g: Gate): + """The chart data page's own modules (catalog.js, ranking.js) over the file in Node.js: music_data_smoke.mjs.""" + if ctx.page is None: + g.note = "skipped" + return + if not (ctx.page / "catalog.js").is_file() or not (ctx.page / "ranking.js").is_file(): + g.fail(f"{ctx.page}: no catalog.js / ranking.js (the chart data page, examples/songs)") + return + r = subprocess.run(["node", str(SMOKE), str(ctx.page), str(ctx.file)], capture_output=True, text=True, + timeout=600) + lines = [x for x in (r.stdout + r.stderr).splitlines() if x.strip()] + if r.returncode: + for x in lines or [f"node exited with {r.returncode}"]: + g.fail(x[:300]) + else: + g.note = lines[-1][:200] if lines else "passed" + + +GATES = (("schema", gate_schema), ("provenance", gate_provenance), ("counts", gate_counts), ("deck", gate_deck), + ("scenarios", gate_scenarios), ("finite", gate_finite), ("references", gate_references), ("bgm", gate_bgm), + ("size", gate_size), ("page", gate_page)) + + +def gates(raw: bytes, ctx: Context, only=None) -> dict: + """The report of the gates (`only`: their names) on the file's bytes.""" + try: + doc = json.loads(raw.decode("utf-8")) + if not isinstance(doc, dict): + raise ValueError("not an object") + except ValueError as e: + doc, results = None, [{"gate": "json", "passed": False, "failures": [f"not JSON: {str(e)[:160]}"], + "failureCount": 1, "warnings": [], "warningCount": 0, "note": ""}] + if doc is not None: + results = [] + for name, fn in GATES: + if only is not None and name not in only: + continue + g = Gate() + try: + fn(doc, ctx, g, raw) if name == "size" else fn(doc, ctx, g) + except Exception as e: # a malformed file the gate did not foresee fails the gate + g.fail(f"the gate stopped: {type(e).__name__}: {str(e)[:160]}") + results.append({"gate": name, "passed": not g.failures, "failures": g.failures[:LISTED], + "failureCount": len(g.failures), "warnings": g.warnings[:LISTED], + "warningCount": len(g.warnings), "note": g.note}) + return {"passed": all(r["passed"] for r in results), "sha256": sha256(raw), "bytes": len(raw), + "gates": results} + + +def report_markdown(report: dict) -> str: + rows = ["| Gate | Result | |", "|---|---|---|"] + for r in report["gates"]: + result = "passed" if r["passed"] else f"**failed** ({r['failureCount']})" + if r["warningCount"]: + result += f", {r['warningCount']} warnings" + rows.append(f"| {r['gate']} | {result} | {r['note']} |") + details = [] + for r in report["gates"]: + for kind, key, n in (("failure", "failures", "failureCount"), ("warning", "warnings", "warningCount")): + for x in r[key]: + details.append(f"- {r['gate']} {kind}: {x}") + if r[n] > len(r[key]): + details.append(f"- {r['gate']}: {r[n] - len(r[key])} more {kind}s") + return "\n".join(["", "#### Gates", ""] + rows + ([""] + details if details else [])) + + +def cmd_check(out: str, master: str, page: str) -> None: + import nnnotes + o, m = Path(out), Path(master) + raw = (o / FILE).read_bytes() + snapshot = json.loads(snapshot_file(m).read_text(encoding="utf-8")) + ctx = Context(region=env("NNNOTES_CATALOG_REGION"), language=os.environ.get("NNNOTES_CATALOG_LANGUAGE"), + master=m, snapshot=snapshot, deck_commit=deck_commit(), nnnotes_version=nnnotes.__version__, + jackets=o / "jackets", schema=SCHEMA, published=published(FILE), page=Path(page), file=o / FILE) + report = gates(raw, ctx) + (o / "check.json").write_text(json.dumps(report, indent=1), encoding="utf-8") + summary(report_markdown(report)) + if not report["passed"]: + fail("a gate failed: nothing is published (the job summary lists why)") + doc = json.loads(raw) + p = doc["provenance"] + version = re.sub(r"[^A-Za-z0-9._-]", "_", str(p["master"]["version"])) + marker = { + "format": BUILD_FORMAT, + "file": FILE, "sha256": report["sha256"], "bytes": report["bytes"], + "archive": f"{ARCHIVE}{version}/{report['sha256']}.json", + "songs": len(doc["songs"]), "charts": sum(len(s["charts"]) for s in doc["songs"]), + "jackets": len(list((o / "jackets").glob("*.webp"))), + "inputs": inputs(snapshot["entry"]), + "provenance": {"region": p["region"], "client": p["client"], "masterVersion": p["master"]["version"], + "deckCommit": p["deck"]["commit"], "exporterVersion": p["exporter"]["version"]}, + "player": {"repository": os.environ.get("PLAYER_REPOSITORY"), "ref": os.environ.get("PLAYER_REF")}, + "gates": {r["gate"]: {"passed": r["passed"], "warnings": r["warningCount"]} for r in report["gates"]}, + "builtAt": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"), + "run": (f"{os.environ['GITHUB_SERVER_URL']}/{os.environ['GITHUB_REPOSITORY']}/actions/runs/" + f"{os.environ['GITHUB_RUN_ID']}" if os.environ.get("GITHUB_RUN_ID") else None), + } + (o / MARKER).write_text(json.dumps(marker, indent=1) + "\n", encoding="utf-8") + + +# ---------------------------------------------------------------- publish +def upload(b, key: str, src: Path, cache: str) -> None: + extra = {"ContentType": story_site.TYPES.get(src.suffix.lower(), "application/octet-stream"), + "CacheControl": cache} + b.s3.upload_file(str(src), b.name, b.prefix + key, ExtraArgs=extra) + + +def read_back(b, key: str, digest: str) -> None: + """The object as stored (S3 GET), else as served (plain HTTP), must have the SHA-256 uploaded.""" + for attempt in range(4): + try: + if sha256(b.s3.get_object(Bucket=b.name, Key=b.prefix + key)["Body"].read()) == digest: + return + except Exception as e: # Cloudflare in front of the store rejects a signed GET at times + print(f"read back {key}: {type(e).__name__}", flush=True) + time.sleep(3 * (attempt + 1)) + try: + if sha256(get(public_url(key), timeout=300)) == digest: + return + except urllib.error.URLError as e: + print(f"read back {key} over HTTP: {e}", flush=True) + fail(f"{key}: the bucket does not serve what was uploaded (SHA-256 {digest[:12]})") + + +def cmd_publish(out: str, dry_run: bool = False) -> None: + o = Path(out) + raw = (o / FILE).read_bytes() + report = json.loads((o / "check.json").read_text(encoding="utf-8")) + if not report.get("passed") or report.get("sha256") != sha256(raw) or not (o / MARKER).is_file(): + fail("the file has not passed the gates (check.json): nothing is published") + marker = json.loads((o / MARKER).read_text(encoding="utf-8")) + b = bucket() + if not b.writable and not dry_run: + fail("publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") + force = os.environ.get("FORCE") == "true" + have = b.keys(JACKETS) + jackets = sorted((o / "jackets").glob("*.webp")) + new = [j for j in jackets if force or have.get(JACKETS + j.name) != j.stat().st_size] + archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) + steps = [(JACKETS + j.name, j, JACKET_CACHE, False) for j in new] + if not archived: + steps.append((marker["archive"], o / FILE, ARCHIVE_CACHE, True)) + # the file before the marker that names it, the jackets and the archive copy before the file + steps += [(FILE, o / FILE, FILE_CACHE, True), (MARKER, o / MARKER, FILE_CACHE, True)] + if dry_run: + for key, _, _, _ in steps: + print(f"would upload {b.prefix}{key}") + else: + story_site.parallel(lambda s: upload(b, s[0], s[1], s[2]), [s for s in steps if not s[3]]) + for key, src, cache, check in steps: + if check: + upload(b, key, src, cache) + read_back(b, key, sha256(src.read_bytes())) + summary(f"- {'would publish' if dry_run else 'published'} {len(new)} jackets (of {len(jackets)}), " + f"{'the archive copy, ' if not archived else ''}{FILE} ({len(raw)} bytes, sha256 " + f"{short(marker['sha256'])}), {MARKER}: {public_url(FILE)}") + + +def main(argv: list[str]) -> None: + if not argv: + sys.exit(__doc__) + cmd, args = argv[0], argv[1:] + if cmd == "plan" and not args: + cmd_plan() + elif cmd == "master" and len(args) == 1: + cmd_master(args[0]) + elif cmd == "build" and len(args) == 1: + cmd_build(args[0]) + elif cmd == "check" and len(args) == 3: + cmd_check(*args) + elif cmd == "publish" and len(args) in (1, 2) and args[1:] in ([], ["--dry-run"]): + cmd_publish(args[0], dry_run=bool(args[1:])) + else: + sys.exit(__doc__) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/.github/scripts/music_data_smoke.mjs b/.github/scripts/music_data_smoke.mjs new file mode 100644 index 0000000..a711edc --- /dev/null +++ b/.github/scripts/music_data_smoke.mjs @@ -0,0 +1,94 @@ +// The chart data page's smoke test (music_data.py check, gate "page"): the page's own pure modules (ournotes-player +// examples/songs: catalog.js, ranking.js) over a music-data.json in Node.js, the way the page reads it. Every chart +// must get a row, every playable chart its figures in every play scenario (Gekisou Live at several ranks and Just +// rates, Free Live, a Great share), finite and positive, and the rankings must work on them. +// +// node music_data_smoke.mjs +// +// Prints one line per problem (at most 40) and exits 1, or a one-line summary. +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { pathToFileURL } from "node:url"; + +const [dir, file] = process.argv.slice(2); +if (!dir || !file) { + console.error("usage: node music_data_smoke.mjs "); + process.exit(2); +} +const catalog = await import(pathToFileURL(path.join(dir, "catalog.js")).href); +const ranking = await import(pathToFileURL(path.join(dir, "ranking.js")).href); +const data = JSON.parse(readFileSync(file, "utf8")); + +const problems = []; +const problem = (m) => problems.push(m); +const finite = (v) => typeof v === "number" && Number.isFinite(v); + +const charts = (data.songs || []).flatMap((s) => (s.charts || []).map((c) => ({ song: s, chart: c }))); +const rows = catalog.chartRows(data); +if (rows.length !== charts.length) problem(`chartRows: ${rows.length} rows for ${charts.length} charts`); +const kind = ranking.plainKind(data); +if (kind === null) problem("plainKind: the file has no plain score-up kind (effect 2000, 5 s, no targets)"); + +// Whether the data has what a scenario needs on a chart (a null rangeWeights or kind leaves a chart without figures in +// the rank and Just scenarios: music_data.py reports those as warnings). +const computable = (deck, scenario) => { + if (!deck) return false; + const has = (s) => s && Array.isArray((s.weights || [])[kind]); + if (scenario && scenario.mode === "free") return (deck.offSeeds || []).length > 0 && deck.offSeeds.every(has); + if (deck.unplayable || !(deck.seeds || []).length) return false; + const rank1 = !scenario || ((scenario.ranks || []).every((r) => r === 1) && (scenario.just ?? 1) >= 1); + return deck.seeds.every((s) => has(s) && (rank1 || Array.isArray((s.rangeWeights || [])[kind]))); +}; +const has = ranking.scenarioData(data); +for (const k of ["free", "ranks", "just"]) if (!has[k]) problem(`scenarioData: no data for the ${k} scenario`); + +const SCENARIOS = [ + ["Gekisou Live, rank 1", null], + ["Gekisou Live, rank 5", { mode: "battle", ranks: [5, 5, 5], just: 1, great: 0 }], + ["Gekisou Live, ranks 2 3 4", { mode: "battle", ranks: [2, 3, 4], just: 1, great: 0 }], + ["Gekisou Live, Just 0", { mode: "battle", ranks: [1, 1, 1], just: 0, great: 0 }], + ["Gekisou Live, rank 3, Just 0.5, Great 0.2", { mode: "battle", ranks: [3, 3, 3], just: 0.5, great: 0.2 }], + ["Free Live", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0 }], + ["Free Live, Great 0.5", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0.5 }], +]; +const SKILLS = [1, 1, 1, 1, 1]; +let figures = 0; +for (const [name, scenario] of SCENARIOS) { + const joined = ranking.joinCharts(data, scenario); + const expected = charts.filter(({ chart }) => computable(chart.deck, scenario)).length; + if (joined.length !== expected) problem(`${name}: ${joined.length} charts with figures, expected ${expected}`); + for (const r of joined) { + figures++; + if (!finite(r.base) || r.base <= 0) problem(`${name}: chart ${r.scoreId}: base ${r.base}`); + if (!Array.isArray(r.weights) || !r.weights.length || !r.weights.every(finite)) { + problem(`${name}: chart ${r.scoreId}: weights are not finite numbers`); + } + } + if (!joined.length) continue; + const ranked = ranking.rank(joined, { skills: SKILLS, source: "bgm", overheadMs: 30000 }); + if (!ranked.some((r) => r.frontier)) problem(`${name}: no chart on the frontier`); + for (const r of ranked) { + if (r.lengthMs === null) problem(`${name}: chart ${r.scoreId}: no play length`); + else if (!finite(r.perMinute) || !finite(r.rate)) problem(`${name}: chart ${r.scoreId}: rate ${r.rate}`); + } + for (const room of [0, 5]) { + ranking.eventDominance(ranked, "bgm", ranking.X_MAX, room); + for (const r of ranked) { + const p = ranking.requiredPower(r, SKILLS, "S", 1, room); + if (p !== null && !(finite(p) && p >= 0)) problem(`${name}: chart ${r.scoreId}: required power ${p}`); + } + } + const one = ranked[0]; + const chance = ranking.reachChance(one, SKILLS, 300000, "S", 1, 0); + if (chance !== null && !(chance >= 0 && chance <= 1)) problem(`${name}: reach chance ${chance}`); + catalog.refigure(rows, data, scenario); +} +catalog.histogram(rows, (r) => r.level); + +if (problems.length) { + for (const m of problems.slice(0, 40)) console.log(m); + if (problems.length > 40) console.log(`${problems.length - 40} more problems`); + process.exit(1); +} +console.log(`page smoke test: ${rows.length} charts, ${SCENARIOS.length} scenarios, ${figures} chart figures; ` + + `plain kind ${kind}, scenarios free/ranks/just`); diff --git a/.github/scripts/songs_page.sh b/.github/scripts/songs_page.sh new file mode 100755 index 0000000..71849aa --- /dev/null +++ b/.github/scripts/songs_page.sh @@ -0,0 +1,18 @@ +#!/usr/bin/env bash +# The chart data page of ournotes-player ($PLAYER_REPOSITORY at $PLAYER_REF: examples/songs, the page that reads +# music-data.json) into $1, for the smoke test of music_data.py check. Only that directory is checked out and nothing +# is built: the test imports the page's pure modules (catalog.js, ranking.js) in Node.js. +set -euo pipefail +dir="$1" +if [ -z "${PLAYER_REF:-}" ]; then + echo "::error::set the repository variable MUSIC_DATA_PLAYER_REF: the ournotes-player commit whose chart data page reads this music data (.github/MUSIC_DATA.md)" + exit 1 +fi +rm -rf "$dir" +git clone --quiet --filter=blob:none --no-checkout "https://github.com/$PLAYER_REPOSITORY.git" "$dir" +git -C "$dir" sparse-checkout set examples/songs +git -C "$dir" checkout --quiet "$PLAYER_REF" +for f in catalog.js ranking.js; do + [ -f "$dir/examples/songs/$f" ] || { echo "::error::$PLAYER_REPOSITORY $PLAYER_REF has no examples/songs/$f"; exit 1; } +done +echo "ournotes-player $(git -C "$dir" rev-parse --short HEAD): examples/songs" diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py new file mode 100644 index 0000000..e601026 --- /dev/null +++ b/.github/scripts/test_music_data.py @@ -0,0 +1,443 @@ +"""Self-test of the music data gates (music_data.py): a synthetic music-data.json passes every gate, and each gate +fails on the defect it is there for. Runs in seconds, without the network: + + python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + +The JSON Schema gate uses docs/schema/music-data.schema.json (or $MUSIC_DATA_SCHEMA) when the checkout has it; the +page smoke test runs when $MUSIC_DATA_PAGE names ournotes-player's examples/songs (and Node.js is installed). Real +files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass) and +$MUSIC_DATA_OLD_SAMPLE (one without them: the scenario gate stops it). +""" +import copy +import hashlib +import importlib.util +import io +import json +import os +import shutil +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import music_data # noqa: E402 +from music_data import Context, gates # noqa: E402 + +ROOT = Path(__file__).resolve().parents[2] +LANGS = ["ja", "en", "zh-Hant", "zh-Hans", "ko"] +TABLES = music_data.SONG_TABLES + ("MasterLiveSkillEffect", "MasterLiveGekisouRankingScoreBonus") +BIN = {t: hashlib.sha256(t.encode()).hexdigest() for t in TABLES} # the manifest's hashes of the served files +DECK = "d4" * 20 +CONTENT = ("deck", "scenarios", "finite", "references", "bgm") + + +def schema_path(): + p = Path(os.environ.get("MUSIC_DATA_SCHEMA") or ROOT / music_data.SCHEMA) + if not p.is_file(): + return None + return p if importlib.util.find_spec("jsonschema") else None + + +def page_path(): + p = os.environ.get("MUSIC_DATA_PAGE") + return Path(p) if p and Path(p, "ranking.js").is_file() and shutil.which("node") else None + + +# ---------------------------------------------------------------- a synthetic file +def text(stem): + return {lang: f"{stem}-{lang}" for lang in LANGS} + + +def check_deck(exact): + return {"deck": [[0, 5000], None], "exact": exact, "predicted": exact + 0.25, "bound": 7.0} + + +def deck_chart(): + seed = {"seed": 0, "score": 120000, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, + "justCount": 0, "lotResults": [0, 0, 0, 0], + "rangeScorePerfect": 4000}], + "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000, + "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} + return {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, + "ranges": [{"index": 0, "mission": 1, "startMs": 1000, "endMs": 5000, "rankBonusPercent": 250, + "rankBonusPercents": [250, 190, 160, 100, 100]}], + "justNotes": 0, "seeds": [seed], + "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], + "unplayable": None} + + +def chart(difficulty, score_id, last=60000): + return {"difficulty": difficulty, "scoreId": score_id, "level": 10, "displayLevel": 10.5, "fullComboCount": 20, + "asset": {"key": f"Live/MusicScore/c/c_{score_id}", "sha256": "ab" * 32}, + "notes": {"judged": 20, "total": 22, "byOperateType": {"1": 20, "120": 2}}, + "bpm": {"main": 120.0, "min": 120.0, "max": 120.0, "changes": [{"timeMs": 0, "bpm": 120.0}]}, + "firstNoteMs": 1000, "lastJudgedNoteMs": last, "lastNoteMs": last, "musicLengthMs": last + 1000, + "skillEventsMs": [1000, 3000], "fevers": [[1000, 5000]], "deck": deck_chart()} + + +def song(i, charts): + ranks = ["D", "C", "B", "A", "S", "SS"] + return {"id": i, "sortOrder": i, "startAt": "2026/01/01 0:00:00", "defaultUnlock": True, "title": text(f"t{i}"), + "ruby": None, "phonetic": text(f"p{i}"), "bandIds": [1], "bandName": None, "vocalCharacterIds": [1], + "lyricist": text("l"), "composer": text("c"), "arranger": None, "musicType": 1, "musicCategories": [1], + "bestMusicTagIds": [1], "jacket": f"jkt_{i}", "gekisouMissions": [1, 2, 3], + "bgm": {"soundId": i, "cueSheet": f"Bgm{i}", "cue": f"song{i}", + "length": {"lengthMs": 90000, "samples": 48000 * 90, "sampleRate": 48000, "durationMs": 90000}}, + "scoreRanks": [{"rank": r, "requiredScore": n * 1000, "battleRequiredScore": n * 2000} + for n, r in enumerate(ranks)], + "charts": charts, "master": {"MasterLiveMusic": {"_id": i, "_rate": 1.5}, "MasterLiveScoreRank": []}} + + +def sample() -> dict: + kind = {"id": 0, "effectType": 2000, "activationTimeSecond": 5.0, "durationMs": 5000, "skillTargetIds": [], + "skillConditionGroup": 0, "skillReleaseConditionGroup": 0, "effectLimitCount": 0, + "effectExecuteLimitCount": 0, "effectExecuteLimitResetConditionGroup": 0, "rows": [1], "values": [10000]} + return { + "format": "nnnotes.music-data/1", + "provenance": { + "region": "tw", "client": {"versionName": "1.0.1", "versionCode": 25}, + "catalog": {"resourceVersion": None, "sha256": "cd" * 32}, + "master": {"source": "api", "version": "v-test", "tables": {t: {"sha256": BIN[t]} for t in TABLES}}, + "exporter": {"name": "nnnotes", "version": "0.1.2", "chartFormat": "nnnotes.live-score/1"}, + "deck": {"name": "ournotes-deck", "version": "0.0.1", + "source": "https://github.com/empty-sekai/ournotes-deck", "commit": DECK, + "format": "ournotes-deck.chart-stats/2"}}, + "languages": LANGS, + "bands": [{"id": 1, "name": text("band"), "mainColor": "#3388BB", "subColor": "#FFFFFF"}], + "characters": [{"id": 1, "bandId": 1, "name": text("ch"), "shortName": text("c"), "mainColor": "#77BBDD"}], + "tags": [{"id": 1, "name": text("tag")}], + "categories": [{"id": 1, "musicCategories": [1], "name": text("cat")}], + "deck": {"model": {"power": 300000, "checkPower": 1000003}, "kinds": [kind]}, + "songs": [song(100001, [chart("easy", 10), chart("expert", 30)]), song(100002, [chart("expert", 40)])], + } + + +def context(tmp_path: Path, doc: dict, **kw) -> tuple[bytes, Context]: + """The file's bytes and what check reads next to it: the decoded master data and its snapshot, the jackets (the + page smoke test only with page=page_path(): it starts Node.js).""" + master = tmp_path / "master" + master.mkdir(exist_ok=True) + files = {} + for t in TABLES: + data = json.dumps({"_allData": []}).encode() + (master / f"{t}.json").write_bytes(data) + files[f"{t}.json"] = hashlib.sha256(data).hexdigest() + listed = [{"name": f"{t}.bin", "hash": BIN[t], "size": 1} for t in TABLES] + (master / music_data.MANIFEST).write_text(json.dumps({"version": "v-test", "files": listed}), encoding="utf-8") + jackets = tmp_path / "jackets" + jackets.mkdir(exist_ok=True) + for s in doc.get("songs") or []: + if s.get("jacket"): + (jackets / f"{s['jacket']}.webp").write_bytes(b"RIFF....WEBP") + # an infinity as nnnotes writes it (1e999: JSON that JavaScript reads too); a NaN stays NaN (nnnotes writes none) + raw = json.dumps(doc, ensure_ascii=False).replace("Infinity", "1e999").encode("utf-8") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(region="tw", language="zh-Hant", master=master, + snapshot={"entry": {"version": "v-test", "client_version": "1.0.1"}, "files": files}, + deck_commit=DECK, nnnotes_version="0.1.2", jackets=jackets, schema=schema_path(), published=None, + page=None, file=tmp_path / "music-data.json") + for k, v in kw.items(): + setattr(ctx, k, v) + return raw, ctx + + +def run(tmp_path, doc, **kw) -> dict: + return gates(*context(tmp_path, doc, **kw)) + + +def gate(report: dict, name: str) -> dict: + return next(g for g in report["gates"] if g["gate"] == name) + + +def failures(report: dict) -> list: + return [(g["gate"], g["failures"]) for g in report["gates"] if not g["passed"]] + + +def seed(doc, song=0, chart=0): + return doc["songs"][song]["charts"][chart]["deck"]["seeds"][0] + + +# ---------------------------------------------------------------- the gates +def test_the_sample_passes_every_gate(tmp_path): + r = run(tmp_path, sample(), page=page_path()) + assert r["passed"], failures(r) + assert [g["gate"] for g in r["gates"]] == [name for name, _ in music_data.GATES] + assert not [g for g in r["gates"] if g["warningCount"]] + assert r["sha256"] == hashlib.sha256((tmp_path / "music-data.json").read_bytes()).hexdigest() + + +def test_schema(tmp_path): + if schema_path() is None: + pytest.skip("no docs/schema/music-data.schema.json (a fork not synced with upstream) or no jsonschema") + doc = sample() + doc["songs"][0]["charts"][0]["level"] = "10" + g = gate(run(tmp_path, doc), "schema") + assert not g["passed"] and any("songs/0/charts/0/level" in f for f in g["failures"]) + + +def test_page_smoke(tmp_path): + if page_path() is None: + pytest.skip("set MUSIC_DATA_PAGE to ournotes-player's examples/songs (and install Node.js)") + assert gate(run(tmp_path, sample(), page=page_path()), "page")["passed"] + doc = sample() + for s in doc["songs"]: + for c in s["charts"]: + c["deck"].pop("offSeeds") + g = gate(run(tmp_path, doc, page=page_path()), "page") + assert not g["passed"] and any("free scenario" in f for f in g["failures"]) + + +def unplayable(doc): + doc["songs"][1]["charts"][0]["deck"]["unplayable"] = "more than three fevers" # its seeds kept + + +@pytest.mark.parametrize("change, name, match", [ + # the play scenario fields + (lambda d: d["songs"][0]["charts"][0]["deck"].pop("offSeeds"), "scenarios", "offSeeds missing"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["offSeeds"].append({}), "scenarios", "offSeeds has 2 entries"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].pop("rankBonusPercents"), "scenarios", + "rankBonusPercents missing"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[250, 190, 160, 100]), + "scenarios", "rankBonusPercents not five ints"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[251, 190, 160, 100, 100]), + "scenarios", "is not rankBonusPercent"), + (lambda d: seed(d).pop("scorePerfect"), "scenarios", "no scorePerfect"), + (lambda d: seed(d).pop("rangeWeights"), "scenarios", "no rangeWeights"), + (lambda d: seed(d).pop("rankCheck"), "scenarios", "no rankCheck"), + (lambda d: seed(d)["ranges"][0].pop("rangeScorePerfect"), "scenarios", "rangeScorePerfect missing"), + (lambda d: seed(d).update(rangeWeights=[[[0.1]]]), "scenarios", "rangeWeights are not"), + (lambda d: seed(d)["rankCheck"].update(exact=99999), "scenarios", "rank check deck is not within"), + # the deck statistics + (lambda d: d.update(deck=None), "deck", "deck is null"), + (lambda d: d["songs"][0]["charts"][0].update(deck=None), "deck", "no deck statistics"), + (lambda d: seed(d)["check"].update(exact=99999), "deck", "check deck is not within"), + (lambda d: seed(d).update(weights=[[0.5]]), "deck", "weights are not"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(seeds=[]), "deck", "no seeds"), + (unplayable, "deck", "unplayable, but has Gekisou on seeds"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(events=[[0, 1000]]), "deck", "1 skill events"), + # numbers + (lambda d: seed(d)["weights"][0].__setitem__(1, float("nan")), "finite", "weights[0][1]: not finite"), + (lambda d: d["songs"][1]["charts"][0]["bpm"].update(main=float("inf")), "finite", "bpm.main: not finite"), + # references and texts + (lambda d: d["songs"][1]["bandIds"].append(9), "references", "band 9 not in"), + (lambda d: d["songs"][1]["vocalCharacterIds"].append(7), "references", "character 7 not in"), + (lambda d: d["songs"][0].update(title=None), "references", "title: no text"), + (lambda d: d["songs"][0].update(title={lang: "" for lang in LANGS}), "references", "empty in every language"), + (lambda d: d["songs"][0]["composer"].pop("ko"), "references", "composer: not a text"), + (lambda d: d["songs"][0].update(jacket=None), "references", "no jacket"), + (lambda d: d["songs"].reverse(), "references", "not sorted"), + (lambda d: d["songs"][1]["charts"][0].update(scoreId=10), "references", "occurs twice"), + (lambda d: d["songs"][0]["charts"].reverse(), "references", "charts ['expert', 'easy']"), + (lambda d: d["songs"][0].update(scoreRanks=[]), "references", "score ranks"), + (lambda d: d["songs"][0].update(bandIds=[]), "references", "neither a band nor a band name"), + # BGM + (lambda d: d["songs"][0]["bgm"].update(length=None), "bgm", "no BGM length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=50000, samples=48000 * 50), "bgm", + "ends before the last note"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=90500), "bgm", "is not samples"), + (lambda d: d["songs"][0]["bgm"]["length"].update(lengthMs=95000), "bgm", "cue length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=1200000, samples=48000 * 1200), "bgm", + "bounds"), + # provenance + (lambda d: d.update(format="nnnotes.music-data/2"), "provenance", "format"), + (lambda d: d["provenance"].update(region="en"), "provenance", "region 'en'"), + (lambda d: d["provenance"]["master"].update(version="other"), "provenance", "the snapshot's 'v-test'"), + (lambda d: d["provenance"]["master"].update(source="embedded"), "provenance", "master.source"), + (lambda d: d["provenance"]["master"]["tables"]["MasterText"].update(sha256="00" * 32), "provenance", + "MasterText.sha256 is not"), + (lambda d: d["provenance"]["master"]["tables"].pop("MasterBand"), "provenance", "lacks MasterBand"), + (lambda d: d["provenance"]["deck"].update(commit="ee" * 20), "provenance", "rust/Cargo.lock pins"), + (lambda d: d["provenance"]["exporter"].update(version="0.0.9"), "provenance", "installed nnnotes"), + (lambda d: d["provenance"]["client"].update(versionName=None), "provenance", "no APK version"), +]) +def test_a_gate_fails(tmp_path, change, name, match): + doc = sample() + change(doc) + r = run(tmp_path, doc) + g = gate(r, name) + assert not r["passed"] and not g["passed"] and any(match in f for f in g["failures"]), g + + +def test_a_table_read_is_not_the_listed_one(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.master / "MasterText.json").write_text('{"_allData": [{}]}', encoding="utf-8") + g = gate(gates(raw, ctx), "provenance") + assert not g["passed"] and g["failures"] == ["MasterText.json read is not the one index.json lists"] + + +def test_a_jacket_file_is_missing(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.jackets / "jkt_100002.webp").unlink() + g = gate(gates(raw, ctx), "references") + assert not g["passed"] and g["failures"] == ["song 100002: no jacket file jackets/jkt_100002.webp"] + + +def test_warnings_do_not_fail(tmp_path): + doc = sample() + seed(doc).update(rangeWeights=None, rankCheck=None) # overlapping ranges + seed(doc, 0, 1)["rangeWeights"][0] = None # a kind reading the confirmed rank + doc["songs"][1]["charts"][0]["deck"]["offSeeds"][0]["weights"][0] = None + doc["songs"][0]["master"]["MasterLiveMusic"]["_rate"] = float("inf") # 1e999 in master data as served + doc["songs"][1]["title"]["zh-Hant"] = "" + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 3 and gate(r, "finite")["warningCount"] == 1 + assert gate(r, "references")["warnings"] == ["songs without a zh-Hant title: 100002"] + + +def test_against_the_published_file(tmp_path): + doc = sample() + raw, ctx = context(tmp_path, doc) + more = copy.deepcopy(doc) + more["songs"].append(song(100003, [chart("expert", 50)])) + ctx.published = json.dumps(more).encode() + r = gates(raw, ctx) + assert not gate(r, "counts")["passed"] and gate(r, "counts")["failures"] == [ + "2 songs, the published file has 3", "3 charts, the published file has 4"] + assert gate(r, "counts")["warnings"] == ["songs no longer in the file: 100003", "charts no longer in the file: 50"] + ctx.published = raw + b" " * (len(raw) * 2) # the same songs in a file three times as big + r = gates(raw, ctx) + assert gate(r, "counts")["passed"] and not gate(r, "size")["passed"] + ctx.published = raw[:len(raw) // 3] # a third: not JSON, and twice exceeded + r = gates(raw, ctx) + assert gate(r, "counts")["warnings"] == ["the published music-data.json is not JSON: skipped"] + assert not gate(r, "size")["passed"] + ctx.published = raw + assert gates(raw, ctx)["passed"] + + +def test_not_json(tmp_path): + r = gates(b"{", Context()) + assert not r["passed"] and r["gates"][0]["gate"] == "json" + + +def test_report_markdown(tmp_path): + doc = sample() + doc["songs"][0]["charts"][0]["deck"].pop("offSeeds") + text_ = music_data.report_markdown(run(tmp_path, doc)) + assert "| scenarios | **failed** (1) |" in text_ + assert "- scenarios failure: chart 10 (100001 easy): offSeeds missing, expected exactly one" in text_ + + +def test_deck_commit_from_cargo_lock(tmp_path): + lock = tmp_path / "rust" / "Cargo.lock" + lock.parent.mkdir() + lock.write_text('[[package]]\nname = "pyo3"\nversion = "0.29.0"\n\n[[package]]\nname = "ournotes-deck"\n' + 'version = "0.0.1"\nsource = "git+https://github.com/empty-sekai/ournotes-deck?rev=' + "a" * 40 + + "#" + "b" * 40 + '"\n', encoding="utf-8") + assert music_data.deck_commit(tmp_path) == "b" * 40 + with pytest.raises(SystemExit, match="sync the fork"): + music_data.deck_commit(tmp_path / "nowhere") + + +# ---------------------------------------------------------------- real files (when named) +def real(name): + p = os.environ.get(name) + if not p or not Path(p).is_file(): + pytest.skip(f"set {name} to a music-data.json") + return Path(p).read_bytes() + + +def test_a_real_file_without_the_scenario_fields(): + raw = real("MUSIC_DATA_OLD_SAMPLE") + r = gates(raw, Context(language="zh-Hant"), only=CONTENT) + assert failures(r) == [("scenarios", gate(r, "scenarios")["failures"])] + g = gate(r, "scenarios") + charts = sum(len(s["charts"]) for s in json.loads(raw)["songs"]) + assert g["failureCount"] >= charts and all("offSeeds missing" in f or "rankBonusPercents missing" in f + or "no scorePerfect" in f for f in g["failures"]) + + +def test_a_real_file_with_the_scenario_fields(tmp_path): + raw = real("MUSIC_DATA_SAMPLE") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(language="zh-Hant", page=page_path(), file=tmp_path / "music-data.json", published=raw) + r = gates(raw, ctx, only=CONTENT + ("counts", "size", "page")) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 0 + + +# ---------------------------------------------------------------- publish (a stand-in bucket) +class FakeS3: + def __init__(self, corrupt=None): + self.store, self.log, self.corrupt = {}, [], corrupt + + def upload_file(self, src, bucket, key, ExtraArgs): + self.store[key] = b"not it" if key == self.corrupt else Path(src).read_bytes() + self.log.append((key, ExtraArgs["CacheControl"], ExtraArgs["ContentType"])) + + def get_object(self, Bucket, Key): + return {"Body": io.BytesIO(self.store[Key])} + + +class FakeBucket: + name, prefix, writable = "moenotes", "music-data/", True + + def __init__(self, s3): + self.s3 = s3 + + def keys(self, sub=""): + return {k[len(self.prefix):]: len(v) for k, v in self.s3.store.items() if k.startswith(self.prefix + sub)} + + +def published_out(tmp_path, monkeypatch, s3): + out = tmp_path / "out" + out.mkdir() + raw, ctx = context(out, sample()) + report = gates(raw, ctx) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "build.json").write_text(json.dumps({"sha256": report["sha256"], + "archive": f"archive/v-test/{report['sha256']}.json"}), + encoding="utf-8") + monkeypatch.setattr(music_data, "bucket", lambda: FakeBucket(s3)) + monkeypatch.setattr(music_data.time, "sleep", lambda s: None) + monkeypatch.setenv("STORY_S3_ENDPOINT", "https://storage.example") + monkeypatch.setenv("STORY_S3_BUCKET", "moenotes") + monkeypatch.delenv("FORCE", raising=False) + return out, report + + +def test_publish_order_and_read_back(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + music_data.cmd_publish(str(out)) + keys = [k for k, _, _ in s3.log] + archive = f"music-data/archive/v-test/{report['sha256']}.json" + assert sorted(keys[:2]) == ["music-data/jackets/jkt_100001.webp", "music-data/jackets/jkt_100002.webp"] + assert keys[2:] == [archive, "music-data/music-data.json", "music-data/build.json"] + caches = {k: c for k, c, _ in s3.log} + assert caches[archive].endswith("immutable") and caches["music-data/music-data.json"] == "no-cache" + assert s3.store["music-data/music-data.json"] == (out / "music-data.json").read_bytes() + s3.log.clear() + music_data.cmd_publish(str(out)) # again: the jackets and the archive copy are there + assert [k for k, _, _ in s3.log] == ["music-data/music-data.json", "music-data/build.json"] + monkeypatch.setenv("FORCE", "true") + s3.log.clear() + music_data.cmd_publish(str(out), dry_run=True) + assert s3.log == [] + + +def test_publish_stops_before_the_marker_when_the_file_does_not_read_back(tmp_path, monkeypatch): + s3 = FakeS3(corrupt="music-data/music-data.json") + out, _ = published_out(tmp_path, monkeypatch, s3) + + def offline(*a, **kw): + raise music_data.urllib.error.URLError("offline") + monkeypatch.setattr(music_data, "get", offline) + with pytest.raises(SystemExit, match="music-data.json: the bucket does not serve what was uploaded"): + music_data.cmd_publish(str(out)) + assert "music-data/build.json" not in s3.store + + +def test_publish_needs_passed_gates(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + (out / "check.json").write_text(json.dumps(dict(report, passed=False)), encoding="utf-8") + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "music-data.json").write_bytes(b"{}") # not the file that was checked + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + assert s3.store == {} diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml new file mode 100644 index 0000000..a5dffd9 --- /dev/null +++ b/.github/workflows/music-data.yml @@ -0,0 +1,146 @@ +name: Music data + +# Publishes the music data file of the current master data for the chart data page (ournotes-player examples/songs): +# `nnnotes music-data` (every song and chart with the deck model's statistics, docs/music-data.md) of +# moenotes-masterdata-sync's decoded master data, into the story site's bucket under music-data/: music-data.json, +# jackets/, archive//.json and the build marker build.json. A run first compares what it would +# build (the master data snapshot, the deck commit, the nnnotes commit) with the published build.json and ends there +# when nothing changed; else it builds, checks the file (quality gates and a smoke test with the page's own modules) +# and uploads it: the jackets and the archive copy first, then music-data.json, the build marker last. A file that +# fails a gate is not published. Nothing is deleted from the bucket. +# +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new +# master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds although nothing changed +# (`force`) or does a dry run. docs: .github/MUSIC_DATA.md + +on: + repository_dispatch: + types: [masterdata-updated] + workflow_dispatch: + inputs: + force: + description: "build and publish although the published file was made from the same inputs" + type: boolean + default: false + dry_run: + description: "build and check, but upload nothing" + type: boolean + default: false + schedule: + - cron: "41 3 * * *" + +permissions: + contents: read + +# One run at a time: a run replaces music-data.json and build.json. +concurrency: + group: music-data + cancel-in-progress: false + +env: + # The story site's bucket (its repository variables) and the key prefix of the music data. + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || 'music-data' }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} + # The ournotes-player whose chart data page reads the file: its modules smoke-test every build. No default yet: the + # page with the play scenarios is not merged; a run stops before building until MUSIC_DATA_PLAYER_REF is set. + PLAYER_REPOSITORY: ${{ vars.MUSIC_DATA_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.MUSIC_DATA_PLAYER_REF || '' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the bundle key and the CDN come from secrets in the build step. + NNNOTES_CATALOG_REGION: tw + NNNOTES_CATALOG_LANGUAGE: zh-Hant + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 10 + outputs: + build: ${{ steps.plan.outputs.build }} + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 # the nnnotes commit: the last one that changed src/, rust/ or pyproject.toml + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - name: Inputs against the published build + id: plan + env: + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py plan + + build: + needs: plan + if: needs.plan.outputs.build == 'true' + runs-on: ubuntu-latest + timeout-minutes: 90 + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: The chart data page + run: .github/scripts/songs_page.sh "$WORK/page" + + - uses: Swatinem/rust-cache@v2 + with: + workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - name: Install nnnotes + # builds the extension module nnnotes._deck: the deck model at the commit rust/Cargo.toml pins + run: uv pip install --system -e ".[test]" boto3 + + - name: Gate self-test + env: + MUSIC_DATA_PAGE: ${{ env.WORK }}/page/examples/songs + run: python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/music_data.py master "$WORK/master" + + - name: Build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + # A new cache on every run, never actions/cache: nnnotes keeps a downloaded catalog file for good (a cached + # one would never see a new resource version), and the cache holds decrypted game files. + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + run: python .github/scripts/music_data.py build "$WORK/out" + + - name: Gates + run: python .github/scripts/music_data.py check "$WORK/out" "$WORK/master" "$WORK/page/examples/songs" + + - name: Publish + # only after every gate passed (a failed step skips it); a dry run lists what it would upload + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py publish "$WORK/out" ${{ github.event.inputs.dry_run == 'true' && '--dry-run' || '' }} From 8877ca8c69e0e89c82e5f94e90e7b7bea808db66 Mon Sep 17 00:00:00 2001 From: nichinichisou Date: Wed, 30 Sep 2026 01:50:36 +0800 Subject: [PATCH 09/14] ci: publish the music data only with MUSIC_DATA_PUBLISH true A publishing switch, the repository variable MUSIC_DATA_PUBLISH (unset: off): unless it is `true` every run of music-data.yml, whatever its trigger, is a dry run. It builds, runs every gate and the page smoke test and lists what it would upload; the publish step gets no bucket key, and `music_data.py publish` itself refuses to upload. MUSIC_DATA.md says how to turn it on. Co-Authored-By: Claude Opus 5.5 --- .github/MUSIC_DATA.md | 13 +++++++++++-- .github/scripts/music_data.py | 24 ++++++++++++++++++++---- .github/scripts/test_music_data.py | 13 +++++++++++++ .github/workflows/music-data.yml | 14 +++++++++----- 4 files changed, 53 insertions(+), 11 deletions(-) diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md index 2f669c4..f8b22ab 100644 --- a/.github/MUSIC_DATA.md +++ b/.github/MUSIC_DATA.md @@ -51,6 +51,13 @@ missed, and `workflow_dispatch`: Runs do not overlap (`concurrency: music-data`). +**Publishing switch.** Nothing is uploaded unless the repository variable `MUSIC_DATA_PUBLISH` is `true` (unset: +off). Off, every run, whatever its trigger (the schedule, `masterdata-updated`, `workflow_dispatch` with or without +`dry_run`), is a dry run: it builds, runs every gate and the smoke test, and lists what it would upload; the publish +step then gets no bucket key (anonymous, it can only read) and `music_data.py publish` itself refuses to upload. +Turn it on (`gh variable set MUSIC_DATA_PUBLISH --body true`) once the published format is final; while it is off, +nothing being published, `plan` finds no `build.json` and every run builds. + ## Gates Every one must pass, else nothing is published. Warnings go to the job summary and `build.json` and do not stop it. @@ -95,6 +102,7 @@ Repository variables: | Variable | Default | | |---|---|---| | `MUSIC_DATA_PLAYER_REF` | none: **required** | the ournotes-player commit whose chart data page reads this file (the page with the play scenarios); a run stops before building without it | +| `MUSIC_DATA_PUBLISH` | none: off | `true`: upload; anything else: every run is a dry run | | `MUSIC_DATA_PLAYER_REPOSITORY` | `empty-sekai/ournotes-player` | | | `MUSIC_DATA_S3_PREFIX` | `music-data` | the key prefix in the bucket | | `MUSIC_DATA_MASTERDATA_REGION` | `hk-tw-mo` | the region of `index.json`; the build reads the TW catalog (`[catalog] region` `tw`) | @@ -103,10 +111,11 @@ Repository variables: ## Before the first run - **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play - scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with - upstream first. Until then `plan` stops naming what is missing. + scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6, `12df2a6`): sync the + fork with upstream first. Until then `plan` stops naming what is missing. - **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play scenario fields, once that page is merged. +- **Publishing.** Set `MUSIC_DATA_PUBLISH` to `true` last, when dry runs pass and the published format is final. ## Notes diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py index 2887336..0c336b5 100755 --- a/.github/scripts/music_data.py +++ b/.github/scripts/music_data.py @@ -18,7 +18,8 @@ publish OUT [--dry-run] the jackets the bucket lacks or has at another size (every one with $FORCE), the file's archive copy, then music-data.json, build.json last; each read back and its SHA-256 checked. - Only a file whose check.json passed; never deletes + Only a file whose check.json passed; never deletes. A dry run unless $MUSIC_DATA_PUBLISH is + `true` (the publishing switch, off by default): it lists what it would upload Bucket: story_site.Bucket ($STORY_S3_ENDPOINT, $STORY_S3_BUCKET, credentials $STORY_S3_ACCESS_KEY / $STORY_S3_SECRET_KEY) with the key prefix $MUSIC_DATA_S3_PREFIX (default music-data). The published files are read @@ -187,6 +188,7 @@ def cmd_plan() -> None: + "\n- this run: " + ("builds" + (" (force)" if force else "") + (f", changed: {', '.join(changed)}" if changed and have else "") if build else "nothing changed, nothing to build")) + summary(f"- publishing: {'on' if publishing() else 'off (MUSIC_DATA_PUBLISH is not `true`: a dry run)'}") output("build", "true" if build else "false") @@ -791,7 +793,15 @@ def read_back(b, key: str, digest: str) -> None: fail(f"{key}: the bucket does not serve what was uploaded (SHA-256 {digest[:12]})") +def publishing() -> bool: + """The publishing switch: uploads only with MUSIC_DATA_PUBLISH `true` (a repository variable, unset: off).""" + return os.environ.get("MUSIC_DATA_PUBLISH") == "true" + + def cmd_publish(out: str, dry_run: bool = False) -> None: + if not publishing() and not dry_run: + summary("- publishing is off (the repository variable MUSIC_DATA_PUBLISH is not `true`): a dry run") + dry_run = True o = Path(out) raw = (o / FILE).read_bytes() report = json.loads((o / "check.json").read_text(encoding="utf-8")) @@ -802,16 +812,22 @@ def cmd_publish(out: str, dry_run: bool = False) -> None: if not b.writable and not dry_run: fail("publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") force = os.environ.get("FORCE") == "true" - have = b.keys(JACKETS) + try: + have = b.keys(JACKETS) + archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) + except Exception as e: # a dry run without a key where the bucket lists to none + if not dry_run: + raise + print(f"cannot list the bucket ({type(e).__name__}): the dry run lists every object", flush=True) + have, archived = {}, False jackets = sorted((o / "jackets").glob("*.webp")) new = [j for j in jackets if force or have.get(JACKETS + j.name) != j.stat().st_size] - archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) steps = [(JACKETS + j.name, j, JACKET_CACHE, False) for j in new] if not archived: steps.append((marker["archive"], o / FILE, ARCHIVE_CACHE, True)) # the file before the marker that names it, the jackets and the archive copy before the file steps += [(FILE, o / FILE, FILE_CACHE, True), (MARKER, o / MARKER, FILE_CACHE, True)] - if dry_run: + if dry_run or not publishing(): for key, _, _, _ in steps: print(f"would upload {b.prefix}{key}") else: diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py index e601026..bc86b88 100644 --- a/.github/scripts/test_music_data.py +++ b/.github/scripts/test_music_data.py @@ -395,6 +395,7 @@ def published_out(tmp_path, monkeypatch, s3): monkeypatch.setenv("STORY_S3_ENDPOINT", "https://storage.example") monkeypatch.setenv("STORY_S3_BUCKET", "moenotes") monkeypatch.delenv("FORCE", raising=False) + monkeypatch.setenv("MUSIC_DATA_PUBLISH", "true") return out, report @@ -430,6 +431,18 @@ def offline(*a, **kw): assert "music-data/build.json" not in s3.store +@pytest.mark.parametrize("switch", [None, "", "false", "True", "1"]) +def test_publishing_is_off_unless_the_switch_is_true(tmp_path, monkeypatch, switch): + s3 = FakeS3() + out, _ = published_out(tmp_path, monkeypatch, s3) + if switch is None: + monkeypatch.delenv("MUSIC_DATA_PUBLISH") + else: + monkeypatch.setenv("MUSIC_DATA_PUBLISH", switch) + music_data.cmd_publish(str(out)) # a dry run, whatever the command line says + assert s3.log == [] and s3.store == {} + + def test_publish_needs_passed_gates(tmp_path, monkeypatch): s3 = FakeS3() out, report = published_out(tmp_path, monkeypatch, s3) diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml index a5dffd9..d0c9478 100644 --- a/.github/workflows/music-data.yml +++ b/.github/workflows/music-data.yml @@ -11,7 +11,8 @@ name: Music data # # Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new # master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds although nothing changed -# (`force`) or does a dry run. docs: .github/MUSIC_DATA.md +# (`force`) or does a dry run. Publishing switch: unless the repository variable MUSIC_DATA_PUBLISH is `true`, every +# run, whatever its trigger, is a dry run (it builds and checks, and uploads nothing). docs: .github/MUSIC_DATA.md on: repository_dispatch: @@ -42,6 +43,8 @@ env: STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || 'music-data' }} + # The publishing switch: anything but 'true' makes every run a dry run. + MUSIC_DATA_PUBLISH: ${{ vars.MUSIC_DATA_PUBLISH || 'false' }} # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} MASTERDATA_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} @@ -138,9 +141,10 @@ jobs: run: python .github/scripts/music_data.py check "$WORK/out" "$WORK/master" "$WORK/page/examples/songs" - name: Publish - # only after every gate passed (a failed step skips it); a dry run lists what it would upload + # Only after every gate passed (a failed step skips it). A dry run (the dry_run input, or the publishing switch + # MUSIC_DATA_PUBLISH off) lists what it would upload; it gets no bucket key: anonymous, it can only read. env: - STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} - STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + STORY_S3_ACCESS_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_ACCESS_KEY || '' }} + STORY_S3_SECRET_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_SECRET_KEY || '' }} FORCE: ${{ github.event.inputs.force }} - run: python .github/scripts/music_data.py publish "$WORK/out" ${{ github.event.inputs.dry_run == 'true' && '--dry-run' || '' }} + run: python .github/scripts/music_data.py publish "$WORK/out" ${{ (github.event.inputs.dry_run == 'true' || vars.MUSIC_DATA_PUBLISH != 'true') && '--dry-run' || '' }} From 9b88ad163a5d17ba7e99a350a2f48fa0d62e7f36 Mon Sep 17 00:00:00 2001 From: luoxiadesu <249937048+luoxiadesu@users.noreply.github.com> Date: Wed, 30 Sep 2026 03:56:13 +0900 Subject: [PATCH 10/14] feat(jp): support Japanese release assets and split APKs --- .github/MUSIC_DATA.md | 8 +- .github/scripts/music_data.py | 10 +- .github/scripts/story_site.py | 29 +++ .github/scripts/test_jp_workflow.py | 52 +++++ .github/workflows/music-data.yml | 14 +- .github/workflows/story-site.yml | 12 +- README.en.md | 4 + README.md | 3 + docs/configuration.md | 6 + docs/jp.md | 123 ++++++++++++ docs/schema/music-data.schema.json | 5 + docs/stages.md | 3 +- src/nnnotes/addressables.py | 55 +++++- src/nnnotes/apkset.py | 77 ++++++++ src/nnnotes/catalog.py | 31 ++- src/nnnotes/catalogdb.py | 23 ++- src/nnnotes/cli.py | 45 ++++- src/nnnotes/cli_assets.py | 74 +++++-- src/nnnotes/config.py | 17 ++ src/nnnotes/configfile.py | 11 +- src/nnnotes/crikey.py | 7 +- src/nnnotes/crilips.py | 3 +- src/nnnotes/cristages.py | 11 +- src/nnnotes/deckdata.py | 12 +- src/nnnotes/gameapi.py | 26 ++- src/nnnotes/jp.py | 283 +++++++++++++++++++++++++++ src/nnnotes/liveaudio.py | 4 +- src/nnnotes/master.py | 28 ++- src/nnnotes/musicdata.py | 3 +- src/nnnotes/nnnotes.example.toml | 10 +- src/nnnotes/player.py | 9 +- src/nnnotes/storysite.py | 2 + src/nnnotes/web.py | 15 +- src/nnnotes/webmodel.py | 2 +- tests/test_configfile.py | 3 +- tests/test_jp.py | 287 ++++++++++++++++++++++++++++ 36 files changed, 1212 insertions(+), 95 deletions(-) create mode 100644 .github/scripts/test_jp_workflow.py create mode 100644 docs/jp.md create mode 100644 src/nnnotes/apkset.py create mode 100644 src/nnnotes/jp.py create mode 100644 tests/test_jp.py diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md index f8b22ab..bb9c26b 100644 --- a/.github/MUSIC_DATA.md +++ b/.github/MUSIC_DATA.md @@ -1,5 +1,9 @@ # Music data workflow (StarMoe) +JP: `MUSIC_DATA_MASTERDATA_REGION=jp` selects the JP package, catalog and client metadata, with +`jp/music-data` as the default output prefix. The marker includes the asset hash and the provenance gate rejects +a catalog from another snapshot. See [Japanese release](../docs/jp.md) for setup and validation limits. + `.github/workflows/music-data.yml` keeps the music data file of the chart data page (ournotes-player `examples/songs`) up to date: `nnnotes music-data` of the current master data ([docs/music-data.md](../docs/music-data.md): every song and chart with the deck model's statistics and the play scenarios), checked by quality gates and @@ -14,8 +18,8 @@ published into the story site's bucket under `music-data/` (`https://storage.bdo A run never deletes anything from the bucket. Its helper steps are `.github/scripts/music_data.py` (with the bucket, HTTP and master data helpers of `story_site.py`), `music_data_smoke.mjs`, `songs_page.sh` and `apk.sh`; the gate -self-test is `test_music_data.py`. Nothing outside `.github/` differs from upstream, so the fork syncs with it as -before. The Cloudflare Pages preview of the page is not part of the workflow. +self-tests are `test_music_data.py` and `test_jp_workflow.py`. The fork also contains JP support in `src/nnnotes`. +The Cloudflare Pages preview of the page is not part of the workflow. ## A run diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py index 0c336b5..b4ada9a 100755 --- a/.github/scripts/music_data.py +++ b/.github/scripts/music_data.py @@ -49,7 +49,7 @@ BUILD_FORMAT = "moenotes.music-data-build/1" # This script's own version of a build: bump it when what it builds or publishes changes, so that the next run builds # although the master data, the deck model and nnnotes are the same. -RECIPE = 1 +RECIPE = 2 FILE, MARKER, JACKETS, ARCHIVE = "music-data.json", "build.json", "jackets/", "archive/" MANIFEST = "MasterManifest.json" SOURCE_PATHS = ("src", "rust", "pyproject.toml") # nnnotes' code: the commit that last changed one of them @@ -58,7 +58,7 @@ ARCHIVE_CACHE = story_site.ASSET_CACHE # archive//.json: content-addressed FILE_CACHE = "no-cache" # music-data.json and build.json change in place JACKET_CACHE = "public, max-age=86400" -SNAPSHOT_KEYS = ("version", "resource_version", "client_version", "verified_at", "manifest_sha256", "table_count") +SNAPSHOT_KEYS = ("version", "resource_version", "resource_hash", "client_version", "verified_at", "manifest_sha256", "table_count") # the gates' bounds (MUSIC_DATA.md) SIZE_RATIO = (0.8, 2.0) # against the published file @@ -159,6 +159,7 @@ def inputs(entry: dict, root: Path = Path(".")) -> dict: """What a build is made of: the master data snapshot, the deck model, nnnotes and this script.""" return {"masterRegion": env("MASTERDATA_REGION"), "masterVersion": entry.get("version"), "resourceVersion": entry.get("resource_version"), "clientVersion": entry.get("client_version"), + "resourceHash": entry.get("resource_hash"), "deckCommit": deck_commit(root), "nnnotesCommit": nnnotes_commit(root), "recipe": RECIPE} @@ -220,6 +221,7 @@ def cmd_master(out: str) -> None: # ---------------------------------------------------------------- build def cmd_build(out: str) -> None: + story_site.configure_region() o = Path(out).resolve() o.mkdir(parents=True, exist_ok=True) nnnotes = [sys.executable, "-m", "nnnotes"] @@ -324,6 +326,10 @@ def gate_provenance(doc, ctx: Context, g: Gate): v = ctx.snapshot["entry"].get("version") if m.get("version") != v: g.fail(f"master.version {m.get('version')!r}, the snapshot's {v!r}") + if ctx.snapshot.get("region") == "jp": + for expected, actual in (("resource_version", "resourceVersion"), ("resource_hash", "resourceHash")): + if (p.get("catalog") or {}).get(actual) != ctx.snapshot["entry"].get(expected): + g.fail(f"JP catalog {actual} differs from the master snapshot") if ctx.master is not None: manifest = json.loads((ctx.master / MANIFEST).read_bytes()) if m.get("version") != manifest.get("version"): diff --git a/.github/scripts/story_site.py b/.github/scripts/story_site.py index 0bdc827..a220a94 100755 --- a/.github/scripts/story_site.py +++ b/.github/scripts/story_site.py @@ -137,6 +137,32 @@ def master_index() -> tuple[str, dict]: return base + region["path"], region +def configure_region() -> None: + """Keep build inputs from one release; JP client/API/CDN are read from the public snapshot.""" + region = env("MASTERDATA_REGION") + names = {"hk-tw-mo": "tw", "en": "en", "kr": "kr", "jp": "jp"} + if region not in names or env("NNNOTES_CATALOG_REGION") != names[region]: + sys.exit("story_site: master region and catalog region differ") + if region != "jp": + return + _, snapshot = master_index() + entry = snapshot.get("entry") or {} + assets = entry.get("assets") or {} + upstream = entry.get("upstream") or {} + api = assets.get("api_root") or upstream.get("api_root") + bundle = assets.get("bundle_root") + from urllib.parse import urlsplit + cdn = upstream.get("cdn_root") + if bundle: + u = urlsplit(bundle) + cdn = f"{u.scheme}://{u.netloc}" + client = entry.get("client_version") + if not all(isinstance(v, str) and v for v in (api, cdn, client)): + sys.exit("story_site: JP snapshot lacks API/CDN/client metadata") + os.environ.update(NNNOTES_SERVERS_JP_API=api, NNNOTES_SERVERS_JP_CDN=cdn, + NNNOTES_SERVERS_JP_CLIENT_VERSION=client, NNNOTES_SERVERS_JP_PROVIDER="jp") + + def fetch_table(url: str, name: str, digest: str, dest: Path) -> None: if not re.fullmatch(r"[A-Za-z0-9_-]+\.json", name): sys.exit(f"story_site: unexpected master file name {name!r}") @@ -215,10 +241,13 @@ def fetch(key: str) -> None: def cmd_build(site_dir: str, ids: list[str]) -> None: + configure_region() site = Path(site_dir).resolve() stories = parse_ids(" ".join(ids)) tmp = site.parent / f"{site.name}.tmp" nnnotes = [sys.executable, "-m", "nnnotes", "web", str(site), "--tmp", str(tmp)] + if env("MASTERDATA_REGION") == "jp": + nnnotes += ["--story-languages", "ja"] status = 0 if stories: cmd = nnnotes + [a for i in stories for a in ("--story", str(i))] diff --git a/.github/scripts/test_jp_workflow.py b/.github/scripts/test_jp_workflow.py new file mode 100644 index 0000000..6a1cf80 --- /dev/null +++ b/.github/scripts/test_jp_workflow.py @@ -0,0 +1,52 @@ +"""The workflow selects one release and derives JP endpoints without copying credentials.""" +import os +from pathlib import Path +import sys + +import pytest + +sys.path.insert(0, str(Path(__file__).parent)) +import story_site +import music_data + + +def test_configure_jp_from_public_snapshot(monkeypatch): + monkeypatch.setenv('MASTERDATA_REGION', 'jp') + monkeypatch.setenv('NNNOTES_CATALOG_REGION', 'jp') + entry = {'client_version': '1.0.4', 'assets': {'api_root': 'https://api.example.test', + 'bundle_root': 'https://cdn.example.test/asset/1.0.0.300/Android/' + 'a' * 32}} + monkeypatch.setattr(story_site, 'master_index', lambda: ('unused', {'entry': entry})) + for key in ('API', 'CDN', 'CLIENT_VERSION', 'PROVIDER'): + monkeypatch.setenv('NNNOTES_SERVERS_JP_' + key, '') + story_site.configure_region() + assert os.environ['NNNOTES_SERVERS_JP_CDN'] == 'https://cdn.example.test' + assert os.environ['NNNOTES_SERVERS_JP_API'] == 'https://api.example.test' + assert os.environ['NNNOTES_SERVERS_JP_CLIENT_VERSION'] == '1.0.4' + + +def test_reject_mixed_master_and_catalog(monkeypatch): + monkeypatch.setenv('MASTERDATA_REGION', 'jp') + monkeypatch.setenv('NNNOTES_CATALOG_REGION', 'tw') + with pytest.raises(SystemExit, match='differ'): + story_site.configure_region() + + +def test_asset_hash_changes_build_inputs(monkeypatch): + monkeypatch.setenv('MASTERDATA_REGION', 'jp') + monkeypatch.setattr(music_data, 'deck_commit', lambda root: 'a' * 40) + monkeypatch.setattr(music_data, 'nnnotes_commit', lambda root: 'b' * 40) + a = music_data.inputs({'version': 'v', 'resource_hash': 'a' * 32}) + b = music_data.inputs({'version': 'v', 'resource_hash': 'b' * 32}) + assert a != b and a['resourceHash'] == 'a' * 32 + + +def test_jp_catalog_provenance_matches_snapshot(): + from test_music_data import sample + doc = sample() + doc['provenance']['catalog']['resourceVersion'] = '1.0.0.300' + doc['provenance']['catalog']['resourceHash'] = 'a' * 32 + context = music_data.Context(snapshot={'region': 'jp', 'entry': { + 'version': doc['provenance']['master']['version'], 'resource_version': '1.0.0.300', 'resource_hash': 'b' * 32}}) + gate = music_data.Gate() + music_data.gate_provenance(doc, context, gate) + assert any('resourceHash differs' in s for s in gate.failures) diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml index d0c9478..7594233 100644 --- a/.github/workflows/music-data.yml +++ b/.github/workflows/music-data.yml @@ -35,14 +35,14 @@ permissions: # One run at a time: a run replaces music-data.json and build.json. concurrency: - group: music-data + group: music-data-${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} cancel-in-progress: false env: # The story site's bucket (its repository variables) and the key prefix of the music data. STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} - MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || 'music-data' }} + MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || (vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'jp/music-data' || 'music-data') }} # The publishing switch: anything but 'true' makes every run a dry run. MUSIC_DATA_PUBLISH: ${{ vars.MUSIC_DATA_PUBLISH || 'false' }} # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). @@ -53,12 +53,14 @@ env: PLAYER_REPOSITORY: ${{ vars.MUSIC_DATA_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} PLAYER_REF: ${{ vars.MUSIC_DATA_PLAYER_REF || '' }} PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} - APK_PACKAGE: ${{ vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + APK_PACKAGE: ${{ vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'com.bushiroad.sirius' || vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} PYTHON_VERSION: "3.13" WORK: ${{ github.workspace }}/../work # nnnotes settings (docs/configuration.md); the bundle key and the CDN come from secrets in the build step. - NNNOTES_CATALOG_REGION: tw - NNNOTES_CATALOG_LANGUAGE: zh-Hant + NNNOTES_CATALOG_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'jp' || 'tw' }} + NNNOTES_CATALOG_LANGUAGE: ${{ vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'ja' || 'zh-Hant' }} + NNNOTES_SERVERS_JP_NAME: JP + NNNOTES_SERVERS_JP_LANGUAGES: ja NNNOTES_SERVERS_TW_NAME: TW/HK/MO NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja @@ -115,7 +117,7 @@ jobs: - name: Gate self-test env: MUSIC_DATA_PAGE: ${{ env.WORK }}/page/examples/songs - run: python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + run: python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py .github/scripts/test_jp_workflow.py - name: APK env: diff --git a/.github/workflows/story-site.yml b/.github/workflows/story-site.yml index 5a933cd..30f90cf 100644 --- a/.github/workflows/story-site.yml +++ b/.github/workflows/story-site.yml @@ -34,14 +34,14 @@ permissions: # One run at a time: every run rewrites the site's indexes from the manifests it fetched. concurrency: - group: story-site + group: story-site-${{ vars.STORY_MASTERDATA_REGION || 'hk-tw-mo' }} cancel-in-progress: false env: # The bucket of the site (repository variables; the defaults are the StarMoe site). STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} - STORY_S3_PREFIX: ${{ vars.STORY_S3_PREFIX || '' }} + STORY_S3_PREFIX: ${{ vars.STORY_S3_PREFIX || (vars.STORY_MASTERDATA_REGION == 'jp' && 'jp' || '') }} # moenotes-masterdata-sync: decoded master data by region (index.json lists every table with its SHA-256). MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} MASTERDATA_REGION: ${{ vars.STORY_MASTERDATA_REGION || 'hk-tw-mo' }} @@ -51,12 +51,14 @@ env: PLAYER_REPOSITORY: ${{ vars.STORY_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} PLAYER_REF: ${{ vars.STORY_PLAYER_REF || '3774d8ac3987' }} PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} - APK_PACKAGE: ${{ vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + APK_PACKAGE: ${{ vars.STORY_MASTERDATA_REGION == 'jp' && 'com.bushiroad.sirius' || vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} PYTHON_VERSION: "3.13" WORK: ${{ github.workspace }}/../work # nnnotes settings (docs/configuration.md); the keys and the CDN come from secrets in the build step. - NNNOTES_CATALOG_REGION: tw - NNNOTES_CATALOG_LANGUAGE: zh-Hant + NNNOTES_CATALOG_REGION: ${{ vars.STORY_MASTERDATA_REGION == 'jp' && 'jp' || 'tw' }} + NNNOTES_CATALOG_LANGUAGE: ${{ vars.STORY_MASTERDATA_REGION == 'jp' && 'ja' || 'zh-Hant' }} + NNNOTES_SERVERS_JP_NAME: JP + NNNOTES_SERVERS_JP_LANGUAGES: ja NNNOTES_SERVERS_TW_NAME: TW/HK/MO NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja diff --git a/README.en.md b/README.en.md index 6906a25..3dd269d 100644 --- a/README.en.md +++ b/README.en.md @@ -1,5 +1,9 @@ # nnnotes?! +Japanese-release support includes anonymous version discovery, authenticated CDN downloads, gzip catalogs, +snapshot-isolated caches and split APKs. See the [JP guide](https://github.com/StarMoe-org/nnnotes/blob/main/docs/jp.md) +for configuration and validation limits. + [简体中文](https://github.com/MetaSekaiLab/nnnotes/blob/main/README.md) | [English](https://github.com/MetaSekaiLab/nnnotes/blob/main/README.en.md) nnnotes is an offline data toolkit for the game files of BanG Dream! Our Notes: it reads Addressables catalogs, diff --git a/README.md b/README.md index 96b94dc..ab28b70 100644 --- a/README.md +++ b/README.md @@ -52,6 +52,9 @@ nnnotes config check # 每项设置的来源和格式是否有效, ## 完成度 +日服现已支持匿名版本查询、CDN 认证下载、gzip catalog、资源路径与缓存隔离,以及 Android 分包读取。 +配置、命令和日服验证范围见 [日服支持](docs/jp.md)。 + 以下结果基于台服 1.0.1(zh-Hant)的全部数据: | 部分 | 状态 | diff --git a/docs/configuration.md b/docs/configuration.md index c98fbd7..af7515c 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -1,5 +1,7 @@ # Configuration +For Japanese-release endpoints, catalog snapshots and split APKs, see [Japanese release](jp.md). + nnnotes has no built-in keys, server addresses or data paths. Every setting comes from one of three sources, and a command that needs a setting nobody gave stops before it does any work. @@ -87,6 +89,10 @@ exist when they are given. | `[catalog] region` | `NNNOTES_CATALOG_REGION` | `--region` | name of one `[servers.]` table | | `[catalog] language` | `NNNOTES_CATALOG_LANGUAGE` | `--language` | catalog language: the `` of `catalog_main_.bin`: `ja`, `en`, `zh-Hant`, `zh-Hans` or `ko`; also the client language of `live`, `story` and `web` (the text field, fonts and line spacing of their UI) and of the model labels of `web --live2d` | | `[servers.] name` | `NNNOTES_SERVERS__NAME` | — | label of the region in `browse` (default: the region name) | +| `[servers.] provider` | `NNNOTES_SERVERS__PROVIDER` | — | `international` or `jp`; empty selects `jp` for the region named jp, international otherwise | +| `[servers.] client_version` | `NNNOTES_SERVERS__CLIENT_VERSION` | — | per-region API client version; overrides `[client] version`, then falls back to the region's APK versionName | +| `[servers.] apk` | `NNNOTES_SERVERS__APK` | `--apk` | per-region APK, APKS/XAPK or directory with base.apk; adjacent splits of base.apk are read automatically; the flag overrides it | +| `[servers.] catalog` | `NNNOTES_SERVERS__CATALOG` | `--catalog` | per-region catalog file; JP needs the matching `.source.json`; the flag overrides it | | `[servers.] cdn` | `NNNOTES_SERVERS__CDN` | — | CDN base URL of the region (a trailing `/` is ignored) | | `[servers.] languages` | `NNNOTES_SERVERS__LANGUAGES` | — | catalog languages `browse` lists: a TOML array of strings; comma-separated in the environment | | `[servers.] api` | `NNNOTES_SERVERS__API` | — | API root of the region: `https://host[:port]` (TLS, port 443 by default), `host[:port]`, or `http://host[:port]` for a plain-text local server; no path | diff --git a/docs/jp.md b/docs/jp.md new file mode 100644 index 0000000..b2e1672 --- /dev/null +++ b/docs/jp.md @@ -0,0 +1,123 @@ +# Japanese release + +JP uses the same catalog, bundle and master parsers, with a separate version discovery and CDN download path. +The anonymous `MasterdataService/Version` response provides the master version, Android asset version/hash and +CDN authentication. No player account or login is needed. + +## Configuration + +Use your own client addresses and encryption settings. As elsewhere in nnnotes, neither addresses nor keys are +built into the package. Add a region table with these settings to your private configuration: + +```toml +[catalog] +region = "jp" +language = "ja" + +[servers.jp] +provider = "jp" +name = "JP" +api = "https://YOUR_API_HOST" +cdn = "https://YOUR_CDN_HOST" +client_version = "YOUR_CURRENT_CLIENT_VERSION" +languages = ["ja"] +apk = "/path/to/jp/base.apk" +master = "/path/to/jp/master" + +[paths] +cache = "/path/to/cache" +``` + +Keep `[bundle] key/nonce_seed` and `[master] key/iv` in your configuration. Master keys are unnecessary when +reading already decoded tables with `music-data --decoded-master`. + +`provider` defaults to `jp` for the region named `jp`, and to `international` otherwise. An alias such as +`servers.japan` needs `provider = "jp"`. `client_version` overrides `[client] version`; when neither is set, +nnnotes reads the APK versionName. Use the current accepted JP client version: an old extracted APK can still +contain readable local data even when its version is no longer accepted by the API. + +The `cdn` setting is also the allowed CDN origin for authentication. A different origin in Version fails before +any authenticated download. The JP transport uses HTTPS, follows no redirects, keeps authentication in memory, +and refreshes it once on 401/403. A 429 does not cause an authentication-refresh loop. Credentials are not written +to catalog metadata or task files. + +## APKs and split data + +`apk` accepts: + +- a normal APK; +- `base.apk` with adjacent `split_*.apk` / `config.*.apk` files; +- a directory containing `base.apk` and its splits; +- an `.apks` or `.xapk` archive containing `base.apk` and its splits. + +JP stores boot data in base.apk and the embedded catalog/master/bundles in the Unity data split. Both are read +through the same APK-set interface. `datapack.unity3d` is loaded alongside `data.unity3d` for external resources +such as materials and shaders. Native libraries are also read from the splits. No repacked APK is required. + +`[servers.] apk` and `catalog` override their `[paths]` equivalents; explicit `--apk` / `--catalog` flags +still win. For mixed-release chart builds, configure separate master directories and APK sets. JP stories and +models must use a separate site directory from international releases, since their model IDs may overlap. + +## Commands + +```sh +nnnotes --region jp master version +nnnotes --region jp master download --latest -o work/jp-master-bin +nnnotes master decode work/jp-master-bin -o work/jp-master +nnnotes --region jp catalog --prefix Live/MusicScore/ --limit 10 +nnnotes --region jp pull Live/MusicScore/0001/0001_00 +nnnotes --region jp catalogs fetch +nnnotes --region jp export --select key:Image/Jacket/jacket_temporary --views none -o out/jp-sample +nnnotes --region jp music-data --decoded-master --no-bgm -o out/jp-music-data.json +``` + +For `--decoded-master`, keep `MasterManifest.json` with the decoded tables (copy it from the download directory, +or use a snapshot from your masterdata service). `--no-bgm` omits song audio length inspection; omit that flag to +include it. `master version` reports `resourceHash` for JP as well as the two version fields. `music-data` records +the JP hash in `provenance.catalog.resourceHash`. + +JP catalogs are `catalog_main.bin` in a version/hash directory, without a language suffix. The HTTP payload can +be gzip; nnnotes keeps its original bytes and SHA-256, and bounds decompression before parsing. Remote bundle and +CRI locations use `{Fwk.Resource.RemoteAssetDir}`. These are indexed, resolved and downloaded through the selected +snapshot, including in `browse`, `catalogs`, `export`, `plan` and `run-stage`. + +The cache lives under `/jp////`. It cannot reuse the +international catalog cache. Each cached `catalog_main.bin` has a `catalog_main.bin.source.json` with its SHA-256 +and public source metadata. For an offline `--catalog` or `catalogs import`, copy the pair together. Credentials +are reacquired only when a missing file must be downloaded. + +Catalog database records include this source metadata, and JP catalog identity includes version/hash even when +the catalog bytes remain the same. Old international records retain their previous IDs. Cached historical files +can be read offline; downloading a missing historical file fails if Version now selects a different snapshot. +Refresh the catalog to use the new snapshot. Historical local bundles require the APK set matching the imported +APK catalog. A JP master download likewise rejects a different current master snapshot. + +## Workflow integration + +For the StarMoe workflows, set `MUSIC_DATA_MASTERDATA_REGION=jp` or `STORY_MASTERDATA_REGION=jp`. This selects the +JP package and Japanese catalog language; the build reads API/CDN/client-version from the JP entry in the public +masterdata index. The default output prefixes become `jp/music-data` and `jp`, with separate concurrency groups. +Custom output prefixes must also be separate from the international outputs. Publishing remains controlled by +the existing workflow switches. No GitHub variable, secret, deployment or bucket is changed by installing nnnotes. + +The music build marker includes resource hash, so hash-only updates rebuild. Its provenance gate checks that the +catalog version/hash matches the master snapshot. A mixed master/catalog region is rejected before building. + +## Validation and current limits + +The synthetic tests use local gRPC and HTTP servers, including gzip catalogs, split packages, authentication +rotation, redirects, 429, source isolation, hash-only updates, offline replay and invalid paths. + +Live validation on 2026-09-30 used JP client 1.0.4 and local JP Android 1.0.3 data: + +- 238 master tables downloaded with manifest hashes and decoded successfully; +- a 73-note chart and a 512×512 jacket read successfully; +- one Live2D runtime model and one complete story exported; the story's 22 FLAC files passed full FFmpeg decoding; +- selected `export` produced six files without failed tasks; +- `music-data --decoded-master --no-bgm` built 85 songs / 340 charts, including the pinned deck model's + statistics, with no unplayable chart; output passed JSON Schema validation. + +This does not establish full JP story/Live2D/movie export coverage, browser playback, every BGM, or native scoring +parity. JP's text tables retain five language columns, but most translated values are empty; story web builds +default to Japanese. Existing Bili chat skin view rules still refer to international `_iconAssetPath` fields; +JP rows with a different schema report no-value and are not represented as verified skin mappings. diff --git a/docs/schema/music-data.schema.json b/docs/schema/music-data.schema.json index dad6a10..3d9221b 100644 --- a/docs/schema/music-data.schema.json +++ b/docs/schema/music-data.schema.json @@ -71,6 +71,11 @@ ], "description": "the resource version recorded for the catalog in the catalog store, null when none is recorded" }, + "resourceHash": { + "type": "string", + "pattern": "^[0-9a-f]{32}$", + "description": "JP Android asset directory hash, when the catalog source is JP" + }, "sha256": { "$ref": "#/$defs/sha256", "description": "SHA-256 of the remote catalog file the charts were read with" diff --git a/docs/stages.md b/docs/stages.md index 9ac98ee..92b437f 100644 --- a/docs/stages.md +++ b/docs/stages.md @@ -13,7 +13,8 @@ The asset export is built from three layers: ## Stages Every stage is at version 1, except `unity.census` (version 2: it lists the scripts animation clip bindings name) -and `cri.movie` (version 2: every stream kind, codec and channel of a USM). +and `cri.movie` (version 2: every stream kind, codec and channel of a USM), and `catalog.index` (version 2: +gzip input and Japanese remote asset placeholders). | Stage | Subject | Inputs | Outputs | |---|---|---|---| diff --git a/src/nnnotes/addressables.py b/src/nnnotes/addressables.py index 23ca07a..19c9dbd 100644 --- a/src/nnnotes/addressables.py +++ b/src/nnnotes/addressables.py @@ -15,6 +15,8 @@ from __future__ import annotations import hashlib +import gzip +import io import struct from dataclasses import dataclass, field from html import escape @@ -32,6 +34,23 @@ CATALOG_MAGIC = 0x0DE38942 CATALOG_VERSION = 2 NONE = 0xFFFFFFFF +REMOTE_PREFIX = "{Fwk.Resource.RemoteAssetDir}/" +MAX_CATALOG = 128 * 1024 * 1024 + + +def catalog_bytes(data: bytes) -> bytes: + """Normalize HTTP gzip or raw binary catalog bytes with a bounded expanded size.""" + if len(data) > MAX_CATALOG: + raise ValueError("catalog exceeds the size limit") + if data.startswith(b"\x1f\x8b"): + try: + with gzip.GzipFile(fileobj=io.BytesIO(data)) as stream: + data = stream.read(MAX_CATALOG + 1) + except (OSError, EOFError): + raise ValueError("invalid gzip catalog") from None + if len(data) > MAX_CATALOG: + raise ValueError("expanded catalog exceeds the size limit") + return data @dataclass(frozen=True) @@ -56,6 +75,9 @@ def decrypt(data: bytes, filename: str, key: BundleKey) -> bytes: def remote_path(internal_id: str) -> str | None: """The path below the CDN base of a remote location (an absolute URL: everything after its host), else None.""" + if internal_id.startswith(REMOTE_PREFIX): + from .jp import relative_path + return "/" + relative_path(internal_id[len(REMOTE_PREFIX):]) scheme, sep, rest = internal_id.partition("://") if not sep or not scheme.isalpha(): return None @@ -65,6 +87,7 @@ def remote_path(internal_id: str) -> str | None: def parse(data: bytes) -> list[dict]: """Every location of a binary catalog: {offset, primary_key, internal_id, dependencies (location offsets)}.""" + data = catalog_bytes(data) def u32(offset): return struct.unpack_from(" dict: def parse_header(data: bytes) -> dict: """The header of a binary catalog: magic, version, keysOffset, locatorId, instanceProvider, sceneProvider ({id, assembly, type, data}), initObjects (the same, the providers the catalog initializes), buildResultHash.""" + data = catalog_bytes(data) if len(data) < HEADER.size: raise ValueError("unsupported catalog format") magic, version, keys, locator, instance, scene, init, build = HEADER.unpack_from(data, 0) @@ -244,6 +268,7 @@ def parse_header(data: bytes) -> dict: def parse_keys(data: bytes) -> list[dict]: """The key table of a binary catalog in stored order (ContentCatalogData.ResourceLocator.KeyData {u32 key object, u32 location set}): {"key": value, "type": the key's type name, "locations": location offsets}.""" + data = catalog_bytes(data) header = parse_header(data) buf = _Buffer(data) out = [] @@ -267,6 +292,7 @@ def parse_locations(data: bytes) -> list[dict]: offsets), dependency hash (i32), extra data (an ObjectTypeData) and resource type (a TypeSerializer.Data). extra_data of bundle and raw file locations is an AssetBundleRequestOptions: {type, hash, bundleName, crc, bundleSize, timeout, redirectLimit, retryCount, flags and the flag bits by name}.""" + data = catalog_bytes(data) header = parse_header(data) buf = _Buffer(data) offsets: set[int] = set() @@ -309,6 +335,7 @@ class Region: label: str cdn: str = field(repr=False) languages: list[str] + config: object = field(default=None, repr=False) class Handler(BaseHTTPRequestHandler): @@ -329,6 +356,12 @@ def listing(self, label: str, rows) -> None: def catalog(self, region: Region, language: str): key = (region.name, language) if key not in self.server.catalogs: + if region.config is not None and region.config.provider(region.name) == "jp": + from .jp import open_catalog + cat = open_catalog(region.config, region.name, bundle_key=self.server.bundle_key) + self.server.catalogs[key] = browse(cat.entries) + self.server.jp_catalogs[key] = cat + return self.server.catalogs[key] catalog = self.server.cache / "catalogs" / region.name / f"catalog_main_{language}.bin" if not catalog.exists(): with urlopen(region.cdn + f"/asset/Android/{catalog.name}", timeout=60) as response: @@ -358,7 +391,13 @@ def do_GET(self): if language not in region.languages: self.send_error(404) return - bundles, files = self.catalog(region, language) + from .gameapi import GameApiError + from .config import ConfigError + try: + bundles, files = self.catalog(region, language) + except (GameApiError, ConfigError, ValueError): + self.send_error(502, "Catalog could not be loaded") + return base = f"/{quote(region.name)}/{quote(language)}/" if len(parts) > 2 and parts[2] == "download": ident = parts[3] if len(parts) == 4 else "" @@ -367,8 +406,17 @@ def do_GET(self): return entry = bundles[int(ident)] name = entry["internal_id"].rsplit("/", 1)[1] - with urlopen(region.cdn + remote_path(entry["internal_id"]), timeout=60) as response: - data = decrypt(response.read(), name, self.server.bundle_key) + try: + cat = getattr(self.server, "jp_catalogs", {}).get((region.name, language)) + if cat is not None: + from .catalog import Bundle + data = cat.fetch(Bundle(int(ident), entry["internal_id"], name, True)).read_bytes() + else: + with urlopen(region.cdn + remote_path(entry["internal_id"]), timeout=60) as response: + data = decrypt(response.read(), name, self.server.bundle_key) + except (GameApiError, ConfigError, ValueError): + self.send_error(502, "Resource could not be downloaded") + return self.send_response(200) self.send_header("Content-Type", "application/octet-stream") self.send_header("Content-Disposition", "attachment; filename*=UTF-8''" + quote(name)) @@ -404,6 +452,7 @@ def serve(regions: list[Region], bundle_key: BundleKey, cache: Path, port: int, server.bundle_key = bundle_key server.cache = Path(cache) server.catalogs = {} + server.jp_catalogs = {} print(f"nnnotes browse: serving on {host}:{port}", flush=True) try: server.serve_forever() diff --git a/src/nnnotes/apkset.py b/src/nnnotes/apkset.py new file mode 100644 index 0000000..1852e98 --- /dev/null +++ b/src/nnnotes/apkset.py @@ -0,0 +1,77 @@ +"""Read a base APK and its splits, an APK directory, or an APKS/XAPK archive. + +The base manifest wins over split manifests. Resource members are resolved across the set; +no merged APK is produced. Nested APKs are spooled so large asset packs need not stay in RAM. +""" +from __future__ import annotations + +import shutil +import tempfile +import zipfile +from contextlib import ExitStack +from pathlib import Path + + +class ApkSet: + def __init__(self, source): + self.source = source + self._stack = ExitStack() + self._members = {} + + def __enter__(self): + try: + self._open() + return self + except BaseException: + self._stack.close() + raise + + def _add(self, source): + archive = self._stack.enter_context(zipfile.ZipFile(source)) + for info in archive.infolist(): + self._members.setdefault(info.filename, (archive, info)) + + def _open(self): + if hasattr(self.source, "read"): + self._add(self.source) + return + path = Path(self.source) + if path.is_dir(): + base = path / "base.apk" + if not base.is_file(): + raise FileNotFoundError("APK directory has no base.apk") + paths = [base] + sorted(p for p in path.glob("*.apk") if p != base) + elif path.suffix.lower() in (".apks", ".xapk"): + outer = self._stack.enter_context(zipfile.ZipFile(path)) + names = [n for n in outer.namelist() if n.lower().endswith(".apk")] + bases = [n for n in names if Path(n).name == "base.apk"] + if len(bases) != 1: + raise ValueError("APK archive must contain exactly one base.apk") + for name in bases + sorted(n for n in names if n not in bases): + tmp = self._stack.enter_context(tempfile.SpooledTemporaryFile(max_size=16 * 1024 * 1024)) + with outer.open(name) as stream: + shutil.copyfileobj(stream, tmp) + tmp.seek(0) + self._add(tmp) + return + else: + paths = [path] + if path.name == "base.apk": + paths += sorted(p for p in path.parent.glob("*.apk") + if p.name.startswith(("split_", "config."))) + for apk in paths: + self._add(apk) + + def namelist(self): + return list(self._members) + + def infolist(self): + return [info for _, info in self._members.values()] + + def read(self, name): + name = name.filename if isinstance(name, zipfile.ZipInfo) else name + archive, info = self._members[name] + return archive.read(info) + + def __exit__(self, *args): + return self._stack.__exit__(*args) diff --git a/src/nnnotes/catalog.py b/src/nnnotes/catalog.py index 82da771..a1df558 100644 --- a/src/nnnotes/catalog.py +++ b/src/nnnotes/catalog.py @@ -11,11 +11,11 @@ import threading import urllib.parse import urllib.request -import zipfile from dataclasses import dataclass from pathlib import Path -from .addressables import BundleKey, decrypt, parse, parse_locations, remote_path +from .apkset import ApkSet +from .addressables import REMOTE_PREFIX, BundleKey, decrypt, parse, parse_locations, remote_path from .cache import write_atomic as _write_atomic from .config import ConfigError, apk_missing @@ -46,6 +46,10 @@ def location_kind(internal_id: str) -> str: def file_name(internal_id: str) -> str: """The name of a bundle or raw file location: its path below the platform directory (`Android/`), which for a bundle is its bare file name.""" + if internal_id.startswith(REMOTE_PREFIX): + from .jp import relative_path + path = relative_path(internal_id[len(REMOTE_PREFIX):]) + return path.rsplit("/", 1)[-1] if path.endswith(".bundle") else path rel = remote_path(internal_id) if rel is None: rel = internal_id[len(LOCAL_PREFIX):] if internal_id.startswith(LOCAL_PREFIX) else internal_id @@ -73,16 +77,18 @@ class Catalog: the first time it is needed (a ConfigError it raises is raised naming the file that needed the setting). """ - def __init__(self, catalog_bytes: bytes, cache_dir: Path, *, cdn=None, bundle_key=None, apk: Path | None = None): + def __init__(self, catalog_bytes: bytes, cache_dir: Path, *, cdn=None, bundle_key=None, apk: Path | None = None, + source=None, session=None): self._settings = {"cdn": cdn, "bundle_key": bundle_key} self.cache_dir = Path(cache_dir) self.cache_dir.mkdir(parents=True, exist_ok=True) self.apk = Path(apk) if apk else None + self.source, self.session = source, session self._sources = {"remote": catalog_bytes} self._locations: list[dict] | None = None self._parsed: tuple | None = None if self.apk is not None: - with zipfile.ZipFile(self.apk) as z: + with ApkSet(self.apk) as z: self._sources["apk"] = z.read(APK_CATALOG) def _parse(self) -> tuple: @@ -259,24 +265,33 @@ def fetch(self, b: Bundle) -> Path: if b.remote: url = self._url(b.internal_id) key = self._setting("bundle_key", b.name) # before the download: a missing key fails first - data = download(url) + data = self._download(url) else: if self.apk is None: raise apk_missing(f"bundle {b.name}") rel = b.internal_id[len(LOCAL_PREFIX):].lstrip("/") - with zipfile.ZipFile(self.apk) as z: + with ApkSet(self.apk) as z: data = z.read(APK_AA_DIR + rel) key = self._setting("bundle_key", b.name) if data[:7] != b"UnityFS" else None _write_atomic(dst, _unityfs(data, b.name, key)) return dst def _url(self, internal_id: str) -> str: + if self.source is not None: + return self.source.url(internal_id) name = internal_id.rsplit('/', 1)[-1] cdn = self._setting("cdn", name) if not cdn: raise RuntimeError(f"{name} not cached and no CDN base given") return cdn.rstrip("/") + remote_path(internal_id) + def _download(self, url): + if self.source is not None: + if self.session is None: + raise ConfigError("JP downloads require a configured JP session") + return self.session.get(url, source=self.source) + return download(url) + def fetch_key(self, key: str) -> list[Path]: """Every bundle of the key's closure; APK-local ones only when an APK is set.""" return [self.fetch(b) for b in self.resolve(key) if b.remote or self.apk] @@ -290,14 +305,14 @@ def fetch_raw(self, e: dict) -> Path: dst = self.cache_dir / "raw" / rel.lstrip("/") dst.parent.mkdir(parents=True, exist_ok=True) if not (dst.exists() and dst.stat().st_size > 0): - _write_atomic(dst, download(self._url(iid))) + _write_atomic(dst, self._download(self._url(iid))) return dst def apk_bundle(self, name_contains: str) -> Path: """An APK-local addressable bundle whose file name contains the substring.""" if self.apk is None: raise apk_missing("an APK bundle") - with zipfile.ZipFile(self.apk) as z: + with ApkSet(self.apk) as z: names = [n for n in z.namelist() if n.startswith(APK_AA_DIR) and n.endswith(".bundle") and name_contains in n.rsplit("/", 1)[1]] diff --git a/src/nnnotes/catalogdb.py b/src/nnnotes/catalogdb.py index 8ff6709..832a87c 100644 --- a/src/nnnotes/catalogdb.py +++ b/src/nnnotes/catalogdb.py @@ -27,9 +27,9 @@ import re import urllib.error import urllib.request -import zipfile from pathlib import Path +from .apkset import ApkSet from . import contract from .addressables import parse_header, parse_keys, parse_locations, remote_path from .catalog import APK_CATALOG, Catalog, file_name, location_kind @@ -294,9 +294,13 @@ def format_diff(d: dict) -> str: # ---------------------------------------------------------------- the store of versions -def version_id(remote_sha: str, apk_sha: str | None) -> str: +def version_id(remote_sha: str, apk_sha: str | None, source: dict | None = None) -> str: """The id of a catalog version: the key hash of its two catalogs' content ids.""" - return contract.digest({"remote": remote_sha, "apk": apk_sha}) + identity = {"remote": remote_sha, "apk": apk_sha} + if source is not None: + from .jp import Source + identity["source"] = Source.from_dict(source).to_dict() + return contract.digest(identity) def default_label(remote_sha: str, resource_version: str | None = None) -> str: @@ -307,7 +311,7 @@ def default_label(remote_sha: str, resource_version: str | None = None) -> str: def apk_catalog(apk) -> bytes: """The local catalog inside an APK.""" - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: return z.read(APK_CATALOG) @@ -361,12 +365,12 @@ def _put(self, data: bytes) -> dict: def add(self, remote: bytes, apk: bytes | None = None, *, label: str | None = None, region: str | None = None, language: str | None = None, hash_text: str | None = None, resource_version: str | None = None, - apk_version_name: str | None = None) -> dict: + apk_version_name: str | None = None, source: dict | None = None) -> dict: """Import a catalog pair (idempotent: a known pair gets the label and the facts it did not have yet); the version. A label names one version per region and language.""" remote_rec = self._put(bytes(remote)) apk_rec = self._put(bytes(apk)) if apk is not None else None - vid = version_id(remote_rec["sha256"], apk_rec["sha256"] if apk_rec else None) + vid = version_id(remote_rec["sha256"], apk_rec["sha256"] if apk_rec else None, source) versions = self._load() v = next((x for x in versions if x["id"] == vid), None) if v is None: @@ -374,11 +378,14 @@ def add(self, remote: bytes, apk: bytes | None = None, *, label: str | None = No "region": region, "language": language, "remote": remote_rec, "apk": apk_rec, "hash": None, "resourceVersion": None, "apkVersionName": None} versions.append(v) + if source is not None: + v["source"] = dict(source) for k, given in (("region", region), ("language", language), ("hash", hash_text), ("resourceVersion", resource_version), ("apkVersionName", apk_version_name)): if given is not None and v[k] is None: v[k] = given - label = label or default_label(remote_rec["sha256"], resource_version) + label = label or (f"{source['version']}/{source['hash']}" if source else + default_label(remote_rec["sha256"], resource_version)) for other in versions: if other is not v and label in other["labels"] and (other["region"], other["language"]) == ( v["region"], v["language"]): @@ -431,7 +438,7 @@ class IndexStage(Stage): """catalog.index: a catalog pair -> its index. Subjects and inputs come from the fact "catalogs": {subject: [Input "remote", Input "apk" (optional)]} (CatalogDB.inputs). Artifact "#index".""" name = "catalog.index" - version = 1 + version = 2 def subjects(self, env) -> list[str]: return sorted(env.fact("catalogs")) diff --git a/src/nnnotes/cli.py b/src/nnnotes/cli.py index b14ead1..4207a87 100644 --- a/src/nnnotes/cli.py +++ b/src/nnnotes/cli.py @@ -62,6 +62,7 @@ from .catalog import Catalog from .compress import DEFAULT_ENCODING, ENCODINGS from .config import DEFAULT_FILE, ENV_CONFIG, Config, ConfigError, config_files, find_file, use, user_file +from .gameapi import GameApiError from .jsonio import dumps, write_json from .webaudio import DEFAULT_AUDIO_FORMAT, WEB_AUDIO @@ -86,8 +87,11 @@ # ---------------------------------------------------------------- settings -> data def load_config(args) -> Config: overrides = {k: getattr(args, dest, None) for k, (dest, _) in FLAG_SETTINGS.items()} - return use(Config.load(getattr(args, "config", None), overrides=overrides, - flags={k: flag for k, (_, flag) in FLAG_SETTINGS.items()})) + cfg = Config.load(getattr(args, "config", None), overrides=overrides, + flags={k: flag for k, (_, flag) in FLAG_SETTINGS.items()}) + if cfg.provider() == "jp" and not cfg.has("catalog", "language"): + cfg = cfg.for_region(cfg.region()) + return use(cfg) def bundle_key(cfg: Config) -> BundleKey: @@ -96,6 +100,8 @@ def bundle_key(cfg: Config) -> BundleKey: def _existing(cfg: Config, section: str, key: str, kind: str = "file") -> Path | None: p = cfg.path(section, key) + if key == "apk" and p is not None and p.is_dir(): + return p if p is not None and not (p.is_file() if kind == "file" else p.is_dir()): raise ConfigError(f"setting {section}.{key}: {kind} {p} not found") return p @@ -107,9 +113,15 @@ def open_catalog(cfg: Config, bundles: bool = True, region: str | None = None) - serves them all. `bundles`: bundles will be fetched; else only the catalog is read. The region, its CDN base and the bundle key are read from the settings only when something must be downloaded (every file in the cache: none of them is needed), then a missing one is a ConfigError naming the setting.""" + if region: + cfg = cfg.for_region(region) cache = cfg.require_path("paths", "cache") catbin = _existing(cfg, "paths", "catalog") apk = _existing(cfg, "paths", "apk") + if cfg.provider() == "jp": + from .jp import open_catalog as jp_catalog + return jp_catalog(cfg, cfg.region(), catalog_file=catbin, + bundle_key=(lambda: bundle_key(cfg)) if bundles else None, apk=apk) language = cfg.require("catalog", "language") if catbin is None else None cdn = (lambda: cfg.cdn(region or cfg.region())) if bundles or catbin is None else None key = (lambda: bundle_key(cfg)) if bundles else None @@ -126,8 +138,10 @@ def master_dir(cfg: Config, region: str | None = None) -> Path: return _existing(cfg, section, key, "directory") -def player_data(cfg: Config): +def player_data(cfg: Config, region: str | None = None): from .player import PlayerData + if region: + cfg = cfg.for_region(region) cfg.require_path("paths", "apk") return PlayerData(_existing(cfg, "paths", "apk")) @@ -183,7 +197,7 @@ def cmd_browse(args, cfg): langs = cfg.get_list(f"servers.{r}", "languages") if not langs: raise cfg.missing(f"servers.{r}", "languages") - regions.append(Region(r, cfg.get(f"servers.{r}", "name") or r, cfg.cdn(r), langs)) + regions.append(Region(r, cfg.get(f"servers.{r}", "name") or r, cfg.cdn(r), langs, cfg)) serve(regions, bundle_key(cfg), cfg.require_path("paths", "cache"), args.port, args.host) @@ -368,7 +382,10 @@ def cmd_master_version(args, cfg): v = gameapi.master_version(cfg, region) except gameapi.GameApiError as e: sys.exit(f"nnnotes: {e}") - _print_json({"region": region, "masterVersion": v.version, "resourceVersion": v.resource_version}) + result = {"region": region, "masterVersion": v.version, "resourceVersion": v.resource_version} + if v.resource_hash is not None: + result["resourceHash"] = v.resource_hash + _print_json(result) def cmd_master_download(args, cfg): @@ -376,8 +393,16 @@ def cmd_master_download(args, cfg): region = cfg.region() cdn = cfg.cdn(region) try: - version = gameapi.master_version(cfg, region).version if args.latest else args.version - r = master.download(cdn, version, Path(args.out), workers=args.workers) + if cfg.provider(region) == "jp": + from .jp import Session, master_version + session = Session(cfg, region) + observation = session.observe() + version = observation.version.version if args.latest else master_version(args.version) + r = master.download(observation.cdn, version, Path(args.out), workers=args.workers, + get=lambda url: session.get(url, master=version), strict=True) + else: + version = gameapi.master_version(cfg, region).version if args.latest else args.version + r = master.download(cdn, version, Path(args.out), workers=args.workers) except (gameapi.GameApiError, master.DownloadError) as e: sys.exit(f"nnnotes: {e}") _print_json(r) @@ -590,6 +615,10 @@ def cmd_web(args, cfg): web.check_player(player) regions = (web.site_regions(cfg, args.web_regions, args.all_regions) if args.web_regions or args.all_regions else None) # None: the one [catalog] region + if (stories or models) and regions and len({cfg.provider(r) for r in regions}) > 1: + raise ConfigError("build JP stories/models in a separate site directory from international releases") + if (stories or models) and regions and cfg.provider(regions[0]) == "jp": + cfg = use(cfg.for_region(regions[0])) base = {"region": regions[0]} if regions else {} unknown = web.unknown_pairs(cfg, args.pair, regions) if args.pair else [] if unknown: @@ -967,6 +996,8 @@ def main(argv=None): except ConfigError as e: print(f"nnnotes: {e}", file=sys.stderr) sys.exit(2) + except GameApiError as e: + sys.exit(f"nnnotes: {e}") if __name__ == "__main__": diff --git a/src/nnnotes/cli_assets.py b/src/nnnotes/cli_assets.py index d89d910..9ae69ce 100644 --- a/src/nnnotes/cli_assets.py +++ b/src/nnnotes/cli_assets.py @@ -49,6 +49,7 @@ import json import sys import threading +import zipfile from collections import Counter from collections.abc import Mapping from concurrent.futures import ThreadPoolExecutor @@ -57,6 +58,7 @@ from pathlib import Path from types import SimpleNamespace +from .apkset import ApkSet from . import contract from .config import ConfigError, describe as describe_setting, usable_cpus, use from .contract import Cost, IncompatibleTask, Input, Task @@ -289,8 +291,22 @@ def catalog(self, vid: str): raise FileNotFoundError(f"catalog version {vid[:12]} is not in the store") remote, apk = db.catalog_bytes(v) cfg, region = self.cfg, v.get("region") - cat = Catalog(remote, self.cache, cdn=lambda: cfg.cdn(region or cfg.region()), - bundle_key=lambda: _bundle_key(cfg), apk=cfg.path("paths", "apk")) + region = region or cfg.get("catalog", "region") + if region: + cfg = cfg.for_region(region) + source, session, cache = None, None, self.cache + if v.get("source") is not None: + from .jp import Source, Session + source = Source.from_dict(v["source"]) + session = Session(cfg, cfg.region()) + cache = source.cache_dir(cache) + cat = Catalog(remote, cache, cdn=lambda: cfg.cdn(cfg.region()), + bundle_key=lambda: _bundle_key(cfg), apk=cfg.path("paths", "apk"), + source=source, session=session) + if source is not None and apk is not None and cat.sources().get("apk") != apk: + raise ConfigError("JP historical catalog needs the APK set it was imported with") + # Replay the APK catalog imported with this version, not today's catalog offsets. + cat._sources = {"remote": remote, **({"apk": apk} if apk is not None else {})} hit = self._catalogs[vid] = (cat, catalogdb.by_id(catalogdb.index(remote, apk))) return hit @@ -305,24 +321,23 @@ def fetch_location(self, vid: str, lid: str) -> Path: if loc["kind"] == "bundle": return cat.fetch(Bundle(0, iid, file_name(iid), remote_path(iid) is not None)) if remote_path(iid) is None: - return self._apk_file(iid) + return self._apk_file(iid, apk=cat.apk, cache=cat.cache_dir) return cat.fetch_raw({"internal_id": iid}) - def _apk_file(self, internal_id: str) -> Path: + def _apk_file(self, internal_id: str, *, apk=None, cache=None) -> Path: """A raw file of the APK as stored, through the cache (raw/, where a CDN file of that path would be); KeyError when the APK does not hold it.""" - import zipfile from .cache import write_atomic from .catalog import APK_AA_DIR if self.cache is None: raise self.cfg.missing("paths", "cache") - apk = self.cfg.path("paths", "apk") + apk = apk or self.cfg.path("paths", "apk") if apk is None: raise self.cfg.missing("paths", "apk") rel = apk_rel(internal_id) - dst = self.cache / "raw" / rel + dst = (cache or self.cache) / "raw" / rel if not (dst.is_file() and dst.stat().st_size > 0): - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: data = z.read(APK_AA_DIR + rel) dst.parent.mkdir(parents=True, exist_ok=True) write_atomic(dst, data) @@ -447,6 +462,9 @@ def __init__(self, args, cfg, common, log=_log): self.calibration = load_calibration(self.root) self.scheduled = calibrated(self.installed.stages, self.calibration) # the stages with calibrated costs self.version, self.remote, self.apk_catalog = self._catalog_version() + if self.version.get("region"): + self.cfg = cfg.for_region(self.version["region"]) + self.apk = self.cfg.path("paths", "apk") self._index = None self._problems: list = [] self._apk_names: set | None = None @@ -473,11 +491,13 @@ def _catalog_version(self): cat = self.common.open_catalog(self.cfg, bundles=False) src = cat.sources() remote, apk = src["remote"], src.get("apk") - vid = catalogdb.version_id(contract.sha256(remote), contract.sha256(apk) if apk is not None else None) + source = cat.source.to_dict() if getattr(cat, "source", None) else None + vid = catalogdb.version_id(contract.sha256(remote), contract.sha256(apk) if apk is not None else None, source) v = next((x for x in self.db.versions() if x["id"] == vid), None) if v is None: v = self.db.add(remote, apk, region=self.cfg.get("catalog", "region"), - language=self.cfg.get("catalog", "language"), apk_version_name=_apk_version(self.apk)) + language=self.cfg.get("catalog", "language"), apk_version_name=_apk_version(self.apk), + source=source, resource_version=source["version"] if source else None) return v, remote, apk def catalogs_fact(self) -> dict: @@ -661,17 +681,22 @@ def problems(self) -> list[dict]: def _cache_rel(self, kind: str, entry: dict, loc: dict) -> str: from .addressables import remote_path + prefix = "" + if self.version.get("source"): + from .jp import Source + prefix = Source.from_dict(self.version["source"]).cache_dir(Path(".")).as_posix() + "/" if kind == "bundle": - return f"bundles/{entry['name']}" + return prefix + f"bundles/{entry['name']}" rel = remote_path(loc["internalId"]) if rel is None: # a raw file of the APK (CatalogFetcher._apk_file) rel = apk_rel(loc["internalId"]) - return "raw/" + rel.lstrip("/") + return prefix + "raw/" + rel.lstrip("/") def _inputs(self, kind: str, entries: list[dict], fetch: bool) -> dict: from .catalogdb import by_id locs = by_id(self.index()) memo_kind = "bundles" if kind == "bundle" else "raw" + memo_prefix = self.version["id"] + ":" if self.version.get("source") else "" out, todo = {}, [] for e in entries: loc = locs[e["location"]] @@ -682,7 +707,7 @@ def _inputs(self, kind: str, entries: list[dict], fetch: bool) -> dict: if path is not None and path.is_file() and path.stat().st_size > 0: todo.append((e, locators, path)) continue - known = self.store.named(memo_kind, e["name"]) if e["name"] != e["stable"] else None + known = self.store.named(memo_kind, memo_prefix + e["name"]) if e["name"] != e["stable"] else None if known is not None: # a hashed file name: its recorded identity holds out[e["stable"]] = Input(kind, known[0], known[1], e["name"], tuple(locators)) elif not e["remote"] and self.apk is None: @@ -701,7 +726,7 @@ def one(job): try: if path is None: path = self.fetcher.fetch_location(self.version["id"], e["location"]) - sha, size = self.store.identify(path, memo_kind, e["name"]) + sha, size = self.store.identify(path, memo_kind, memo_prefix + e["name"]) except ConfigError: raise except KeyError as x: # not in the APK @@ -726,9 +751,8 @@ def _in_apk(self, internal_id: str) -> bool: """Whether the APK holds the file of a location (read from its directory, not fetched).""" from .catalog import APK_AA_DIR if self._apk_names is None: - import zipfile try: - with zipfile.ZipFile(self.apk) as z: + with ApkSet(self.apk) as z: self._apk_names = set(z.namelist()) except (OSError, zipfile.BadZipFile): raise ConfigError(f"setting paths.apk: {self.apk} is not a readable APK") from None @@ -1309,7 +1333,7 @@ def _apk_catalog(args, cfg) -> bytes | None: apk = cfg.path("paths", "apk") if apk is None: return None - if not apk.is_file(): + if not apk.exists(): raise ConfigError(f"setting paths.apk: file {apk} not found") return apk_catalog(apk) @@ -1320,10 +1344,14 @@ def cmd_catalogs_import(args, cfg, common): except OSError as e: args.usage(f"{args.file}: {e.strerror or e}") try: + source = None + if cfg.provider() == "jp": + from .jp import read_source + source = read_source(args.file, remote).to_dict() v = _db(args, cfg).add(remote, _apk_catalog(args, cfg), label=args.label, region=cfg.get("catalog", "region"), language=cfg.get("catalog", "language"), resource_version=args.resource_version, - apk_version_name=_apk_version(cfg.path("paths", "apk"))) + apk_version_name=_apk_version(cfg.path("paths", "apk")), source=source) except ValueError as e: args.usage(str(e)) print(_version_line(v)) @@ -1332,6 +1360,16 @@ def cmd_catalogs_import(args, cfg, common): def cmd_catalogs_fetch(args, cfg, common): from . import catalogdb region = cfg.region() + if cfg.provider(region) == "jp": + from .jp import open_catalog + cat = open_catalog(cfg, region, apk=cfg.path("paths", "apk")) + src = cat.sources() + v = _db(args, cfg).add(src["remote"], src.get("apk"), label=args.label, region=region, + language=cfg.get("catalog", "language") or "ja", source=cat.source.to_dict(), + resource_version=cat.source.version, + apk_version_name=_apk_version(cfg.path("paths", "apk"))) + print(_version_line(v)) + return cdn = cfg.cdn(region) language = cfg.require("catalog", "language") try: diff --git a/src/nnnotes/config.py b/src/nnnotes/config.py index 50077bd..c0da951 100644 --- a/src/nnnotes/config.py +++ b/src/nnnotes/config.py @@ -202,6 +202,10 @@ def get_list(self, section: str, key: str) -> list[str]: return list(v) def path(self, section: str, key: str) -> Path | None: + if section == "paths" and key in ("apk", "catalog") and self.origin(section, key) != "flag": + region = self.get("catalog", "region") + if region and self.has(f"servers.{region}", key): + section = f"servers.{region}" v, origin = self._raw(section, key) if v is None: return None @@ -245,6 +249,19 @@ def regions(self) -> list[str]: def region(self) -> str: return self.require("catalog", "region") + def provider(self, region: str | None = None) -> str: + region = region or self.get("catalog", "region") + value = self.get(f"servers.{region}", "provider") or ("jp" if region == "jp" else "international") + if value not in ("jp", "international"): + raise ConfigError(f"setting servers.{region}.provider: must be jp or international") + return value + + def for_region(self, region: str) -> "Config": + overrides = {**self._over, ("catalog", "region"): region} + if self.provider(region) == "jp": + overrides[("catalog", "language")] = "ja" + return Config(self._data, self._base, self.source, self._env, overrides, self._flags) + def cdn(self, region: str) -> str: """CDN base of a region, without a trailing slash.""" return self.require(f"servers.{region}", "cdn").rstrip("/") diff --git a/src/nnnotes/configfile.py b/src/nnnotes/configfile.py index a879f55..2bcb8e2 100644 --- a/src/nnnotes/configfile.py +++ b/src/nnnotes/configfile.py @@ -28,6 +28,7 @@ LINKS = ("auto", "clone", "hard", "copy") FILES = {("paths", "catalog"), ("paths", "apk"), ("paths", "ffmpeg"), ("paths", "vgmstream"), ("paths", "node")} DIRS = {("paths", "master"), ("paths", "player"), (REGION, "master")} +REGION_PATHS = {"apk", "catalog"} CHECK = "nnnotes.config-check/1" PATHS = "nnnotes.config-path/1" PROBLEMS = ("invalid", "not found", "unknown") @@ -59,7 +60,8 @@ def _template_settings() -> list[Setting]: elif (m := _KEYVAL.match(line)) and section: sec = REGION if section == f"servers.{TEMPLATE_REGION}" else section key, value = m.group(2), m.group(3).strip() - kind = "list" if value.startswith("[") else "path" if sec == "paths" or (sec, key) in DIRS else "string" + kind = "list" if value.startswith("[") else "path" if (sec == "paths" or (sec, key) in DIRS + or sec == REGION and key in REGION_PATHS) else "string" out.append(Setting(sec, key, kind, (sec, key) in SECRETS, " ".join(comment))) comment = [] for k in FONT_KEYS: @@ -120,6 +122,8 @@ def validate(setting: Setting, value) -> None: u = urlsplit(value) if u.scheme not in ("https", "http") or not u.netloc: raise ValueError("must be an http(s):// URL") + elif k == (REGION, "provider") and value not in ("jp", "international"): + raise ValueError("must be jp or international") elif k in ((REGION, "api"), ("bootstrap", "api")): from .gameapi import channel_target channel_target(value) @@ -287,7 +291,10 @@ def _status(cfg: Config, s: Setting, section: str) -> tuple[str, str | None]: validate(s, cfg.get_list(section, s.key)) elif s.kind == "path": p = cfg.path(section, s.key) - if (s.section, s.key) in FILES or s.section == FONTS: + if s.key == "apk": + if not p.exists(): + return "not found", "no such APK or directory" + elif (s.section, s.key) in FILES or s.section == FONTS or (s.section, s.key) == (REGION, "catalog"): if not p.is_file(): return "not found", "no such file" elif (s.section, s.key) in DIRS and not p.is_dir(): diff --git a/src/nnnotes/crikey.py b/src/nnnotes/crikey.py index bff86c1..c03157b 100644 --- a/src/nnnotes/crikey.py +++ b/src/nnnotes/crikey.py @@ -13,17 +13,18 @@ import io import re import struct -import zipfile from pathlib import Path +from .apkset import ApkSet + DATA_IN_APK = "assets/bin/Data/data.unity3d" def _load_data(src: Path): import UnityPy # here, not at import: the modules that import this one seldom call it src = Path(src) - if src.suffix.lower() in (".apk", ".zip"): - with zipfile.ZipFile(src) as z: + if src.is_dir() or src.suffix.lower() in (".apk", ".zip", ".apks", ".xapk"): + with ApkSet(src) as z: return UnityPy.load(io.BytesIO(z.read(DATA_IN_APK))) return UnityPy.load(str(src)) diff --git a/src/nnnotes/crilips.py b/src/nnnotes/crilips.py index 15000ff..1516903 100644 --- a/src/nnnotes/crilips.py +++ b/src/nnnotes/crilips.py @@ -24,6 +24,7 @@ from dataclasses import dataclass, field from pathlib import Path +from .apkset import ApkSet from .jsonio import write_json LIB_ENTRY_SUFFIX = "lib/arm64-v8a/libcri_lips_unity.so" @@ -190,7 +191,7 @@ def _scan_const(rodata: bytes, c: Const) -> bytes: def _library_in(apk: Path) -> bytes | None: try: - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: for info in z.infolist(): if info.filename.endswith(LIB_ENTRY_SUFFIX): return z.read(info.filename) diff --git a/src/nnnotes/cristages.py b/src/nnnotes/cristages.py index c627f8c..1f8f1ef 100644 --- a/src/nnnotes/cristages.py +++ b/src/nnnotes/cristages.py @@ -37,10 +37,10 @@ import subprocess import tempfile import threading -import zipfile from collections import defaultdict from pathlib import Path +from .apkset import ApkSet from . import contract from .atoms import impl_id from .contract import Cost, IncompatibleTask, Input @@ -72,14 +72,19 @@ def boot_input(store, apk) -> Input | None: if apk is None: return None apk = Path(apk) - if apk.suffix.lower() not in (".apk", ".zip"): + if not apk.is_dir() and apk.suffix.lower() not in (".apk", ".zip", ".apks", ".xapk"): sha, size = store.identify(apk) return Input("boot", sha, size, apk.name, ({"kind": "file", "path": str(apk.resolve())},)) + if apk.is_dir(): + with ApkSet(apk) as z: + data = z.read(BOOT_IN_APK) + sha = store.put(data) + return Input("boot", sha, len(data), "data.unity3d", ({"kind": "store"},)) apk_sha, _ = store.identify(apk) name = f"boot:{apk_sha}" known = store.named("files", name) if known is None or not store.has(*known): - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: data = z.read(BOOT_IN_APK) known = store.identify(store.path(store.put(data)), "files", name) return Input("boot", known[0], known[1], "data.unity3d", ({"kind": "store"},)) diff --git a/src/nnnotes/deckdata.py b/src/nnnotes/deckdata.py index 4cd6905..180db4b 100644 --- a/src/nnnotes/deckdata.py +++ b/src/nnnotes/deckdata.py @@ -26,6 +26,8 @@ from collections import Counter from dataclasses import dataclass from pathlib import Path + +from .apkset import ApkSet from typing import Callable DECK_FORMAT = "nnnotes.deck-data/1" # the deck model's input format (ournotes-deck data::FORMAT) @@ -188,7 +190,7 @@ def apk_master(apk) -> MasterSource: """The master data files the APK ships (`assets/Master/`: MasterManifest.json + .bin).""" apk = Path(apk) try: - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: raw = z.read(APK_MASTER + MANIFEST) except KeyError: raise DeckDataError(f"{apk}: no {APK_MASTER}{MANIFEST}") from None @@ -197,7 +199,7 @@ def apk_master(apk) -> MasterSource: version, hashes = _manifest(raw, f"{apk} {APK_MASTER}") def read(names): - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: have = set(z.namelist()) missing = [n for n in names if APK_MASTER + n not in have] if missing: @@ -369,7 +371,7 @@ def apk_client(apk) -> dict: """{versionName, versionCode} of an APK's AndroidManifest.xml (None when absent).""" from .player import MANIFEST_IN_APK, manifest_version_code, manifest_version_name try: - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: data = z.read(MANIFEST_IN_APK) if MANIFEST_IN_APK in z.namelist() else None except (zipfile.BadZipFile, OSError): raise DeckDataError(f"{apk}: not a readable APK") from None @@ -381,7 +383,9 @@ def apk_client(apk) -> dict: def catalog_info(cat, store_root=None) -> dict: """{resourceVersion, sha256} of a catalog's remote catalog file (resource_version).""" sha = hashlib.sha256(cat.sources()["remote"]).hexdigest() - return {"resourceVersion": resource_version(store_root, sha), "sha256": sha} + source = getattr(cat, "source", None) + return {"resourceVersion": source.version if source else resource_version(store_root, sha), "sha256": sha, + **({"resourceHash": source.hash} if source else {})} def resource_version(store_root, remote_sha: str) -> str | None: diff --git a/src/nnnotes/gameapi.py b/src/nnnotes/gameapi.py index 45cc300..77b61f7 100644 --- a/src/nnnotes/gameapi.py +++ b/src/nnnotes/gameapi.py @@ -28,6 +28,7 @@ import grpc +from .apkset import ApkSet from .config import Config, ConfigError, describe VERSION_METHOD = "/app.masterdata.MasterdataService/Version" @@ -111,6 +112,7 @@ def strings(buf: bytes, names: dict[int, str]) -> dict[str, str]: class MasterVersion: version: str # master data version resource_version: str + resource_hash: str | None = None @dataclass(frozen=True) @@ -180,7 +182,7 @@ def channel_target(root: str) -> tuple[str, bool]: def call(root: str, method: str, request: bytes, client_version: str, *, timeout: float = TIMEOUT, - attempts: int = ATTEMPTS, setting: str = "the API root") -> bytes: + attempts: int = ATTEMPTS, setting: str = "the API root", response_metadata: dict | None = None) -> bytes: """One unary call with raw bytes; returns the response message's bytes. `setting` names the root in errors. Raises GameApiError.""" target, tls = channel_target(root) @@ -192,7 +194,11 @@ def call(root: str, method: str, request: bytes, client_version: str, *, timeout for i in range(attempts): last = i + 1 == attempts try: - body = stub(request, timeout=timeout, metadata=metadata(client_version)) + body, response = stub.with_call(request, timeout=timeout, metadata=metadata(client_version)) + if response_metadata is not None: + response_metadata.clear() + response_metadata.update(response.initial_metadata() or ()) + response_metadata.update(response.trailing_metadata() or ()) except grpc.RpcError as e: code = e.code() if code in RETRY_STATUS and not last: @@ -256,23 +262,26 @@ def apk_version_name(apk: Path) -> str | None: """`versionName` of the APK's AndroidManifest.xml (None when the APK has none).""" from .player import MANIFEST_IN_APK, manifest_version_name try: - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: data = z.read(MANIFEST_IN_APK) if MANIFEST_IN_APK in z.namelist() else None except (zipfile.BadZipFile, OSError): raise ConfigError(f"setting paths.apk: {apk} is not a readable APK") from None return manifest_version_name(data) if data is not None else None -def client_version(cfg: Config) -> str: +def client_version(cfg: Config, region: str | None = None) -> str: """The `x-client-version` of calls: `[client] version`, else the versionName of `[paths] apk`.""" - v = cfg.get("client", "version") + region = region or cfg.get("catalog", "region") + if region: + cfg = cfg.for_region(region) + v = cfg.get(f"servers.{region}", "client_version") or cfg.get("client", "version") what = "setting client.version" if v is None: apk = cfg.path("paths", "apk") if apk is None: raise ConfigError(f"the game API needs the client version: give it as {describe('client', 'version')}, " f"or give base.apk as {describe('paths', 'apk', '--apk')} to use its versionName") - if not apk.is_file(): + if not apk.exists(): raise ConfigError(f"setting paths.apk: file {apk} not found") v = apk_version_name(apk) if not v: @@ -288,8 +297,11 @@ def client_version(cfg: Config) -> str: def master_version(cfg: Config, region: str, *, timeout: float = TIMEOUT) -> MasterVersion: """The master data version region `region` serves now (its `[servers.] api`).""" section = f"servers.{region}" + if cfg.provider(region) == "jp": + from .jp import Session + return Session(cfg, region).observe(timeout=timeout).version root = api_root(cfg, section) - return fetch_master_version(root, client_version(cfg), timeout=timeout, setting=f"[{section}] api") + return fetch_master_version(root, client_version(cfg, region), timeout=timeout, setting=f"[{section}] api") def server_list(cfg: Config, *, timeout: float = TIMEOUT) -> list[Server]: diff --git a/src/nnnotes/jp.py b/src/nnnotes/jp.py new file mode 100644 index 0000000..ebdcbde --- /dev/null +++ b/src/nnnotes/jp.py @@ -0,0 +1,283 @@ +"""Japanese release: anonymous Version discovery and snapshot-bound CDN downloads. + +Addresses and client versions come from Config. Credentials remain in a Session, never in +source metadata, cache records, task descriptions or exception messages. +""" +from __future__ import annotations + +import base64 +import hashlib +import http.client +import json +import re +import threading +import time +from dataclasses import asdict, dataclass, field +from pathlib import Path +from urllib.parse import quote, unquote, urlsplit + +from . import gameapi +from .config import ConfigError + +PLACEHOLDER = "{Fwk.Resource.RemoteAssetDir}/" +CATALOG_LIMIT = 64 * 1024 * 1024 +FILE_LIMIT = 1024 * 1024 * 1024 + + +def numeric_version(value): + if not isinstance(value, str) or not re.fullmatch(r"[0-9]+(?:\.[0-9]+){1,3}", value): + raise ValueError("invalid numeric version") + parts = tuple(map(int, value.split("."))) + if max(parts) > 2147483647: + raise ValueError("numeric version component is too large") + return parts + (0,) * (4 - len(parts)) + + +def master_version(value): + if not isinstance(value, str) or not re.fullmatch(r"[0-9]+(?:\.[0-9]+){1,3}/[0-9a-f]{32}", value): + raise ValueError("invalid JP master version") + return value + + +def select_asset(raw, client): + """Select the largest applicable live minimum; no top-level fallback when live has no match.""" + if not raw: + return None + try: + data = json.loads(raw) + if not isinstance(data, dict): + raise ValueError() + live = data.get("live") + if live: + if not isinstance(live, list): + raise ValueError() + current = numeric_version(client) + candidates = [] + for entry in live: + if not isinstance(entry, dict): + continue + try: + minimum = numeric_version(entry.get("minClientVersion")) + except ValueError: + continue + if minimum <= current: + candidates.append((minimum, entry)) + if not candidates: + return None + data = max(candidates, key=lambda x: x[0])[1] + version, digest = data.get("version"), data.get("Android") + numeric_version(version) + if not isinstance(digest, str) or not re.fullmatch(r"[0-9a-f]{32}", digest): + raise ValueError() + return version, digest + except (ValueError, TypeError): + raise gameapi.GameApiError("JP Version: invalid asset metadata") from None + + +def origin(value): + try: + u = urlsplit(value) + if (not u.hostname or u.username or u.password or u.query or u.fragment + or u.path not in ("", "/") or u.scheme not in ("https", "http") + or u.scheme == "http" and u.hostname not in ("localhost", "127.0.0.1", "::1")): + raise ValueError() + port = u.port + host = f"[{u.hostname}]" if ":" in u.hostname else u.hostname + return f"{u.scheme}://{host}" + (f":{port}" if port and port != (443 if u.scheme == "https" else 80) else "") + except (ValueError, TypeError, AttributeError): + raise ConfigError("JP origin must be HTTPS with no path or credentials (HTTP loopback is allowed for tests)") from None + + +def relative_path(value): + decoded = unquote(value) + if (not value or value.startswith("/") or "\\" in decoded or "%" in decoded + or any(ord(c) < 32 for c in decoded) or any(p in ("", ".", "..") for p in decoded.split("/")) + or "?" in value or "#" in value or ":" in value): + raise ValueError("invalid JP resource path") + return quote(decoded, safe="/!$&'()+,;=@[]-._~") + + +@dataclass(frozen=True) +class Source: + provider: str + version: str + hash: str + cdn: str = field(repr=False) + platform: str = "Android" + + def __post_init__(self): + if self.provider != "jp" or self.platform != "Android": + raise ValueError("unsupported JP asset source") + numeric_version(self.version) + if not re.fullmatch(r"[0-9a-f]{32}", self.hash): + raise ValueError("invalid JP asset hash") + if origin(self.cdn) != self.cdn: + raise ValueError("JP source CDN is not a canonical origin") + + @property + def root(self): + return f"{self.cdn}/asset/{self.version}/Android/{self.hash}" + + @property + def catalog_url(self): + return self.root + "/catalog_main.bin" + + def url(self, internal_id): + if not internal_id.startswith(PLACEHOLDER): + raise ValueError("JP resource has no remote asset placeholder") + return self.root + "/" + relative_path(internal_id[len(PLACEHOLDER):]) + + def to_dict(self): + return asdict(self) + + @classmethod + def from_dict(cls, data): + if not isinstance(data, dict) or set(data) != {"provider", "version", "hash", "cdn", "platform"}: + raise ValueError("invalid JP source metadata") + return cls(**data) + + def cache_dir(self, cache): + site = hashlib.sha256(self.cdn.encode()).hexdigest()[:16] + return Path(cache) / "jp" / site / self.version / self.hash + + +@dataclass(frozen=True) +class Observation: + version: gameapi.MasterVersion + source: Source | None + cdn: str = field(repr=False) + authorization: str = field(repr=False) + user_agent: str = field(repr=False) + + +class Session: + def __init__(self, cfg, region): + self.cfg, self.region = cfg.for_region(region), region + self._observation = None + self._created = 0 + self._lock = threading.RLock() + + def observe(self, *, timeout=gameapi.TIMEOUT): + with self._lock: + section = f"servers.{self.region}" + api = gameapi.api_root(self.cfg, section) + origin(api) + allowed = origin(self.cfg.cdn(self.region)) + client = gameapi.client_version(self.cfg, self.region) + headers = {} + body = gameapi.call(api, gameapi.VERSION_METHOD, b"", client, timeout=timeout, + setting=f"[{section}] api", response_metadata=headers) + try: + version = master_version(gameapi.parse_version(body).version) + cdn = origin(headers.get("x-sirius-env", "")) + if cdn != allowed: + raise gameapi.GameApiError("JP Version: CDN differs from the configured CDN origin") + credential = headers.get("x-sirius-cred") + if not isinstance(credential, str) or not credential or len(credential) > 8192: + raise ValueError("missing CDN credential") + asset = select_asset(headers.get("x-asset-version"), client) + except (ValueError, ConfigError): + raise gameapi.GameApiError("JP Version: invalid master version or CDN metadata") from None + source = Source("jp", *asset, cdn) if asset else None + authorization = "Basic " + base64.b64encode(("sirius:" + credential).encode()).decode() + result = Observation(gameapi.MasterVersion(version, asset[0] if asset else "", asset[1] if asset else None), + source, cdn, authorization, "OurNotes/" + client) + self._observation, self._created = result, time.monotonic() + return result + + def _current(self): + with self._lock: + if self._observation is None or time.monotonic() - self._created > 300: + return self.observe() + return self._observation + + def get(self, url, *, source=None, master=None, limit=FILE_LIMIT): + """Download against one immutable source. Auth refresh never changes the requested snapshot.""" + for attempt in range(2): + observation = self._current() + if source is not None and observation.source != source: + raise gameapi.GameApiError("JP assets changed: refresh the catalog before downloading missing resources") + if master is not None and observation.version.version != master: + raise gameapi.GameApiError("JP master changed: start a new master download") + parsed = urlsplit(url) + base = origin(f"{parsed.scheme}://{parsed.netloc}") + prefix = source.root + "/" if source else observation.cdn + "/master/" + master_version(master) + "/" + if base != observation.cdn or not url.startswith(prefix): + raise gameapi.GameApiError("JP download URL is outside the selected source") + relative_path(url[len(prefix):]) + conn_type = http.client.HTTPSConnection if parsed.scheme == "https" else http.client.HTTPConnection + connection = conn_type(parsed.netloc, timeout=120) + try: + connection.request("GET", parsed.path, headers={"Authorization": observation.authorization, + "User-Agent": observation.user_agent}) + response = connection.getresponse() + status = response.status + if status == 200: + data = response.read(limit + 1) + if len(data) > limit: + raise gameapi.GameApiError("JP download exceeds the size limit") + return data + except (OSError, http.client.HTTPException): + raise gameapi.GameApiError("JP CDN download failed: connection error") from None + finally: + connection.close() + if status in (401, 403) and attempt == 0: + with self._lock: + if self._observation is observation: + self.observe() + continue + raise gameapi.GameApiError(f"JP CDN download failed: HTTP {status}") + raise AssertionError("unreachable") + + +def sidecar(path): + return Path(str(path) + ".source.json") + + +def read_source(path, data): + try: + doc = json.loads(sidecar(path).read_bytes()) + if doc["catalog_sha256"] != hashlib.sha256(data).hexdigest(): + raise ValueError() + return Source.from_dict(doc["source"]) + except (OSError, ValueError, KeyError, TypeError): + raise ConfigError("JP catalog needs its matching .source.json sidecar (written by catalogs fetch)") from None + + +def write_source(path, data, source): + from .cache import write_atomic + doc = {"catalog_sha256": hashlib.sha256(data).hexdigest(), "source": source.to_dict()} + write_atomic(sidecar(path), (json.dumps(doc, sort_keys=True) + "\n").encode()) + + +def open_catalog(cfg, region, *, catalog_file=None, bundle_key=None, apk=None): + from .addressables import parse_header + from .cache import write_atomic + from .catalog import Catalog + session = Session(cfg, region) + if catalog_file is not None: + data = Path(catalog_file).read_bytes() + source = read_source(catalog_file, data) + else: + source = session.observe().source + if source is None: + raise gameapi.GameApiError("JP Version: no Android asset version applies to this client") + path = source.cache_dir(cfg.require_path("paths", "cache")) / "catalog_main.bin" + if path.exists(): + data = path.read_bytes() + if not sidecar(path).exists(): + # Interrupted between the two atomic writes: fetch the same snapshot again. + data = session.get(source.catalog_url, source=source, limit=CATALOG_LIMIT) + parse_header(data) + write_atomic(path, data) + write_source(path, data, source) + elif read_source(path, data) != source: + raise ConfigError("JP cached catalog source differs from the selected version") + else: + data = session.get(source.catalog_url, source=source, limit=CATALOG_LIMIT) + parse_header(data) + path.parent.mkdir(parents=True, exist_ok=True) + write_atomic(path, data) + write_source(path, data, source) + return Catalog(data, source.cache_dir(cfg.require_path("paths", "cache")), bundle_key=bundle_key, + apk=apk, source=source, session=session) diff --git a/src/nnnotes/liveaudio.py b/src/nnnotes/liveaudio.py index c0e08a6..8b0b754 100644 --- a/src/nnnotes/liveaudio.py +++ b/src/nnnotes/liveaudio.py @@ -25,9 +25,9 @@ from __future__ import annotations import struct -import zipfile from pathlib import Path +from .apkset import ApkSet from .acb import commands as _commands, tables as _tables, u16s as _u16s from .catalog import Catalog from .jsonio import write_json @@ -129,7 +129,7 @@ def cmd(table, index): def acf_info(apk: Path) -> dict: """Categories, buses and REACT entries of the game's ACF (`assets/Cri/Sound/Sirius.acf` in the APK).""" - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: top, T = _tables(z.read(ACF_IN_APK)) cat_names = {r["Index"]: r["Name"] for r in T["CategoryNameTable"]} cats = {} diff --git a/src/nnnotes/master.py b/src/nnnotes/master.py index 67e272b..da9186f 100644 --- a/src/nnnotes/master.py +++ b/src/nnnotes/master.py @@ -18,6 +18,7 @@ import hashlib import json import os +import re import time import urllib.error import urllib.request @@ -218,25 +219,39 @@ def _get(url: str, retries: int = 3, timeout: int = 60) -> bytes: raise AssertionError("unreachable") -def download(cdn: str, version: str, out_dir: Path, workers: int = 16) -> dict: +def download(cdn: str, version: str, out_dir: Path, workers: int = 16, *, get=None, strict=False) -> dict: """Master data `version` from the CDN into / (the manifest and every listed file, SHA-256 checked; files already present with the right hash are kept).""" out_dir = Path(out_dir) out_dir.mkdir(parents=True, exist_ok=True) base = f"{cdn.rstrip('/')}/master/{version}" - manifest_raw = _get(f"{base}/MasterManifest.json") + get = get or _get + manifest_raw = get(f"{base}/MasterManifest.json") manifest = json.loads(manifest_raw.decode("utf-8")) - (out_dir / "MasterManifest.json").write_bytes(manifest_raw) + if strict and manifest.get("version") != version: + raise DownloadError("master manifest version differs from the requested version") files = manifest.get("files", []) + if strict: + seen = set() + for entry in files: + name = entry.get("name") + if (not isinstance(name, str) or not re.fullmatch(r"[A-Za-z0-9_-]+\.bin", name) + or name in seen or not re.fullmatch(r"[0-9a-fA-F]{64}", str(entry.get("hash", ""))) + or type(entry.get("size")) is not int or not 64 < entry["size"] <= 128 * 1024 * 1024): + raise DownloadError("invalid JP master manifest entry") + seen.add(name) def one(f: dict): name, sha = f["name"], (f.get("hash") or "").lower() if "/" in name or "\\" in name or name in ("", ".", ".."): return {"file": name, "error": "unexpected file name"} dst = out_dir / name - if sha and dst.is_file() and hashlib.sha256(dst.read_bytes()).hexdigest() == sha: + if (sha and dst.is_file() and (not strict or dst.stat().st_size == f["size"]) + and hashlib.sha256(dst.read_bytes()).hexdigest() == sha): return "kept" - data = _get(f"{base}/{name}") + data = get(f"{base}/{name}") + if strict and (len(data) != f.get("size") or not sha): + return {"file": name, "error": "size or hash missing/different from the manifest"} if sha and hashlib.sha256(data).hexdigest() != sha: return {"file": name, "error": "sha256 differs from the manifest"} tmp = dst.with_name(f"{name}.{os.getpid()}.part") @@ -246,6 +261,9 @@ def one(f: dict): with ThreadPoolExecutor(max(1, workers)) as ex: results = list(ex.map(one, files)) + if not strict or not any(isinstance(r, dict) for r in results): + from .cache import write_atomic + write_atomic(out_dir / "MasterManifest.json", manifest_raw) return {"version": manifest.get("version", version), "files": len(files), "downloaded": results.count("downloaded"), "kept": results.count("kept"), "failed": [r for r in results if isinstance(r, dict)], "out": str(out_dir)} diff --git a/src/nnnotes/musicdata.py b/src/nnnotes/musicdata.py index 24157dd..29fab44 100644 --- a/src/nnnotes/musicdata.py +++ b/src/nnnotes/musicdata.py @@ -496,7 +496,8 @@ def record(sid: int) -> dict: "provenance": { "region": region, "client": {"versionName": client.get("versionName"), "versionCode": client.get("versionCode")}, - "catalog": {"resourceVersion": catalog.get("resourceVersion"), "sha256": catalog.get("sha256")}, + "catalog": {"resourceVersion": catalog.get("resourceVersion"), "sha256": catalog.get("sha256"), + **({"resourceHash": catalog["resourceHash"]} if catalog.get("resourceHash") else {})}, "master": {"source": master_source, "version": master_version, "tables": {t: {"sha256": table_sha[t]} for t in read}}, "exporter": exporter, diff --git a/src/nnnotes/nnnotes.example.toml b/src/nnnotes/nnnotes.example.toml index 90fd280..f0b5c39 100644 --- a/src/nnnotes/nnnotes.example.toml +++ b/src/nnnotes/nnnotes.example.toml @@ -26,6 +26,14 @@ language = "" # One table per region; add more tables ([servers.]) for other regions. [servers.tw] +# protocol: international or jp (empty: jp for the region named jp, international otherwise) +provider = "" +# optional per-region client version; empty: [client] version, else this region's APK versionName +client_version = "" +# optional per-region APK, base.apk with adjacent splits, APKS/XAPK, or directory containing base.apk +apk = "" +# optional per-region catalog file; JP files use a .source.json sidecar written by catalogs fetch +catalog = "" # label of the region in the local catalog browser (`nnnotes browse`) name = "" # CDN base URL of the region: catalogs, bundles, raw audio data and master data are fetched below it @@ -46,7 +54,7 @@ api = "" version = "" [paths] -# catalog .bin file to read instead of downloading catalog_main_.bin into the cache (flag --catalog) +# catalog .bin file to read instead of downloading; JP source metadata is stored in .source.json catalog = "" # cache directory: downloaded catalogs, decrypted bundles, raw CDN files (flag --cache) cache = "" diff --git a/src/nnnotes/player.py b/src/nnnotes/player.py index d04365e..a05781f 100644 --- a/src/nnnotes/player.py +++ b/src/nnnotes/player.py @@ -18,13 +18,13 @@ import json import re import struct -import zipfile from importlib import resources from pathlib import Path import UnityPy from UnityPy.helpers.TypeTreeNode import TypeTreeNode +from .apkset import ApkSet from .config import ConfigError from .unity import DEFAULT_RESOURCES, deref, external_path, is_pptr @@ -149,8 +149,13 @@ def manifest_version_code(data: bytes) -> int | None: class PlayerData: def __init__(self, apk: Path): - with zipfile.ZipFile(apk) as z: + with ApkSet(apk) as z: self.env = UnityPy.load(io.BytesIO(z.read(DATA_IN_APK))) + # Split Unity builds keep resources.assets (materials, shaders and settings) in + # datapack.unity3d; level0 in data.unity3d refers to it by external file id. + datapack = "assets/bin/Data/datapack.unity3d" + if datapack in z.namelist(): + self.env.load_file(io.BytesIO(z.read(datapack)), name="datapack.unity3d") self.defaults = UnityPy.load(io.BytesIO(z.read(DEFAULT_RESOURCES_IN_APK))) self.game_version = (manifest_version_name(z.read(MANIFEST_IN_APK)) if MANIFEST_IN_APK in z.namelist() else None) diff --git a/src/nnnotes/storysite.py b/src/nnnotes/storysite.py index 7251578..2ed2114 100644 --- a/src/nnnotes/storysite.py +++ b/src/nnnotes/storysite.py @@ -658,6 +658,8 @@ def build(out_dir, adv_ids, cfg: Config, player_dir, audio_format: str = "aac", if fonts not in FONT_SOURCES: raise ValueError(f"fonts {fonts!r}: one of {', '.join(FONT_SOURCES)}") require_extra() + if story_languages is None and cfg.provider((regions or [cfg.get("catalog", "region")])[0]) == "jp": + story_languages = ["ja"] langs = check_languages(story_languages) base = languages.check(cfg.require("catalog", "language")) default = base if base in langs else langs[0] diff --git a/src/nnnotes/web.py b/src/nnnotes/web.py index 29bc010..5989669 100644 --- a/src/nnnotes/web.py +++ b/src/nnnotes/web.py @@ -437,7 +437,7 @@ def open_data(cfg: Config, region: str | None = None): """Catalog (fetching from the CDN of `region`), the master dir of `region` and PlayerData from the settings (as the command line opens them; `region` None: [catalog] region).""" from .cli import master_dir, open_catalog, player_data - return open_catalog(cfg, region=region), master_dir(cfg, region), player_data(cfg) + return open_catalog(cfg, region=region), master_dir(cfg, region), player_data(cfg, region) def _lock_fetches(cat, lock) -> None: @@ -723,6 +723,8 @@ def ingest(store: Store, site: Path, music_id: int, difficulty: str, live_dir: P def _worker_init(cfg: Config, lock, job: dict) -> None: + if job.get("region"): + cfg = cfg.for_region(job["region"]) use(cfg) configure_caches(job) cat, master, player = open_data(cfg, job.get("region")) @@ -1511,6 +1513,9 @@ def _build_group(site: Path, tmp_root: Path, pairs, cfg: Config, region: str, pr reads: "ReadSets | None" = None) -> tuple[list, list, int]: """The charts `pairs` of one region group into charts/.json (data of `region`); manifests that exist are skipped unless `force` and gain the group's `regions`. -> (results, skipped ids, workers).""" + cfg = cfg.for_region(region) + if cfg.provider(region) == "jp": + job = {**job, "language": "ja"} by_music: dict[int, list[str]] = {} skipped = [] for music_id, difficulty in pairs: @@ -1637,7 +1642,13 @@ def build(out_dir, pairs, cfg: Config, player_dir: Path, audio_format: str = DEF masters = region_masters(cfg, regions) for m in masters.values(): liveoptions.resolve(live_options, m) - groups = region_groups(masters) + # Equal master tables alone do not imply equal assets or player data across releases. + groups = [] + for group in region_groups(masters): + by_provider = {} + for region in group: + by_provider.setdefault(cfg.provider(region), []).append(region) + groups.extend(by_provider.values()) site = Path(out_dir).resolve() (site / "charts").mkdir(parents=True, exist_ok=True) site_store(site, encoding) diff --git a/src/nnnotes/webmodel.py b/src/nnnotes/webmodel.py index c2922a4..0711506 100644 --- a/src/nnnotes/webmodel.py +++ b/src/nnnotes/webmodel.py @@ -456,7 +456,7 @@ def write_models_index(site: Path) -> tuple[int, set[str]]: def _open(cfg: Config, region: str | None = None): from .cli import open_catalog, player_data - return open_catalog(cfg, region=region), player_data(cfg) + return open_catalog(cfg, region=region), player_data(cfg, region) def _worker_init(cfg: Config, lock, job: dict) -> None: diff --git a/tests/test_configfile.py b/tests/test_configfile.py index 1ab6373..131574f 100644 --- a/tests/test_configfile.py +++ b/tests/test_configfile.py @@ -73,7 +73,8 @@ def test_init_with_values(tmp_path, capsys): data = tomllib.loads(f.read_text(encoding="utf-8")) assert data["bundle"]["key"] == KEY and data["catalog"]["region"] == "en" assert data["servers"] == {"en": {"name": "", "cdn": "https://cdn.test/", "api": "", "languages": ["en", "ja"], - "master": ""}} # the example region table renamed + "master": "", "provider": "", "client_version": "", "apk": "", + "catalog": ""}} # the example region table renamed assert data["paths"]["cache"] == str((tmp_path / "cache").absolute()) assert "# CDN base URL of the region" in f.read_text(encoding="utf-8") # the comments stay if os.name != "nt": diff --git a/tests/test_jp.py b/tests/test_jp.py new file mode 100644 index 0000000..4b4faec --- /dev/null +++ b/tests/test_jp.py @@ -0,0 +1,287 @@ +"""JP transport, snapshot identity and split APK regressions; all servers and data are synthetic.""" +import base64 +import gzip +import io +import json +import threading +import zipfile +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer +from types import SimpleNamespace + +import pytest + +import synth +from test_gameapi import server as grpc_server, version_response +from nnnotes import addressables, catalogdb, cli, configfile, deckdata, gameapi, jp +from nnnotes.apkset import ApkSet +from nnnotes.catalog import Catalog, file_name, location_kind +from nnnotes.cli_assets import CatalogFetcher +from nnnotes.config import Config, ConfigError + +HASH = "a" * 32 +MASTER = "1.0.0.100/" + "b" * 32 +server = grpc_server + + +def cat_bytes(): + return synth.CatalogWriter().build([ + ("Live/MusicScore/0001/0001_00", "Assets/chart.asset", [1]), + ("chart.bundle", jp.PLACEHOLDER + "chart.bundle", []), + ("Cri/Sound/test", jp.PLACEHOLDER + "audio/test.acb", []), + ]) + + +@pytest.fixture +def upstream(server): + state = SimpleNamespace(hash=HASH, credential="synthetic-secret", asset=True, seen=[], status=200, + body=gzip.compress(cat_bytes()), cdn_override=None) + + class Handler(BaseHTTPRequestHandler): + def do_GET(self): + state.seen.append((self.path, dict(self.headers))) + expected = "Basic " + base64.b64encode(("sirius:" + state.credential).encode()).decode() + status = state.status if self.headers.get("Authorization") == expected else 401 + self.send_response(status) + if status == 302: + self.send_header("Location", "http://127.0.0.1:1/leak") + self.end_headers() + if status == 200: + self.wfile.write(state.body) + + def log_message(self, *args): + pass + + http = ThreadingHTTPServer(("127.0.0.1", 0), Handler) + thread = threading.Thread(target=http.serve_forever, daemon=True) + thread.start() + state.cdn = f"http://127.0.0.1:{http.server_port}" + + def version(req, ctx): + headers = [("x-sirius-env", state.cdn_override or state.cdn), ("x-sirius-cred", state.credential)] + if state.asset: + headers.append(("x-asset-version", json.dumps({"live": [ + {"minClientVersion": "1.0.4", "version": "1.0.0.300", "Android": state.hash}]}))) + ctx.send_initial_metadata(headers) + return version_response(MASTER) + + server.handlers[gameapi.VERSION_METHOD] = version + state.grpc = server + state.cfg = Config({"catalog": {"region": "jp", "language": "ja"}, + "client": {"version": "1.0.1"}, + "servers": {"jp": {"cdn": state.cdn, "api": server.root, "client_version": "1.0.4"}}}, + environ={}) + try: + yield state + finally: + http.shutdown() + http.server_close() + thread.join() + + +def test_version_headers_and_master_without_asset(upstream): + result = gameapi.master_version(upstream.cfg, "jp") + assert (result.version, result.resource_version, result.resource_hash) == (MASTER, "1.0.0.300", HASH) + assert upstream.grpc.seen[0].metadata["x-client-version"] == "1.0.4" + assert upstream.grpc.seen[0].method == gameapi.VERSION_METHOD + upstream.asset = False + result = gameapi.master_version(upstream.cfg, "jp") + assert result.version == MASTER and result.resource_version == "" and result.resource_hash is None + + +def test_live_selection_numeric_and_no_fallback(): + def entry(client, version): + return {"minClientVersion": client, "version": version, "Android": HASH} + raw = {"version": "9.9", "Android": HASH, + "live": [entry("1.0.10", "1.0.10"), entry("1.0.2", "1.0.2"), entry("8.0", "8.0")]} + assert jp.select_asset(json.dumps(raw), "1.0.12") == ("1.0.10", HASH) + assert jp.select_asset(json.dumps(raw), "1.0.1") is None + assert jp.select_asset(json.dumps({"version": "1.0", "Android": HASH, "live": []}), "1.0") == ("1.0", HASH) + with pytest.raises(gameapi.GameApiError, match="invalid asset"): + jp.select_asset('{"version":"../secret"}', "1.0") + + +def test_authenticated_download_rotation_and_redaction(upstream): + session = jp.Session(upstream.cfg, "jp") + obs = session.observe() + assert upstream.credential not in repr(obs) + upstream.credential = "synthetic-rotated" + assert session.get(obs.source.catalog_url, source=obs.source) == upstream.body + assert len(upstream.seen) == 2 and len(upstream.grpc.seen) == 2 + assert upstream.seen[-1][1]["User-Agent"] == "OurNotes/1.0.4" + assert "Authorization" not in json.dumps(obs.source.to_dict()) + + +@pytest.mark.parametrize("status", [302, 429, 500, 403]) +def test_download_does_not_follow_redirects_or_loop(upstream, status): + session = jp.Session(upstream.cfg, "jp") + obs = session.observe() + upstream.status = status + with pytest.raises(gameapi.GameApiError, match=f"HTTP {status}") as error: + session.get(obs.source.catalog_url, source=obs.source) + expected = 2 if status == 403 else 1 + assert len(upstream.seen) == expected and len(upstream.grpc.seen) == expected + assert upstream.credential not in str(error.value) and upstream.cdn not in str(error.value) + + +def test_unapproved_origin_fails_before_download(upstream): + upstream.cdn_override = "https://elsewhere.invalid" + with pytest.raises(gameapi.GameApiError, match="CDN differs"): + jp.Session(upstream.cfg, "jp").observe() + assert not upstream.seen + + +def test_hash_only_change_does_not_rebind_old_catalog(upstream): + session = jp.Session(upstream.cfg, "jp") + old = session.observe().source + upstream.hash = "c" * 32 + session.observe() + with pytest.raises(gameapi.GameApiError, match="assets changed"): + session.get(old.catalog_url, source=old) + assert not upstream.seen + + +def test_gzip_catalog_classification_and_dependency_closure(tmp_path): + raw = cat_bytes() + compressed = gzip.compress(raw) + assert addressables.parse(raw) == addressables.parse(compressed) + assert addressables.parse_keys(raw) == addressables.parse_keys(compressed) + cat = Catalog(compressed, tmp_path) + assert len(cat.resolve("Live/MusicScore/0001/0001_00")) == 1 + assert cat.resolve("Live/MusicScore/0001/0001_00")[0].remote + assert len(cat.raw_files()) == 1 + assert catalogdb.summary(catalogdb.index(compressed))["remoteBundles"] == 1 + assert file_name(jp.PLACEHOLDER + "nested/chart.bundle") == "chart.bundle" + assert file_name(jp.PLACEHOLDER + "audio/test.acb") == "audio/test.acb" + + +def test_gzip_limit(monkeypatch): + monkeypatch.setattr(addressables, "MAX_CATALOG", 256) + with pytest.raises(ValueError, match="expanded catalog"): + addressables.parse(gzip.compress(b"x" * 257)) + + +@pytest.mark.parametrize("path", ["../secret", "%2e%2e/secret", "%252e%252e/secret", "/secret", "x\\y", "a?b", "a#b"]) +def test_placeholder_rejects_unsafe_paths(path): + with pytest.raises(ValueError): + location_kind(jp.PLACEHOLDER + path) + + +def test_catalog_cache_source_and_offline_replay(upstream, tmp_path, monkeypatch): + cfg = upstream.cfg + cfg._data["paths"] = {"cache": str(tmp_path / "cache")} + cat = cli.open_catalog(cfg, bundles=False) + path = cat.cache_dir / "catalog_main.bin" + assert path.read_bytes() == upstream.body + assert jp.read_source(path, upstream.body) == cat.source + assert cat.cache_dir != tmp_path / "cache" + assert cat.keys("Live/") == ["Live/MusicScore/0001/0001_00"] + record = catalogdb.CatalogDB(tmp_path / "store").add(upstream.body, source=cat.source.to_dict(), region="jp") + assert upstream.credential not in json.dumps(record) + cfg._data["paths"]["catalog"] = str(path) + monkeypatch.setattr(jp.Session, "observe", lambda *a, **k: pytest.fail("offline read made an API call")) + assert cli.open_catalog(cfg, bundles=False).keys() == cat.keys() + path.write_bytes(b"changed") + with pytest.raises(ConfigError, match="matching"): + cli.open_catalog(cfg, bundles=False) + + +def test_catalogdb_source_identity_and_fetcher(upstream, tmp_path): + cfg = upstream.cfg + cfg._data["paths"] = {"cache": str(tmp_path / "cache")} + session = jp.Session(cfg, "jp") + source = session.observe().source + db = catalogdb.CatalogDB(tmp_path / "store") + old = db.add(cat_bytes(), source=source.to_dict(), region="jp") + new_source = {**source.to_dict(), "hash": "c" * 32} + new = db.add(cat_bytes(), source=new_source, region="jp") + assert old["id"] != new["id"] and old["labels"] != new["labels"] + assert old["remote"] == new["remote"] + fetcher = CatalogFetcher(tmp_path / "store", tmp_path / "cache", cfg) + cat, _ = fetcher.catalog(old["id"]) + assert cat.source == source and cat.cache_dir == source.cache_dir(tmp_path / "cache") + + +def zip_bytes(files): + output = io.BytesIO() + with zipfile.ZipFile(output, "w") as z: + for name, data in files.items(): + z.writestr(name, data) + return output.getvalue() + + +@pytest.mark.parametrize("kind", ["base", "directory", "apks"]) +def test_split_apk_member_routing_and_manifest_precedence(tmp_path, kind): + base = zip_bytes({"AndroidManifest.xml": synth.axml(["versionName", "manifest", "1.0.4"], 2), + "assets/bin/Data/data.unity3d": b"boot"}) + assets = zip_bytes({"AndroidManifest.xml": b"split manifest", "assets/aa/catalog.bin": cat_bytes(), + "assets/Master/MasterManifest.json": b'{"version":"embedded", "files":[]}'}) + (tmp_path / "base.apk").write_bytes(base) + (tmp_path / "split_UnityDataAssetPack.apk").write_bytes(assets) + source = tmp_path / "base.apk" if kind == "base" else tmp_path + if kind == "apks": + source = tmp_path / "game.apks" + source.write_bytes(zip_bytes({"base.apk": base, "split_UnityDataAssetPack.apk": assets})) + with ApkSet(source) as archive: + assert archive.read("assets/bin/Data/data.unity3d") == b"boot" + assert archive.read("assets/aa/catalog.bin") == cat_bytes() + assert gameapi.apk_version_name(source) == "1.0.4" + assert deckdata.apk_master(source).version == "embedded" + assert catalogdb.apk_catalog(source) == cat_bytes() + + +def test_per_region_apk_and_client_settings(tmp_path): + cfg = Config({"catalog": {"region": "tw"}, "paths": {"apk": "global.apk"}, + "servers": {"jp": {"apk": "jp.apks", "client_version": "1.0.4"}}, + "client": {"version": "1.0.1"}}, base=tmp_path, environ={}) + assert cfg.path("paths", "apk") == tmp_path / "global.apk" + assert cfg.for_region("jp").path("paths", "apk") == tmp_path / "jp.apks" + assert gameapi.client_version(cfg, "jp") == "1.0.4" and gameapi.client_version(cfg, "tw") == "1.0.1" + assert configfile.resolve("servers.jp.apk")[0].kind == "path" + assert configfile.resolve("servers.jp.client_version")[0].kind == "string" + + +def test_datapack_is_loaded_in_the_boot_environment(tmp_path, monkeypatch): + from nnnotes import player + boot = SimpleNamespace(objects=[SimpleNamespace(type=SimpleNamespace(name="GraphicsSettings"), + assets_file=SimpleNamespace(unity_version="6000.3.12f1"))]) + loaded = [] + boot.load_file = lambda stream, **kw: loaded.append((stream.read(), kw["name"])) + monkeypatch.setattr(player.UnityPy, "load", lambda stream: boot if stream.read() == b"boot" else object()) + path = tmp_path / "base.apk" + path.write_bytes(zip_bytes({player.DATA_IN_APK: b"boot", player.DEFAULT_RESOURCES_IN_APK: b"defaults", + "assets/bin/Data/datapack.unity3d": b"resources"})) + player.PlayerData(path) + assert loaded == [(b"resources", "datapack.unity3d")] + + +def test_jp_master_manifest_checks_before_writing(tmp_path): + from nnnotes import master + manifest = {"version": MASTER, "files": [{"name": "../secret.bin", "hash": "a" * 64, "size": 128}]} + seen = [] + def get(url): + seen.append(url) + return json.dumps(manifest).encode() + with pytest.raises(master.DownloadError, match="manifest entry"): + master.download("https://cdn.invalid", MASTER, tmp_path, get=get, strict=True) + assert len(seen) == 1 and not (tmp_path / "MasterManifest.json").exists() + + +def test_export_cache_locator_uses_source_namespace(tmp_path): + from nnnotes.cli_assets import Workspace + source = jp.Source("jp", "1.0.0.300", HASH, "https://cdn.invalid") + workspace = Workspace.__new__(Workspace) + workspace.version = {"source": source.to_dict()} + rel = workspace._cache_rel("bundle", {"name": "same.bundle"}, {}) + assert tmp_path / rel == source.cache_dir(tmp_path) / "bundles/same.bundle" + rel = workspace._cache_rel("raw", {}, {"internalId": jp.PLACEHOLDER + "audio/sample.acb"}) + assert tmp_path / rel == source.cache_dir(tmp_path) / "raw/audio/sample.acb" + + +def test_music_data_preserves_asset_hash(tmp_path): + from test_musicdata import CHARTS, bgm, master_dir, KEY, PROV + from nnnotes import musicdata + output = tmp_path / "music.json" + musicdata.export(output, deckdata.master_files(master_dir(tmp_path)), KEY, CHARTS.__getitem__, bgm, + **{**PROV, "catalog": {**PROV["catalog"], "resourceHash": HASH}}) + doc = json.loads(output.read_bytes()) + assert doc["provenance"]["catalog"]["resourceHash"] == HASH From a41b371b0eb55d188eb2e28ed83237d516a35f36 Mon Sep 17 00:00:00 2001 From: luoxiadesu <249937048+luoxiadesu@users.noreply.github.com> Date: Wed, 30 Sep 2026 04:06:32 +0900 Subject: [PATCH 11/14] fix(jp): reject incomplete downloads and isolate APK updates --- docs/jp.md | 6 ++-- src/nnnotes/catalog.py | 13 ++++++-- src/nnnotes/catalogdb.py | 2 +- src/nnnotes/cli_assets.py | 4 ++- src/nnnotes/cri.py | 3 +- src/nnnotes/jp.py | 11 +++++++ src/nnnotes/master.py | 10 +++++- tests/test_jp.py | 68 ++++++++++++++++++++++++++++++++++++++- tests/test_master.py | 11 +++++++ 9 files changed, 118 insertions(+), 10 deletions(-) diff --git a/docs/jp.md b/docs/jp.md index b2e1672..5362452 100644 --- a/docs/jp.md +++ b/docs/jp.md @@ -82,7 +82,8 @@ CRI locations use `{Fwk.Resource.RemoteAssetDir}`. These are indexed, resolved a snapshot, including in `browse`, `catalogs`, `export`, `plan` and `run-stage`. The cache lives under `/jp////`. It cannot reuse the -international catalog cache. Each cached `catalog_main.bin` has a `catalog_main.bin.source.json` with its SHA-256 +international catalog cache. Embedded files are further isolated under `apk//`, so an APK +update cannot reuse old local bundles when the CDN snapshot stays unchanged. Each cached `catalog_main.bin` has a `catalog_main.bin.source.json` with its SHA-256 and public source metadata. For an offline `--catalog` or `catalogs import`, copy the pair together. Credentials are reacquired only when a missing file must be downloaded. @@ -106,7 +107,8 @@ catalog version/hash matches the master snapshot. A mixed master/catalog region ## Validation and current limits The synthetic tests use local gRPC and HTTP servers, including gzip catalogs, split packages, authentication -rotation, redirects, 429, source isolation, hash-only updates, offline replay and invalid paths. +rotation, redirects, 429, truncated responses, source isolation, APK-only and hash-only updates, offline replay +and invalid paths. Live validation on 2026-09-30 used JP client 1.0.4 and local JP Android 1.0.3 data: diff --git a/src/nnnotes/catalog.py b/src/nnnotes/catalog.py index a1df558..73a9c45 100644 --- a/src/nnnotes/catalog.py +++ b/src/nnnotes/catalog.py @@ -7,6 +7,7 @@ from __future__ import annotations import http.client +import hashlib import sys import threading import urllib.parse @@ -91,6 +92,12 @@ def __init__(self, catalog_bytes: bytes, cache_dir: Path, *, cdn=None, bundle_ke with ApkSet(self.apk) as z: self._sources["apk"] = z.read(APK_CATALOG) + def local_cache_dir(self) -> Path: + """JP embedded files also depend on the APK catalog, independently of the CDN snapshot.""" + if self.source is not None and "apk" in self._sources: + return self.cache_dir / "apk" / hashlib.sha256(self._sources["apk"]).hexdigest() + return self.cache_dir + def _parse(self) -> tuple: """(entries, by offset, by primary key), parsed the first time a lookup needs them (a command that only wants the catalog files, sources(), does not pay for it).""" @@ -247,7 +254,7 @@ def resolve(self, key: str) -> list[Bundle]: # --- fetch ------------------------------------------------------------- def cached(self, b: Bundle) -> Path | None: """The file fetch(b) returns when the bundle is in the cache already, else None.""" - dst = self.cache_dir / "bundles" / b.name + dst = (self.cache_dir if b.remote else self.local_cache_dir()) / "bundles" / b.name return dst if dst.is_file() and dst.stat().st_size > 0 else None def cached_raw(self, e: dict) -> Path | None: @@ -258,7 +265,7 @@ def cached_raw(self, e: dict) -> Path | None: def fetch(self, b: Bundle) -> Path: """Local path to the decrypted bundle (CDN download or APK read).""" - dst = self.cache_dir / "bundles" / b.name + dst = (self.cache_dir if b.remote else self.local_cache_dir()) / "bundles" / b.name dst.parent.mkdir(parents=True, exist_ok=True) if dst.exists() and dst.stat().st_size > 0: return dst @@ -319,7 +326,7 @@ def apk_bundle(self, name_contains: str) -> Path: if len(names) != 1: raise KeyError(f"{name_contains!r}: {len(names)} APK bundles match") name = names[0].rsplit("/", 1)[1] - dst = self.cache_dir / "bundles" / name + dst = self.local_cache_dir() / "bundles" / name if not (dst.exists() and dst.stat().st_size > 0): dst.parent.mkdir(parents=True, exist_ok=True) data = z.read(names[0]) diff --git a/src/nnnotes/catalogdb.py b/src/nnnotes/catalogdb.py index 832a87c..2436f98 100644 --- a/src/nnnotes/catalogdb.py +++ b/src/nnnotes/catalogdb.py @@ -384,7 +384,7 @@ def add(self, remote: bytes, apk: bytes | None = None, *, label: str | None = No ("resourceVersion", resource_version), ("apkVersionName", apk_version_name)): if given is not None and v[k] is None: v[k] = given - label = label or (f"{source['version']}/{source['hash']}" if source else + label = label or (f"{source['version']}/{source['hash']}/{vid[:12]}" if source else default_label(remote_rec["sha256"], resource_version)) for other in versions: if other is not v and label in other["labels"] and (other["region"], other["language"]) == ( diff --git a/src/nnnotes/cli_assets.py b/src/nnnotes/cli_assets.py index 9ae69ce..11b89f4 100644 --- a/src/nnnotes/cli_assets.py +++ b/src/nnnotes/cli_assets.py @@ -321,7 +321,7 @@ def fetch_location(self, vid: str, lid: str) -> Path: if loc["kind"] == "bundle": return cat.fetch(Bundle(0, iid, file_name(iid), remote_path(iid) is not None)) if remote_path(iid) is None: - return self._apk_file(iid, apk=cat.apk, cache=cat.cache_dir) + return self._apk_file(iid, apk=cat.apk, cache=cat.local_cache_dir()) return cat.fetch_raw({"internal_id": iid}) def _apk_file(self, internal_id: str, *, apk=None, cache=None) -> Path: @@ -685,6 +685,8 @@ def _cache_rel(self, kind: str, entry: dict, loc: dict) -> str: if self.version.get("source"): from .jp import Source prefix = Source.from_dict(self.version["source"]).cache_dir(Path(".")).as_posix() + "/" + if not entry.get("remote", True) and self.version.get("apk"): + prefix += "apk/" + self.version["apk"]["sha256"] + "/" if kind == "bundle": return prefix + f"bundles/{entry['name']}" rel = remote_path(loc["internalId"]) diff --git a/src/nnnotes/cri.py b/src/nnnotes/cri.py index 0154d35..5926506 100644 --- a/src/nnnotes/cri.py +++ b/src/nnnotes/cri.py @@ -60,7 +60,8 @@ def hca_key(apk) -> int: """crikey.find_key of an APK, read once per process (keyed by the file's path, size and modification time).""" - k = cache.file_id(apk) + path = Path(apk) + k = cache.file_id(path / "base.apk" if path.is_dir() else path) with _keys_lock: if k in _keys: return _keys[k] diff --git a/src/nnnotes/jp.py b/src/nnnotes/jp.py index ebdcbde..842d93f 100644 --- a/src/nnnotes/jp.py +++ b/src/nnnotes/jp.py @@ -213,9 +213,20 @@ def get(self, url, *, source=None, master=None, limit=FILE_LIMIT): response = connection.getresponse() status = response.status if status == 200: + length = response.getheader("Content-Length") + if length is not None and not response.chunked: + if not length.isdecimal(): + raise gameapi.GameApiError("JP download has invalid Content-Length") + length = int(length) + if length > limit: + raise gameapi.GameApiError("JP download exceeds the size limit") + else: + length = None data = response.read(limit + 1) if len(data) > limit: raise gameapi.GameApiError("JP download exceeds the size limit") + if length is not None and len(data) != length: + raise gameapi.GameApiError("JP download is truncated") return data except (OSError, http.client.HTTPException): raise gameapi.GameApiError("JP CDN download failed: connection error") from None diff --git a/src/nnnotes/master.py b/src/nnnotes/master.py index da9186f..bf412e1 100644 --- a/src/nnnotes/master.py +++ b/src/nnnotes/master.py @@ -228,18 +228,26 @@ def download(cdn: str, version: str, out_dir: Path, workers: int = 16, *, get=No get = get or _get manifest_raw = get(f"{base}/MasterManifest.json") manifest = json.loads(manifest_raw.decode("utf-8")) - if strict and manifest.get("version") != version: + if strict and (not isinstance(manifest, dict) or manifest.get("version") != version): raise DownloadError("master manifest version differs from the requested version") files = manifest.get("files", []) if strict: + if not isinstance(files, list) or not files: + raise DownloadError("JP master manifest has no file list") seen = set() for entry in files: + if not isinstance(entry, dict): + raise DownloadError("invalid JP master manifest entry") name = entry.get("name") if (not isinstance(name, str) or not re.fullmatch(r"[A-Za-z0-9_-]+\.bin", name) or name in seen or not re.fullmatch(r"[0-9a-fA-F]{64}", str(entry.get("hash", ""))) or type(entry.get("size")) is not int or not 64 < entry["size"] <= 128 * 1024 * 1024): raise DownloadError("invalid JP master manifest entry") seen.add(name) + else: + # Preserve the original partial-download receipt even if a worker raises. + from .cache import write_atomic + write_atomic(out_dir / "MasterManifest.json", manifest_raw) def one(f: dict): name, sha = f["name"], (f.get("hash") or "").lower() diff --git a/tests/test_jp.py b/tests/test_jp.py index 4b4faec..e67e3e6 100644 --- a/tests/test_jp.py +++ b/tests/test_jp.py @@ -34,7 +34,7 @@ def cat_bytes(): @pytest.fixture def upstream(server): state = SimpleNamespace(hash=HASH, credential="synthetic-secret", asset=True, seen=[], status=200, - body=gzip.compress(cat_bytes()), cdn_override=None) + body=gzip.compress(cat_bytes()), cdn_override=None, content_length=None) class Handler(BaseHTTPRequestHandler): def do_GET(self): @@ -42,6 +42,8 @@ def do_GET(self): expected = "Basic " + base64.b64encode(("sirius:" + state.credential).encode()).decode() status = state.status if self.headers.get("Authorization") == expected else 401 self.send_response(status) + if state.content_length is not None: + self.send_header("Content-Length", str(state.content_length)) if status == 302: self.send_header("Location", "http://127.0.0.1:1/leak") self.end_headers() @@ -111,6 +113,24 @@ def test_authenticated_download_rotation_and_redaction(upstream): assert "Authorization" not in json.dumps(obs.source.to_dict()) +def test_truncated_download_is_not_cached(upstream, tmp_path): + session = jp.Session(upstream.cfg, "jp") + source = session.observe().source + upstream.content_length = len(upstream.body) + 5 + cat = Catalog(cat_bytes(), tmp_path, source=source, session=session) + with pytest.raises(gameapi.GameApiError, match="truncated"): + cat.fetch_raw({"internal_id": jp.PLACEHOLDER + "audio/test.acb"}) + assert cat.cached_raw({"internal_id": jp.PLACEHOLDER + "audio/test.acb"}) is None + + +def test_declared_download_limit_is_checked(upstream): + session = jp.Session(upstream.cfg, "jp") + source = session.observe().source + upstream.content_length = 10000 + with pytest.raises(gameapi.GameApiError, match="size limit"): + session.get(source.catalog_url, source=source, limit=1000) + + @pytest.mark.parametrize("status", [302, 429, 500, 403]) def test_download_does_not_follow_redirects_or_loop(upstream, status): session = jp.Session(upstream.cfg, "jp") @@ -201,6 +221,33 @@ def test_catalogdb_source_identity_and_fetcher(upstream, tmp_path): assert cat.source == source and cat.cache_dir == source.cache_dir(tmp_path / "cache") +def test_apk_update_with_same_asset_snapshot_does_not_reuse_embedded_cache(tmp_path): + from nnnotes.catalog import APK_CATALOG, APK_AA_DIR + source = jp.Source("jp", "1.0", HASH, "https://cdn.invalid") + apk = tmp_path / "base.apk" + cache = tmp_path / "cache" + db = catalogdb.CatalogDB(tmp_path / "store") + cfg = Config({"catalog": {"region": "jp"}, "paths": {"apk": str(apk)}}, environ={}) + from nnnotes.cli_assets import Workspace + for revision in (1, 2): + local = synth.CatalogWriter().build([("Local", "Assets/local", [1]), + ("local.bundle", synth.local("local.bundle"), [])], + build_hash=str(revision) * 32) + content = b"UnityFS\0" + bytes([revision]) + apk.write_bytes(zip_bytes({APK_CATALOG: local, APK_AA_DIR + "Android/local.bundle": content})) + cat = Catalog(cat_bytes(), source.cache_dir(cache), source=source, apk=apk) + v = db.add(cat_bytes(), local, source=source.to_dict(), region="jp") + workspace = Workspace.__new__(Workspace) + workspace.version = v + rel = workspace._cache_rel("bundle", {"name": "local.bundle", "remote": False}, {}) + assert cat.fetch(cat.resolve("Local")[0]).read_bytes() == content + assert (cache / rel).read_bytes() == content + fetcher = CatalogFetcher(tmp_path / "store", cache, cfg) + replay, _ = fetcher.catalog(v["id"]) + assert replay.fetch(replay.resolve("Local")[0]).read_bytes() == content + assert len(db.versions()) == 2 + + def zip_bytes(files): output = io.BytesIO() with zipfile.ZipFile(output, "w") as z: @@ -266,6 +313,15 @@ def get(url): assert len(seen) == 1 and not (tmp_path / "MasterManifest.json").exists() +@pytest.mark.parametrize("files", [None, [], {}, [None]]) +def test_jp_master_rejects_missing_or_malformed_file_lists(tmp_path, files): + from nnnotes import master + with pytest.raises(master.DownloadError, match="manifest"): + master.download("https://cdn.invalid", MASTER, tmp_path, strict=True, + get=lambda _: json.dumps({"version": MASTER, "files": files}).encode()) + assert not (tmp_path / "MasterManifest.json").exists() + + def test_export_cache_locator_uses_source_namespace(tmp_path): from nnnotes.cli_assets import Workspace source = jp.Source("jp", "1.0.0.300", HASH, "https://cdn.invalid") @@ -285,3 +341,13 @@ def test_music_data_preserves_asset_hash(tmp_path): **{**PROV, "catalog": {**PROV["catalog"], "resourceHash": HASH}}) doc = json.loads(output.read_bytes()) assert doc["provenance"]["catalog"]["resourceHash"] == HASH + + +def test_directory_hca_key_tracks_replaced_base_apk(tmp_path, monkeypatch): + from nnnotes import cri + apk = tmp_path / "base.apk" + apk.write_bytes(b"first") + monkeypatch.setattr(cri.crikey, "find_key", lambda directory: len((directory / "base.apk").read_bytes())) + assert cri.hca_key(tmp_path) == 5 + apk.write_bytes(b"replacement") + assert cri.hca_key(tmp_path) == 11 diff --git a/tests/test_master.py b/tests/test_master.py index a236ced..9301e77 100644 --- a/tests/test_master.py +++ b/tests/test_master.py @@ -95,6 +95,17 @@ def test_download_rejects_path_names(tmp_path): assert not (tmp_path / "evil.bin").exists() +def test_international_manifest_survives_worker_exception(tmp_path): + manifest = json.dumps({"version": "v", "files": [{"name": "MasterA.bin", "hash": ""}]}).encode() + def get(url): + if url.endswith("MasterManifest.json"): + return manifest + raise master.DownloadError("connection failed") + with pytest.raises(master.DownloadError): + master.download("https://cdn.invalid", "v", tmp_path, get=get) + assert (tmp_path / "MasterManifest.json").read_bytes() == manifest + + # ---------------------------------------------------------------- the stored JSON def test_table_json_is_the_json_rule(): finite = {"_allData": [{"_id": 1, "_name": "テスト\n\"q\"", "_v": 0.1, "_w": -2.5e-12, "_big": 1e300, "_n": None, From 96ae15e8ac151009e522e877125bb7fcae9e7f21 Mon Sep 17 00:00:00 2001 From: nichinichisou Date: Wed, 30 Sep 2026 03:17:00 +0800 Subject: [PATCH 12/14] ci: check the music data seeds, rank bonuses and luck points The deck gate checks the seeds (the one seed 0 on a chart without a luck range, else two or more different seeds, the same on every luck chart: their number is the file's, not fixed), every seed range's rank 1 bonus (trunc(rangeScore * rankBonusPercent / 100)) and its luckPoints, which nnnotes and ournotes-deck add with the Gekisou skills. RECIPE 2. Co-Authored-By: Claude Opus 5.5 --- .github/MUSIC_DATA.md | 9 +++-- .github/scripts/music_data.py | 41 +++++++++++++++++++- .github/scripts/test_music_data.py | 62 ++++++++++++++++++++++-------- 3 files changed, 91 insertions(+), 21 deletions(-) diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md index f8b22ab..9aca4f8 100644 --- a/.github/MUSIC_DATA.md +++ b/.github/MUSIC_DATA.md @@ -68,7 +68,7 @@ Every one must pass, else nothing is published. Warnings go to the job summary a | `schema` | the file against `docs/schema/music-data.schema.json` of the checkout (JSON Schema 2020-12) | | `provenance` | `format`; `region` `tw`; `master.source` `api`; `master.version` equal to the snapshot's and its `MasterManifest.json`'s; every table's SHA-256 the manifest's, every decoded table read the one `index.json` lists; the song tables and the deck model's present; `deck.commit` the one `rust/Cargo.lock` pins; `exporter.version` the installed nnnotes; an APK version; a catalog SHA-256 (warning: the APK is another client version than the snapshot's) | | `counts` | no fewer songs and charts than the published file (warning: ids no longer in it) | -| `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning), `weights[kind][position]` numbers, every check deck within its bound | +| `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning): the one seed 0 on a chart without a luck range, else two or more different seeds, the same on every luck chart (their number is the file's, not fixed), `weights[kind][position]` numbers, every seed range's `rankBonus` = trunc(`rangeScore` x `rankBonusPercent` / 100) and its `luckPoints` an int, every check deck within its bound | | `scenarios` | the play scenario fields: `offSeeds` exactly one entry (seed 0, score, weights, check within its bound), every range's `rankBonusPercents` five ints (the first `rankBonusPercent`), every seed's `scorePerfect`, `rangeWeights` (`[kind][position][range]`) and `rankCheck` (within its bound), every seed range's `rangeScorePerfect` (warnings, none in TW: a null `rangeWeights`, a null kind in it or in `offSeeds`' weights) | | `finite` | no NaN or infinity (warning: one inside master data rows, `songs[].master`, which the format writes as `1e999`) | | `references` | texts in every language of `languages` (names and titles not empty); unique ids; songs sorted; the songs' bands, vocal characters and tags in the file; a band or a band name; a jacket, and its file in `jackets/`; a BGM cue; score ranks; charts in difficulty order, score ids unique (warnings: a title without a `zh-Hant` text, a music category on no tab, a character of no band) | @@ -89,7 +89,8 @@ python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py with, optionally, `MUSIC_DATA_SCHEMA` (a schema file when the checkout has none), `MUSIC_DATA_PAGE` (an `examples/songs` directory: the smoke test), `MUSIC_DATA_SAMPLE` (a real file with the play scenario fields: its -content gates pass) and `MUSIC_DATA_OLD_SAMPLE` (one without them: the scenario gate stops it). +content gates pass; one made before the ranges' `luckPoints`: the deck gate stops it on those alone) and +`MUSIC_DATA_OLD_SAMPLE` (one without the play scenario fields: the scenario gate stops it). ## Settings @@ -112,7 +113,9 @@ Repository variables: - **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6, `12df2a6`): sync the - fork with upstream first. Until then `plan` stops naming what is missing. + fork with upstream first. Until then `plan` stops naming what is missing. The deck gate also needs the ranges' + `luckPoints`, which come with nnnotes' and ournotes-deck's Gekisou skill changes: until the fork has them every + build stops there. - **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play scenario fields, once that page is merged. - **Publishing.** Set `MUSIC_DATA_PUBLISH` to `true` last, when dry runs pass and the published format is final. diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py index 0c336b5..05af9ff 100755 --- a/.github/scripts/music_data.py +++ b/.github/scripts/music_data.py @@ -49,7 +49,7 @@ BUILD_FORMAT = "moenotes.music-data-build/1" # This script's own version of a build: bump it when what it builds or publishes changes, so that the next run builds # although the master data, the deck model and nnnotes are the same. -RECIPE = 1 +RECIPE = 2 FILE, MARKER, JACKETS, ARCHIVE = "music-data.json", "build.json", "jackets/", "archive/" MANIFEST = "MasterManifest.json" SOURCE_PATHS = ("src", "rust", "pyproject.toml") # nnnotes' code: the commit that last changed one of them @@ -66,6 +66,7 @@ BGM_CUE_SLACK_MS = 1000 # |durationMs - lengthMs| BGM_TAIL_MS = 60_000 # BGM after the last note: more is reported RANKS = 5 +LUCK_MISSION = 2 # a range of it draws lots: its chart has several seeds DIFFICULTIES = ("easy", "normal", "hard", "expert") SONG_TABLES = ("MasterLiveMusic", "MasterLiveMusicScore", "MasterText", "MasterBand", "MasterCharacter", "MasterTag", "MasterLiveMusicCategory", "MasterSound", "MasterSoundCueSheet", "MasterLiveScoreRank") @@ -290,6 +291,16 @@ def within(c) -> bool: and abs(c["exact"] - c["predicted"]) <= c["bound"]) +def rank_bonus(range_score: int, percent: int) -> int: + """trunc(rangeScore * percent / 100): a range's rank bonus.""" + q = abs(range_score * percent) // 100 + return q if range_score * percent >= 0 else -q + + +def luck_chart(deck: dict) -> bool: + return any(isinstance(r, dict) and r.get("mission") == LUCK_MISSION for r in deck.get("ranges") or []) + + def gate_schema(doc, ctx: Context, g: Gate): if ctx.schema is None: g.note = "skipped" @@ -410,6 +421,9 @@ def weights_shape(w, kinds: int, positions: int, nullable: bool) -> bool: def gate_deck(doc, ctx: Context, g: Gate): + """The deck statistics. Seeds (deck.model.seeds): the one seed 0 on a chart without a luck range, else two or more + seeds, the same on every luck chart (their number is the file's); every range's rankBonus the rank 1 bonus and + its luckPoints (the range's luck points without skills).""" deck = doc.get("deck") if not isinstance(deck, dict): g.fail("deck is null: no deck statistics (made with --no-deck?)") @@ -421,6 +435,7 @@ def gate_deck(doc, ctx: Context, g: Gate): if not (is_num(power) and power > 0): g.fail(f"deck.model.power {power!r}") n = unplayable = seeds = 0 + luck_seeds = None for song, chart in charts_of(doc): n += 1 w, d = where(song, chart), chart.get("deck") @@ -440,6 +455,17 @@ def gate_deck(doc, ctx: Context, g: Gate): g.fail(f"{w}: unplayable, but has Gekisou on seeds") elif not d.get("seeds"): g.fail(f"{w}: no seeds") + values = [s.get("seed") for s in d.get("seeds") or []] + if values and not luck_chart(d): + if len(values) != 1 or not is_int(values[0]) or values[0] != 0: + g.fail(f"{w}: seeds {values[:4]!r} without a luck range, expected the one seed 0") + elif values: + if not all(map(is_int, values)) or len(values) < 2 or len(set(values)) != len(values): + g.fail(f"{w}: {len(values)} seeds on a luck chart, expected two or more different int seeds") + elif luck_seeds is None: + luck_seeds = values + elif values != luck_seeds: + g.fail(f"{w}: its {len(values)} seeds are not the {len(luck_seeds)} of the first luck chart") for seed in d.get("seeds") or []: seeds += 1 s = f"{w} seed {seed.get('seed')}" @@ -449,9 +475,20 @@ def gate_deck(doc, ctx: Context, g: Gate): g.fail(f"{s}: weights are not [kind][position] numbers") if len(seed.get("ranges") or []) != len(d.get("ranges") or []): g.fail(f"{s}: {len(seed.get('ranges') or [])} range results for {len(d.get('ranges') or [])} ranges") + for i, (r, rr) in enumerate(zip(seed.get("ranges") or [], d.get("ranges") or [], strict=False)): + r, rr = (r if isinstance(r, dict) else {}), (rr if isinstance(rr, dict) else {}) + if not (is_int(r.get("rangeScore")) and is_int(r.get("rankBonus"))): + g.fail(f"{s} range {i}: rangeScore {r.get('rangeScore')!r}, rankBonus {r.get('rankBonus')!r}") + elif is_int(rr.get("rankBonusPercent")) and r["rankBonus"] != rank_bonus(r["rangeScore"], + rr["rankBonusPercent"]): + g.fail(f"{s} range {i}: rankBonus {r['rankBonus']} is not trunc({r['rangeScore']} * " + f"{rr['rankBonusPercent']} / 100)") + if not is_int(r.get("luckPoints")): + g.fail(f"{s} range {i}: luckPoints {'missing' if 'luckPoints' not in r else 'not an int'}") if not within(seed.get("check")): g.fail(f"{s}: the check deck is not within its bound") - g.note = f"{n} charts, {kinds} kinds, {seeds} seeds, {unplayable} unplayable" + g.note = (f"{n} charts, {kinds} kinds, {seeds} seeds, {unplayable} unplayable, " + f"{len(luck_seeds or [])} seeds per luck chart") def gate_scenarios(doc, ctx: Context, g: Gate): diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py index bc86b88..c90400f 100644 --- a/.github/scripts/test_music_data.py +++ b/.github/scripts/test_music_data.py @@ -5,8 +5,9 @@ The JSON Schema gate uses docs/schema/music-data.schema.json (or $MUSIC_DATA_SCHEMA) when the checkout has it; the page smoke test runs when $MUSIC_DATA_PAGE names ournotes-player's examples/songs (and Node.js is installed). Real -files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass) and -$MUSIC_DATA_OLD_SAMPLE (one without them: the scenario gate stops it). +files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass; one made before +the ranges' luckPoints, the deck gate stops on those alone) and $MUSIC_DATA_OLD_SAMPLE (one without the play scenario +fields: the scenario gate stops it). """ import copy import hashlib @@ -53,27 +54,31 @@ def check_deck(exact): return {"deck": [[0, 5000], None], "exact": exact, "predicted": exact + 0.25, "bound": 7.0} -def deck_chart(): - seed = {"seed": 0, "score": 120000, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, - "justCount": 0, "lotResults": [0, 0, 0, 0], - "rangeScorePerfect": 4000}], - "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000, - "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} +LUCK_SEEDS = [11, 22] # a luck chart's seeds; else the one seed 0 + + +def deck_chart(luck=False): + def one(n): + return {"seed": n, "score": 120000 + n, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, + "justCount": 0, "luckPoints": 3 if luck else 0, + "lotResults": [0, 0, 0, 0], "rangeScorePerfect": 4000}], + "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000 + n, + "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} return {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, - "ranges": [{"index": 0, "mission": 1, "startMs": 1000, "endMs": 5000, "rankBonusPercent": 250, - "rankBonusPercents": [250, 190, 160, 100, 100]}], - "justNotes": 0, "seeds": [seed], + "ranges": [{"index": 0, "mission": 2 if luck else 1, "startMs": 1000, "endMs": 5000, + "rankBonusPercent": 250, "rankBonusPercents": [250, 190, 160, 100, 100]}], + "justNotes": 0, "seeds": [one(n) for n in (LUCK_SEEDS if luck else [0])], "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], "unplayable": None} -def chart(difficulty, score_id, last=60000): +def chart(difficulty, score_id, last=60000, luck=False): return {"difficulty": difficulty, "scoreId": score_id, "level": 10, "displayLevel": 10.5, "fullComboCount": 20, "asset": {"key": f"Live/MusicScore/c/c_{score_id}", "sha256": "ab" * 32}, "notes": {"judged": 20, "total": 22, "byOperateType": {"1": 20, "120": 2}}, "bpm": {"main": 120.0, "min": 120.0, "max": 120.0, "changes": [{"timeMs": 0, "bpm": 120.0}]}, "firstNoteMs": 1000, "lastJudgedNoteMs": last, "lastNoteMs": last, "musicLengthMs": last + 1000, - "skillEventsMs": [1000, 3000], "fevers": [[1000, 5000]], "deck": deck_chart()} + "skillEventsMs": [1000, 3000], "fevers": [[1000, 5000]], "deck": deck_chart(luck)} def song(i, charts): @@ -109,7 +114,8 @@ def sample() -> dict: "tags": [{"id": 1, "name": text("tag")}], "categories": [{"id": 1, "musicCategories": [1], "name": text("cat")}], "deck": {"model": {"power": 300000, "checkPower": 1000003}, "kinds": [kind]}, - "songs": [song(100001, [chart("easy", 10), chart("expert", 30)]), song(100002, [chart("expert", 40)])], + "songs": [song(100001, [chart("easy", 10), chart("expert", 30, luck=True)]), + song(100002, [chart("expert", 40, luck=True)])], } @@ -216,6 +222,14 @@ def unplayable(doc): (lambda d: d["songs"][0]["charts"][0]["deck"].update(seeds=[]), "deck", "no seeds"), (unplayable, "deck", "unplayable, but has Gekisou on seeds"), (lambda d: d["songs"][0]["charts"][0]["deck"].update(events=[[0, 1000]]), "deck", "1 skill events"), + (lambda d: seed(d).update(seed=5), "deck", "seeds [5] without a luck range, expected the one seed 0"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["seeds"].pop(), "deck", "1 seeds on a luck chart"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["seeds"][1].update(seed=11), "deck", "2 seeds on a luck chart"), + (lambda d: d["songs"][1]["charts"][0]["deck"]["seeds"][1].update(seed=33), "deck", + "its 2 seeds are not the 2 of the first luck chart"), + (lambda d: seed(d)["ranges"][0].update(rankBonus=9999), "deck", "rankBonus 9999 is not trunc(4000 * 250 / 100)"), + (lambda d: seed(d)["ranges"][0].pop("luckPoints"), "deck", "seed 0 range 0: luckPoints missing"), + (lambda d: seed(d, 0, 1)["ranges"][0].update(luckPoints=2.5), "deck", "seed 11 range 0: luckPoints not an int"), # numbers (lambda d: seed(d)["weights"][0].__setitem__(1, float("nan")), "finite", "weights[0][1]: not finite"), (lambda d: d["songs"][1]["charts"][0]["bpm"].update(main=float("inf")), "finite", "bpm.main: not finite"), @@ -339,10 +353,23 @@ def real(name): return Path(p).read_bytes() +def luck_points(raw: bytes) -> bool: + """Whether a file's deck.seeds ranges have luckPoints (a file made before them has none).""" + return any("luckPoints" in r for _, c in music_data.charts_of(json.loads(raw)) + for s in (c.get("deck") or {}).get("seeds") or [] for r in s.get("ranges") or []) + + +def before_luck_points(r: dict) -> bool: + """The deck gate stopped only on the ranges' missing luckPoints.""" + g = gate(r, "deck") + return not g["passed"] and all(f.endswith("luckPoints missing") for f in g["failures"]) + + def test_a_real_file_without_the_scenario_fields(): raw = real("MUSIC_DATA_OLD_SAMPLE") r = gates(raw, Context(language="zh-Hant"), only=CONTENT) - assert failures(r) == [("scenarios", gate(r, "scenarios")["failures"])] + assert [n for n, _ in failures(r)] == (["scenarios"] if luck_points(raw) else ["deck", "scenarios"]) + assert luck_points(raw) or before_luck_points(r) g = gate(r, "scenarios") charts = sum(len(s["charts"]) for s in json.loads(raw)["songs"]) assert g["failureCount"] >= charts and all("offSeeds missing" in f or "rankBonusPercents missing" in f @@ -354,7 +381,10 @@ def test_a_real_file_with_the_scenario_fields(tmp_path): (tmp_path / "music-data.json").write_bytes(raw) ctx = Context(language="zh-Hant", page=page_path(), file=tmp_path / "music-data.json", published=raw) r = gates(raw, ctx, only=CONTENT + ("counts", "size", "page")) - assert r["passed"], failures(r) + if luck_points(raw): + assert r["passed"], failures(r) + else: + assert [n for n, _ in failures(r)] == ["deck"] and before_luck_points(r), failures(r) assert gate(r, "scenarios")["warningCount"] == 0 From 4c22bfe2ed93d1cd1410425a8ff574a7200211a6 Mon Sep 17 00:00:00 2001 From: nichinichisou Date: Wed, 30 Sep 2026 04:49:01 +0800 Subject: [PATCH 13/14] ci: validate Gekisou aptitude and smoke test the page API Check shape references, mission coverage, band variants, mean/error pairs, seed rules, dimensions, tail identities and simulation bounds. Cap gzip size at 2 MB overall and 1.2 MB for aptitude. Exercise page lookups, gains, uncertainty and missing cross terms. Default figures must remain unchanged. Reconstruct only deterministic check seeds with positional master skill factors, never stochastic means. Add a real chart-stats sample test and a necessary rounding bound for stochastic tailPerfect. Keep production PLAYER_REF and publishing unchanged. Co-Authored-By: Claude Code --- .github/MUSIC_DATA.md | 26 +- .github/scripts/music_data.py | 380 ++++++++++++++++++++++++++- .github/scripts/music_data_smoke.mjs | 83 +++++- .github/scripts/test_music_data.py | 255 ++++++++++++++++-- 4 files changed, 719 insertions(+), 25 deletions(-) diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md index 9aca4f8..6c24ae5 100644 --- a/.github/MUSIC_DATA.md +++ b/.github/MUSIC_DATA.md @@ -2,7 +2,8 @@ `.github/workflows/music-data.yml` keeps the music data file of the chart data page (ournotes-player `examples/songs`) up to date: `nnnotes music-data` of the current master data ([docs/music-data.md](../docs/music-data.md): -every song and chart with the deck model's statistics and the play scenarios), checked by quality gates and +every song and chart with the deck model's statistics, the play scenarios and the chart's Gekisou skill aptitude), +checked by quality gates and published into the story site's bucket under `music-data/` (`https://storage.bdon.moe/moenotes/music-data/`): | Object | Content | Cache-Control | @@ -70,10 +71,12 @@ Every one must pass, else nothing is published. Warnings go to the job summary a | `counts` | no fewer songs and charts than the published file (warning: ids no longer in it) | | `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning): the one seed 0 on a chart without a luck range, else two or more different seeds, the same on every luck chart (their number is the file's, not fixed), `weights[kind][position]` numbers, every seed range's `rankBonus` = trunc(`rangeScore` x `rankBonusPercent` / 100) and its `luckPoints` an int, every check deck within its bound | | `scenarios` | the play scenario fields: `offSeeds` exactly one entry (seed 0, score, weights, check within its bound), every range's `rankBonusPercents` five ints (the first `rankBonusPercent`), every seed's `scorePerfect`, `rangeWeights` (`[kind][position][range]`) and `rankCheck` (within its bound), every seed range's `rangeScorePerfect` (warnings, none in TW: a null `rangeWeights`, a null kind in it or in `offSeeds`' weights) | +| `aptitude` | the Gekisou skill aptitude (every shape alone on a chart). `deck.model.gekisouAptitude` a text; `deck.gekisouAptitude`: every key, `plainKind` the page's plain kind, a `host` text, the `seedRule` (a deterministic test, increasing batches, the targets, the cross seeds), `shapes` numbered 0, 1, 2, ... (source `member` or `support`, mission 1 to 4, `bandCondition` a support skill's alone and exactly when an effect has condition 5000, effect rows with every key and their condition groups, condition 5000 without targets, skills with a level and, with a band condition alone, member targets and bands). Every chart's `deck.gekisouAptitude`: null exactly when the chart is unplayable with Gekisou on, has no Gekisou range or there is no shape; else `factors` one per range (counts; no Just or Perfect notes outside a Just range; `lotteries` `[0, 0]` outside a luck range, else the mean of `deck.seeds`' `lotResults`) and `variants` one per shape of the chart's missions (or mission 4) in shape order, a band condition shape's `bandMatch` true then false: every `[mean, se]` two finite numbers with se >= 0 (every se 0 when deterministic), ranges one per range, `tail` = `score` less the ranges' `rangeScore` and `rankBonus` (allowing 0.0005 per rounded term plus 1e-6), deterministic point deltas integers and `tailPerfect` checked against the baseline Perfect range bonuses, 1 seed when deterministic else a batch of the seed rule (the last one when `seTargetMet` is false), `crossSeeds` min(seeds, the rule's), `weights` one per position and `rangeWeights` per position and range where the plain kind and `deck.seeds[0].rangeWeights` are, else null, the `check` on `deck.seeds[0]`'s seed, a rank per range (1 where the ranks are not linear), a plain kind value or null per position, within its bound (warning: variants that missed the standard error target) | | `finite` | no NaN or infinity (warning: one inside master data rows, `songs[].master`, which the format writes as `1e999`) | | `references` | texts in every language of `languages` (names and titles not empty); unique ids; songs sorted; the songs' bands, vocal characters and tags in the file; a band or a band name; a jacket, and its file in `jackets/`; a BGM cue; score ranks; charts in difficulty order, score ids unique (warnings: a title without a `zh-Hant` text, a music category on no tab, a character of no band) | | `bgm` | every song's BGM length: `durationMs = samples * 1000 // sampleRate`, 30 s to 10 min, within 1 s of the cue's `lengthMs`, not ending before a chart's last note (warning: more than a minute after it) | | `size` | 0.8 to 2 times the published file | +| `gzip` | the file gzipped at most 2 MB (0.37 MB before the aptitude), its Gekisou skill aptitude gzipped at most 1.2 MB (about 0.4 MB expected) | | `page` | `music_data_smoke.mjs`: the page's `catalog.js` and `ranking.js` in Node.js over the file: a row per chart, a plain score-up kind, data for the free, rank and Just scenarios, finite positive figures for every chart the data covers in seven scenarios (Gekisou Live at several ranks, Just rates and a Great share, Free Live), the ranking, frontier and event figures | | (publish) | read back after upload, SHA-256 checked | @@ -89,9 +92,22 @@ python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py with, optionally, `MUSIC_DATA_SCHEMA` (a schema file when the checkout has none), `MUSIC_DATA_PAGE` (an `examples/songs` directory: the smoke test), `MUSIC_DATA_SAMPLE` (a real file with the play scenario fields: its -content gates pass; one made before the ranges' `luckPoints`: the deck gate stops it on those alone) and +content gates pass; one made before the ranges' `luckPoints` and the aptitude: the deck and aptitude gates stop it +on those alone) and `MUSIC_DATA_OLD_SAMPLE` (one without the play scenario fields: the scenario gate stops it). +### Aptitude page smoke + +With the aptitude API (ournotes-player PR #11, `1522c24`), the same Node smoke also checks shape/skill/band +lookups, chart variants, all five battle scenarios, Free Live exclusion, finite gains, raw standard errors, +missing cross terms and the absence of standard errors for transformed or combined figures. Removing aptitude +must not change the default chart figures: default ranking still has no card Gekisou skills. + +Only deterministic variants are reconstructed against their individual `check` seed, using positional cards and +`masterSkillFactor` for the game's float32 conversion. Stochastic means are never used to reconstruct a check. +Older pinned page modules explicitly report `API unavailable (skipped)`; moving `MUSIC_DATA_PLAYER_REF` remains a +separate rollout decision. No browser or page build is needed. + ## Settings Repository secrets: the story site's (`.github/STORY_SITE.md`), no new one: `NNNOTES_BUNDLE_KEY`, @@ -113,9 +129,9 @@ Repository variables: - **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6, `12df2a6`): sync the - fork with upstream first. Until then `plan` stops naming what is missing. The deck gate also needs the ranges' - `luckPoints`, which come with nnnotes' and ournotes-deck's Gekisou skill changes: until the fork has them every - build stops there. + fork with upstream first. Until then `plan` stops naming what is missing. The deck and aptitude gates also need + the ranges' `luckPoints` and the Gekisou skill aptitude, which come with nnnotes' and ournotes-deck's Gekisou skill + changes: until the fork has them every build stops there. - **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play scenario fields, once that page is merged. - **Publishing.** Set `MUSIC_DATA_PUBLISH` to `true` last, when dry runs pass and the published format is final. diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py index 05af9ff..a9116a0 100755 --- a/.github/scripts/music_data.py +++ b/.github/scripts/music_data.py @@ -27,6 +27,7 @@ """ from __future__ import annotations +import gzip import hashlib import json import math @@ -62,11 +63,18 @@ # the gates' bounds (MUSIC_DATA.md) SIZE_RATIO = (0.8, 2.0) # against the published file +FILE_GZIP_MAX = 2_000_000 # the file gzipped, the download (0.37 MB before the aptitude) +APTITUDE_GZIP_MAX = 1_200_000 # the Gekisou skill aptitude gzipped (about 0.4 MB expected) BGM_MS = (30_000, 600_000) # a song's BGM length BGM_CUE_SLACK_MS = 1000 # |durationMs - lengthMs| BGM_TAIL_MS = 60_000 # BGM after the last note: more is reported RANKS = 5 LUCK_MISSION = 2 # a range of it draws lots: its chart has several seeds +JUST_MISSION = 3 # a range of it judges Just +MISSIONS = (1, 2, 3, 4) # a Gekisou skill's: combo, luck, Just, every one +PLAIN_KIND = (2000, 5000) # the page's plain kind (ranking.js plainKind): effect type, ms +BAND_CONDITION = 5000 # the skill condition on the paired member (its band) +APTITUDE_SLACK = 1e-6 # relative: the aptitude's identities on means of integers DIFFICULTIES = ("easy", "normal", "hard", "expert") SONG_TABLES = ("MasterLiveMusic", "MasterLiveMusicScore", "MasterText", "MasterBand", "MasterCharacter", "MasterTag", "MasterLiveMusicCategory", "MasterSound", "MasterSoundCueSheet", "MasterLiveScoreRank") @@ -297,6 +305,18 @@ def rank_bonus(range_score: int, percent: int) -> int: return q if range_score * percent >= 0 else -q +def plain_kind(doc: dict): + """The id of the page's plain score-up kind in deck.kinds (ranking.js plainKind): effect type 2000 on the whole + deck for 5 s, without targets, conditions or limits; None for none.""" + for k in ((doc.get("deck") or {}).get("kinds")) or []: + if (isinstance(k, dict) and k.get("effectType") == PLAIN_KIND[0] and not k.get("skillTargetIds") + and not any(k.get(x) for x in ("skillConditionGroup", "skillReleaseConditionGroup", + "effectLimitCount", "effectExecuteLimitCount")) + and (PLAIN_KIND[1] if k.get("durationMs") is None else k["durationMs"]) == PLAIN_KIND[1]): + return k.get("id") + return None + + def luck_chart(deck: dict) -> bool: return any(isinstance(r, dict) and r.get("mission") == LUCK_MISSION for r in deck.get("ranges") or []) @@ -561,6 +581,359 @@ def gate_scenarios(doc, ctx: Context, g: Gate): g.note = "offSeeds, rankBonusPercents, scorePerfect, rangeWeights, rankCheck, rangeScorePerfect" +APTITUDE_KEYS = ("plainKind", "host", "seedRule", "shapes") +SEED_RULE_KEYS = ("deterministicTest", "batches", "relative", "baseline", "crossSeeds") +SHAPE_KEYS = ("id", "source", "mission", "bandCondition", "effects", "skills") +EFFECT_INTS = ("effectType", "triggerType", "effectValue", "maxEffectValue", "effectLimitCount", + "effectExecuteLimitCount") +EFFECT_GROUPS = ("trigger", "condition", "release", "reset") +FACTOR_INTS = ("judgedNotes", "justNotes", "perfectNotes", "tailNotes", "comboAtStart") +VARIANT_KEYS = ("shape", "bandMatch", "deterministic", "seeds", "seTargetMet", "crossSeeds", "score", "scorePerfect", + "tail", "tailPerfect", "converted", "ranges", "weights", "rangeWeights", "check") +VARIANT_PAIRS = ("score", "scorePerfect", "tail", "tailPerfect", "converted") +VARIANT_RANGE_PAIRS = ("rangeScore", "rankBonus", "rangeScorePerfect", "maxCombo", "justCount", "luckPoints") +APTITUDE_CHECK_KEYS = ("seed", "ranks", "deck", "exact", "predicted", "bound") + + +def pair(v) -> bool: + """A [mean, standard error]: two finite numbers, the error not negative.""" + return isinstance(v, list) and len(v) == 2 and is_num(v[0]) and is_num(v[1]) and v[1] >= 0 + + +def close(a: float, b: float) -> bool: + return abs(a - b) <= APTITUDE_SLACK * max(1.0, abs(a), abs(b)) + + +def gate_aptitude(doc, ctx: Context, g: Gate): + """The charts' Gekisou skill aptitude: deck.gekisouAptitude (its plain kind the page's, the seed rule, shapes + numbered from 0: source, mission, band condition, effect rows, skills); every chart's deck.gekisouAptitude, null + exactly when the chart is unplayable with Gekisou on, has no Gekisou range or there is no shape; else factors per + range and one variant per shape of the chart's missions (or mission 4) in shape order, a band condition shape's + bandMatch true then false: every [mean, se] two finite numbers with se >= 0 (0 when deterministic), ranges per + range, tail = score - the ranges' rangeScore and rankBonus, the seeds and the cross seeds by the seed rule, weights + and rangeWeights where the plain kind and the chart's rank weights are, the check within its bound. Warnings: + variants whose standard error missed the seed rule's target.""" + deck = doc.get("deck") + if not isinstance(deck, dict): + g.fail("deck is null: no Gekisou skill aptitude") + return + plain = plain_kind(doc) + if not (isinstance((deck.get("model") or {}).get("gekisouAptitude"), str) and deck["model"]["gekisouAptitude"]): + g.fail("deck.model.gekisouAptitude: no text") + head = deck.get("gekisouAptitude") + if not isinstance(head, dict): + g.fail("deck.gekisouAptitude missing" if head is None else f"deck.gekisouAptitude {head!r}") + head = {} + elif [k for k in APTITUDE_KEYS if k not in head]: + g.fail(f"deck.gekisouAptitude: no {', '.join(k for k in APTITUDE_KEYS if k not in head)}") + if head: + pk = head.get("plainKind", "missing") + if not (pk is None or is_int(pk)) or pk != plain: + g.fail(f"deck.gekisouAptitude.plainKind {pk!r}, the page's plain kind is {plain!r}") + if not (isinstance(head.get("host"), str) and head["host"].strip()): + g.fail("deck.gekisouAptitude.host: no text") + rule = head.get("seedRule") if isinstance(head.get("seedRule"), dict) else {} + batches = rule.get("batches") + if head and not (all(k in rule for k in SEED_RULE_KEYS) and is_int(rule["deterministicTest"]) + and rule["deterministicTest"] >= 1 and isinstance(batches, list) and batches + and all(map(is_int, batches)) and batches == sorted(set(batches)) and batches[0] >= 1 + and is_num(rule["relative"]) and rule["relative"] >= 0 and is_num(rule["baseline"]) + and rule["baseline"] >= 0 and is_int(rule["crossSeeds"]) and rule["crossSeeds"] >= 1): + g.fail(f"deck.gekisouAptitude.seedRule {rule!r}"[:200]) + rule, batches = {}, None + + # the shapes + shapes = head.get("shapes") if isinstance(head.get("shapes"), list) else [] + if head and not isinstance(head.get("shapes"), list): + g.fail("deck.gekisouAptitude.shapes is not a list") + if any(not isinstance(s, dict) or not is_int(s.get("id")) or s["id"] != i + for i, s in enumerate(shapes)): + g.fail("deck.gekisouAptitude.shapes: ids are not 0, 1, 2, ... in order") + by_id: dict = {} + for s in shapes: + if not isinstance(s, dict): + continue + w = f"shape {s.get('id')}" + absent = [k for k in SHAPE_KEYS if k not in s] + if absent: + g.fail(f"{w}: no {', '.join(absent)}") + continue + if is_int(s["id"]): + by_id[s["id"]] = s + if s["source"] not in ("member", "support"): + g.fail(f"{w}: source {s['source']!r}") + if not is_int(s["mission"]) or s["mission"] not in MISSIONS: + g.fail(f"{w}: mission {s['mission']!r}") + band = s["bandCondition"] + if not isinstance(band, bool) or (band and s["source"] != "support"): + g.fail(f"{w}: bandCondition {band!r} (a support skill's alone)") + effects = s["effects"] + fives = 0 + if not isinstance(effects, list): + g.fail(f"{w}: effects are not a list") + effects = [] + for i, e in enumerate(effects): + if not isinstance(e, dict): + g.fail(f"{w} effect {i}: not an object") + continue + bad = [k for k in EFFECT_INTS if not is_int(e.get(k))] + if not is_num(e.get("activationTimeSecond")): + bad.append("activationTimeSecond") + if not (isinstance(e.get("skillTargetIds"), list) and all(map(is_int, e["skillTargetIds"]))): + bad.append("skillTargetIds") + for k in EFFECT_GROUPS: + sets = e.get(k) + if not (isinstance(sets, list) and all(isinstance(x, list) for x in sets)): + bad.append(k) + continue + for c in (c for x in sets for c in x): + if not (isinstance(c, dict) and is_int(c.get("type")) and isinstance(c.get("values"), list) + and isinstance(c.get("positive"), bool) and "targetIds" in c): + bad.append(k) + elif c["type"] == BAND_CONDITION: + fives += 1 + if c["targetIds"] is not None: + bad.append(f"{k} (condition {BAND_CONDITION} targetIds not null)") + elif not (isinstance(c["targetIds"], list) and all(map(is_int, c["targetIds"]))): + bad.append(k) + cu = e.get("cumulative", "missing") + if cu is not None and not (isinstance(cu, dict) and all( + k in cu for k in ("type", "values", "targetIds", "maxCumulativeCount"))): + bad.append("cumulative") + if bad: + g.fail(f"{w} effect {i}: {', '.join(dict.fromkeys(bad))} missing or malformed") + if isinstance(band, bool) and band != (fives > 0): + g.fail(f"{w}: bandCondition {band}, its effects have {fives} condition {BAND_CONDITION}") + skills = s["skills"] + if not (isinstance(skills, list) and skills): + g.fail(f"{w}: no skills") + continue + for k in skills: + ok = (isinstance(k, dict) and is_int(k.get("id")) and is_int(k.get("level")) and k["level"] >= 1 + and "memberTargetIds" in k and "bandIds" in k) + targets, band_ids = (k.get("memberTargetIds"), k.get("bandIds")) if isinstance(k, dict) else (0, 0) + if band is True: + ok = (ok and isinstance(targets, list) and bool(targets) and all(map(is_int, targets)) + and targets == sorted(set(targets)) and isinstance(band_ids, list) + and all(map(is_int, band_ids))) + else: + ok = ok and targets is None and band_ids is None + if not ok: + g.fail(f"{w}: skill {k!r} is not an id, a level and (with a band condition alone) member targets " + f"and bands"[:240]) + + # the charts + aptitudes = nulls = variants = deterministic = 0 + missed = [] + for song, chart in charts_of(doc): + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + continue # the deck gate fails it + if "gekisouAptitude" not in d: + g.fail(f"{w}: no deck.gekisouAptitude") + continue + a, ranges, dseeds = d["gekisouAptitude"], d.get("ranges") or [], d.get("seeds") or [] + why = ("unplayable with Gekisou on" if d.get("unplayable") else "without a Gekisou range" if not ranges + else None if shapes else "without a Gekisou skill shape") + if a is None: + nulls += 1 + if why is None: + g.fail(f"{w}: deck.gekisouAptitude is null, but the chart is playable with Gekisou on") + continue + if why is not None: + g.fail(f"{w}: a Gekisou skill aptitude on a chart {why}") + continue + if not (isinstance(a, dict) and isinstance(a.get("factors"), list) and isinstance(a.get("variants"), list)): + g.fail(f"{w}: deck.gekisouAptitude has no factors and variants lists") + continue + aptitudes += 1 + positions = d.get("positions") + missions = {r.get("mission") for r in ranges if isinstance(r, dict)} + linear = bool(dseeds) and dseeds[0].get("rangeWeights") is not None + + # factors + if len(a["factors"]) != len(ranges): + g.fail(f"{w}: {len(a['factors'])} factors for {len(ranges)} ranges") + for j, (f, r) in enumerate(zip(a["factors"], ranges, strict=False)): + f, r = (f if isinstance(f, dict) else {}), (r if isinstance(r, dict) else {}) + bad = [k for k in FACTOR_INTS if not (is_int(f.get(k)) and f[k] >= 0)] + if not pair(f.get("lotteries")): + bad.append("lotteries") + if bad: + g.fail(f"{w} factors {j}: {', '.join(bad)} missing or not counts") + continue + if r.get("mission") != JUST_MISSION and (f["justNotes"] or f["perfectNotes"]): + g.fail(f"{w} factors {j}: Just or Perfect notes in a range without the Just mission") + if f["justNotes"] + f["perfectNotes"] > f["judgedNotes"]: + g.fail(f"{w} factors {j}: {f['justNotes']} Just and {f['perfectNotes']} Perfect notes of " + f"{f['judgedNotes']} judged") + at = [s["ranges"][j] for s in dseeds if isinstance(s.get("ranges"), list) and len(s["ranges"]) > j + and isinstance(s["ranges"][j], dict)] + lots = [sum(x["lotResults"]) for x in at + if isinstance(x.get("lotResults"), list) and all(map(is_int, x["lotResults"]))] + if r.get("mission") != LUCK_MISSION and f["lotteries"] != [0, 0]: + g.fail(f"{w} factors {j}: lotteries {f['lotteries']} in a range without the luck mission") + elif lots and len(lots) == len(dseeds) and not close(f["lotteries"][0], sum(lots) / len(lots)): + g.fail(f"{w} factors {j}: lotteries {f['lotteries'][0]}, deck.seeds' lotResults give " + f"{sum(lots) / len(lots)}") + + # the variants: which, in order + want = [(s["id"], match) for s in sorted(by_id.values(), key=lambda s: s["id"]) + if s["mission"] == 4 or s["mission"] in missions + for match in ((True, False) if s["bandCondition"] is True else (None,))] + got = [(v.get("shape"), v.get("bandMatch")) if isinstance(v, dict) else None for v in a["variants"]] + unknown = [v[0] for v in got if v is not None and not (is_int(v[0]) and v[0] in by_id)] + other = [v[0] for v in got if v is not None and is_int(v[0]) and v[0] in by_id + and by_id[v[0]]["mission"] != 4 and by_id[v[0]]["mission"] not in missions] + if unknown: + g.fail(f"{w}: variants of shapes {sorted(set(map(repr, unknown)))} not in deck.gekisouAptitude.shapes") + if other: + g.fail(f"{w}: variants of shapes {sorted(set(other))} of a mission the chart does not play") + if got != want and not unknown and not other: + lacking = [x for x in want if x not in got] + g.fail(f"{w}: variants {'lack ' + repr(lacking[:6]) if lacking else 'not in shape order, true first'}" + f" ({len(got)} for {len(want)})") + + for v in a["variants"]: + if not isinstance(v, dict): + g.fail(f"{w}: a variant {v!r}") + continue + s = f"{w} shape {v.get('shape')}" + ("" if v.get("bandMatch") is None else f" {v['bandMatch']}") + absent = [k for k in VARIANT_KEYS if k not in v] + if absent: + g.fail(f"{s}: no {', '.join(absent)}") + continue + variants += 1 + shape = by_id.get(v["shape"]) if is_int(v["shape"]) else None + det = v["deterministic"] + if not isinstance(det, bool): + g.fail(f"{s}: deterministic {det!r}") + det = False + deterministic += det + if v["bandMatch"] is not None and not isinstance(v["bandMatch"], bool): + g.fail(f"{s}: bandMatch {v['bandMatch']!r}") + elif shape is not None and (v["bandMatch"] is None) == (shape.get("bandCondition") is True): + g.fail(f"{s}: bandMatch {v['bandMatch']!r} for a shape " + f"{'with' if v['bandMatch'] is None else 'without'} a band condition") + n, cross, met = v["seeds"], v["crossSeeds"], v["seTargetMet"] + if not (is_int(n) and n >= 1 and isinstance(met, bool)): + g.fail(f"{s}: seeds {n!r}, seTargetMet {met!r}") + elif det and (n != 1 or not met): + g.fail(f"{s}: deterministic, but {n} seeds and seTargetMet {met}") + elif not det and batches and (n not in batches or (not met and n != batches[-1])): + g.fail(f"{s}: {n} seeds (seTargetMet {met}), not a batch of the seed rule {batches}") + elif not det and not met: + missed.append(f"{chart.get('scoreId')} shape {v['shape']}") + if (not is_int(cross) or cross < 1 or (is_int(n) and is_int(rule.get("crossSeeds")) + and cross != min(n, rule["crossSeeds"]))): + g.fail(f"{s}: crossSeeds {cross!r}, expected min({n}, {rule['crossSeeds']})") + + # every [mean, se] + values = {k: v[k] for k in VARIANT_PAIRS} + rs = v["ranges"] + if not isinstance(rs, list) or len(rs) != len(ranges): + g.fail(f"{s}: {len(rs) if isinstance(rs, list) else repr(rs)} range results for {len(ranges)} ranges") + rs = [] + for j, r in enumerate(rs): + for k in VARIANT_RANGE_PAIRS: + values[f"ranges[{j}].{k}"] = r.get(k) if isinstance(r, dict) else None + wt, rw = v["weights"], v["rangeWeights"] + if plain is None: + if wt is not None or rw is not None: + g.fail(f"{s}: weights or rangeWeights without a plain kind") + else: + if not (isinstance(wt, list) and len(wt) == positions): + g.fail(f"{s}: weights are not one [mean, se] per position") + else: + values.update({f"weights[{k}]": x for k, x in enumerate(wt)}) + if not linear: + if rw is not None: + g.fail(f"{s}: rangeWeights, but deck.seeds[0].rangeWeights is null") + elif not (isinstance(rw, list) and len(rw) == positions and all( + isinstance(x, list) and len(x) == len(ranges) for x in rw)): + g.fail(f"{s}: rangeWeights are not [position][range] [mean, se]") + else: + values.update({f"rangeWeights[{k}][{j}]": y for k, x in enumerate(rw) for j, y in enumerate(x)}) + bad = [k for k, x in values.items() if not pair(x)] + if bad: + g.fail(f"{s}: {', '.join(bad[:6])}{' ...' if len(bad) > 6 else ''} not [mean, se] (finite, se >= 0)") + elif det and any(x[1] for x in values.values()): + g.fail(f"{s}: deterministic, but a standard error is not 0: " + f"{', '.join(k for k, x in values.items() if x[1])[:160]}") + elif rs and all(pair(r.get(k)) for r in rs for k in ("rangeScore", "rankBonus")): + inside = sum(r["rangeScore"][0] + r["rankBonus"][0] for r in rs) + if abs(v["tail"][0] - (v["score"][0] - inside)) > ( + 0.0005 * (2 + 2 * len(rs)) + APTITUDE_SLACK): + g.fail(f"{s}: tail {v['tail'][0]} is not score {v['score'][0]} less the ranges' rangeScore and " + f"rankBonus {inside}") + + # Perfect rank-bonus deltas are not exported. Only a necessary rounding bound can be checked + # for stochastic means: each difference of two truncated bonuses is within 2 points of delta*p/100. + if not bad and len(rs) == len(ranges): + perfect_terms = [r["rangeScorePerfect"][0] * (1 + info["rankBonusPercent"] / 100) + for r, info in zip(rs, ranges, strict=True)] + rounding = 0.001 + sum(0.0005 * abs(1 + info["rankBonusPercent"] / 100) for info in ranges) + residual = v["scorePerfect"][0] - v["tailPerfect"][0] - sum(perfect_terms) + if abs(residual) > 2 * len(ranges) + rounding + APTITUDE_SLACK: + g.fail(f"{s}: tailPerfect violates the necessary Perfect rank-bonus rounding bound") + + # Deterministic point deltas are integers. Perfect rank bonuses can then be recovered exactly + # from the first baseline seed (stochastic Perfect rank-bonus deltas are not exported). + if det and not bad and rs and dseeds: + points = [x[0] for k, x in values.items() if not k.startswith(("weights", "rangeWeights"))] + if not all(float(x).is_integer() for x in points): + g.fail(f"{s}: deterministic, but a point delta is not an integer") + else: + perfect = 0 + for r, info, base in zip(rs, ranges, dseeds[0]["ranges"], strict=True): + delta, baseline = int(r["rangeScorePerfect"][0]), base["rangeScorePerfect"] + percent = info["rankBonusPercent"] + perfect += delta + rank_bonus(baseline + delta, percent) - rank_bonus(baseline, percent) + if v["tailPerfect"][0] != v["scorePerfect"][0] - perfect: + g.fail(f"{s}: tailPerfect is not scorePerfect less the Perfect range scores and bonuses") + + # the check + c = v["check"] + if not (isinstance(c, dict) and all(k in c for k in APTITUDE_CHECK_KEYS)): + g.fail(f"{s}: the check has no {', '.join(APTITUDE_CHECK_KEYS)}") + continue + if dseeds and c["seed"] != dseeds[0].get("seed"): + g.fail(f"{s}: the check's seed {c['seed']!r} is not deck.seeds[0]'s {dseeds[0].get('seed')!r}") + ranks = c["ranks"] + if not (isinstance(ranks, list) and len(ranks) == len(ranges) + and all(is_int(x) and 1 <= x <= RANKS for x in ranks)): + g.fail(f"{s}: the check's ranks {ranks!r} are not one rank per range") + elif not linear and any(x != 1 for x in ranks): + g.fail(f"{s}: the check's ranks {ranks} on a chart whose ranks are not linear (rank 1 alone)") + cd = c["deck"] + if not (isinstance(cd, list) and len(cd) == positions and all( + x is None or (plain is not None and isinstance(x, list) and len(x) == 2 and x[0] == plain + and is_num(x[1])) for x in cd)): + g.fail(f"{s}: the check deck is not a [plain kind, value] or null per position") + if not within(c): + g.fail(f"{s}: the check is not within its bound") + if missed: + g.warn(f"{len(missed)} variants missed the seed rule's standard error target: {', '.join(missed[:10])}" + + (" ..." if len(missed) > 10 else "")) + g.note = (f"{len(shapes)} shapes; {aptitudes} charts with an aptitude, {nulls} null; {variants} variants, " + f"{deterministic} deterministic; plain kind {plain}") + + +def gate_gzip(doc, ctx: Context, g: Gate, raw: bytes): + """The file's gzip size (what the page downloads) and that of its Gekisou skill aptitude, against loose caps.""" + whole = len(gzip.compress(raw, 6)) + deck = doc.get("deck") if isinstance(doc.get("deck"), dict) else {} + part = {"deck": deck.get("gekisouAptitude"), + "charts": [(c.get("deck") or {}).get("gekisouAptitude") if isinstance(c.get("deck"), dict) else None + for _, c in charts_of(doc)]} + aptitude = len(gzip.compress(json.dumps(part, ensure_ascii=False, separators=(",", ":")).encode("utf-8"), 6)) + if whole > FILE_GZIP_MAX: + g.fail(f"{whole} bytes gzipped, more than {FILE_GZIP_MAX}") + if aptitude > APTITUDE_GZIP_MAX: + g.fail(f"the Gekisou skill aptitude: {aptitude} bytes gzipped, more than {APTITUDE_GZIP_MAX}") + g.note = f"{whole} bytes gzipped, the aptitude {aptitude}" + + def nonfinite(v, path: str, out: list): if isinstance(v, float): if not math.isfinite(v): @@ -725,8 +1098,9 @@ def gate_page(doc, ctx: Context, g: Gate): GATES = (("schema", gate_schema), ("provenance", gate_provenance), ("counts", gate_counts), ("deck", gate_deck), - ("scenarios", gate_scenarios), ("finite", gate_finite), ("references", gate_references), ("bgm", gate_bgm), - ("size", gate_size), ("page", gate_page)) + ("scenarios", gate_scenarios), ("aptitude", gate_aptitude), ("finite", gate_finite), + ("references", gate_references), ("bgm", gate_bgm), ("size", gate_size), ("gzip", gate_gzip), + ("page", gate_page)) def gates(raw: bytes, ctx: Context, only=None) -> dict: @@ -745,7 +1119,7 @@ def gates(raw: bytes, ctx: Context, only=None) -> dict: continue g = Gate() try: - fn(doc, ctx, g, raw) if name == "size" else fn(doc, ctx, g) + fn(doc, ctx, g, raw) if name in ("size", "gzip") else fn(doc, ctx, g) except Exception as e: # a malformed file the gate did not foresee fails the gate g.fail(f"the gate stopped: {type(e).__name__}: {str(e)[:160]}") results.append({"gate": name, "passed": not g.failures, "failures": g.failures[:LISTED], diff --git a/.github/scripts/music_data_smoke.mjs b/.github/scripts/music_data_smoke.mjs index a711edc..635b456 100644 --- a/.github/scripts/music_data_smoke.mjs +++ b/.github/scripts/music_data_smoke.mjs @@ -85,10 +85,91 @@ for (const [name, scenario] of SCENARIOS) { } catalog.histogram(rows, (r) => r.level); +// Aptitude is a separate view, never folded into the default song ranking. Older pinned pages may lack this API; +// report that explicitly until PLAYER_REF is deliberately moved to a page with the aptitude UI. +let aptitudeFigures = 0; +const aptitudeApi = ["aptitudeShapes", "chartVariants", "aptitudeFigures", "aptitudeRate", "aptitudeSe", + "masterSkillFactor"].every((k) => typeof ranking[k] === "function") + && ["gekisouSkill", "shapeSkills", "shapeBands"].every((k) => typeof catalog[k] === "function"); +const near = (a, b) => finite(a) && finite(b) && Math.abs(a - b) <= 1e-8 * Math.max(1, Math.abs(a), Math.abs(b)); +if (aptitudeApi && data.deck?.gekisouAptitude) { + const shapes = ranking.aptitudeShapes(data); + if (shapes.size !== data.deck.gekisouAptitude.shapes.length) problem("aptitudeShapes: missing shapes"); + for (const shape of shapes.values()) { + const skills = catalog.shapeSkills(data, shape, "zh-Hant"); + if (skills.length !== new Set(shape.skills.map((s) => `${s.id}:${s.level}`)).size) { + problem(`shape ${shape.id}: shapeSkills count`); + } + const table = shape.source === "support" ? "supportSkills" : "skills"; + for (const skill of skills) { + const named = catalog.gekisouSkill(data, table, skill.id, "zh-Hant"); + if (!named.name || named.name !== skill.name) problem(`shape ${shape.id}: skill name lookup`); + } + const bands = catalog.shapeBands(data, shape, "zh-Hant"); + if (bands.length !== new Set(shape.skills.flatMap((s) => s.bandIds || [])).size || bands.some((s) => !s)) { + problem(`shape ${shape.id}: shapeBands lookup`); + } + } + const power = data.deck.model.power; + for (const { chart } of charts) { + const d = chart.deck; + if (!d) continue; + const variants = ranking.chartVariants(d); + if (variants.length !== (d.unplayable ? 0 : d.gekisouAptitude?.variants.length || 0)) { + problem(`chart ${chart.scoreId}: chartVariants count`); + } + // Removing aptitude must not alter any default figure. + const baseline = ranking.chartFigures(d, kind, power); + const without = ranking.chartFigures({ ...d, gekisouAptitude: null }, kind, power); + if (JSON.stringify(baseline) !== JSON.stringify(without)) problem(`chart ${chart.scoreId}: aptitude changes default`); + for (const variant of variants) { + for (const [name, scenario] of SCENARIOS) { + const tag = `aptitude chart ${chart.scoreId} shape ${variant.shape}, ${name}`; + const f = ranking.aptitudeFigures(variant, d.ranges, power, scenario); + if (scenario?.mode === "free") { + if (f !== null) problem(`${tag}: Free Live has aptitude`); + continue; + } + aptitudeFigures++; + if (!f || !finite(f.base)) { problem(`${tag}: missing or nonfinite base`); continue; } + if (f.weights !== null && (f.weights.length !== d.positions || !f.weights.every(finite))) { + problem(`${tag}: invalid weights`); + } + const zero = Array(d.positions).fill(0); + if (!near(ranking.aptitudeRate(f, zero), f.base)) problem(`${tag}: no-skill rate`); + const rate = ranking.aptitudeRate(f, SKILLS); + if (f.weights === null ? rate !== null : !finite(rate)) problem(`${tag}: missing cross-term handling`); + if (ranking.aptitudeSe(f, SKILLS) !== null) problem(`${tag}: SE assigned without covariance`); + if (scenario === null) { + if (!near(f.base, variant.score[0] / power) + || !near(ranking.aptitudeSe(f, zero), variant.score[1] / power)) problem(`${tag}: raw mean/SE`); + } else if (ranking.aptitudeSe(f, zero) !== null) problem(`${tag}: transformed SE is not null`); + } + // Only deterministic deltas describe the individual check seed. Never reconstruct a stochastic check + // from sampled means. Ordinary check cards are positional, unlike the UI's random-order expectation. + if (!variant.deterministic || kind === null) continue; + const c = variant.check; + const seed = d.seeds.find((s) => s.seed === c.seed); + const sc = { mode: "battle", ranks: c.ranks, just: 1, great: 0 }; + const base = ranking.scenarioSeed(seed, d.ranges, kind, sc); + const gain = ranking.aptitudeFigures(variant, d.ranges, power, sc); + if (!base || !gain?.weights) continue; + const predicted = data.deck.model.checkPower * ((base.score / power) + gain.base + + c.deck.reduce((sum, card, k) => sum + (card ? ranking.masterSkillFactor(card[1]) + * (base.weights[k] + gain.weights[k]) : 0), 0)); + // Exported weight rounding adds a small reconstruction error on top of the engine's bound. + const rounding = data.deck.model.checkPower * 1e-7 * (1 + d.ranges.length) * d.positions; + if (!finite(predicted) || Math.abs(predicted - c.exact) > c.bound + rounding) { + problem(`aptitude chart ${chart.scoreId} shape ${variant.shape}: deterministic check reconstruction`); + } + } + } +} + if (problems.length) { for (const m of problems.slice(0, 40)) console.log(m); if (problems.length > 40) console.log(`${problems.length - 40} more problems`); process.exit(1); } console.log(`page smoke test: ${rows.length} charts, ${SCENARIOS.length} scenarios, ${figures} chart figures; ` - + `plain kind ${kind}, scenarios free/ranks/just`); + + `plain kind ${kind}, scenarios free/ranks/just; aptitude ${aptitudeApi ? aptitudeFigures + " figures" : "API unavailable (skipped)"}`); diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py index c90400f..3df390c 100644 --- a/.github/scripts/test_music_data.py +++ b/.github/scripts/test_music_data.py @@ -6,8 +6,8 @@ The JSON Schema gate uses docs/schema/music-data.schema.json (or $MUSIC_DATA_SCHEMA) when the checkout has it; the page smoke test runs when $MUSIC_DATA_PAGE names ournotes-player's examples/songs (and Node.js is installed). Real files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass; one made before -the ranges' luckPoints, the deck gate stops on those alone) and $MUSIC_DATA_OLD_SAMPLE (one without the play scenario -fields: the scenario gate stops it). +the ranges' luckPoints and the Gekisou skill aptitude, the deck and aptitude gates stop it on those alone) and +$MUSIC_DATA_OLD_SAMPLE (one without the play scenario fields: the scenario gate stops it). """ import copy import hashlib @@ -55,21 +55,83 @@ def check_deck(exact): LUCK_SEEDS = [11, 22] # a luck chart's seeds; else the one seed 0 +# the Gekisou skill aptitude's shapes: id, source, mission, band condition +SHAPES = [(0, "member", 1, False), (1, "support", 2, True), (2, "member", 3, False), (3, "support", 4, False)] +SEED_RULE = {"deterministicTest": 4, "batches": [32, 64, 128, 256, 512, 1024], "relative": 0.01, "baseline": 0.001, + "crossSeeds": 64} + + +def effect(band=False): + return {"effectType": 2000, "triggerType": 7010, "activationTimeSecond": 5.0, "effectValue": 1000, + "maxEffectValue": 0, "effectLimitCount": 0, "effectExecuteLimitCount": 0, "skillTargetIds": [], + "trigger": [[{"type": 7010, "values": [1], "positive": True, "targetIds": []}]], + "condition": [[{"type": 5000, "values": [1], "positive": True, "targetIds": None}]] if band else [], + "release": [], "reset": [], "cumulative": None} + + +def aptitude_head(): + return {"plainKind": 0, "host": "one performer with a synthetic empty Gekisou skill", "seedRule": SEED_RULE, + "shapes": [{"id": i, "source": src, "mission": m, "bandCondition": band, "effects": [effect(band)], + "skills": [{"id": 10 + i, "level": 5, "memberTargetIds": [41] if band else None, + "bandIds": [1] if band else None}]} for i, src, m, band in SHAPES]} + + +def variant(deck, shape, match, deterministic): + """One shape's aptitude on a chart: deterministic on a chart without a luck range, else on 32 seeds.""" + se = 0 if deterministic else 12.5 + + def p(m): + return [m, se] + ranges = [{"rangeScore": p(300), "rankBonus": p(750), "rangeScorePerfect": p(300), "maxCombo": p(0), + "justCount": p(0), "luckPoints": p(2)} for _ in deck["ranges"]] + n = 1 if deterministic else 32 + linear = deck["seeds"][0]["rangeWeights"] is not None + result = {"shape": shape, "bandMatch": match, "deterministic": deterministic, "seeds": n, "seTargetMet": True, + "crossSeeds": min(n, SEED_RULE["crossSeeds"]), "score": p(1500), "scorePerfect": p(1500), + "tail": p(1500 - 1050 * len(ranges)), "tailPerfect": p(1500 - 1050 * len(ranges)), + "converted": p(0), "ranges": ranges, + "weights": [p(0.01), p(0.02)], + "rangeWeights": [[p(0.001)] * len(ranges), [p(0.002)] * len(ranges)] if linear else None, + "check": {"seed": deck["seeds"][0]["seed"], "ranks": [2] * len(ranges) if linear else [1] * len(ranges), + "deck": [[0, 7000], None], "exact": 5000, "predicted": 5000.5, "bound": 3.0}} + base = deck["seeds"][0] + score = base["score"] + 1500 + weight = base["weights"][0][0] + 0.01 + if linear: + for r, info in zip(base["ranges"], deck["ranges"], strict=True): + percent = info["rankBonusPercents"][1] + score += int(r["rangeScore"] * percent / 100) - r["rankBonus"] + 300 * (percent - 250) / 100 + weight += (190 - 250) / 100 * (0.1 + 0.001) + predicted = 1000003 * (score / 300000 + 0.7 * weight) + result["check"].update(exact=round(predicted), predicted=predicted) + return result + + +def aptitude(deck, luck): + missions = {r["mission"] for r in deck["ranges"]} + variants = [variant(deck, i, match, deterministic=not luck) for i, _, m, band in SHAPES + if m == 4 or m in missions for match in ((True, False) if band else (None,))] + factors = [{"judgedNotes": 12, "justNotes": 0, "perfectNotes": 0, "tailNotes": 2, "comboAtStart": 5, + "lotteries": [4.0, 0.0] if r["mission"] == 2 else [0, 0]} for r in deck["ranges"]] + return {"factors": factors, "variants": variants} def deck_chart(luck=False): def one(n): return {"seed": n, "score": 120000 + n, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, "justCount": 0, "luckPoints": 3 if luck else 0, - "lotResults": [0, 0, 0, 0], "rangeScorePerfect": 4000}], + "lotResults": [1, 2, 0, 1] if luck else [0, 0, 0, 0], + "rangeScorePerfect": 4000}], "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000 + n, "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} - return {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, - "ranges": [{"index": 0, "mission": 2 if luck else 1, "startMs": 1000, "endMs": 5000, - "rankBonusPercent": 250, "rankBonusPercents": [250, 190, 160, 100, 100]}], - "justNotes": 0, "seeds": [one(n) for n in (LUCK_SEEDS if luck else [0])], - "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], - "unplayable": None} + d = {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, + "ranges": [{"index": 0, "mission": 2 if luck else 1, "startMs": 1000, "endMs": 5000, + "rankBonusPercent": 250, "rankBonusPercents": [250, 190, 160, 100, 100]}], + "justNotes": 0, "seeds": [one(n) for n in (LUCK_SEEDS if luck else [0])], + "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], + "unplayable": None} + d["gekisouAptitude"] = aptitude(d, luck) + return d def chart(difficulty, score_id, last=60000, luck=False): @@ -113,7 +175,8 @@ def sample() -> dict: "characters": [{"id": 1, "bandId": 1, "name": text("ch"), "shortName": text("c"), "mainColor": "#77BBDD"}], "tags": [{"id": 1, "name": text("tag")}], "categories": [{"id": 1, "musicCategories": [1], "name": text("cat")}], - "deck": {"model": {"power": 300000, "checkPower": 1000003}, "kinds": [kind]}, + "deck": {"model": {"power": 300000, "checkPower": 1000003, "gekisouAptitude": "one skill at a time"}, + "kinds": [kind], "gekisouAptitude": aptitude_head()}, "songs": [song(100001, [chart("easy", 10), chart("expert", 30, luck=True)]), song(100002, [chart("expert", 40, luck=True)])], } @@ -164,6 +227,30 @@ def seed(doc, song=0, chart=0): return doc["songs"][song]["charts"][chart]["deck"]["seeds"][0] +def apt(doc, song=0, chart=0): + return doc["songs"][song]["charts"][chart]["deck"]["gekisouAptitude"] + + +def var(doc, i=0, song=0, chart=0): + return apt(doc, song, chart)["variants"][i] + + +def shape(doc, i): + return doc["deck"]["gekisouAptitude"]["shapes"][i] + + +def nonlinear(doc, song=0, chart=0): + """A chart whose ranks are not linear (overlapping ranges): no rangeWeights, the checks at rank 1.""" + d = doc["songs"][song]["charts"][chart]["deck"] + for s in d["seeds"]: + s.update(rangeWeights=None, rankCheck=None) + for v in d["gekisouAptitude"]["variants"]: + v["rangeWeights"] = None + v["check"]["ranks"] = [1] * len(d["ranges"]) + predicted = 1000003 * ((d["seeds"][0]["score"] + v["score"][0]) / 300000 + 0.7 * 0.51) + v["check"].update(exact=round(predicted), predicted=predicted) + + # ---------------------------------------------------------------- the gates def test_the_sample_passes_every_gate(tmp_path): r = run(tmp_path, sample(), page=page_path()) @@ -194,6 +281,20 @@ def test_page_smoke(tmp_path): assert not g["passed"] and any("free scenario" in f for f in g["failures"]) +def test_page_aptitude_check_reconstruction(tmp_path): + page = page_path() + if page is None or "aptitudeFigures" not in (page / "ranking.js").read_text(encoding="utf-8"): + pytest.skip("page does not yet expose the aptitude API") + doc = sample() + # Stochastic check values are not reconstructible from means; changing them must not trigger reconstruction. + var(doc, 0, 0, 1)["check"].update(exact=99999999, predicted=99999999) + assert gate(run(tmp_path, doc, page=page), "page")["passed"] + # Even a self-consistent exported check is independently rejected when deterministic deltas disagree. + var(doc)["check"].update(exact=99999999, predicted=99999999) + g = gate(run(tmp_path, doc, page=page), "page") + assert any("deterministic check reconstruction" in f for f in g["failures"]), g + + def unplayable(doc): doc["songs"][1]["charts"][0]["deck"]["unplayable"] = "more than three fevers" # its seeds kept @@ -214,6 +315,70 @@ def unplayable(doc): (lambda d: seed(d)["ranges"][0].pop("rangeScorePerfect"), "scenarios", "rangeScorePerfect missing"), (lambda d: seed(d).update(rangeWeights=[[[0.1]]]), "scenarios", "rangeWeights are not"), (lambda d: seed(d)["rankCheck"].update(exact=99999), "scenarios", "rank check deck is not within"), + # the Gekisou skill aptitude + (lambda d: d["deck"].pop("gekisouAptitude"), "aptitude", "deck.gekisouAptitude missing"), + (lambda d: d["deck"]["model"].pop("gekisouAptitude"), "aptitude", "deck.model.gekisouAptitude: no text"), + (lambda d: d["deck"]["gekisouAptitude"].update(plainKind=1), "aptitude", "the page's plain kind is 0"), + (lambda d: d["deck"]["kinds"][0].update(durationMs=6000), "aptitude", "plainKind 0, the page's plain kind is None"), + (lambda d: d["deck"]["kinds"][0].update(durationMs=6000), "aptitude", "weights or rangeWeights without a plain"), + (lambda d: d["deck"]["gekisouAptitude"].pop("host"), "aptitude", "deck.gekisouAptitude: no host"), + (lambda d: d["deck"]["gekisouAptitude"].update(seedRule=dict(SEED_RULE, batches=[64, 32])), "aptitude", + "deck.gekisouAptitude.seedRule"), + (lambda d: shape(d, 1).update(id=5), "aptitude", "ids are not 0, 1, 2, ... in order"), + (lambda d: shape(d, 0).update(source="card"), "aptitude", "shape 0: source 'card'"), + (lambda d: shape(d, 2).update(mission=5), "aptitude", "shape 2: mission 5"), + (lambda d: shape(d, 0).update(bandCondition=True), "aptitude", "bandCondition True (a support skill's alone)"), + (lambda d: shape(d, 3).update(bandCondition=True), "aptitude", + "shape 3: bandCondition True, its effects have 0 condition 5000"), + (lambda d: shape(d, 1)["effects"][0]["condition"][0][0].update(targetIds=[41]), "aptitude", + "condition 5000 targetIds not null"), + (lambda d: shape(d, 0)["effects"][0].pop("cumulative"), "aptitude", "effect 0: cumulative missing or malformed"), + (lambda d: shape(d, 0)["effects"][0].update(effectValue=1.5), "aptitude", "effectValue missing or malformed"), + (lambda d: shape(d, 0)["effects"][0].update(reset=None), "aptitude", "reset missing or malformed"), + (lambda d: shape(d, 1)["skills"][0].update(memberTargetIds=None), "aptitude", "is not an id, a level"), + (lambda d: shape(d, 0)["skills"][0].update(bandIds=[1]), "aptitude", "is not an id, a level"), + (lambda d: shape(d, 0).update(skills=[]), "aptitude", "shape 0: no skills"), + (lambda d: d["songs"][0]["charts"][0]["deck"].pop("gekisouAptitude"), "aptitude", "no deck.gekisouAptitude"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(gekisouAptitude=None), "aptitude", + "deck.gekisouAptitude is null, but the chart is playable"), + (unplayable, "aptitude", "a Gekisou skill aptitude on a chart unplayable with Gekisou on"), + (lambda d: apt(d)["factors"].pop(), "aptitude", "0 factors for 1 ranges"), + (lambda d: apt(d)["factors"][0].update(justNotes=3), "aptitude", "Just or Perfect notes in a range without"), + (lambda d: apt(d)["factors"][0].update(tailNotes=-1), "aptitude", "tailNotes missing or not counts"), + (lambda d: apt(d)["factors"][0].update(lotteries=[1, 0]), "aptitude", "in a range without the luck mission"), + (lambda d: apt(d, 0, 1)["factors"][0].update(lotteries=[3.0, 0]), "aptitude", "deck.seeds' lotResults give 4.0"), + (lambda d: apt(d)["variants"].pop(0), "aptitude", "variants lack [(0, None)]"), + (lambda d: apt(d)["variants"].reverse(), "aptitude", "not in shape order, true first"), + (lambda d: apt(d, 0, 1)["variants"].pop(1), "aptitude", "variants lack [(1, False)]"), + (lambda d: var(d).update(shape=9), "aptitude", "variants of shapes ['9'] not in deck.gekisouAptitude.shapes"), + (lambda d: var(d).update(shape=2), "aptitude", "variants of shapes [2] of a mission the chart does not play"), + (lambda d: var(d, 0, 0, 1).update(bandMatch=None), "aptitude", "bandMatch None for a shape with a band"), + (lambda d: var(d).update(bandMatch=True), "aptitude", "bandMatch True for a shape without a band"), + (lambda d: var(d).pop("tail"), "aptitude", "shape 0: no tail"), + (lambda d: var(d).update(score=[1500]), "aptitude", "score not [mean, se]"), + (lambda d: var(d).update(score=[1500, -1]), "aptitude", "score not [mean, se]"), + (lambda d: var(d).update(converted=[float("nan"), 0]), "aptitude", "converted not [mean, se]"), + (lambda d: var(d)["ranges"][0].update(luckPoints=None), "aptitude", "ranges[0].luckPoints not [mean, se]"), + (lambda d: var(d).update(score=[1500, 1]), "aptitude", "deterministic, but a standard error is not 0: score"), + (lambda d: var(d).update(seeds=2), "aptitude", "deterministic, but 2 seeds"), + (lambda d: var(d, 0, 0, 1).update(seeds=33), "aptitude", "33 seeds (seTargetMet True), not a batch"), + (lambda d: var(d, 0, 0, 1).update(seTargetMet=False), "aptitude", "32 seeds (seTargetMet False), not a batch"), + (lambda d: var(d, 0, 0, 1).update(crossSeeds=64), "aptitude", "crossSeeds 64, expected min(32, 64)"), + (lambda d: var(d)["ranges"].append(dict(var(d)["ranges"][0])), "aptitude", "2 range results for 1 ranges"), + (lambda d: var(d).update(tail=[451, 0]), "aptitude", "tail 451 is not score 1500 less the ranges' rangeScore"), + (lambda d: var(d)["weights"].append([0, 0]), "aptitude", "weights are not one [mean, se] per position"), + (lambda d: seed(d).update(rangeWeights=None, rankCheck=None), "aptitude", + "rangeWeights, but deck.seeds[0].rangeWeights is null"), + (lambda d: var(d).update(rangeWeights=[[[0, 0]]]), "aptitude", "rangeWeights are not [position][range]"), + (lambda d: (nonlinear(d), var(d)["check"].update(ranks=[2])), "aptitude", "ranks are not linear (rank 1 alone)"), + (lambda d: var(d)["check"].update(seed=5), "aptitude", "the check's seed 5 is not deck.seeds[0]'s 0"), + (lambda d: var(d, 0, 0, 1)["check"].update(seed=22), "aptitude", "the check's seed 22 is not deck.seeds[0]'s 11"), + (lambda d: var(d)["check"].update(ranks=[6]), "aptitude", "the check's ranks [6] are not one rank per range"), + (lambda d: var(d)["check"].update(deck=[[1, 7000], None]), "aptitude", "not a [plain kind, value] or null"), + (lambda d: var(d)["check"].update(exact=99999999), "aptitude", "the check is not within its bound"), + (lambda d: var(d)["check"].pop("bound"), "aptitude", "the check has no seed, ranks"), + (lambda d: var(d).update(tailPerfect=[449, 0]), "aptitude", "tailPerfect is not scorePerfect"), + (lambda d: var(d).update(converted=[0.5, 0]), "aptitude", "a point delta is not an integer"), # the deck statistics (lambda d: d.update(deck=None), "deck", "deck is null"), (lambda d: d["songs"][0]["charts"][0].update(deck=None), "deck", "no deck statistics"), @@ -289,7 +454,7 @@ def test_a_jacket_file_is_missing(tmp_path): def test_warnings_do_not_fail(tmp_path): doc = sample() - seed(doc).update(rangeWeights=None, rankCheck=None) # overlapping ranges + nonlinear(doc) # overlapping ranges seed(doc, 0, 1)["rangeWeights"][0] = None # a kind reading the confirmed rank doc["songs"][1]["charts"][0]["deck"]["offSeeds"][0]["weights"][0] = None doc["songs"][0]["master"]["MasterLiveMusic"]["_rate"] = float("inf") # 1e999 in master data as served @@ -300,6 +465,50 @@ def test_warnings_do_not_fail(tmp_path): assert gate(r, "references")["warnings"] == ["songs without a zh-Hant title: 100002"] +def test_aptitude_warnings(tmp_path): + doc = sample() + var(doc, 0, 0, 1).update(seeds=1024, seTargetMet=False, crossSeeds=64) # the last batch, not the target + r = run(tmp_path, doc) + assert r["passed"], failures(r) + assert gate(r, "aptitude")["warnings"] == [ + "1 variants missed the seed rule's standard error target: 30 shape 1"] + assert gate(r, "aptitude")["note"] == ("4 shapes; 3 charts with an aptitude, 0 null; 8 variants, " + "2 deterministic; plain kind 0") + + +def test_no_gekisou_skill_shapes(tmp_path): + doc = sample() + doc["deck"]["gekisouAptitude"]["shapes"] = [] + kept = apt(doc) + for _, c in music_data.charts_of(doc): + c["deck"]["gekisouAptitude"] = None + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + doc["songs"][0]["charts"][0]["deck"]["gekisouAptitude"] = kept + g = gate(run(tmp_path, doc), "aptitude") + assert g["failures"] == ["chart 10 (100001 easy): a Gekisou skill aptitude on a chart without a Gekisou skill " + "shape"] + + +def test_an_unplayable_chart_has_no_aptitude(tmp_path): + doc = sample() + doc["songs"][1]["charts"][0]["deck"].update(unplayable="more than three fevers", seeds=[], gekisouAptitude=None) + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + assert gate(r, "aptitude")["note"].startswith("4 shapes; 2 charts with an aptitude, 1 null") + + +def test_gzip_caps(tmp_path, monkeypatch): + raw, ctx = context(tmp_path, sample()) + g = gate(gates(raw, ctx, only=("gzip",)), "gzip") + assert g["passed"] and g["note"].startswith(f"{len(music_data.gzip.compress(raw, 6))} bytes gzipped") + monkeypatch.setattr(music_data, "APTITUDE_GZIP_MAX", 100) + g = gate(gates(raw, ctx, only=("gzip",)), "gzip") + assert not g["passed"] and g["failures"][0].startswith("the Gekisou skill aptitude: ") + monkeypatch.setattr(music_data, "FILE_GZIP_MAX", 100) + assert len(gate(gates(raw, ctx, only=("gzip",)), "gzip")["failures"]) == 2 + + def test_against_the_published_file(tmp_path): doc = sample() raw, ctx = context(tmp_path, doc) @@ -380,12 +589,26 @@ def test_a_real_file_with_the_scenario_fields(tmp_path): raw = real("MUSIC_DATA_SAMPLE") (tmp_path / "music-data.json").write_bytes(raw) ctx = Context(language="zh-Hant", page=page_path(), file=tmp_path / "music-data.json", published=raw) - r = gates(raw, ctx, only=CONTENT + ("counts", "size", "page")) - if luck_points(raw): - assert r["passed"], failures(r) - else: - assert [n for n, _ in failures(r)] == ["deck"] and before_luck_points(r), failures(r) + r = gates(raw, ctx, only=CONTENT + ("aptitude", "counts", "size", "gzip", "page")) + new = [n for n, ok in (("deck", luck_points(raw)), ("aptitude", "gekisouAptitude" in json.loads(raw)["deck"])) + if not ok] # the gates of fields made after the file + assert [n for n, _ in failures(r)] == new, failures(r) + assert "deck" not in new or before_luck_points(r) + g = gate(r, "aptitude") + assert "aptitude" not in new or all(f.endswith(("gekisouAptitude missing", "gekisouAptitude: no text", + ": no deck.gekisouAptitude")) for f in g["failures"]), g assert gate(r, "scenarios")["warningCount"] == 0 + print(f"aptitude: {g['note']}; warnings: {g['warnings']}; gzip: {gate(r, 'gzip')['note']}") + + +def test_real_aptitude_deck_sample(): + """Optional real chart-stats output, wrapped without modifying its statistics or touching the source file.""" + stats = json.loads(real("MUSIC_DATA_APTITUDE_DECK_SAMPLE")) + doc = {"deck": {k: stats[k] for k in ("model", "kinds", "gekisouAptitude")}, + "songs": [{"id": c["musicId"], "charts": [{"scoreId": c["scoreId"], "difficulty": c["difficulty"], + "skillEventsMs": [e[1] for e in c["events"]], "deck": c}]} for c in stats["charts"]]} + report = gates(json.dumps(doc).encode(), Context(), only=("deck", "aptitude", "gzip")) + assert report["passed"], failures(report) # ---------------------------------------------------------------- publish (a stand-in bucket) From 56918cc24c55b12e76f08ae92fe1f05d1ee8eee6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Exmeaning=20=28=E6=9D=B1=E9=9B=AA=29?= Date: Wed, 30 Sep 2026 05:06:42 +0800 Subject: [PATCH 14/14] =?UTF-8?q?=E5=A2=9E=E5=8A=A0=E5=A4=9A=E6=9C=8D?= =?UTF-8?q?=E5=89=A7=E6=83=85=E6=9E=84=E5=BB=BA?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .github/STORY_SITE.md | 15 +- .github/scripts/story_site.py | 44 +++++- .github/scripts/test_jp_workflow.py | 33 ++++ .github/workflows/story-site-region.yml | 174 +++++++++++++++++++++ .github/workflows/story-site.yml | 195 +++++------------------- docs/jp.md | 3 +- 6 files changed, 305 insertions(+), 159 deletions(-) create mode 100644 .github/workflows/story-site-region.yml diff --git a/.github/STORY_SITE.md b/.github/STORY_SITE.md index 466dbbd..6056712 100644 --- a/.github/STORY_SITE.md +++ b/.github/STORY_SITE.md @@ -3,11 +3,17 @@ `.github/workflows/story-site.yml` keeps the StarMoe story site up to date: it builds the stories the published site lacks with this repository's `nnnotes web --story` and uploads them to the bucket that serves the site (`https://storage.bdon.moe/moenotes/`, the layout `nnnotes web` writes: `stories.json`, `stories/`, `models/`, -`assets/`, `story/`). It only adds: a run never deletes anything from the bucket. Its helper steps are in +`assets/`, `story/`), one site per game region: `hk-tw-mo` at the bucket root, `jp` under `jp/` +(`https://storage.bdon.moe/moenotes/jp/`; JP Live2D model ids overlap the international ones). It only adds: a run +never deletes anything from the bucket. Its helper steps are in `.github/scripts/`; nothing outside `.github/` differs from upstream, so the fork syncs with it as before. ## A run +`story-site.yml` picks the regions (`story_site.py regions`: those of `STORY_REGIONS` that the dispatch's +`client_payload.regions` names, the `region` input of a manual run, every one on the schedule) and calls +`story-site-region.yml` once per region, in parallel. Each region's run: + 1. **plan** (a few seconds): the MasterAdv ids of the decoded master data of moenotes-masterdata-sync (`MasterAdv.json`, SHA-256 checked against its `index.json`) against the `stories/.json` objects of the bucket. The ids without a manifest, at most `STORY_LIMIT` (40) in id order, are the run's stories; none: the run ends here. @@ -32,9 +38,10 @@ snapshot: `dispatch_repositories`), a daily schedule (03:23 UTC) in case a dispa |---|---| | `stories` | MasterAdv ids to build (spaces or commas); empty: every story the site lacks | | `force` | rebuild the given stories and their Live2D models although their manifests exist | +| `region` | `all` (every region of `STORY_REGIONS`), `hk-tw-mo` or `jp` | | `dry_run` | build, then list what would be uploaded instead of uploading | -Runs do not overlap (`concurrency: story-site`). +Runs of one region do not overlap (`concurrency: story-site-`); the regions build side by side. ## Settings @@ -49,8 +56,8 @@ Repository secrets (Settings → Secrets and variables → Actions → Secrets): | `STORY_S3_ACCESS_KEY`, `STORY_S3_SECRET_KEY` | an S3 key that can list, read and write the bucket | Repository variables (optional; the defaults are the StarMoe site): `STORY_S3_ENDPOINT` (`https://storage.bdon.moe`), -`STORY_S3_BUCKET` (`moenotes`), `STORY_S3_PREFIX` (empty: the bucket root), `MASTERDATA_BASE_URL` -(`https://metadata.bdon.moe`), `STORY_MASTERDATA_REGION` (`hk-tw-mo`), `STORY_LIMIT` (`40`), `STORY_PLAYER_REPOSITORY` +`STORY_S3_BUCKET` (`moenotes`), `STORY_S3_PREFIX` (empty: the bucket root; the hk-tw-mo site), `STORY_S3_PREFIX_JP` (`jp`), `MASTERDATA_BASE_URL` +(`https://metadata.bdon.moe`), `STORY_REGIONS` (`hk-tw-mo jp`: the regions with a site), `STORY_LIMIT` (`40`), `STORY_PLAYER_REPOSITORY` (`empty-sekai/ournotes-player`), `STORY_PLAYER_REF` (`3774d8ac3987`), `PLAYFETCH_VERSION` (`v0.92`), `STORY_APK_PACKAGE` (`com.bilibili.sirius`). diff --git a/.github/scripts/story_site.py b/.github/scripts/story_site.py index a220a94..c923f0b 100755 --- a/.github/scripts/story_site.py +++ b/.github/scripts/story_site.py @@ -2,6 +2,9 @@ """The story site's CI steps (.github/workflows/story-site.yml): the published site lives in an S3 bucket, a run adds the stories it lacks. + regions the regions this run builds (GitHub outputs `regions`, a JSON list, and `count`): those of + $STORY_REGIONS (default hk-tw-mo) that the repository_dispatch payload names, or + $REQUESTED_REGION of a workflow_dispatch (empty / "all": every one), or all on a schedule plan the stories to build: $REQUESTED, else every MasterAdv id without a manifest in the bucket (at most $STORY_LIMIT, in id order); GitHub outputs `stories` (space separated) and `count` master OUT the decoded master data of $MASTERDATA_REGION from moenotes-masterdata-sync (index.json: @@ -137,6 +140,43 @@ def master_index() -> tuple[str, dict]: return base + region["path"], region +# Regions with a story site (story-site-region.yml gives each its own bucket prefix: JP model ids overlap the +# international ones). +SITE_REGIONS = ("hk-tw-mo", "jp") + + +def split_list(text: str) -> list[str]: + return [part for part in re.split(r"[\s,]+", text.strip()) if part] + + +def select_regions(enabled: str, event: str, payload: list[str] | None, requested: str) -> list[str]: + """The regions a run builds: of $STORY_REGIONS, the dispatch's regions, the requested one, or all (schedule).""" + allowed = [r for r in split_list(enabled) if r in SITE_REGIONS] + unknown = [r for r in split_list(enabled) if r not in SITE_REGIONS] + if unknown: + sys.exit(f"story_site: no story site for region(s) {', '.join(unknown)}") + if event == "repository_dispatch": + # Older dispatches carry no regions: build every enabled one (a run without new stories ends at plan). + wanted = payload if payload else allowed + elif event == "workflow_dispatch" and requested and requested != "all": + wanted = [requested] + else: + wanted = allowed + return [r for r in allowed if r in wanted] + + +def cmd_regions() -> None: + payload = None + event_path = os.environ.get("GITHUB_EVENT_PATH") + if event_path and Path(event_path).is_file(): + regions = (json.loads(Path(event_path).read_text(encoding="utf-8")).get("client_payload") or {}).get("regions") + payload = [r for r in regions if isinstance(r, str)] if isinstance(regions, list) else None + regions = select_regions(env("STORY_REGIONS", "hk-tw-mo"), env("GITHUB_EVENT_NAME", ""), payload, + os.environ.get("REQUESTED_REGION", "")) + output("regions", json.dumps(regions)) + output("count", str(len(regions))) + + def configure_region() -> None: """Keep build inputs from one release; JP client/API/CDN are read from the public snapshot.""" region = env("MASTERDATA_REGION") @@ -298,7 +338,9 @@ def main(argv: list[str]) -> None: if not argv: sys.exit(__doc__) cmd, args = argv[0], argv[1:] - if cmd == "plan" and not args: + if cmd == "regions" and not args: + cmd_regions() + elif cmd == "plan" and not args: cmd_plan() elif cmd == "master" and len(args) == 1: cmd_master(args[0]) diff --git a/.github/scripts/test_jp_workflow.py b/.github/scripts/test_jp_workflow.py index 6a1cf80..e6eeccd 100644 --- a/.github/scripts/test_jp_workflow.py +++ b/.github/scripts/test_jp_workflow.py @@ -50,3 +50,36 @@ def test_jp_catalog_provenance_matches_snapshot(): gate = music_data.Gate() music_data.gate_provenance(doc, context, gate) assert any('resourceHash differs' in s for s in gate.failures) + + +def test_regions_follow_the_dispatch_payload(): + assert story_site.select_regions('hk-tw-mo jp', 'repository_dispatch', ['jp'], '') == ['jp'] + assert story_site.select_regions('hk-tw-mo jp', 'repository_dispatch', ['hk-tw-mo', 'en', 'kr'], '') == ['hk-tw-mo'] + # an older dispatch without regions: every enabled region + assert story_site.select_regions('hk-tw-mo,jp', 'repository_dispatch', None, '') == ['hk-tw-mo', 'jp'] + + +def test_regions_of_schedule_and_manual_runs(): + assert story_site.select_regions('hk-tw-mo jp', 'schedule', None, '') == ['hk-tw-mo', 'jp'] + assert story_site.select_regions('hk-tw-mo jp', 'workflow_dispatch', None, 'all') == ['hk-tw-mo', 'jp'] + assert story_site.select_regions('hk-tw-mo jp', 'workflow_dispatch', None, 'jp') == ['jp'] + # a region STORY_REGIONS leaves out is not built + assert story_site.select_regions('hk-tw-mo', 'workflow_dispatch', None, 'jp') == [] + + +def test_regions_reject_a_region_without_a_story_site(): + with pytest.raises(SystemExit, match='no story site'): + story_site.select_regions('hk-tw-mo kr', 'schedule', None, '') + + +def test_regions_command_reads_the_event(tmp_path, monkeypatch): + import json + event = tmp_path / 'event.json' + event.write_text(json.dumps({'client_payload': {'regions': ['jp']}}), encoding='utf-8') + out = tmp_path / 'out' + monkeypatch.setenv('GITHUB_EVENT_PATH', str(event)) + monkeypatch.setenv('GITHUB_EVENT_NAME', 'repository_dispatch') + monkeypatch.setenv('GITHUB_OUTPUT', str(out)) + monkeypatch.setenv('STORY_REGIONS', 'hk-tw-mo jp') + story_site.cmd_regions() + assert out.read_text(encoding='utf-8').splitlines() == ['regions=["jp"]', 'count=1'] diff --git a/.github/workflows/story-site-region.yml b/.github/workflows/story-site-region.yml new file mode 100644 index 0000000..aa9203d --- /dev/null +++ b/.github/workflows/story-site-region.yml @@ -0,0 +1,174 @@ +name: Story site (one region) + +# One region's story site: plan the stories its site lacks, build them, publish. Called by story-site.yml once per +# region (a matrix); every region's site is separate (its own bucket prefix, master data, APK and catalog). + +on: + workflow_call: + inputs: + region: + description: "moenotes-masterdata-sync region: hk-tw-mo or jp" + type: string + required: true + stories: + type: string + default: "" + force: + type: boolean + default: false + dry_run: + type: boolean + default: false + +permissions: + contents: read + +env: + # The bucket of the site (repository variables; the defaults are the StarMoe site). JP is published under jp/: + # its Live2D model ids overlap the international ones. + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + STORY_S3_PREFIX: ${{ inputs.region == 'jp' && (vars.STORY_S3_PREFIX_JP || 'jp') || (vars.STORY_S3_PREFIX || '') }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every table with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ inputs.region }} + # At most this many stories per run (a run must end within the job limit); the next run builds the rest. + STORY_LIMIT: ${{ vars.STORY_LIMIT || '40' }} + # The ournotes-player the site is written for (its page files are part of the site). + PLAYER_REPOSITORY: ${{ vars.STORY_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.STORY_PLAYER_REF || '3774d8ac3987' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ inputs.region == 'jp' && 'com.bushiroad.sirius' || vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the keys and the CDN come from secrets in the build step, JP's API/CDN/ + # client version from its public master snapshot (story_site.py configure_region). + NNNOTES_CATALOG_REGION: ${{ inputs.region == 'jp' && 'jp' || 'tw' }} + NNNOTES_CATALOG_LANGUAGE: ${{ inputs.region == 'jp' && 'ja' || 'zh-Hant' }} + NNNOTES_SERVERS_JP_NAME: JP + NNNOTES_SERVERS_JP_LANGUAGES: ja + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 15 + outputs: + stories: ${{ steps.plan.outputs.stories }} + count: ${{ steps.plan.outputs.count }} + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - run: uv pip install --system --quiet boto3 + - name: Stories to build + id: plan + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + REQUESTED: ${{ inputs.stories }} + run: python .github/scripts/story_site.py plan + + build: + needs: plan + if: needs.plan.outputs.count != '0' + runs-on: ubuntu-latest + timeout-minutes: 340 + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: Free disk space + # The runner's preinstalled SDKs take most of its disk; a story build keeps its bundles and work files. + run: sudo rm -rf /usr/local/lib/android /usr/share/dotnet /opt/ghc /usr/local/.ghcup /opt/hostedtoolcache/CodeQL && df -h "$GITHUB_WORKSPACE" + + - uses: Swatinem/rust-cache@v2 # the install builds the extension module nnnotes._deck (rust/) + with: + workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - name: Install nnnotes + # UnityPy 1.25.x requires etcpak, which has no manylinux x86_64 wheel on PyPI (only PyPy / cp37), so the + # runner's cp313 environment cannot resolve it. pyproject.toml already pins it to the 1.24 line. + run: uv pip install --system -e ".[fonts]" boto3 + + - name: Tools + run: | + sudo apt-get update -qq && sudo apt-get install -y -qq ffmpeg > /dev/null + .github/scripts/tools.sh "$WORK/tools" + + - name: Fonts + uses: actions/cache@v6 + with: + path: ${{ env.WORK }}/fonts + key: story-fonts-${{ hashFiles('.github/scripts/fonts.sh') }} + - run: .github/scripts/fonts.sh "$WORK/fonts" + + - name: ournotes-player + uses: actions/cache@v6 + with: + path: ${{ env.WORK }}/player + key: story-player-${{ env.PLAYER_REPOSITORY }}-${{ env.PLAYER_REF }} + - run: .github/scripts/player.sh "$WORK/player" + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/story_site.py master "$WORK/master" + + - name: The published site + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + run: python .github/scripts/story_site.py fetch "$WORK/site" + + - name: Build + id: build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + NNNOTES_PATHS_PLAYER: ${{ env.WORK }}/player + NNNOTES_PATHS_VGMSTREAM: ${{ env.WORK }}/tools/vgmstream-cli + NNNOTES_PATHS_FONTS_JA: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf + NNNOTES_PATHS_FONTS_EN: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf + NNNOTES_PATHS_FONTS_ZH_HANT: ${{ env.WORK }}/fonts/NotoSansCJKtc-Regular.otf + NNNOTES_PATHS_FONTS_ZH_HANS: ${{ env.WORK }}/fonts/NotoSansCJKsc-Regular.otf + NNNOTES_PATHS_FONTS_KO: ${{ env.WORK }}/fonts/Pretendard-SemiBold.otf + NNNOTES_PATHS_FONTS_EMOJI: ${{ env.WORK }}/fonts/NotoColorEmoji.ttf + FORCE: ${{ inputs.force }} + run: python .github/scripts/story_site.py build "$WORK/site" ${{ needs.plan.outputs.stories }} + + - name: Publish + # Also after a partial build: the stories that were built are published, the failed ones are built again by + # the next run (they have no manifest). A dry run lists what it would upload. + if: always() && steps.build.outcome != 'skipped' && steps.build.outcome != 'cancelled' + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + run: python .github/scripts/story_site.py publish "$WORK/site" ${{ inputs.dry_run && '--dry-run' || '' }} + + - name: Failures + if: always() + run: | + f="$WORK/site.story-failures.json" + if [ -f "$f" ]; then echo "::error::stories failed (${{ inputs.region }}), see the job summary"; { echo '### Failed stories (${{ inputs.region }})'; echo '```json'; cat "$f"; echo '```'; } >> "$GITHUB_STEP_SUMMARY"; fi diff --git a/.github/workflows/story-site.yml b/.github/workflows/story-site.yml index 30f90cf..05dfe07 100644 --- a/.github/workflows/story-site.yml +++ b/.github/workflows/story-site.yml @@ -1,19 +1,26 @@ name: Story site -# Adds the stories that the published story site does not have yet to it. The site (stories.json, stories/, models/, -# assets/, story/: the layout `nnnotes web` writes) lives in an S3 bucket; a run fetches its manifests (not its assets), -# builds the missing stories with `nnnotes web --story` on top of them, and uploads what the build added or changed: -# new assets first, then the story and model manifests, the indexes last. Nothing is deleted from the bucket. +# Adds the stories that the published story sites do not have yet to them, one site per game region (hk-tw-mo at the +# bucket root, jp under jp/). A site (stories.json, stories/, models/, assets/, story/: the layout `nnnotes web` +# writes) lives in an S3 bucket; a region's run (story-site-region.yml) fetches its manifests (not its assets), builds +# the missing stories with `nnnotes web --story` on top of them, and uploads what the build added or changed: new +# assets first, then the story and model manifests, the indexes last. Nothing is deleted from the bucket. # -# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new -# master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds given stories (with `force`: -# again) or does a dry run. docs: .github/STORY_SITE.md +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` (client_payload.regions) when a +# region starts serving a new master snapshot: those regions build; a daily schedule catches a missed dispatch (every +# region); workflow_dispatch builds one region or all, given stories (with `force`: again) or does a dry run. +# The regions: repository variable STORY_REGIONS (default "hk-tw-mo jp"). docs: .github/STORY_SITE.md on: repository_dispatch: types: [masterdata-updated] workflow_dispatch: inputs: + region: + description: "the region to build (all: every region of STORY_REGIONS)" + type: choice + options: [all, hk-tw-mo, jp] + default: all stories: description: "MasterAdv ids to build, separated by spaces or commas (empty: every story the site lacks)" required: false @@ -32,155 +39,37 @@ on: permissions: contents: read -# One run at a time: every run rewrites the site's indexes from the manifests it fetched. -concurrency: - group: story-site-${{ vars.STORY_MASTERDATA_REGION || 'hk-tw-mo' }} - cancel-in-progress: false - -env: - # The bucket of the site (repository variables; the defaults are the StarMoe site). - STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} - STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} - STORY_S3_PREFIX: ${{ vars.STORY_S3_PREFIX || (vars.STORY_MASTERDATA_REGION == 'jp' && 'jp' || '') }} - # moenotes-masterdata-sync: decoded master data by region (index.json lists every table with its SHA-256). - MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} - MASTERDATA_REGION: ${{ vars.STORY_MASTERDATA_REGION || 'hk-tw-mo' }} - # At most this many stories per run (a run must end within the job limit); the next run builds the rest. - STORY_LIMIT: ${{ vars.STORY_LIMIT || '40' }} - # The ournotes-player the site is written for (its page files are part of the site). - PLAYER_REPOSITORY: ${{ vars.STORY_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} - PLAYER_REF: ${{ vars.STORY_PLAYER_REF || '3774d8ac3987' }} - PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} - APK_PACKAGE: ${{ vars.STORY_MASTERDATA_REGION == 'jp' && 'com.bushiroad.sirius' || vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} - PYTHON_VERSION: "3.13" - WORK: ${{ github.workspace }}/../work - # nnnotes settings (docs/configuration.md); the keys and the CDN come from secrets in the build step. - NNNOTES_CATALOG_REGION: ${{ vars.STORY_MASTERDATA_REGION == 'jp' && 'jp' || 'tw' }} - NNNOTES_CATALOG_LANGUAGE: ${{ vars.STORY_MASTERDATA_REGION == 'jp' && 'ja' || 'zh-Hant' }} - NNNOTES_SERVERS_JP_NAME: JP - NNNOTES_SERVERS_JP_LANGUAGES: ja - NNNOTES_SERVERS_TW_NAME: TW/HK/MO - NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja - jobs: - plan: + regions: runs-on: ubuntu-latest - timeout-minutes: 15 + timeout-minutes: 5 outputs: - stories: ${{ steps.plan.outputs.stories }} - count: ${{ steps.plan.outputs.count }} + regions: ${{ steps.regions.outputs.regions }} + count: ${{ steps.regions.outputs.count }} steps: - uses: actions/checkout@v7 - - uses: actions/setup-python@v7 - with: - python-version: ${{ env.PYTHON_VERSION }} - - uses: astral-sh/setup-uv@v7 - with: - enable-cache: true - - run: uv pip install --system --quiet boto3 - - name: Stories to build - id: plan + - name: Regions to build + id: regions env: - STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} - STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} - REQUESTED: ${{ github.event.inputs.stories }} - run: python .github/scripts/story_site.py plan - - build: - needs: plan - if: needs.plan.outputs.count != '0' - runs-on: ubuntu-latest - timeout-minutes: 340 - steps: - - uses: actions/checkout@v7 - - uses: actions/setup-python@v7 - with: - python-version: ${{ env.PYTHON_VERSION }} - cache: pip - cache-dependency-path: pyproject.toml - - uses: actions/setup-node@v7 - with: - node-version: "22" - - - name: Free disk space - # The runner's preinstalled SDKs take most of its disk; a story build keeps its bundles and work files. - run: sudo rm -rf /usr/local/lib/android /usr/share/dotnet /opt/ghc /usr/local/.ghcup /opt/hostedtoolcache/CodeQL && df -h "$GITHUB_WORKSPACE" - - - uses: Swatinem/rust-cache@v2 # the install builds the extension module nnnotes._deck (rust/) - with: - workspaces: rust - - uses: astral-sh/setup-uv@v7 - with: - enable-cache: true - - name: Install nnnotes - # UnityPy 1.25.x requires etcpak, which has no manylinux x86_64 wheel on PyPI (only PyPy / cp37), so the - # runner's cp313 environment cannot resolve it. pyproject.toml already pins it to the 1.24 line. - run: uv pip install --system -e ".[fonts]" boto3 - - - name: Tools - run: | - sudo apt-get update -qq && sudo apt-get install -y -qq ffmpeg > /dev/null - .github/scripts/tools.sh "$WORK/tools" - - - name: Fonts - uses: actions/cache@v6 - with: - path: ${{ env.WORK }}/fonts - key: story-fonts-${{ hashFiles('.github/scripts/fonts.sh') }} - - run: .github/scripts/fonts.sh "$WORK/fonts" - - - name: ournotes-player - uses: actions/cache@v6 - with: - path: ${{ env.WORK }}/player - key: story-player-${{ env.PLAYER_REPOSITORY }}-${{ env.PLAYER_REF }} - - run: .github/scripts/player.sh "$WORK/player" - - - name: APK - env: - PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} - run: .github/scripts/apk.sh "$WORK/apk" - - - name: Master data - run: python .github/scripts/story_site.py master "$WORK/master" - - - name: The published site - env: - STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} - STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} - run: python .github/scripts/story_site.py fetch "$WORK/site" - - - name: Build - id: build - env: - NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} - NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} - NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} - NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache - NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master - NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk - NNNOTES_PATHS_PLAYER: ${{ env.WORK }}/player - NNNOTES_PATHS_VGMSTREAM: ${{ env.WORK }}/tools/vgmstream-cli - NNNOTES_PATHS_FONTS_JA: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf - NNNOTES_PATHS_FONTS_EN: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf - NNNOTES_PATHS_FONTS_ZH_HANT: ${{ env.WORK }}/fonts/NotoSansCJKtc-Regular.otf - NNNOTES_PATHS_FONTS_ZH_HANS: ${{ env.WORK }}/fonts/NotoSansCJKsc-Regular.otf - NNNOTES_PATHS_FONTS_KO: ${{ env.WORK }}/fonts/Pretendard-SemiBold.otf - NNNOTES_PATHS_FONTS_EMOJI: ${{ env.WORK }}/fonts/NotoColorEmoji.ttf - FORCE: ${{ github.event.inputs.force }} - run: python .github/scripts/story_site.py build "$WORK/site" ${{ needs.plan.outputs.stories }} - - - name: Publish - # Also after a partial build: the stories that were built are published, the failed ones are built again by - # the next run (they have no manifest). A dry run lists what it would upload. - if: always() && steps.build.outcome != 'skipped' && steps.build.outcome != 'cancelled' - env: - STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} - STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} - run: python .github/scripts/story_site.py publish "$WORK/site" ${{ github.event.inputs.dry_run == 'true' && '--dry-run' || '' }} - - - name: Failures - if: always() - run: | - f="$WORK/site.story-failures.json" - if [ -f "$f" ]; then echo "::error::stories failed, see the job summary"; { echo '### Failed stories'; echo '```json'; cat "$f"; echo '```'; } >> "$GITHUB_STEP_SUMMARY"; fi + STORY_REGIONS: ${{ vars.STORY_REGIONS || 'hk-tw-mo jp' }} + REQUESTED_REGION: ${{ github.event.inputs.region }} + run: python3 .github/scripts/story_site.py regions + + site: + needs: regions + if: needs.regions.outputs.count != '0' + strategy: + fail-fast: false + matrix: + region: ${{ fromJSON(needs.regions.outputs.regions) }} + # One run per region at a time: every run rewrites its site's indexes from the manifests it fetched. + concurrency: + group: story-site-${{ matrix.region }} + cancel-in-progress: false + uses: ./.github/workflows/story-site-region.yml + with: + region: ${{ matrix.region }} + stories: ${{ github.event.inputs.stories || '' }} + force: ${{ github.event.inputs.force == 'true' }} + dry_run: ${{ github.event.inputs.dry_run == 'true' }} + secrets: inherit diff --git a/docs/jp.md b/docs/jp.md index 5362452..be4110a 100644 --- a/docs/jp.md +++ b/docs/jp.md @@ -95,7 +95,8 @@ APK catalog. A JP master download likewise rejects a different current master sn ## Workflow integration -For the StarMoe workflows, set `MUSIC_DATA_MASTERDATA_REGION=jp` or `STORY_MASTERDATA_REGION=jp`. This selects the +For the StarMoe workflows, set `MUSIC_DATA_MASTERDATA_REGION=jp`; the story site builds JP beside hk-tw-mo (`STORY_REGIONS`, +default `hk-tw-mo jp`, one run per region). This selects the JP package and Japanese catalog language; the build reads API/CDN/client-version from the JP entry in the public masterdata index. The default output prefixes become `jp/music-data` and `jp`, with separate concurrency groups. Custom output prefixes must also be separate from the international outputs. Publishing remains controlled by