diff --git a/.github/MUSIC_DATA.md b/.github/MUSIC_DATA.md new file mode 100644 index 0000000..220f5c0 --- /dev/null +++ b/.github/MUSIC_DATA.md @@ -0,0 +1,154 @@ +# Music data workflow (StarMoe) + +JP: `MUSIC_DATA_MASTERDATA_REGION=jp` selects the JP package, catalog and client metadata, with +`jp/music-data` as the default output prefix. The marker includes the asset hash and the provenance gate rejects +a catalog from another snapshot. See [Japanese release](../docs/jp.md) for setup and validation limits. + +`.github/workflows/music-data.yml` keeps the music data file of the chart data page (ournotes-player +`examples/songs`) up to date: `nnnotes music-data` of the current master data ([docs/music-data.md](../docs/music-data.md): +every song and chart with the deck model's statistics, the play scenarios and the chart's Gekisou skill aptitude), +checked by quality gates and +published into the story site's bucket under `music-data/` (`https://storage.bdon.moe/moenotes/music-data/`): + +| Object | Content | Cache-Control | +|---|---|---| +| `music-data.json` | the current file | `no-cache` | +| `jackets/.webp` | every song's jacket (`--jackets`), where the page looks for them | `public, max-age=86400` | +| `archive//.json` | every published file, kept | `public, max-age=31536000, immutable` | +| `build.json` | the build marker: the file's SHA-256, size and counts, what it was made from, the gate results, the run | `no-cache` | + +A run never deletes anything from the bucket. Its helper steps are `.github/scripts/music_data.py` (with the bucket, +HTTP and master data helpers of `story_site.py`), `music_data_smoke.mjs`, `songs_page.sh` and `apk.sh`; the gate +self-tests are `test_music_data.py` and `test_jp_workflow.py`. The fork also contains JP support in `src/nnnotes`. +The Cloudflare Pages preview of the page is not part of the workflow. + +## A run + +1. **plan** (seconds): the inputs of a build, from `index.json` of moenotes-masterdata-sync (the snapshot's master + data version, resource version and client version of `MUSIC_DATA_MASTERDATA_REGION`) and the checkout (the + ournotes-deck commit `rust/Cargo.lock` pins, the last nnnotes commit that changed `src/`, `rust/` or + `pyproject.toml`, and `RECIPE` of `music_data.py`), against `inputs` of the published `build.json`. The same + inputs (and a published `music-data.json`): the run ends here. `force` builds anyway. +2. **build**: + - the chart data page's modules (`examples/songs` of `MUSIC_DATA_PLAYER_REF`, not built), nnnotes with its deck + model (the install builds `nnnotes._deck`), the gate self-test, the APK (playfetch with `PLAYFETCH_CREDENTIALS`: + `provenance.client`; nnnotes also reads the bundles the APK carries, as in the story site's builds), the decoded + master data of moenotes-masterdata-sync (every file SHA-256 checked against `index.json`, `MasterManifest.json` + included); + - `nnnotes music-data --decoded-master --jackets jackets -o music-data.json`: the master data as decoded (no + master key), `provenance.master` the manifest's version and SHA-256 of the files as served; the charts, cue + sheets and jackets from the TW catalog, downloaded afresh on every run (never `actions/cache`: nnnotes keeps a + downloaded catalog for good, and the cache holds decrypted game files); + - the gates (below); a failed gate stops the run, the job summary lists why; + - upload: the jackets the bucket lacks (or has at another size; every one with `force`), the archive copy, then + `music-data.json`, `build.json` last. The archive copy, the file and the marker are each read back and checked + against their SHA-256 before the next is written: a failure leaves the previous `build.json`, so the next run + builds again. + +Triggers: `repository_dispatch` `masterdata-updated` (moenotes-masterdata-sync's `dispatch_repositories` already +names this repository for the story site: both workflows run), a daily schedule (03:41 UTC) in case a dispatch was +missed, and `workflow_dispatch`: + +| Input | Meaning | +|---|---| +| `force` | build and publish although the published file was made from the same inputs; upload every jacket again | +| `dry_run` | build and check, then list what would be uploaded instead of uploading | + +Runs do not overlap (`concurrency: music-data`). + +**Publishing switch.** Nothing is uploaded unless the repository variable `MUSIC_DATA_PUBLISH` is `true` (unset: +off). Off, every run, whatever its trigger (the schedule, `masterdata-updated`, `workflow_dispatch` with or without +`dry_run`), is a dry run: it builds, runs every gate and the smoke test, and lists what it would upload; the publish +step then gets no bucket key (anonymous, it can only read) and `music_data.py publish` itself refuses to upload. +Turn it on (`gh variable set MUSIC_DATA_PUBLISH --body true`) once the published format is final; while it is off, +nothing being published, `plan` finds no `build.json` and every run builds. + +## Gates + +Every one must pass, else nothing is published. Warnings go to the job summary and `build.json` and do not stop it. + +| Gate | Checks | +|---|---| +| (build) | nnnotes' own checks: every table, chart and cue sheet read, every chart measured, the deck statistics cross-checked against the chart facts and the master data (the command writes no file otherwise) | +| `schema` | the file against `docs/schema/music-data.schema.json` of the checkout (JSON Schema 2020-12) | +| `provenance` | `format`; `region` `tw`; `master.source` `api`; `master.version` equal to the snapshot's and its `MasterManifest.json`'s; every table's SHA-256 the manifest's, every decoded table read the one `index.json` lists; the song tables and the deck model's present; `deck.commit` the one `rust/Cargo.lock` pins; `exporter.version` the installed nnnotes; an APK version; a catalog SHA-256 (warning: the APK is another client version than the snapshot's) | +| `counts` | no fewer songs and charts than the published file (warning: ids no longer in it) | +| `deck` | deck statistics on every chart: kinds, a positive power, events and positions matching the chart, seeds unless unplayable (a warning): the one seed 0 on a chart without a luck range, else two or more different seeds, the same on every luck chart (their number is the file's, not fixed), `weights[kind][position]` numbers, every seed range's `rankBonus` = trunc(`rangeScore` x `rankBonusPercent` / 100) and its `luckPoints` an int, every check deck within its bound | +| `scenarios` | the play scenario fields: `offSeeds` exactly one entry (seed 0, score, weights, check within its bound), every range's `rankBonusPercents` five ints (the first `rankBonusPercent`), every seed's `scorePerfect`, `rangeWeights` (`[kind][position][range]`) and `rankCheck` (within its bound), every seed range's `rangeScorePerfect` (warnings, none in TW: a null `rangeWeights`, a null kind in it or in `offSeeds`' weights) | +| `aptitude` | the Gekisou skill aptitude (every shape alone on a chart). `deck.model.gekisouAptitude` a text; `deck.gekisouAptitude`: every key, `plainKind` the page's plain kind, a `host` text, the `seedRule` (a deterministic test, increasing batches, the targets, the cross seeds), `shapes` numbered 0, 1, 2, ... (source `member` or `support`, mission 1 to 4, `bandCondition` a support skill's alone and exactly when an effect has condition 5000, effect rows with every key and their condition groups, condition 5000 without targets, skills with a level and, with a band condition alone, member targets and bands). Every chart's `deck.gekisouAptitude`: null exactly when the chart is unplayable with Gekisou on, has no Gekisou range or there is no shape; else `factors` one per range (counts; no Just or Perfect notes outside a Just range; `lotteries` `[0, 0]` outside a luck range, else the mean of `deck.seeds`' `lotResults`) and `variants` one per shape of the chart's missions (or mission 4) in shape order, a band condition shape's `bandMatch` true then false: every `[mean, se]` two finite numbers with se >= 0 (every se 0 when deterministic), ranges one per range, `tail` = `score` less the ranges' `rangeScore` and `rankBonus` (allowing 0.0005 per rounded term plus 1e-6), deterministic point deltas integers and `tailPerfect` checked against the baseline Perfect range bonuses, 1 seed when deterministic else a batch of the seed rule (the last one when `seTargetMet` is false), `crossSeeds` min(seeds, the rule's), `weights` one per position and `rangeWeights` per position and range where the plain kind and `deck.seeds[0].rangeWeights` are, else null, the `check` on `deck.seeds[0]`'s seed, a rank per range (1 where the ranks are not linear), a plain kind value or null per position, within its bound (warning: variants that missed the standard error target) | +| `finite` | no NaN or infinity (warning: one inside master data rows, `songs[].master`, which the format writes as `1e999`) | +| `references` | texts in every language of `languages` (names and titles not empty); unique ids; songs sorted; the songs' bands, vocal characters and tags in the file; a band or a band name; a jacket, and its file in `jackets/`; a BGM cue; score ranks; charts in difficulty order, score ids unique (warnings: a title without a `zh-Hant` text, a music category on no tab, a character of no band) | +| `bgm` | every song's BGM length: `durationMs = samples * 1000 // sampleRate`, 30 s to 10 min, within 1 s of the cue's `lengthMs`, not ending before a chart's last note (warning: more than a minute after it) | +| `size` | 0.8 to 2 times the published file | +| `gzip` | the file gzipped at most 2 MB (0.37 MB before the aptitude), its Gekisou skill aptitude gzipped at most 1.2 MB (about 0.4 MB expected) | +| `page` | `music_data_smoke.mjs`: the page's `catalog.js` and `ranking.js` in Node.js over the file: a row per chart, a plain score-up kind, data for the free, rank and Just scenarios, finite positive figures for every chart the data covers in seven scenarios (Gekisou Live at several ranks, Just rates and a Great share, Free Live), the ranking, frontier and event figures | +| (publish) | read back after upload, SHA-256 checked | + +`counts` and `size` compare with the published `music-data.json` and are skipped while nothing is published. A +legitimate drop (a song the game removed) stops the run: a person checks it, then moves the published +`music-data.json` away (the archive keeps it) or changes the gate in a pull request. + +The self-test runs in each build and locally in seconds, without the network: + +``` +python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py +``` + +with, optionally, `MUSIC_DATA_SCHEMA` (a schema file when the checkout has none), `MUSIC_DATA_PAGE` (an +`examples/songs` directory: the smoke test), `MUSIC_DATA_SAMPLE` (a real file with the play scenario fields: its +content gates pass; one made before the ranges' `luckPoints` and the aptitude: the deck and aptitude gates stop it +on those alone) and +`MUSIC_DATA_OLD_SAMPLE` (one without the play scenario fields: the scenario gate stops it). + +### Aptitude page smoke + +With the aptitude API (ournotes-player PR #11, `1522c24`), the same Node smoke also checks shape/skill/band +lookups, chart variants, all five battle scenarios, Free Live exclusion, finite gains, raw standard errors, +missing cross terms and the absence of standard errors for transformed or combined figures. Removing aptitude +must not change the default chart figures: default ranking still has no card Gekisou skills. + +Only deterministic variants are reconstructed against their individual `check` seed, using positional cards and +`masterSkillFactor` for the game's float32 conversion. Stochastic means are never used to reconstruct a check. +Older pinned page modules explicitly report `API unavailable (skipped)`; moving `MUSIC_DATA_PLAYER_REF` remains a +separate rollout decision. No browser or page build is needed. + +## Settings + +Repository secrets: the story site's (`.github/STORY_SITE.md`), no new one: `NNNOTES_BUNDLE_KEY`, +`NNNOTES_BUNDLE_NONCE_SEED`, `NNNOTES_SERVERS_TW_CDN`, `PLAYFETCH_CREDENTIALS`, `STORY_S3_ACCESS_KEY`, +`STORY_S3_SECRET_KEY`. The master key is not needed: the master data comes decoded. + +Repository variables: + +| Variable | Default | | +|---|---|---| +| `MUSIC_DATA_PLAYER_REF` | none: **required** | the ournotes-player commit whose chart data page reads this file (the page with the play scenarios); a run stops before building without it | +| `MUSIC_DATA_PUBLISH` | none: off | `true`: upload; anything else: every run is a dry run | +| `MUSIC_DATA_PLAYER_REPOSITORY` | `empty-sekai/ournotes-player` | | +| `MUSIC_DATA_S3_PREFIX` | `music-data` | the key prefix in the bucket | +| `MUSIC_DATA_MASTERDATA_REGION` | `hk-tw-mo` | the region of `index.json`; the build reads the TW catalog (`[catalog] region` `tw`) | +| `STORY_S3_ENDPOINT`, `STORY_S3_BUCKET`, `MASTERDATA_BASE_URL`, `PLAYFETCH_VERSION`, `STORY_APK_PACKAGE` | the story site's | shared with it | + +## Before the first run + +- **nnnotes.** The workflow runs this fork's nnnotes. It needs upstream's `music-data` command with the play + scenarios (MetaSekaiLab/nnnotes `a03591e`) and `--decoded-master` (MetaSekaiLab/nnnotes#6, `12df2a6`): sync the + fork with upstream first. Until then `plan` stops naming what is missing. The deck and aptitude gates also need + the ranges' `luckPoints` and the Gekisou skill aptitude, which come with nnnotes' and ournotes-deck's Gekisou skill + changes: until the fork has them every build stops there. +- **The page.** Set `MUSIC_DATA_PLAYER_REF` to the ournotes-player commit of the chart data page that reads the play + scenario fields, once that page is merged. +- **Publishing.** Set `MUSIC_DATA_PUBLISH` to `true` last, when dry runs pass and the published format is final. + +## Notes + +- **Versions.** A new ournotes-deck pin (`rust/Cargo.toml`, `rust/Cargo.lock`) or a new nnnotes commit in `src/`, + `rust/` or `pyproject.toml` reaches this fork with a sync, and the next run builds a new file. The deck + statistics may then differ: `provenance.deck.commit` and `build.json` name the commit. +- **The page and the data.** The page's modules are pinned by `MUSIC_DATA_PLAYER_REF`: after a page release that + reads new fields, move it (a new ref alone does not start a build; run with `force` to check the published data + against the new page). +- **Byte identity.** A build from the same inputs gives the same bytes (the file is canonical); the jackets' + WebP bytes depend on the Pillow version. +- **Logs.** The steps print counts, ids, SHA-256 and field names, not game content; nothing decrypted is cached or + uploaded as an artifact. diff --git a/.github/STORY_SITE.md b/.github/STORY_SITE.md new file mode 100644 index 0000000..6056712 --- /dev/null +++ b/.github/STORY_SITE.md @@ -0,0 +1,76 @@ +# Story site workflow (StarMoe) + +`.github/workflows/story-site.yml` keeps the StarMoe story site up to date: it builds the stories the published site +lacks with this repository's `nnnotes web --story` and uploads them to the bucket that serves the site +(`https://storage.bdon.moe/moenotes/`, the layout `nnnotes web` writes: `stories.json`, `stories/`, `models/`, +`assets/`, `story/`), one site per game region: `hk-tw-mo` at the bucket root, `jp` under `jp/` +(`https://storage.bdon.moe/moenotes/jp/`; JP Live2D model ids overlap the international ones). It only adds: a run +never deletes anything from the bucket. Its helper steps are in +`.github/scripts/`; nothing outside `.github/` differs from upstream, so the fork syncs with it as before. + +## A run + +`story-site.yml` picks the regions (`story_site.py regions`: those of `STORY_REGIONS` that the dispatch's +`client_payload.regions` names, the `region` input of a manual run, every one on the schedule) and calls +`story-site-region.yml` once per region, in parallel. Each region's run: + +1. **plan** (a few seconds): the MasterAdv ids of the decoded master data of moenotes-masterdata-sync + (`MasterAdv.json`, SHA-256 checked against its `index.json`) against the `stories/.json` objects of the bucket. + The ids without a manifest, at most `STORY_LIMIT` (40) in id order, are the run's stories; none: the run ends here. +2. **build** (only when there is something to build): + - fonts (pinned by SHA-256: the files the published stories record in `ui/fonts.json`), vgmstream, ffmpeg, the + built ournotes-player (`STORY_PLAYER_REF`), the APK (playfetch with the account in `PLAYFETCH_CREDENTIALS`), the + decoded master data; + - every object of the site except `assets/` (the manifests and indexes, about 150 MB), over plain HTTP like the + player's browser: the bucket serves public read, and Cloudflare's S3-signed ranged downloads were rejected + intermittently with `SignatureDoesNotMatch`; + - `nnnotes web site --story ...` (with the Live2D models these stories load that the site lacks), then + `nnnotes web site --player-only`, which rewrites `stories.json`, `models.json`, `charts.json` and the player + pages from every manifest present, also after a failed build; + - upload: the assets the bucket lacks, then the new or changed manifests and player files, the indexes last. + Stories that failed have no manifest, so the next run builds them again; their errors are in the job summary. + +Triggers: `repository_dispatch` `masterdata-updated` (moenotes-masterdata-sync sends it when a region serves a new +snapshot: `dispatch_repositories`), a daily schedule (03:23 UTC) in case a dispatch was missed, and +`workflow_dispatch`: + +| Input | Meaning | +|---|---| +| `stories` | MasterAdv ids to build (spaces or commas); empty: every story the site lacks | +| `force` | rebuild the given stories and their Live2D models although their manifests exist | +| `region` | `all` (every region of `STORY_REGIONS`), `hk-tw-mo` or `jp` | +| `dry_run` | build, then list what would be uploaded instead of uploading | + +Runs of one region do not overlap (`concurrency: story-site-`); the regions build side by side. + +## Settings + +Repository secrets (Settings → Secrets and variables → Actions → Secrets): + +| Secret | Value | +|---|---| +| `NNNOTES_BUNDLE_KEY` | `[bundle] key` (32 hex digits) | +| `NNNOTES_BUNDLE_NONCE_SEED` | `[bundle] nonce_seed` | +| `NNNOTES_SERVERS_TW_CDN` | `[servers.tw] cdn`: the TW CDN base URL | +| `PLAYFETCH_CREDENTIALS` | the whole `credentials.json` of `playfetch login` (the account that can pull `com.bilibili.sirius`) | +| `STORY_S3_ACCESS_KEY`, `STORY_S3_SECRET_KEY` | an S3 key that can list, read and write the bucket | + +Repository variables (optional; the defaults are the StarMoe site): `STORY_S3_ENDPOINT` (`https://storage.bdon.moe`), +`STORY_S3_BUCKET` (`moenotes`), `STORY_S3_PREFIX` (empty: the bucket root; the hk-tw-mo site), `STORY_S3_PREFIX_JP` (`jp`), `MASTERDATA_BASE_URL` +(`https://metadata.bdon.moe`), `STORY_REGIONS` (`hk-tw-mo jp`: the regions with a site), `STORY_LIMIT` (`40`), `STORY_PLAYER_REPOSITORY` +(`empty-sekai/ournotes-player`), `STORY_PLAYER_REF` (`3774d8ac3987`), `PLAYFETCH_VERSION` (`v0.92`), +`STORY_APK_PACKAGE` (`com.bilibili.sirius`). + +## Notes + +- **The player version.** The site's player pages come from `STORY_PLAYER_REF`, and its data must be what that player + reads. After a player release that changes the data format, sync this fork with upstream nnnotes and set + `STORY_PLAYER_REF` to the matching player; new stories are then written for it. Stories already published are + not rebuilt (run with `stories` and `force` for that). +- **The APK.** Google Play serves the current game version. When nnnotes' type trees do not match it, the build stops + naming the class and the Unity version: sync the fork with upstream once nnnotes supports that version. +- **Bytes.** A rebuild on the runner is not byte-identical to a build elsewhere: the AAC encoder (ffmpeg) and the PNG + encoder (zlib) differ between machines. The pixels and every JSON value are the same, so new stories and the + published ones fit together; only a forced rebuild re-uploads such files. +- **Master data.** The build reads moenotes-masterdata-sync's current snapshot; stories published from another master + source keep their texts until they are rebuilt. diff --git a/.github/scripts/apk.sh b/.github/scripts/apk.sh new file mode 100755 index 0000000..9aa5f52 --- /dev/null +++ b/.github/scripts/apk.sh @@ -0,0 +1,26 @@ +#!/usr/bin/env bash +# The game's split APKs (base.apk; split_config.arm64_v8a.apk holds the CRI Lips library) into $1, with playfetch +# ($PLAYFETCH_VERSION) and the account store in the secret PLAYFETCH_CREDENTIALS (the credentials.json of +# `playfetch login`). Google Play serves the current version; nnnotes stops, naming the class, when its type trees do not +# match the APK's Unity version. +set -euo pipefail +dir="$1" +if [ -z "${PLAYFETCH_CREDENTIALS_JSON:-}" ]; then + echo "::error::set the repository secret PLAYFETCH_CREDENTIALS (the credentials.json of playfetch login)" + exit 1 +fi +bin="$RUNNER_TEMP/playfetch" +curl -fsSL -o "$bin" "https://github.com/Exmeaning/playfetch/releases/download/$PLAYFETCH_VERSION/playfetch-$PLAYFETCH_VERSION-linux-amd64" +curl -fsSL -o "$RUNNER_TEMP/SHA256SUMS" "https://github.com/Exmeaning/playfetch/releases/download/$PLAYFETCH_VERSION/SHA256SUMS" +(cd "$RUNNER_TEMP" && grep " playfetch-$PLAYFETCH_VERSION-linux-amd64\$" SHA256SUMS | sed "s| playfetch-$PLAYFETCH_VERSION-linux-amd64| playfetch|" | sha256sum -c -) +chmod 755 "$bin" +export PLAYFETCH_CREDENTIALS="$RUNNER_TEMP/playfetch-credentials.json" +umask 077 +printf '%s' "$PLAYFETCH_CREDENTIALS_JSON" > "$PLAYFETCH_CREDENTIALS" +"$bin" pull "$APK_PACKAGE" -out-root "$RUNNER_TEMP/apk-downloads" -mode split +got="$(find "$RUNNER_TEMP/apk-downloads/$APK_PACKAGE" -name base.apk | sort | tail -1)" +[ -n "$got" ] || { echo "::error::playfetch pulled no base.apk"; exit 1; } +mkdir -p "$dir" +cp "$(dirname "$got")"/*.apk "$dir/" +rm -f "$PLAYFETCH_CREDENTIALS" +ls -l "$dir" diff --git a/.github/scripts/fonts.sh b/.github/scripts/fonts.sh new file mode 100755 index 0000000..40dd4f8 --- /dev/null +++ b/.github/scripts/fonts.sh @@ -0,0 +1,27 @@ +#!/usr/bin/env bash +# The open fonts the site's story text is drawn with (the files the published stories record in ui/fonts.json), into +# $1, each checked against its SHA-256: Noto Sans CJK 2.004 (ja and en: JP, zh-Hant: TC, zh-Hans: SC), Pretendard 1.3.9 +# SemiBold (ko), Noto Color Emoji 2.051 (the emoji sprites). All under the SIL Open Font License 1.1. +set -euo pipefail +dir="$1" +mkdir -p "$dir" +cd "$dir" +cjk=https://github.com/notofonts/noto-cjk/raw/165c01b46ea533872e002e0785ff17e44f6d97d8/Sans/OTF +fetch() { # fetch FILE URL SHA256 + if [ ! -f "$1" ] || ! echo "$3 $1" | sha256sum -c --quiet - 2>/dev/null; then + curl -fsSL -o "$1" "$2" + echo "$3 $1" | sha256sum -c - + fi +} +fetch NotoSansCJKjp-Regular.otf "$cjk/Japanese/NotoSansCJKjp-Regular.otf" 68a3fc98800b2a27b371f2fb79991daf3633bd89309d4ffaa6946fd587f375b5 +fetch NotoSansCJKtc-Regular.otf "$cjk/TraditionalChinese/NotoSansCJKtc-Regular.otf" dce08bd4fd91aa8aa76ed8fea4b694c2dfb8550f67871e326843212ddbeb88b4 +fetch NotoSansCJKsc-Regular.otf "$cjk/SimplifiedChinese/NotoSansCJKsc-Regular.otf" 2c76254f6fc379fddfce0a7e84fb5385bb135d3e399294f6eeb6680d0365b74b +fetch NotoColorEmoji.ttf https://github.com/googlefonts/noto-emoji/raw/v2.051/fonts/NotoColorEmoji.ttf 72a635cb3d2f3524c51620cdde406b217204e8a6a06c6a096ff8ed4b5fd6e27b +if [ ! -f Pretendard-SemiBold.otf ] || ! echo "c89bc43027dc7cde5726e96223376f8eec09302b2fc1f8147fd5b57cfc376118 Pretendard-SemiBold.otf" | sha256sum -c --quiet - 2>/dev/null; then + curl -fsSL -o pretendard.zip https://github.com/orioncactus/pretendard/releases/download/v1.3.9/Pretendard-1.3.9.zip + echo "04be351a74d6bf7d60c480a3087e51d185485d35a52023142af1df19eb8c428a pretendard.zip" | sha256sum -c - + unzip -q -o -j pretendard.zip public/static/Pretendard-SemiBold.otf + rm pretendard.zip + echo "c89bc43027dc7cde5726e96223376f8eec09302b2fc1f8147fd5b57cfc376118 Pretendard-SemiBold.otf" | sha256sum -c - +fi +ls -l diff --git a/.github/scripts/music_data.py b/.github/scripts/music_data.py new file mode 100755 index 0000000..2623513 --- /dev/null +++ b/.github/scripts/music_data.py @@ -0,0 +1,1280 @@ +#!/usr/bin/env python3 +"""The music data CI steps (.github/workflows/music-data.yml): `nnnotes music-data` of moenotes-masterdata-sync's +decoded master data, checked by quality gates and published into the story site's bucket under $MUSIC_DATA_S3_PREFIX +(music-data.json, jackets/, archive/, build.json) for the chart data page of ournotes-player (examples/songs). + + plan the inputs of a build (the master data snapshot of $MASTERDATA_REGION in index.json, the + deck commit rust/Cargo.lock pins, the last nnnotes commit of src/, rust/ and pyproject.toml, + RECIPE) against those of the published build.json; GitHub output `build`: true when they + differ, nothing is published or $FORCE is true + master OUT every file of the snapshot of $MASTERDATA_REGION into OUT (SHA-256 checked against + index.json, MasterManifest.json included) and its index entry as OUT.snapshot.json + build OUT `nnnotes music-data --decoded-master --jackets OUT/jackets -o OUT/music-data.json` ([paths] + master: the master step's OUT), its printed summary in OUT/music-data.summary.json + check OUT MASTER PAGE the quality gates (.github/MUSIC_DATA.md) on OUT/music-data.json, with the published file + as the baseline and PAGE the chart data page's modules (examples/songs) for the smoke test; + the report in OUT/check.json, and OUT/build.json (the build marker) when every gate passed; + exits 1 when one failed + publish OUT [--dry-run] + the jackets the bucket lacks or has at another size (every one with $FORCE), the file's + archive copy, then music-data.json, build.json last; each read back and its SHA-256 checked. + Only a file whose check.json passed; never deletes. A dry run unless $MUSIC_DATA_PUBLISH is + `true` (the publishing switch, off by default): it lists what it would upload + +Bucket: story_site.Bucket ($STORY_S3_ENDPOINT, $STORY_S3_BUCKET, credentials $STORY_S3_ACCESS_KEY / +$STORY_S3_SECRET_KEY) with the key prefix $MUSIC_DATA_S3_PREFIX (default music-data). The published files are read +over plain HTTP, as the page reads them (the bucket serves public read). +""" +from __future__ import annotations + +import gzip +import hashlib +import json +import math +import os +import re +import subprocess +import sys +import time +import urllib.error +import urllib.request +from dataclasses import dataclass +from datetime import datetime, timezone +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import story_site # noqa: E402 (the bucket, HTTP and master data helpers) +from story_site import env, get, output, summary # noqa: E402 + +FORMAT = "nnnotes.music-data/1" +BUILD_FORMAT = "moenotes.music-data-build/1" +# This script's own version of a build: bump it when what it builds or publishes changes, so that the next run builds +# although the master data, the deck model and nnnotes are the same. +RECIPE = 2 +FILE, MARKER, JACKETS, ARCHIVE = "music-data.json", "build.json", "jackets/", "archive/" +MANIFEST = "MasterManifest.json" +SOURCE_PATHS = ("src", "rust", "pyproject.toml") # nnnotes' code: the commit that last changed one of them +SCHEMA = Path("docs/schema/music-data.schema.json") +SMOKE = Path(__file__).resolve().parent / "music_data_smoke.mjs" +ARCHIVE_CACHE = story_site.ASSET_CACHE # archive//.json: content-addressed +FILE_CACHE = "no-cache" # music-data.json and build.json change in place +JACKET_CACHE = "public, max-age=86400" +SNAPSHOT_KEYS = ("version", "resource_version", "resource_hash", "client_version", "verified_at", "manifest_sha256", "table_count") + +# the gates' bounds (MUSIC_DATA.md) +SIZE_RATIO = (0.8, 2.0) # against the published file +FILE_GZIP_MAX = 2_000_000 # the file gzipped, the download (0.37 MB before the aptitude) +APTITUDE_GZIP_MAX = 1_200_000 # the Gekisou skill aptitude gzipped (about 0.4 MB expected) +BGM_MS = (30_000, 600_000) # a song's BGM length +BGM_CUE_SLACK_MS = 1000 # |durationMs - lengthMs| +BGM_TAIL_MS = 60_000 # BGM after the last note: more is reported +RANKS = 5 +LUCK_MISSION = 2 # a range of it draws lots: its chart has several seeds +JUST_MISSION = 3 # a range of it judges Just +MISSIONS = (1, 2, 3, 4) # a Gekisou skill's: combo, luck, Just, every one +PLAIN_KIND = (2000, 5000) # the page's plain kind (ranking.js plainKind): effect type, ms +BAND_CONDITION = 5000 # the skill condition on the paired member (its band) +APTITUDE_SLACK = 1e-6 # relative: the aptitude's identities on means of integers +DIFFICULTIES = ("easy", "normal", "hard", "expert") +SONG_TABLES = ("MasterLiveMusic", "MasterLiveMusicScore", "MasterText", "MasterBand", "MasterCharacter", "MasterTag", + "MasterLiveMusicCategory", "MasterSound", "MasterSoundCueSheet", "MasterLiveScoreRank") +SHA256 = re.compile(r"[0-9a-f]{64}") +COMMIT = re.compile(r"[0-9a-f]{40}") +LISTED = 20 # failures and warnings listed per gate + + +def fail(message: str): + sys.exit(f"music_data: {message}") + + +def sha256(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +# ---------------------------------------------------------------- bucket +def s3_prefix() -> str: + p = (os.environ.get("MUSIC_DATA_S3_PREFIX") or "music-data").strip("/") + return f"{p}/" if p else "" + + +def bucket() -> "story_site.Bucket": + b = story_site.Bucket() + b.prefix = s3_prefix() + return b + + +def public_url(key: str) -> str: + return f"{env('STORY_S3_ENDPOINT').rstrip('/')}/{env('STORY_S3_BUCKET')}/{s3_prefix()}{key}" + + +def published(key: str) -> bytes | None: + """A published file (plain HTTP), None when the bucket has none.""" + try: + return get(public_url(key), timeout=300) + except urllib.error.HTTPError as e: + if e.code == 404: + return None + raise + + +def exists(key: str) -> bool: + request = urllib.request.Request(public_url(key), method="HEAD", + headers={"User-Agent": "moenotes-music-data (GitHub Actions)", + "Cache-Control": "no-cache"}) + try: + with urllib.request.urlopen(request, timeout=60): + return True + except urllib.error.HTTPError as e: + if e.code == 404: + return False + raise + + +# ---------------------------------------------------------------- the inputs of a build +def deck_commit(root: Path = Path(".")) -> str: + """The ournotes-deck commit nnnotes builds its deck model with (rust/Cargo.lock).""" + lock = root / "rust" / "Cargo.lock" + if not lock.is_file(): + fail("no rust/Cargo.lock: this nnnotes has no music-data command with the deck model (sync the fork with " + "upstream, .github/MUSIC_DATA.md)") + for block in lock.read_text(encoding="utf-8").split("[[package]]"): + if re.search(r'^name = "ournotes-deck"$', block, re.M): + m = re.search(r'^source = "git\+[^"#]*#([0-9a-f]{40})"$', block, re.M) + if m: + return m.group(1) + fail("rust/Cargo.lock has no ournotes-deck git commit") + + +def nnnotes_commit(root: Path = Path(".")) -> str: + """The commit that last changed nnnotes' code (SOURCE_PATHS); the checkout needs its history.""" + def git(*args): + return subprocess.run(["git", "-C", str(root), *args], capture_output=True, text=True) + if git("rev-parse", "--is-shallow-repository").stdout.strip() != "false": + fail("the checkout is shallow: the nnnotes commit needs the history (actions/checkout fetch-depth: 0)") + commit = git("log", "-1", "--format=%H", "--", *SOURCE_PATHS).stdout.strip() + if not COMMIT.fullmatch(commit): + fail("no commit changed src/, rust/ or pyproject.toml") + return commit + + +def require_decoded_master(root: Path = Path(".")) -> None: + cli = root / "src" / "nnnotes" / "cli.py" + if "--decoded-master" not in cli.read_text(encoding="utf-8"): + fail("this nnnotes has no `music-data --decoded-master` (MetaSekaiLab/nnnotes#6): sync the fork with " + "upstream (.github/MUSIC_DATA.md)") + + +def inputs(entry: dict, root: Path = Path(".")) -> dict: + """What a build is made of: the master data snapshot, the deck model, nnnotes and this script.""" + return {"masterRegion": env("MASTERDATA_REGION"), "masterVersion": entry.get("version"), + "resourceVersion": entry.get("resource_version"), "clientVersion": entry.get("client_version"), + "resourceHash": entry.get("resource_hash"), + "deckCommit": deck_commit(root), "nnnotesCommit": nnnotes_commit(root), "recipe": RECIPE} + + +def short(v) -> str: + return str(v)[:12] if v is not None else "?" + + +# ---------------------------------------------------------------- plan +def cmd_plan() -> None: + _, region = story_site.master_index() + require_decoded_master() + now = inputs(region.get("entry") or {}) + raw = published(MARKER) + marker = json.loads(raw) if raw else None + have = marker.get("inputs") if isinstance(marker, dict) else None + changed = [k for k in now if not isinstance(have, dict) or have.get(k) != now[k]] + if have and not exists(FILE): + changed.append("music-data.json (missing)") + force = os.environ.get("FORCE") == "true" + build = force or bool(changed) + summary(f"### Music data\n\n- master: {now['masterVersion']} of {now['masterRegion']} (resource " + f"{now['resourceVersion']}, client {now['clientVersion']}); deck {short(now['deckCommit'])}; nnnotes " + f"{short(now['nnnotesCommit'])}; recipe {RECIPE}\n- published: " + + (f"master {have.get('masterVersion')}, deck {short(have.get('deckCommit'))}, nnnotes " + f"{short(have.get('nnnotesCommit'))}, built {marker.get('builtAt')}" if isinstance(have, dict) + else "nothing") + + "\n- this run: " + ("builds" + (" (force)" if force else "") + (f", changed: {', '.join(changed)}" + if changed and have else "") + if build else "nothing changed, nothing to build")) + summary(f"- publishing: {'on' if publishing() else 'off (MUSIC_DATA_PUBLISH is not `true`: a dry run)'}") + output("build", "true" if build else "false") + + +# ---------------------------------------------------------------- master data +def snapshot_file(master: Path) -> Path: + return master.parent / f"{master.name}.snapshot.json" + + +def cmd_master(out: str) -> None: + url, region = story_site.master_index() + files = region["files"] + entry = region.get("entry") or {} + if MANIFEST not in files: + fail(f"index.json lists no {MANIFEST} for {env('MASTERDATA_REGION')}") + if entry.get("manifest_sha256") and entry["manifest_sha256"] != files[MANIFEST]: + fail(f"index.json: the entry's manifest_sha256 is not the SHA-256 of its {MANIFEST}") + d = Path(out) + d.mkdir(parents=True, exist_ok=True) + story_site.parallel(lambda item: story_site.fetch_table(url, item[0], item[1], d / item[0]), files.items(), 8) + version = json.loads((d / MANIFEST).read_bytes()).get("version") + if entry.get("version") is not None and str(version) != str(entry["version"]): + fail(f"{MANIFEST} has version {version}, index.json {entry['version']} (the service changed snapshots? " + f"run again)") + snapshot = {"region": env("MASTERDATA_REGION"), "entry": {k: entry.get(k) for k in SNAPSHOT_KEYS}, + "files": files} + snapshot_file(d).write_text(json.dumps(snapshot, indent=1, sort_keys=True), encoding="utf-8") + print(f"master data: {version}, {len(files)} files into {d}") + + +# ---------------------------------------------------------------- build +def cmd_build(out: str) -> None: + story_site.configure_region() + o = Path(out).resolve() + o.mkdir(parents=True, exist_ok=True) + nnnotes = [sys.executable, "-m", "nnnotes"] + usage = subprocess.run(nnnotes + ["music-data", "--help"], capture_output=True, text=True).stdout + if "--decoded-master" not in usage: + fail("the installed nnnotes has no `music-data --decoded-master` (sync the fork with upstream)") + cmd = nnnotes + ["music-data", "--decoded-master", "--jackets", str(o / "jackets"), "-o", str(o / FILE)] + print("+ " + " ".join(cmd[1:]), flush=True) + with open(o / "music-data.summary.json", "wb") as f: + status = subprocess.run(cmd, stdout=f).returncode + if status: + fail(f"nnnotes music-data exited with {status}") + r = json.loads((o / "music-data.summary.json").read_text(encoding="utf-8")) + summary(f"- built: {r.get('songs')} songs, {r.get('charts')} charts, {r.get('jackets')} jackets, deck " + f"{short(r.get('deck'))}, {r.get('bytes')} bytes, sha256 {short(r.get('sha256'))}") + + +# ---------------------------------------------------------------- the gates +@dataclass +class Context: + """What the gates check the file against; None: the gate (or that part of it) is skipped (the self-test).""" + region: str | None = None # [catalog] region: provenance.region + language: str | None = None # [catalog] language: a song title should have it + master: Path | None = None # the decoded master data read (MasterManifest.json, .json) + snapshot: dict | None = None # its index entry and files (the master step's OUT.snapshot.json) + deck_commit: str | None = None + nnnotes_version: str | None = None + jackets: Path | None = None + schema: Path | None = None + published: bytes | None = None # the published music-data.json (None: nothing is published) + page: Path | None = None # examples/songs of ournotes-player + file: Path | None = None # the file on disk, for the smoke test + + +class Gate: + def __init__(self): + self.failures: list[str] = [] + self.warnings: list[str] = [] + self.note = "" + + def fail(self, message: str): + self.failures.append(message) + + def warn(self, message: str): + self.warnings.append(message) + + +def charts_of(doc: dict): + for song in doc.get("songs") or []: + for chart in song.get("charts") or []: + yield song, chart + + +def where(song: dict, chart: dict) -> str: + return f"chart {chart.get('scoreId')} ({song.get('id')} {chart.get('difficulty')})" + + +def is_int(v) -> bool: + return isinstance(v, int) and not isinstance(v, bool) + + +def is_num(v) -> bool: + return (isinstance(v, (int, float)) and not isinstance(v, bool)) and math.isfinite(v) + + +def within(c) -> bool: + return (isinstance(c, dict) and is_int(c.get("exact")) and is_num(c.get("predicted")) and is_num(c.get("bound")) + and abs(c["exact"] - c["predicted"]) <= c["bound"]) + + +def rank_bonus(range_score: int, percent: int) -> int: + """trunc(rangeScore * percent / 100): a range's rank bonus.""" + q = abs(range_score * percent) // 100 + return q if range_score * percent >= 0 else -q + + +def plain_kind(doc: dict): + """The id of the page's plain score-up kind in deck.kinds (ranking.js plainKind): effect type 2000 on the whole + deck for 5 s, without targets, conditions or limits; None for none.""" + for k in ((doc.get("deck") or {}).get("kinds")) or []: + if (isinstance(k, dict) and k.get("effectType") == PLAIN_KIND[0] and not k.get("skillTargetIds") + and not any(k.get(x) for x in ("skillConditionGroup", "skillReleaseConditionGroup", + "effectLimitCount", "effectExecuteLimitCount")) + and (PLAIN_KIND[1] if k.get("durationMs") is None else k["durationMs"]) == PLAIN_KIND[1]): + return k.get("id") + return None + + +def luck_chart(deck: dict) -> bool: + return any(isinstance(r, dict) and r.get("mission") == LUCK_MISSION for r in deck.get("ranges") or []) + + +def gate_schema(doc, ctx: Context, g: Gate): + if ctx.schema is None: + g.note = "skipped" + return + if not ctx.schema.is_file(): + g.fail(f"no JSON Schema {ctx.schema}") + return + import jsonschema + schema = json.loads(ctx.schema.read_text(encoding="utf-8")) + validator = jsonschema.Draft202012Validator(schema) + for e in sorted(validator.iter_errors(doc), key=lambda e: list(map(str, e.absolute_path))): + g.fail(f"{'/'.join(map(str, e.absolute_path)) or '(root)'}: {e.message[:160]}") + g.note = ctx.schema.as_posix() + + +def gate_provenance(doc, ctx: Context, g: Gate): + if doc.get("format") != FORMAT: + g.fail(f"format {doc.get('format')!r}, expected {FORMAT}") + p = doc.get("provenance") or {} + if ctx.region is not None and p.get("region") != ctx.region: + g.fail(f"region {p.get('region')!r}, expected {ctx.region!r}") + m = p.get("master") or {} + if m.get("source") != "api": + g.fail(f"master.source {m.get('source')!r}, expected 'api'") + tables = m.get("tables") or {} + missing = [t for t in SONG_TABLES if t not in tables] + if missing: + g.fail(f"master.tables lacks {', '.join(missing)}") + if doc.get("deck") is not None and len(tables) <= len(SONG_TABLES): + g.fail("master.tables has only the song tables, but the file has deck statistics") + if ctx.snapshot is not None: + v = ctx.snapshot["entry"].get("version") + if m.get("version") != v: + g.fail(f"master.version {m.get('version')!r}, the snapshot's {v!r}") + if ctx.snapshot.get("region") == "jp": + for expected, actual in (("resource_version", "resourceVersion"), ("resource_hash", "resourceHash")): + if (p.get("catalog") or {}).get(actual) != ctx.snapshot["entry"].get(expected): + g.fail(f"JP catalog {actual} differs from the master snapshot") + if ctx.master is not None: + manifest = json.loads((ctx.master / MANIFEST).read_bytes()) + if m.get("version") != manifest.get("version"): + g.fail(f"master.version {m.get('version')!r}, {MANIFEST}'s {manifest.get('version')!r}") + listed = {f.get("name"): str(f.get("hash") or "").lower() for f in manifest.get("files") or []} + files = (ctx.snapshot or {}).get("files") or {} + for t, v in sorted(tables.items()): + if (v or {}).get("sha256") != listed.get(f"{t}.bin"): + g.fail(f"master.tables.{t}.sha256 is not {MANIFEST}'s {t}.bin") + decoded = ctx.master / f"{t}.json" + if not decoded.is_file(): + g.fail(f"no decoded table {t}.json") + elif files and sha256(decoded.read_bytes()) != files.get(f"{t}.json"): + g.fail(f"{t}.json read is not the one index.json lists") + deck = p.get("deck") + if doc.get("deck") is not None and not isinstance(deck, dict): + g.fail("provenance.deck is null, but the file has deck statistics") + if isinstance(deck, dict): + if deck.get("name") != "ournotes-deck" or not COMMIT.fullmatch(str(deck.get("commit"))): + g.fail("provenance.deck names no ournotes-deck commit") + elif ctx.deck_commit is not None and deck["commit"] != ctx.deck_commit: + g.fail(f"deck commit {deck['commit'][:12]}, rust/Cargo.lock pins {ctx.deck_commit[:12]}") + ex = p.get("exporter") or {} + if ex.get("name") != "nnnotes": + g.fail(f"exporter {ex.get('name')!r}") + elif ctx.nnnotes_version is not None and ex.get("version") != ctx.nnnotes_version: + g.fail(f"exporter version {ex.get('version')!r}, installed nnnotes {ctx.nnnotes_version!r}") + client = p.get("client") or {} + if not isinstance(client.get("versionName"), str) or not is_int(client.get("versionCode")): + g.fail("provenance.client has no APK version (was the APK read?)") + elif ctx.snapshot is not None and ctx.snapshot["entry"].get("client_version") not in (None, + client["versionName"]): + g.warn(f"the APK is {client['versionName']}, the snapshot names client " + f"{ctx.snapshot['entry']['client_version']}") + if not SHA256.fullmatch(str((p.get("catalog") or {}).get("sha256"))): + g.fail("provenance.catalog has no SHA-256") + g.note = f"master {m.get('version')}, deck {short((deck or {}).get('commit'))}, nnnotes {ex.get('version')}" + + +def baseline(ctx: Context, g: Gate): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return None + try: + return json.loads(ctx.published) + except ValueError: + g.warn("the published music-data.json is not JSON: skipped") + return None + + +def gate_counts(doc, ctx: Context, g: Gate): + old = baseline(ctx, g) + if old is None: + return + songs = {s.get("id") for s in doc.get("songs") or []} + charts = {c.get("scoreId") for _, c in charts_of(doc)} + old_songs = {s.get("id") for s in old.get("songs") or []} + old_charts = {c.get("scoreId") for _, c in charts_of(old)} + if len(songs) < len(old_songs): + g.fail(f"{len(songs)} songs, the published file has {len(old_songs)}") + if len(charts) < len(old_charts): + g.fail(f"{len(charts)} charts, the published file has {len(old_charts)}") + if old_songs - songs: + g.warn(f"songs no longer in the file: {', '.join(map(str, sorted(old_songs - songs)))}") + if old_charts - charts: + g.warn(f"charts no longer in the file: {', '.join(map(str, sorted(old_charts - charts)))}") + g.note = f"songs {len(old_songs)} -> {len(songs)}, charts {len(old_charts)} -> {len(charts)}" + + +def gate_size(doc, ctx: Context, g: Gate, raw: bytes): + if ctx.published is None: + g.note = "skipped: nothing is published yet" + return + ratio = len(raw) / max(1, len(ctx.published)) + if not SIZE_RATIO[0] <= ratio <= SIZE_RATIO[1]: + g.fail(f"{len(raw)} bytes, {ratio:.2f} times the published {len(ctx.published)} (bounds {SIZE_RATIO[0]} to " + f"{SIZE_RATIO[1]})") + g.note = f"{len(ctx.published)} -> {len(raw)} bytes ({ratio:.2f})" + + +def weights_shape(w, kinds: int, positions: int, nullable: bool) -> bool: + return isinstance(w, list) and len(w) == kinds and all( + (nullable and k is None) or (isinstance(k, list) and len(k) == positions and all(map(is_num, k))) for k in w) + + +def gate_deck(doc, ctx: Context, g: Gate): + """The deck statistics. Seeds (deck.model.seeds): the one seed 0 on a chart without a luck range, else two or more + seeds, the same on every luck chart (their number is the file's); every range's rankBonus the rank 1 bonus and + its luckPoints (the range's luck points without skills).""" + deck = doc.get("deck") + if not isinstance(deck, dict): + g.fail("deck is null: no deck statistics (made with --no-deck?)") + return + kinds = len(deck.get("kinds") or []) + if not kinds: + g.fail("deck.kinds is empty") + power = (deck.get("model") or {}).get("power") + if not (is_num(power) and power > 0): + g.fail(f"deck.model.power {power!r}") + n = unplayable = seeds = 0 + luck_seeds = None + for song, chart in charts_of(doc): + n += 1 + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, events = d.get("positions"), d.get("events") or [] + if len(events) != len(chart.get("skillEventsMs") or []): + g.fail(f"{w}: {len(events)} skill events, the chart has {len(chart.get('skillEventsMs') or [])}") + if not is_int(positions) or (events and positions != max(e[0] for e in events) + 1): + g.fail(f"{w}: positions {positions!r} do not match the events") + continue + if d.get("unplayable"): + unplayable += 1 + g.warn(f"{w}: unplayable with Gekisou on ({d['unplayable']})") + if d.get("seeds"): + g.fail(f"{w}: unplayable, but has Gekisou on seeds") + elif not d.get("seeds"): + g.fail(f"{w}: no seeds") + values = [s.get("seed") for s in d.get("seeds") or []] + if values and not luck_chart(d): + if len(values) != 1 or not is_int(values[0]) or values[0] != 0: + g.fail(f"{w}: seeds {values[:4]!r} without a luck range, expected the one seed 0") + elif values: + if not all(map(is_int, values)) or len(values) < 2 or len(set(values)) != len(values): + g.fail(f"{w}: {len(values)} seeds on a luck chart, expected two or more different int seeds") + elif luck_seeds is None: + luck_seeds = values + elif values != luck_seeds: + g.fail(f"{w}: its {len(values)} seeds are not the {len(luck_seeds)} of the first luck chart") + for seed in d.get("seeds") or []: + seeds += 1 + s = f"{w} seed {seed.get('seed')}" + if not is_int(seed.get("score")): + g.fail(f"{s}: score {seed.get('score')!r}") + if not weights_shape(seed.get("weights"), kinds, positions, nullable=False): + g.fail(f"{s}: weights are not [kind][position] numbers") + if len(seed.get("ranges") or []) != len(d.get("ranges") or []): + g.fail(f"{s}: {len(seed.get('ranges') or [])} range results for {len(d.get('ranges') or [])} ranges") + for i, (r, rr) in enumerate(zip(seed.get("ranges") or [], d.get("ranges") or [], strict=False)): + r, rr = (r if isinstance(r, dict) else {}), (rr if isinstance(rr, dict) else {}) + if not (is_int(r.get("rangeScore")) and is_int(r.get("rankBonus"))): + g.fail(f"{s} range {i}: rangeScore {r.get('rangeScore')!r}, rankBonus {r.get('rankBonus')!r}") + elif is_int(rr.get("rankBonusPercent")) and r["rankBonus"] != rank_bonus(r["rangeScore"], + rr["rankBonusPercent"]): + g.fail(f"{s} range {i}: rankBonus {r['rankBonus']} is not trunc({r['rangeScore']} * " + f"{rr['rankBonusPercent']} / 100)") + if not is_int(r.get("luckPoints")): + g.fail(f"{s} range {i}: luckPoints {'missing' if 'luckPoints' not in r else 'not an int'}") + if not within(seed.get("check")): + g.fail(f"{s}: the check deck is not within its bound") + g.note = (f"{n} charts, {kinds} kinds, {seeds} seeds, {unplayable} unplayable, " + f"{len(luck_seeds or [])} seeds per luck chart") + + +def gate_scenarios(doc, ctx: Context, g: Gate): + """The play scenario fields (Gekisou off, every rank, the Perfect play): offSeeds exactly one, every range's + rankBonusPercents five ints, every seed scorePerfect, rangeWeights and rankCheck, every seed range + rangeScorePerfect. A null rangeWeights, a null kind in it or in offSeeds' weights is a warning (none in TW).""" + kinds = len(((doc.get("deck") or {}).get("kinds")) or []) + null_rw = null_kind = null_off = 0 + for song, chart in charts_of(doc): + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + g.fail(f"{w}: no deck statistics") + continue + positions, ranges = d.get("positions"), d.get("ranges") or [] + off = d.get("offSeeds") + if not isinstance(off, list) or len(off) != 1: + g.fail(f"{w}: offSeeds {'missing' if off is None else f'has {len(off)} entries'}, expected exactly one") + else: + o = off[0] + if o.get("seed") != 0 or not is_int(o.get("score")): + g.fail(f"{w}: Gekisou off seed {o.get('seed')!r} score {o.get('score')!r}") + if not weights_shape(o.get("weights"), kinds, positions, nullable=True): + g.fail(f"{w}: Gekisou off weights are not [kind][position] numbers") + else: + null_off += sum(k is None for k in o["weights"]) + if not within(o.get("check")): + g.fail(f"{w}: the Gekisou off check deck is not within its bound") + for i, r in enumerate(ranges): + p = r.get("rankBonusPercents") + if not (isinstance(p, list) and len(p) == RANKS and all(map(is_int, p))): + g.fail(f"{w} range {i}: rankBonusPercents {'missing' if p is None else 'not five ints'}") + elif p[0] != r.get("rankBonusPercent"): + g.fail(f"{w} range {i}: rankBonusPercents[0] {p[0]} is not rankBonusPercent " + f"{r.get('rankBonusPercent')}") + for seed in d.get("seeds") or []: + s = f"{w} seed {seed.get('seed')}" + absent = [k for k in ("scorePerfect", "rangeWeights", "rankCheck") if k not in seed] + if absent: + g.fail(f"{s}: no {', '.join(absent)}") + continue + if not is_int(seed["scorePerfect"]): + g.fail(f"{s}: scorePerfect {seed['scorePerfect']!r}") + for i, r in enumerate(seed.get("ranges") or []): + if not is_int(r.get("rangeScorePerfect")): + why = "missing" if "rangeScorePerfect" not in r else "not an int" + g.fail(f"{s} range {i}: rangeScorePerfect {why}") + rw = seed["rangeWeights"] + if rw is None: + null_rw += 1 + elif not (isinstance(rw, list) and len(rw) == kinds and all( + k is None or (isinstance(k, list) and len(k) == positions and all( + isinstance(x, list) and len(x) == len(ranges) and all(map(is_num, x)) for x in k)) + for k in rw)): + g.fail(f"{s}: rangeWeights are not [kind][position][range] numbers") + else: + null_kind += sum(k is None for k in rw) + rc = seed["rankCheck"] + if rc is not None: + if len(rc.get("ranks") or []) != len(ranges) or not all( + is_int(x) and 1 <= x <= RANKS for x in rc.get("ranks") or []): + g.fail(f"{s}: rankCheck ranks are not one rank per range") + elif not within(rc): + g.fail(f"{s}: the rank check deck is not within its bound") + if null_rw: + g.warn(f"{null_rw} seeds have rangeWeights null (overlapping ranges: no rank scenarios on those charts)") + if null_kind: + g.warn(f"{null_kind} rangeWeights kinds are null (conditions on the confirmed rank)") + if null_off: + g.warn(f"{null_off} Gekisou off weight kinds are null (conditions on the Gekisou state)") + g.note = "offSeeds, rankBonusPercents, scorePerfect, rangeWeights, rankCheck, rangeScorePerfect" + + +APTITUDE_KEYS = ("plainKind", "host", "seedRule", "shapes") +SEED_RULE_KEYS = ("deterministicTest", "batches", "relative", "baseline", "crossSeeds") +SHAPE_KEYS = ("id", "source", "mission", "bandCondition", "effects", "skills") +EFFECT_INTS = ("effectType", "triggerType", "effectValue", "maxEffectValue", "effectLimitCount", + "effectExecuteLimitCount") +EFFECT_GROUPS = ("trigger", "condition", "release", "reset") +FACTOR_INTS = ("judgedNotes", "justNotes", "perfectNotes", "tailNotes", "comboAtStart") +VARIANT_KEYS = ("shape", "bandMatch", "deterministic", "seeds", "seTargetMet", "crossSeeds", "score", "scorePerfect", + "tail", "tailPerfect", "converted", "ranges", "weights", "rangeWeights", "check") +VARIANT_PAIRS = ("score", "scorePerfect", "tail", "tailPerfect", "converted") +VARIANT_RANGE_PAIRS = ("rangeScore", "rankBonus", "rangeScorePerfect", "maxCombo", "justCount", "luckPoints") +APTITUDE_CHECK_KEYS = ("seed", "ranks", "deck", "exact", "predicted", "bound") + + +def pair(v) -> bool: + """A [mean, standard error]: two finite numbers, the error not negative.""" + return isinstance(v, list) and len(v) == 2 and is_num(v[0]) and is_num(v[1]) and v[1] >= 0 + + +def close(a: float, b: float) -> bool: + return abs(a - b) <= APTITUDE_SLACK * max(1.0, abs(a), abs(b)) + + +def gate_aptitude(doc, ctx: Context, g: Gate): + """The charts' Gekisou skill aptitude: deck.gekisouAptitude (its plain kind the page's, the seed rule, shapes + numbered from 0: source, mission, band condition, effect rows, skills); every chart's deck.gekisouAptitude, null + exactly when the chart is unplayable with Gekisou on, has no Gekisou range or there is no shape; else factors per + range and one variant per shape of the chart's missions (or mission 4) in shape order, a band condition shape's + bandMatch true then false: every [mean, se] two finite numbers with se >= 0 (0 when deterministic), ranges per + range, tail = score - the ranges' rangeScore and rankBonus, the seeds and the cross seeds by the seed rule, weights + and rangeWeights where the plain kind and the chart's rank weights are, the check within its bound. Warnings: + variants whose standard error missed the seed rule's target.""" + deck = doc.get("deck") + if not isinstance(deck, dict): + g.fail("deck is null: no Gekisou skill aptitude") + return + plain = plain_kind(doc) + if not (isinstance((deck.get("model") or {}).get("gekisouAptitude"), str) and deck["model"]["gekisouAptitude"]): + g.fail("deck.model.gekisouAptitude: no text") + head = deck.get("gekisouAptitude") + if not isinstance(head, dict): + g.fail("deck.gekisouAptitude missing" if head is None else f"deck.gekisouAptitude {head!r}") + head = {} + elif [k for k in APTITUDE_KEYS if k not in head]: + g.fail(f"deck.gekisouAptitude: no {', '.join(k for k in APTITUDE_KEYS if k not in head)}") + if head: + pk = head.get("plainKind", "missing") + if not (pk is None or is_int(pk)) or pk != plain: + g.fail(f"deck.gekisouAptitude.plainKind {pk!r}, the page's plain kind is {plain!r}") + if not (isinstance(head.get("host"), str) and head["host"].strip()): + g.fail("deck.gekisouAptitude.host: no text") + rule = head.get("seedRule") if isinstance(head.get("seedRule"), dict) else {} + batches = rule.get("batches") + if head and not (all(k in rule for k in SEED_RULE_KEYS) and is_int(rule["deterministicTest"]) + and rule["deterministicTest"] >= 1 and isinstance(batches, list) and batches + and all(map(is_int, batches)) and batches == sorted(set(batches)) and batches[0] >= 1 + and is_num(rule["relative"]) and rule["relative"] >= 0 and is_num(rule["baseline"]) + and rule["baseline"] >= 0 and is_int(rule["crossSeeds"]) and rule["crossSeeds"] >= 1): + g.fail(f"deck.gekisouAptitude.seedRule {rule!r}"[:200]) + rule, batches = {}, None + + # the shapes + shapes = head.get("shapes") if isinstance(head.get("shapes"), list) else [] + if head and not isinstance(head.get("shapes"), list): + g.fail("deck.gekisouAptitude.shapes is not a list") + if any(not isinstance(s, dict) or not is_int(s.get("id")) or s["id"] != i + for i, s in enumerate(shapes)): + g.fail("deck.gekisouAptitude.shapes: ids are not 0, 1, 2, ... in order") + by_id: dict = {} + for s in shapes: + if not isinstance(s, dict): + continue + w = f"shape {s.get('id')}" + absent = [k for k in SHAPE_KEYS if k not in s] + if absent: + g.fail(f"{w}: no {', '.join(absent)}") + continue + if is_int(s["id"]): + by_id[s["id"]] = s + if s["source"] not in ("member", "support"): + g.fail(f"{w}: source {s['source']!r}") + if not is_int(s["mission"]) or s["mission"] not in MISSIONS: + g.fail(f"{w}: mission {s['mission']!r}") + band = s["bandCondition"] + if not isinstance(band, bool) or (band and s["source"] != "support"): + g.fail(f"{w}: bandCondition {band!r} (a support skill's alone)") + effects = s["effects"] + fives = 0 + if not isinstance(effects, list): + g.fail(f"{w}: effects are not a list") + effects = [] + for i, e in enumerate(effects): + if not isinstance(e, dict): + g.fail(f"{w} effect {i}: not an object") + continue + bad = [k for k in EFFECT_INTS if not is_int(e.get(k))] + if not is_num(e.get("activationTimeSecond")): + bad.append("activationTimeSecond") + if not (isinstance(e.get("skillTargetIds"), list) and all(map(is_int, e["skillTargetIds"]))): + bad.append("skillTargetIds") + for k in EFFECT_GROUPS: + sets = e.get(k) + if not (isinstance(sets, list) and all(isinstance(x, list) for x in sets)): + bad.append(k) + continue + for c in (c for x in sets for c in x): + if not (isinstance(c, dict) and is_int(c.get("type")) and isinstance(c.get("values"), list) + and isinstance(c.get("positive"), bool) and "targetIds" in c): + bad.append(k) + elif c["type"] == BAND_CONDITION: + fives += 1 + if c["targetIds"] is not None: + bad.append(f"{k} (condition {BAND_CONDITION} targetIds not null)") + elif not (isinstance(c["targetIds"], list) and all(map(is_int, c["targetIds"]))): + bad.append(k) + cu = e.get("cumulative", "missing") + if cu is not None and not (isinstance(cu, dict) and all( + k in cu for k in ("type", "values", "targetIds", "maxCumulativeCount"))): + bad.append("cumulative") + if bad: + g.fail(f"{w} effect {i}: {', '.join(dict.fromkeys(bad))} missing or malformed") + if isinstance(band, bool) and band != (fives > 0): + g.fail(f"{w}: bandCondition {band}, its effects have {fives} condition {BAND_CONDITION}") + skills = s["skills"] + if not (isinstance(skills, list) and skills): + g.fail(f"{w}: no skills") + continue + for k in skills: + ok = (isinstance(k, dict) and is_int(k.get("id")) and is_int(k.get("level")) and k["level"] >= 1 + and "memberTargetIds" in k and "bandIds" in k) + targets, band_ids = (k.get("memberTargetIds"), k.get("bandIds")) if isinstance(k, dict) else (0, 0) + if band is True: + ok = (ok and isinstance(targets, list) and bool(targets) and all(map(is_int, targets)) + and targets == sorted(set(targets)) and isinstance(band_ids, list) + and all(map(is_int, band_ids))) + else: + ok = ok and targets is None and band_ids is None + if not ok: + g.fail(f"{w}: skill {k!r} is not an id, a level and (with a band condition alone) member targets " + f"and bands"[:240]) + + # the charts + aptitudes = nulls = variants = deterministic = 0 + missed = [] + for song, chart in charts_of(doc): + w, d = where(song, chart), chart.get("deck") + if not isinstance(d, dict): + continue # the deck gate fails it + if "gekisouAptitude" not in d: + g.fail(f"{w}: no deck.gekisouAptitude") + continue + a, ranges, dseeds = d["gekisouAptitude"], d.get("ranges") or [], d.get("seeds") or [] + why = ("unplayable with Gekisou on" if d.get("unplayable") else "without a Gekisou range" if not ranges + else None if shapes else "without a Gekisou skill shape") + if a is None: + nulls += 1 + if why is None: + g.fail(f"{w}: deck.gekisouAptitude is null, but the chart is playable with Gekisou on") + continue + if why is not None: + g.fail(f"{w}: a Gekisou skill aptitude on a chart {why}") + continue + if not (isinstance(a, dict) and isinstance(a.get("factors"), list) and isinstance(a.get("variants"), list)): + g.fail(f"{w}: deck.gekisouAptitude has no factors and variants lists") + continue + aptitudes += 1 + positions = d.get("positions") + missions = {r.get("mission") for r in ranges if isinstance(r, dict)} + linear = bool(dseeds) and dseeds[0].get("rangeWeights") is not None + + # factors + if len(a["factors"]) != len(ranges): + g.fail(f"{w}: {len(a['factors'])} factors for {len(ranges)} ranges") + for j, (f, r) in enumerate(zip(a["factors"], ranges, strict=False)): + f, r = (f if isinstance(f, dict) else {}), (r if isinstance(r, dict) else {}) + bad = [k for k in FACTOR_INTS if not (is_int(f.get(k)) and f[k] >= 0)] + if not pair(f.get("lotteries")): + bad.append("lotteries") + if bad: + g.fail(f"{w} factors {j}: {', '.join(bad)} missing or not counts") + continue + if r.get("mission") != JUST_MISSION and (f["justNotes"] or f["perfectNotes"]): + g.fail(f"{w} factors {j}: Just or Perfect notes in a range without the Just mission") + if f["justNotes"] + f["perfectNotes"] > f["judgedNotes"]: + g.fail(f"{w} factors {j}: {f['justNotes']} Just and {f['perfectNotes']} Perfect notes of " + f"{f['judgedNotes']} judged") + at = [s["ranges"][j] for s in dseeds if isinstance(s.get("ranges"), list) and len(s["ranges"]) > j + and isinstance(s["ranges"][j], dict)] + lots = [sum(x["lotResults"]) for x in at + if isinstance(x.get("lotResults"), list) and all(map(is_int, x["lotResults"]))] + if r.get("mission") != LUCK_MISSION and f["lotteries"] != [0, 0]: + g.fail(f"{w} factors {j}: lotteries {f['lotteries']} in a range without the luck mission") + elif lots and len(lots) == len(dseeds) and not close(f["lotteries"][0], sum(lots) / len(lots)): + g.fail(f"{w} factors {j}: lotteries {f['lotteries'][0]}, deck.seeds' lotResults give " + f"{sum(lots) / len(lots)}") + + # the variants: which, in order + want = [(s["id"], match) for s in sorted(by_id.values(), key=lambda s: s["id"]) + if s["mission"] == 4 or s["mission"] in missions + for match in ((True, False) if s["bandCondition"] is True else (None,))] + got = [(v.get("shape"), v.get("bandMatch")) if isinstance(v, dict) else None for v in a["variants"]] + unknown = [v[0] for v in got if v is not None and not (is_int(v[0]) and v[0] in by_id)] + other = [v[0] for v in got if v is not None and is_int(v[0]) and v[0] in by_id + and by_id[v[0]]["mission"] != 4 and by_id[v[0]]["mission"] not in missions] + if unknown: + g.fail(f"{w}: variants of shapes {sorted(set(map(repr, unknown)))} not in deck.gekisouAptitude.shapes") + if other: + g.fail(f"{w}: variants of shapes {sorted(set(other))} of a mission the chart does not play") + if got != want and not unknown and not other: + lacking = [x for x in want if x not in got] + g.fail(f"{w}: variants {'lack ' + repr(lacking[:6]) if lacking else 'not in shape order, true first'}" + f" ({len(got)} for {len(want)})") + + for v in a["variants"]: + if not isinstance(v, dict): + g.fail(f"{w}: a variant {v!r}") + continue + s = f"{w} shape {v.get('shape')}" + ("" if v.get("bandMatch") is None else f" {v['bandMatch']}") + absent = [k for k in VARIANT_KEYS if k not in v] + if absent: + g.fail(f"{s}: no {', '.join(absent)}") + continue + variants += 1 + shape = by_id.get(v["shape"]) if is_int(v["shape"]) else None + det = v["deterministic"] + if not isinstance(det, bool): + g.fail(f"{s}: deterministic {det!r}") + det = False + deterministic += det + if v["bandMatch"] is not None and not isinstance(v["bandMatch"], bool): + g.fail(f"{s}: bandMatch {v['bandMatch']!r}") + elif shape is not None and (v["bandMatch"] is None) == (shape.get("bandCondition") is True): + g.fail(f"{s}: bandMatch {v['bandMatch']!r} for a shape " + f"{'with' if v['bandMatch'] is None else 'without'} a band condition") + n, cross, met = v["seeds"], v["crossSeeds"], v["seTargetMet"] + if not (is_int(n) and n >= 1 and isinstance(met, bool)): + g.fail(f"{s}: seeds {n!r}, seTargetMet {met!r}") + elif det and (n != 1 or not met): + g.fail(f"{s}: deterministic, but {n} seeds and seTargetMet {met}") + elif not det and batches and (n not in batches or (not met and n != batches[-1])): + g.fail(f"{s}: {n} seeds (seTargetMet {met}), not a batch of the seed rule {batches}") + elif not det and not met: + missed.append(f"{chart.get('scoreId')} shape {v['shape']}") + if (not is_int(cross) or cross < 1 or (is_int(n) and is_int(rule.get("crossSeeds")) + and cross != min(n, rule["crossSeeds"]))): + g.fail(f"{s}: crossSeeds {cross!r}, expected min({n}, {rule['crossSeeds']})") + + # every [mean, se] + values = {k: v[k] for k in VARIANT_PAIRS} + rs = v["ranges"] + if not isinstance(rs, list) or len(rs) != len(ranges): + g.fail(f"{s}: {len(rs) if isinstance(rs, list) else repr(rs)} range results for {len(ranges)} ranges") + rs = [] + for j, r in enumerate(rs): + for k in VARIANT_RANGE_PAIRS: + values[f"ranges[{j}].{k}"] = r.get(k) if isinstance(r, dict) else None + wt, rw = v["weights"], v["rangeWeights"] + if plain is None: + if wt is not None or rw is not None: + g.fail(f"{s}: weights or rangeWeights without a plain kind") + else: + if not (isinstance(wt, list) and len(wt) == positions): + g.fail(f"{s}: weights are not one [mean, se] per position") + else: + values.update({f"weights[{k}]": x for k, x in enumerate(wt)}) + if not linear: + if rw is not None: + g.fail(f"{s}: rangeWeights, but deck.seeds[0].rangeWeights is null") + elif not (isinstance(rw, list) and len(rw) == positions and all( + isinstance(x, list) and len(x) == len(ranges) for x in rw)): + g.fail(f"{s}: rangeWeights are not [position][range] [mean, se]") + else: + values.update({f"rangeWeights[{k}][{j}]": y for k, x in enumerate(rw) for j, y in enumerate(x)}) + bad = [k for k, x in values.items() if not pair(x)] + if bad: + g.fail(f"{s}: {', '.join(bad[:6])}{' ...' if len(bad) > 6 else ''} not [mean, se] (finite, se >= 0)") + elif det and any(x[1] for x in values.values()): + g.fail(f"{s}: deterministic, but a standard error is not 0: " + f"{', '.join(k for k, x in values.items() if x[1])[:160]}") + elif rs and all(pair(r.get(k)) for r in rs for k in ("rangeScore", "rankBonus")): + inside = sum(r["rangeScore"][0] + r["rankBonus"][0] for r in rs) + if abs(v["tail"][0] - (v["score"][0] - inside)) > ( + 0.0005 * (2 + 2 * len(rs)) + APTITUDE_SLACK): + g.fail(f"{s}: tail {v['tail'][0]} is not score {v['score'][0]} less the ranges' rangeScore and " + f"rankBonus {inside}") + + # Perfect rank-bonus deltas are not exported. Only a necessary rounding bound can be checked + # for stochastic means: each difference of two truncated bonuses is within 2 points of delta*p/100. + if not bad and len(rs) == len(ranges): + perfect_terms = [r["rangeScorePerfect"][0] * (1 + info["rankBonusPercent"] / 100) + for r, info in zip(rs, ranges, strict=True)] + rounding = 0.001 + sum(0.0005 * abs(1 + info["rankBonusPercent"] / 100) for info in ranges) + residual = v["scorePerfect"][0] - v["tailPerfect"][0] - sum(perfect_terms) + if abs(residual) > 2 * len(ranges) + rounding + APTITUDE_SLACK: + g.fail(f"{s}: tailPerfect violates the necessary Perfect rank-bonus rounding bound") + + # Deterministic point deltas are integers. Perfect rank bonuses can then be recovered exactly + # from the first baseline seed (stochastic Perfect rank-bonus deltas are not exported). + if det and not bad and rs and dseeds: + points = [x[0] for k, x in values.items() if not k.startswith(("weights", "rangeWeights"))] + if not all(float(x).is_integer() for x in points): + g.fail(f"{s}: deterministic, but a point delta is not an integer") + else: + perfect = 0 + for r, info, base in zip(rs, ranges, dseeds[0]["ranges"], strict=True): + delta, baseline = int(r["rangeScorePerfect"][0]), base["rangeScorePerfect"] + percent = info["rankBonusPercent"] + perfect += delta + rank_bonus(baseline + delta, percent) - rank_bonus(baseline, percent) + if v["tailPerfect"][0] != v["scorePerfect"][0] - perfect: + g.fail(f"{s}: tailPerfect is not scorePerfect less the Perfect range scores and bonuses") + + # the check + c = v["check"] + if not (isinstance(c, dict) and all(k in c for k in APTITUDE_CHECK_KEYS)): + g.fail(f"{s}: the check has no {', '.join(APTITUDE_CHECK_KEYS)}") + continue + if dseeds and c["seed"] != dseeds[0].get("seed"): + g.fail(f"{s}: the check's seed {c['seed']!r} is not deck.seeds[0]'s {dseeds[0].get('seed')!r}") + ranks = c["ranks"] + if not (isinstance(ranks, list) and len(ranks) == len(ranges) + and all(is_int(x) and 1 <= x <= RANKS for x in ranks)): + g.fail(f"{s}: the check's ranks {ranks!r} are not one rank per range") + elif not linear and any(x != 1 for x in ranks): + g.fail(f"{s}: the check's ranks {ranks} on a chart whose ranks are not linear (rank 1 alone)") + cd = c["deck"] + if not (isinstance(cd, list) and len(cd) == positions and all( + x is None or (plain is not None and isinstance(x, list) and len(x) == 2 and x[0] == plain + and is_num(x[1])) for x in cd)): + g.fail(f"{s}: the check deck is not a [plain kind, value] or null per position") + if not within(c): + g.fail(f"{s}: the check is not within its bound") + if missed: + g.warn(f"{len(missed)} variants missed the seed rule's standard error target: {', '.join(missed[:10])}" + + (" ..." if len(missed) > 10 else "")) + g.note = (f"{len(shapes)} shapes; {aptitudes} charts with an aptitude, {nulls} null; {variants} variants, " + f"{deterministic} deterministic; plain kind {plain}") + + +def gate_gzip(doc, ctx: Context, g: Gate, raw: bytes): + """The file's gzip size (what the page downloads) and that of its Gekisou skill aptitude, against loose caps.""" + whole = len(gzip.compress(raw, 6)) + deck = doc.get("deck") if isinstance(doc.get("deck"), dict) else {} + part = {"deck": deck.get("gekisouAptitude"), + "charts": [(c.get("deck") or {}).get("gekisouAptitude") if isinstance(c.get("deck"), dict) else None + for _, c in charts_of(doc)]} + aptitude = len(gzip.compress(json.dumps(part, ensure_ascii=False, separators=(",", ":")).encode("utf-8"), 6)) + if whole > FILE_GZIP_MAX: + g.fail(f"{whole} bytes gzipped, more than {FILE_GZIP_MAX}") + if aptitude > APTITUDE_GZIP_MAX: + g.fail(f"the Gekisou skill aptitude: {aptitude} bytes gzipped, more than {APTITUDE_GZIP_MAX}") + g.note = f"{whole} bytes gzipped, the aptitude {aptitude}" + + +def nonfinite(v, path: str, out: list): + if isinstance(v, float): + if not math.isfinite(v): + out.append(path) + elif isinstance(v, list): + for i, x in enumerate(v): + nonfinite(x, f"{path}[{i}]", out) + elif isinstance(v, dict): + for k, x in v.items(): + nonfinite(x, f"{path}.{k}" if path else k, out) + + +def gate_finite(doc, ctx: Context, g: Gate): + """No NaN or infinity; one inside master data rows as served (songs[].master, master: `1e999` is how the format + writes a binary32 infinity) is reported, not failed.""" + found: list[str] = [] + nonfinite(doc, "", found) + for path in found: + if re.match(r"(songs\[\d+\]\.master|master)\b", path): + g.warn(f"{path}: not finite (master data as served)") + else: + g.fail(f"{path}: not finite") + g.note = f"{len(found)} non-finite numbers" + + +def gate_references(doc, ctx: Context, g: Gate): + langs = doc.get("languages") or [] + if not langs: + g.fail("no languages") + + def text(t, what: str, required: bool) -> bool: + if t is None: + if required: + g.fail(f"{what}: no text") + return False + if not isinstance(t, dict) or set(t) != set(langs) or not all(isinstance(v, str) for v in t.values()): + g.fail(f"{what}: not a text in {', '.join(langs)}") + return False + if required and not any(t.values()): + g.fail(f"{what}: empty in every language") + return False + return True + + def ids(rows, what: str) -> set: + seen = [r.get("id") for r in rows] + if len(set(seen)) != len(seen) or not all(map(is_int, seen)): + g.fail(f"{what}: ids are not unique ints") + return set(seen) + + bands = ids(doc.get("bands") or [], "bands") + characters = ids(doc.get("characters") or [], "characters") + tags = ids(doc.get("tags") or [], "tags") + ids(doc.get("categories") or [], "categories") + for b in doc.get("bands") or []: + text(b.get("name"), f"band {b.get('id')} name", True) + for c in doc.get("characters") or []: + text(c.get("name"), f"character {c.get('id')} name", True) + text(c.get("shortName"), f"character {c.get('id')} short name", False) + if c.get("bandId") not in bands: + g.warn(f"character {c.get('id')}: band {c.get('bandId')} is not in bands") + for t in doc.get("tags") or []: + text(t.get("name"), f"tag {t.get('id')} name", True) + categories = set() + for c in doc.get("categories") or []: + text(c.get("name"), f"category {c.get('id')} name", True) + categories.update(c.get("musicCategories") or []) + songs = doc.get("songs") or [] + if not songs: + g.fail("no songs") + ids(songs, "songs") + if [s.get("id") for s in songs] != sorted(s.get("id") for s in songs if is_int(s.get("id"))): + g.fail("songs are not sorted by id") + score_ids: set = set() + untitled = [] + for s in songs: + sid = f"song {s.get('id')}" + if text(s.get("title"), f"{sid} title", True) and ctx.language and not s["title"].get(ctx.language): + untitled.append(s.get("id")) + for k in ("ruby", "phonetic", "bandName", "lyricist", "composer", "arranger"): + text(s.get(k), f"{sid} {k}", False) + for key, known, what in (("bandIds", bands, "band"), ("vocalCharacterIds", characters, "character"), + ("bestMusicTagIds", tags, "tag")): + unknown = [x for x in s.get(key) or [] if x not in known] + if unknown: + g.fail(f"{sid}: {what} {', '.join(map(str, unknown))} not in the file's {what}s") + if not s.get("bandIds") and s.get("bandName") is None: + g.fail(f"{sid}: neither a band nor a band name") + other = [x for x in s.get("musicCategories") or [] if x not in categories] + if other: + g.warn(f"{sid}: music categories {', '.join(map(str, other))} are on no category tab") + jacket = s.get("jacket") + if not isinstance(jacket, str) or not jacket: + g.fail(f"{sid}: no jacket") + elif ctx.jackets is not None: + f = ctx.jackets / f"{jacket}.webp" + if not f.is_file() or not f.stat().st_size: + g.fail(f"{sid}: no jacket file jackets/{jacket}.webp") + bgm = s.get("bgm") or {} + if not (is_int(bgm.get("soundId")) and bgm.get("cueSheet") and bgm.get("cue")): + g.fail(f"{sid}: no BGM cue") + ranks = [r.get("rank") for r in s.get("scoreRanks") or []] + if not ranks or len(set(ranks)) != len(ranks): + g.fail(f"{sid}: score ranks {ranks!r}") + charts = s.get("charts") or [] + order = [c.get("difficulty") for c in charts] + if not charts or order != [d for d in DIFFICULTIES if d in order] or len(set(order)) != len(order): + g.fail(f"{sid}: charts {order!r}") + for c in charts: + if c.get("scoreId") in score_ids: + g.fail(f"{where(s, c)}: the score id occurs twice") + score_ids.add(c.get("scoreId")) + if untitled: + g.warn(f"songs without a {ctx.language} title: {', '.join(map(str, untitled))}") + g.note = (f"{len(songs)} songs, {len(bands)} bands, {len(characters)} characters, {len(tags)} tags" + + ("" if ctx.jackets is None else ", jackets present")) + + +def gate_bgm(doc, ctx: Context, g: Gate): + lengths = [] + for s in doc.get("songs") or []: + sid = f"song {s.get('id')}" + L = (s.get("bgm") or {}).get("length") + if not isinstance(L, dict): + g.fail(f"{sid}: no BGM length") + continue + dur, samples, rate = L.get("durationMs"), L.get("samples"), L.get("sampleRate") + if not (is_int(dur) and is_int(samples) and is_int(rate) and rate > 0 and samples > 0): + g.fail(f"{sid}: BGM length {dur!r} ms, {samples!r} samples at {rate!r} Hz") + continue + if dur != samples * 1000 // rate: + g.fail(f"{sid}: BGM durationMs {dur} is not samples * 1000 // sampleRate") + if not BGM_MS[0] <= dur <= BGM_MS[1]: + g.fail(f"{sid}: BGM of {dur} ms (bounds {BGM_MS[0]} to {BGM_MS[1]})") + if is_int(L.get("lengthMs")) and abs(dur - L["lengthMs"]) > BGM_CUE_SLACK_MS: + g.fail(f"{sid}: BGM stream {dur} ms, cue length {L['lengthMs']} ms") + last = max((c.get("lastNoteMs") or 0 for c in s.get("charts") or []), default=0) + if dur < last: + g.fail(f"{sid}: BGM of {dur} ms ends before the last note at {last} ms") + elif dur - last > BGM_TAIL_MS: + g.warn(f"{sid}: BGM plays {dur - last} ms after the last note") + lengths.append(dur) + if lengths: + g.note = f"{min(lengths) / 1000:.1f} s to {max(lengths) / 1000:.1f} s" + + +def gate_page(doc, ctx: Context, g: Gate): + """The chart data page's own modules (catalog.js, ranking.js) over the file in Node.js: music_data_smoke.mjs.""" + if ctx.page is None: + g.note = "skipped" + return + if not (ctx.page / "catalog.js").is_file() or not (ctx.page / "ranking.js").is_file(): + g.fail(f"{ctx.page}: no catalog.js / ranking.js (the chart data page, examples/songs)") + return + r = subprocess.run(["node", str(SMOKE), str(ctx.page), str(ctx.file)], capture_output=True, text=True, + timeout=600) + lines = [x for x in (r.stdout + r.stderr).splitlines() if x.strip()] + if r.returncode: + for x in lines or [f"node exited with {r.returncode}"]: + g.fail(x[:300]) + else: + g.note = lines[-1][:200] if lines else "passed" + + +GATES = (("schema", gate_schema), ("provenance", gate_provenance), ("counts", gate_counts), ("deck", gate_deck), + ("scenarios", gate_scenarios), ("aptitude", gate_aptitude), ("finite", gate_finite), + ("references", gate_references), ("bgm", gate_bgm), ("size", gate_size), ("gzip", gate_gzip), + ("page", gate_page)) + + +def gates(raw: bytes, ctx: Context, only=None) -> dict: + """The report of the gates (`only`: their names) on the file's bytes.""" + try: + doc = json.loads(raw.decode("utf-8")) + if not isinstance(doc, dict): + raise ValueError("not an object") + except ValueError as e: + doc, results = None, [{"gate": "json", "passed": False, "failures": [f"not JSON: {str(e)[:160]}"], + "failureCount": 1, "warnings": [], "warningCount": 0, "note": ""}] + if doc is not None: + results = [] + for name, fn in GATES: + if only is not None and name not in only: + continue + g = Gate() + try: + fn(doc, ctx, g, raw) if name in ("size", "gzip") else fn(doc, ctx, g) + except Exception as e: # a malformed file the gate did not foresee fails the gate + g.fail(f"the gate stopped: {type(e).__name__}: {str(e)[:160]}") + results.append({"gate": name, "passed": not g.failures, "failures": g.failures[:LISTED], + "failureCount": len(g.failures), "warnings": g.warnings[:LISTED], + "warningCount": len(g.warnings), "note": g.note}) + return {"passed": all(r["passed"] for r in results), "sha256": sha256(raw), "bytes": len(raw), + "gates": results} + + +def report_markdown(report: dict) -> str: + rows = ["| Gate | Result | |", "|---|---|---|"] + for r in report["gates"]: + result = "passed" if r["passed"] else f"**failed** ({r['failureCount']})" + if r["warningCount"]: + result += f", {r['warningCount']} warnings" + rows.append(f"| {r['gate']} | {result} | {r['note']} |") + details = [] + for r in report["gates"]: + for kind, key, n in (("failure", "failures", "failureCount"), ("warning", "warnings", "warningCount")): + for x in r[key]: + details.append(f"- {r['gate']} {kind}: {x}") + if r[n] > len(r[key]): + details.append(f"- {r['gate']}: {r[n] - len(r[key])} more {kind}s") + return "\n".join(["", "#### Gates", ""] + rows + ([""] + details if details else [])) + + +def cmd_check(out: str, master: str, page: str) -> None: + import nnnotes + o, m = Path(out), Path(master) + raw = (o / FILE).read_bytes() + snapshot = json.loads(snapshot_file(m).read_text(encoding="utf-8")) + ctx = Context(region=env("NNNOTES_CATALOG_REGION"), language=os.environ.get("NNNOTES_CATALOG_LANGUAGE"), + master=m, snapshot=snapshot, deck_commit=deck_commit(), nnnotes_version=nnnotes.__version__, + jackets=o / "jackets", schema=SCHEMA, published=published(FILE), page=Path(page), file=o / FILE) + report = gates(raw, ctx) + (o / "check.json").write_text(json.dumps(report, indent=1), encoding="utf-8") + summary(report_markdown(report)) + if not report["passed"]: + fail("a gate failed: nothing is published (the job summary lists why)") + doc = json.loads(raw) + p = doc["provenance"] + version = re.sub(r"[^A-Za-z0-9._-]", "_", str(p["master"]["version"])) + marker = { + "format": BUILD_FORMAT, + "file": FILE, "sha256": report["sha256"], "bytes": report["bytes"], + "archive": f"{ARCHIVE}{version}/{report['sha256']}.json", + "songs": len(doc["songs"]), "charts": sum(len(s["charts"]) for s in doc["songs"]), + "jackets": len(list((o / "jackets").glob("*.webp"))), + "inputs": inputs(snapshot["entry"]), + "provenance": {"region": p["region"], "client": p["client"], "masterVersion": p["master"]["version"], + "deckCommit": p["deck"]["commit"], "exporterVersion": p["exporter"]["version"]}, + "player": {"repository": os.environ.get("PLAYER_REPOSITORY"), "ref": os.environ.get("PLAYER_REF")}, + "gates": {r["gate"]: {"passed": r["passed"], "warnings": r["warningCount"]} for r in report["gates"]}, + "builtAt": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"), + "run": (f"{os.environ['GITHUB_SERVER_URL']}/{os.environ['GITHUB_REPOSITORY']}/actions/runs/" + f"{os.environ['GITHUB_RUN_ID']}" if os.environ.get("GITHUB_RUN_ID") else None), + } + (o / MARKER).write_text(json.dumps(marker, indent=1) + "\n", encoding="utf-8") + + +# ---------------------------------------------------------------- publish +def upload(b, key: str, src: Path, cache: str) -> None: + extra = {"ContentType": story_site.TYPES.get(src.suffix.lower(), "application/octet-stream"), + "CacheControl": cache} + b.s3.upload_file(str(src), b.name, b.prefix + key, ExtraArgs=extra) + + +def read_back(b, key: str, digest: str) -> None: + """The object as stored (S3 GET), else as served (plain HTTP), must have the SHA-256 uploaded.""" + for attempt in range(4): + try: + if sha256(b.s3.get_object(Bucket=b.name, Key=b.prefix + key)["Body"].read()) == digest: + return + except Exception as e: # Cloudflare in front of the store rejects a signed GET at times + print(f"read back {key}: {type(e).__name__}", flush=True) + time.sleep(3 * (attempt + 1)) + try: + if sha256(get(public_url(key), timeout=300)) == digest: + return + except urllib.error.URLError as e: + print(f"read back {key} over HTTP: {e}", flush=True) + fail(f"{key}: the bucket does not serve what was uploaded (SHA-256 {digest[:12]})") + + +def publishing() -> bool: + """The publishing switch: uploads only with MUSIC_DATA_PUBLISH `true` (a repository variable, unset: off).""" + return os.environ.get("MUSIC_DATA_PUBLISH") == "true" + + +def cmd_publish(out: str, dry_run: bool = False) -> None: + if not publishing() and not dry_run: + summary("- publishing is off (the repository variable MUSIC_DATA_PUBLISH is not `true`): a dry run") + dry_run = True + o = Path(out) + raw = (o / FILE).read_bytes() + report = json.loads((o / "check.json").read_text(encoding="utf-8")) + if not report.get("passed") or report.get("sha256") != sha256(raw) or not (o / MARKER).is_file(): + fail("the file has not passed the gates (check.json): nothing is published") + marker = json.loads((o / MARKER).read_text(encoding="utf-8")) + b = bucket() + if not b.writable and not dry_run: + fail("publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") + force = os.environ.get("FORCE") == "true" + try: + have = b.keys(JACKETS) + archived = b.keys(marker["archive"]).get(marker["archive"]) == len(raw) + except Exception as e: # a dry run without a key where the bucket lists to none + if not dry_run: + raise + print(f"cannot list the bucket ({type(e).__name__}): the dry run lists every object", flush=True) + have, archived = {}, False + jackets = sorted((o / "jackets").glob("*.webp")) + new = [j for j in jackets if force or have.get(JACKETS + j.name) != j.stat().st_size] + steps = [(JACKETS + j.name, j, JACKET_CACHE, False) for j in new] + if not archived: + steps.append((marker["archive"], o / FILE, ARCHIVE_CACHE, True)) + # the file before the marker that names it, the jackets and the archive copy before the file + steps += [(FILE, o / FILE, FILE_CACHE, True), (MARKER, o / MARKER, FILE_CACHE, True)] + if dry_run or not publishing(): + for key, _, _, _ in steps: + print(f"would upload {b.prefix}{key}") + else: + story_site.parallel(lambda s: upload(b, s[0], s[1], s[2]), [s for s in steps if not s[3]]) + for key, src, cache, check in steps: + if check: + upload(b, key, src, cache) + read_back(b, key, sha256(src.read_bytes())) + summary(f"- {'would publish' if dry_run else 'published'} {len(new)} jackets (of {len(jackets)}), " + f"{'the archive copy, ' if not archived else ''}{FILE} ({len(raw)} bytes, sha256 " + f"{short(marker['sha256'])}), {MARKER}: {public_url(FILE)}") + + +def main(argv: list[str]) -> None: + if not argv: + sys.exit(__doc__) + cmd, args = argv[0], argv[1:] + if cmd == "plan" and not args: + cmd_plan() + elif cmd == "master" and len(args) == 1: + cmd_master(args[0]) + elif cmd == "build" and len(args) == 1: + cmd_build(args[0]) + elif cmd == "check" and len(args) == 3: + cmd_check(*args) + elif cmd == "publish" and len(args) in (1, 2) and args[1:] in ([], ["--dry-run"]): + cmd_publish(args[0], dry_run=bool(args[1:])) + else: + sys.exit(__doc__) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/.github/scripts/music_data_smoke.mjs b/.github/scripts/music_data_smoke.mjs new file mode 100644 index 0000000..635b456 --- /dev/null +++ b/.github/scripts/music_data_smoke.mjs @@ -0,0 +1,175 @@ +// The chart data page's smoke test (music_data.py check, gate "page"): the page's own pure modules (ournotes-player +// examples/songs: catalog.js, ranking.js) over a music-data.json in Node.js, the way the page reads it. Every chart +// must get a row, every playable chart its figures in every play scenario (Gekisou Live at several ranks and Just +// rates, Free Live, a Great share), finite and positive, and the rankings must work on them. +// +// node music_data_smoke.mjs +// +// Prints one line per problem (at most 40) and exits 1, or a one-line summary. +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { pathToFileURL } from "node:url"; + +const [dir, file] = process.argv.slice(2); +if (!dir || !file) { + console.error("usage: node music_data_smoke.mjs "); + process.exit(2); +} +const catalog = await import(pathToFileURL(path.join(dir, "catalog.js")).href); +const ranking = await import(pathToFileURL(path.join(dir, "ranking.js")).href); +const data = JSON.parse(readFileSync(file, "utf8")); + +const problems = []; +const problem = (m) => problems.push(m); +const finite = (v) => typeof v === "number" && Number.isFinite(v); + +const charts = (data.songs || []).flatMap((s) => (s.charts || []).map((c) => ({ song: s, chart: c }))); +const rows = catalog.chartRows(data); +if (rows.length !== charts.length) problem(`chartRows: ${rows.length} rows for ${charts.length} charts`); +const kind = ranking.plainKind(data); +if (kind === null) problem("plainKind: the file has no plain score-up kind (effect 2000, 5 s, no targets)"); + +// Whether the data has what a scenario needs on a chart (a null rangeWeights or kind leaves a chart without figures in +// the rank and Just scenarios: music_data.py reports those as warnings). +const computable = (deck, scenario) => { + if (!deck) return false; + const has = (s) => s && Array.isArray((s.weights || [])[kind]); + if (scenario && scenario.mode === "free") return (deck.offSeeds || []).length > 0 && deck.offSeeds.every(has); + if (deck.unplayable || !(deck.seeds || []).length) return false; + const rank1 = !scenario || ((scenario.ranks || []).every((r) => r === 1) && (scenario.just ?? 1) >= 1); + return deck.seeds.every((s) => has(s) && (rank1 || Array.isArray((s.rangeWeights || [])[kind]))); +}; +const has = ranking.scenarioData(data); +for (const k of ["free", "ranks", "just"]) if (!has[k]) problem(`scenarioData: no data for the ${k} scenario`); + +const SCENARIOS = [ + ["Gekisou Live, rank 1", null], + ["Gekisou Live, rank 5", { mode: "battle", ranks: [5, 5, 5], just: 1, great: 0 }], + ["Gekisou Live, ranks 2 3 4", { mode: "battle", ranks: [2, 3, 4], just: 1, great: 0 }], + ["Gekisou Live, Just 0", { mode: "battle", ranks: [1, 1, 1], just: 0, great: 0 }], + ["Gekisou Live, rank 3, Just 0.5, Great 0.2", { mode: "battle", ranks: [3, 3, 3], just: 0.5, great: 0.2 }], + ["Free Live", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0 }], + ["Free Live, Great 0.5", { mode: "free", ranks: [1, 1, 1], just: 1, great: 0.5 }], +]; +const SKILLS = [1, 1, 1, 1, 1]; +let figures = 0; +for (const [name, scenario] of SCENARIOS) { + const joined = ranking.joinCharts(data, scenario); + const expected = charts.filter(({ chart }) => computable(chart.deck, scenario)).length; + if (joined.length !== expected) problem(`${name}: ${joined.length} charts with figures, expected ${expected}`); + for (const r of joined) { + figures++; + if (!finite(r.base) || r.base <= 0) problem(`${name}: chart ${r.scoreId}: base ${r.base}`); + if (!Array.isArray(r.weights) || !r.weights.length || !r.weights.every(finite)) { + problem(`${name}: chart ${r.scoreId}: weights are not finite numbers`); + } + } + if (!joined.length) continue; + const ranked = ranking.rank(joined, { skills: SKILLS, source: "bgm", overheadMs: 30000 }); + if (!ranked.some((r) => r.frontier)) problem(`${name}: no chart on the frontier`); + for (const r of ranked) { + if (r.lengthMs === null) problem(`${name}: chart ${r.scoreId}: no play length`); + else if (!finite(r.perMinute) || !finite(r.rate)) problem(`${name}: chart ${r.scoreId}: rate ${r.rate}`); + } + for (const room of [0, 5]) { + ranking.eventDominance(ranked, "bgm", ranking.X_MAX, room); + for (const r of ranked) { + const p = ranking.requiredPower(r, SKILLS, "S", 1, room); + if (p !== null && !(finite(p) && p >= 0)) problem(`${name}: chart ${r.scoreId}: required power ${p}`); + } + } + const one = ranked[0]; + const chance = ranking.reachChance(one, SKILLS, 300000, "S", 1, 0); + if (chance !== null && !(chance >= 0 && chance <= 1)) problem(`${name}: reach chance ${chance}`); + catalog.refigure(rows, data, scenario); +} +catalog.histogram(rows, (r) => r.level); + +// Aptitude is a separate view, never folded into the default song ranking. Older pinned pages may lack this API; +// report that explicitly until PLAYER_REF is deliberately moved to a page with the aptitude UI. +let aptitudeFigures = 0; +const aptitudeApi = ["aptitudeShapes", "chartVariants", "aptitudeFigures", "aptitudeRate", "aptitudeSe", + "masterSkillFactor"].every((k) => typeof ranking[k] === "function") + && ["gekisouSkill", "shapeSkills", "shapeBands"].every((k) => typeof catalog[k] === "function"); +const near = (a, b) => finite(a) && finite(b) && Math.abs(a - b) <= 1e-8 * Math.max(1, Math.abs(a), Math.abs(b)); +if (aptitudeApi && data.deck?.gekisouAptitude) { + const shapes = ranking.aptitudeShapes(data); + if (shapes.size !== data.deck.gekisouAptitude.shapes.length) problem("aptitudeShapes: missing shapes"); + for (const shape of shapes.values()) { + const skills = catalog.shapeSkills(data, shape, "zh-Hant"); + if (skills.length !== new Set(shape.skills.map((s) => `${s.id}:${s.level}`)).size) { + problem(`shape ${shape.id}: shapeSkills count`); + } + const table = shape.source === "support" ? "supportSkills" : "skills"; + for (const skill of skills) { + const named = catalog.gekisouSkill(data, table, skill.id, "zh-Hant"); + if (!named.name || named.name !== skill.name) problem(`shape ${shape.id}: skill name lookup`); + } + const bands = catalog.shapeBands(data, shape, "zh-Hant"); + if (bands.length !== new Set(shape.skills.flatMap((s) => s.bandIds || [])).size || bands.some((s) => !s)) { + problem(`shape ${shape.id}: shapeBands lookup`); + } + } + const power = data.deck.model.power; + for (const { chart } of charts) { + const d = chart.deck; + if (!d) continue; + const variants = ranking.chartVariants(d); + if (variants.length !== (d.unplayable ? 0 : d.gekisouAptitude?.variants.length || 0)) { + problem(`chart ${chart.scoreId}: chartVariants count`); + } + // Removing aptitude must not alter any default figure. + const baseline = ranking.chartFigures(d, kind, power); + const without = ranking.chartFigures({ ...d, gekisouAptitude: null }, kind, power); + if (JSON.stringify(baseline) !== JSON.stringify(without)) problem(`chart ${chart.scoreId}: aptitude changes default`); + for (const variant of variants) { + for (const [name, scenario] of SCENARIOS) { + const tag = `aptitude chart ${chart.scoreId} shape ${variant.shape}, ${name}`; + const f = ranking.aptitudeFigures(variant, d.ranges, power, scenario); + if (scenario?.mode === "free") { + if (f !== null) problem(`${tag}: Free Live has aptitude`); + continue; + } + aptitudeFigures++; + if (!f || !finite(f.base)) { problem(`${tag}: missing or nonfinite base`); continue; } + if (f.weights !== null && (f.weights.length !== d.positions || !f.weights.every(finite))) { + problem(`${tag}: invalid weights`); + } + const zero = Array(d.positions).fill(0); + if (!near(ranking.aptitudeRate(f, zero), f.base)) problem(`${tag}: no-skill rate`); + const rate = ranking.aptitudeRate(f, SKILLS); + if (f.weights === null ? rate !== null : !finite(rate)) problem(`${tag}: missing cross-term handling`); + if (ranking.aptitudeSe(f, SKILLS) !== null) problem(`${tag}: SE assigned without covariance`); + if (scenario === null) { + if (!near(f.base, variant.score[0] / power) + || !near(ranking.aptitudeSe(f, zero), variant.score[1] / power)) problem(`${tag}: raw mean/SE`); + } else if (ranking.aptitudeSe(f, zero) !== null) problem(`${tag}: transformed SE is not null`); + } + // Only deterministic deltas describe the individual check seed. Never reconstruct a stochastic check + // from sampled means. Ordinary check cards are positional, unlike the UI's random-order expectation. + if (!variant.deterministic || kind === null) continue; + const c = variant.check; + const seed = d.seeds.find((s) => s.seed === c.seed); + const sc = { mode: "battle", ranks: c.ranks, just: 1, great: 0 }; + const base = ranking.scenarioSeed(seed, d.ranges, kind, sc); + const gain = ranking.aptitudeFigures(variant, d.ranges, power, sc); + if (!base || !gain?.weights) continue; + const predicted = data.deck.model.checkPower * ((base.score / power) + gain.base + + c.deck.reduce((sum, card, k) => sum + (card ? ranking.masterSkillFactor(card[1]) + * (base.weights[k] + gain.weights[k]) : 0), 0)); + // Exported weight rounding adds a small reconstruction error on top of the engine's bound. + const rounding = data.deck.model.checkPower * 1e-7 * (1 + d.ranges.length) * d.positions; + if (!finite(predicted) || Math.abs(predicted - c.exact) > c.bound + rounding) { + problem(`aptitude chart ${chart.scoreId} shape ${variant.shape}: deterministic check reconstruction`); + } + } + } +} + +if (problems.length) { + for (const m of problems.slice(0, 40)) console.log(m); + if (problems.length > 40) console.log(`${problems.length - 40} more problems`); + process.exit(1); +} +console.log(`page smoke test: ${rows.length} charts, ${SCENARIOS.length} scenarios, ${figures} chart figures; ` + + `plain kind ${kind}, scenarios free/ranks/just; aptitude ${aptitudeApi ? aptitudeFigures + " figures" : "API unavailable (skipped)"}`); diff --git a/.github/scripts/player.sh b/.github/scripts/player.sh new file mode 100755 index 0000000..fb9d3b9 --- /dev/null +++ b/.github/scripts/player.sh @@ -0,0 +1,11 @@ +#!/usr/bin/env bash +# A built ournotes-player checkout ($PLAYER_REPOSITORY at $PLAYER_REF) in $1: nnnotes writes its page files into the site. +set -euo pipefail +dir="$1" +if [ ! -f "$dir/dist/ournotes-player.element.min.js" ] || [ "$(git -C "$dir" rev-parse HEAD 2>/dev/null)" != "$(git -C "$dir" rev-parse "$PLAYER_REF^{commit}" 2>/dev/null)" ]; then + rm -rf "$dir" + git clone --quiet --filter=blob:none "https://github.com/$PLAYER_REPOSITORY.git" "$dir" + git -C "$dir" checkout --quiet "$PLAYER_REF" + (cd "$dir" && npm ci --ignore-scripts --no-audit --no-fund && npm run build) +fi +echo "ournotes-player $(git -C "$dir" rev-parse --short HEAD) ($(node -p "require('$dir/package.json').version"))" diff --git a/.github/scripts/songs_page.sh b/.github/scripts/songs_page.sh new file mode 100755 index 0000000..71849aa --- /dev/null +++ b/.github/scripts/songs_page.sh @@ -0,0 +1,18 @@ +#!/usr/bin/env bash +# The chart data page of ournotes-player ($PLAYER_REPOSITORY at $PLAYER_REF: examples/songs, the page that reads +# music-data.json) into $1, for the smoke test of music_data.py check. Only that directory is checked out and nothing +# is built: the test imports the page's pure modules (catalog.js, ranking.js) in Node.js. +set -euo pipefail +dir="$1" +if [ -z "${PLAYER_REF:-}" ]; then + echo "::error::set the repository variable MUSIC_DATA_PLAYER_REF: the ournotes-player commit whose chart data page reads this music data (.github/MUSIC_DATA.md)" + exit 1 +fi +rm -rf "$dir" +git clone --quiet --filter=blob:none --no-checkout "https://github.com/$PLAYER_REPOSITORY.git" "$dir" +git -C "$dir" sparse-checkout set examples/songs +git -C "$dir" checkout --quiet "$PLAYER_REF" +for f in catalog.js ranking.js; do + [ -f "$dir/examples/songs/$f" ] || { echo "::error::$PLAYER_REPOSITORY $PLAYER_REF has no examples/songs/$f"; exit 1; } +done +echo "ournotes-player $(git -C "$dir" rev-parse --short HEAD): examples/songs" diff --git a/.github/scripts/story_site.py b/.github/scripts/story_site.py new file mode 100755 index 0000000..c923f0b --- /dev/null +++ b/.github/scripts/story_site.py @@ -0,0 +1,358 @@ +#!/usr/bin/env python3 +"""The story site's CI steps (.github/workflows/story-site.yml): the published site lives in an S3 bucket, a run adds +the stories it lacks. + + regions the regions this run builds (GitHub outputs `regions`, a JSON list, and `count`): those of + $STORY_REGIONS (default hk-tw-mo) that the repository_dispatch payload names, or + $REQUESTED_REGION of a workflow_dispatch (empty / "all": every one), or all on a schedule + plan the stories to build: $REQUESTED, else every MasterAdv id without a manifest in the bucket + (at most $STORY_LIMIT, in id order); GitHub outputs `stories` (space separated) and `count` + master OUT the decoded master data of $MASTERDATA_REGION from moenotes-masterdata-sync (index.json: + every table with its SHA-256), for `nnnotes web` + fetch SITE every object of the site except assets/ into SITE (the manifests `nnnotes web` skips and the + indexes it rewrites from), with their SHA-256 in SITE.fetched.json + build SITE ID... `nnnotes web SITE --story ID ...` ($FORCE: --force), then `--player-only`, which rewrites the + player files and the indexes from every manifest present even after a failed build + publish SITE [--dry-run] + the new assets (those the bucket lacks), then every other file that is new or changed since + `fetch`, the indexes (stories.json, models.json, charts.json) last; never deletes + +Bucket: $STORY_S3_ENDPOINT, $STORY_S3_BUCKET, $STORY_S3_PREFIX (a key prefix, may be empty), credentials +$STORY_S3_ACCESS_KEY / $STORY_S3_SECRET_KEY (unset: anonymous, reads only). Objects are stored as nnnotes writes them: +a compressed asset (.gz / .br) as it is, without Content-Encoding (the player decodes it). +""" +from __future__ import annotations + +import hashlib +import json +import os +import re +import subprocess +import sys +import urllib.request +from concurrent.futures import ThreadPoolExecutor +from pathlib import Path + +INDEXES = ("stories.json", "models.json", "charts.json") +ASSETS = "assets/" +STORY_MANIFEST = re.compile(r"stories/(\d+)\.json") +ASSET_CACHE = "public, max-age=31536000, immutable" # content-addressed: never changes +OTHER_CACHE = "public, max-age=300" +TYPES = {".json": "application/json", ".gz": "application/gzip", ".br": "application/octet-stream", + ".m4a": "audio/mp4", ".mp4": "video/mp4", ".webm": "video/webm", ".png": "image/png", ".jpg": "image/jpeg", + ".glsl": "text/plain; charset=utf-8", ".html": "text/html; charset=utf-8", + ".js": "text/javascript; charset=utf-8", ".mjs": "text/javascript; charset=utf-8", + ".css": "text/css; charset=utf-8", ".map": "application/json", ".txt": "text/plain; charset=utf-8", + ".svg": "image/svg+xml", ".wasm": "application/wasm", ".bin": "application/octet-stream"} + + +def get(url: str, timeout: int = 120) -> bytes: + # an explicit User-Agent: Cloudflare in front of the services refuses Python-urllib's + request = urllib.request.Request(url, headers={"User-Agent": "moenotes-story-site (GitHub Actions)", + "Cache-Control": "no-cache"}) + with urllib.request.urlopen(request, timeout=timeout) as r: + return r.read() + + +def env(name: str, default: str | None = None) -> str: + value = os.environ.get(name, "") + if value: + return value + if default is None: + sys.exit(f"story_site: set {name}") + return default + + +def output(name: str, value: str) -> None: + path = os.environ.get("GITHUB_OUTPUT") + if path: + with open(path, "a", encoding="utf-8") as f: + f.write(f"{name}={value}\n") + print(f"{name}={value}") + + +def summary(text: str) -> None: + path = os.environ.get("GITHUB_STEP_SUMMARY") + if path: + with open(path, "a", encoding="utf-8") as f: + f.write(text + "\n") + print(text) + + +# ---------------------------------------------------------------- bucket +class Bucket: + def __init__(self): + import boto3 + from botocore import UNSIGNED + from botocore.config import Config + key, secret = os.environ.get("STORY_S3_ACCESS_KEY"), os.environ.get("STORY_S3_SECRET_KEY") + # S3-compatible stores (SeaweedFS here) take path-style requests and no default CRC checksums + config = Config(s3={"addressing_style": "path"}, retries={"max_attempts": 8, "mode": "adaptive"}, + request_checksum_calculation="when_required", response_checksum_validation="when_required", + max_pool_connections=32, **({} if key and secret else {"signature_version": UNSIGNED})) + self.writable = bool(key and secret) + self.s3 = boto3.client("s3", endpoint_url=env("STORY_S3_ENDPOINT"), aws_access_key_id=key or None, + aws_secret_access_key=secret or None, region_name="us-east-1", config=config) + self.name = env("STORY_S3_BUCKET") + prefix = os.environ.get("STORY_S3_PREFIX", "").strip("/") + self.prefix = f"{prefix}/" if prefix else "" + + def keys(self, sub: str = "") -> dict[str, int]: + """{site path: size} of the objects under `sub` (a site path prefix).""" + out = {} + for page in self.s3.get_paginator("list_objects_v2").paginate(Bucket=self.name, Prefix=self.prefix + sub): + for o in page.get("Contents", []): + out[o["Key"][len(self.prefix):]] = o["Size"] + return out + + def download(self, path: str, dest: Path) -> None: + dest.parent.mkdir(parents=True, exist_ok=True) + self.s3.download_file(self.name, self.prefix + path, str(dest)) + + def upload(self, path: str, src: Path) -> None: + suffix = Path(path).suffix.lower() + extra = {"ContentType": TYPES.get(suffix, "application/octet-stream"), + "CacheControl": ASSET_CACHE if path.startswith(ASSETS) else OTHER_CACHE} + self.s3.upload_file(str(src), self.name, self.prefix + path, ExtraArgs=extra) + + +def parallel(fn, items, workers: int = 16) -> None: + with ThreadPoolExecutor(workers) as ex: + for _ in ex.map(fn, items): + pass + + +def sha256(path: Path) -> str: + h = hashlib.sha256() + with open(path, "rb") as f: + for block in iter(lambda: f.read(1 << 20), b""): + h.update(block) + return h.hexdigest() + + +# ---------------------------------------------------------------- master data +def master_index() -> tuple[str, dict]: + base = env("MASTERDATA_BASE_URL").rstrip("/") + index = json.loads(get(f"{base}/index.json", 60)) + region = index["regions"][env("MASTERDATA_REGION")] + if not re.fullmatch(r"/([a-z]+/)?master/", region["path"]): + sys.exit(f"story_site: unexpected master path {region['path']!r}") + return base + region["path"], region + + +# Regions with a story site (story-site-region.yml gives each its own bucket prefix: JP model ids overlap the +# international ones). +SITE_REGIONS = ("hk-tw-mo", "jp") + + +def split_list(text: str) -> list[str]: + return [part for part in re.split(r"[\s,]+", text.strip()) if part] + + +def select_regions(enabled: str, event: str, payload: list[str] | None, requested: str) -> list[str]: + """The regions a run builds: of $STORY_REGIONS, the dispatch's regions, the requested one, or all (schedule).""" + allowed = [r for r in split_list(enabled) if r in SITE_REGIONS] + unknown = [r for r in split_list(enabled) if r not in SITE_REGIONS] + if unknown: + sys.exit(f"story_site: no story site for region(s) {', '.join(unknown)}") + if event == "repository_dispatch": + # Older dispatches carry no regions: build every enabled one (a run without new stories ends at plan). + wanted = payload if payload else allowed + elif event == "workflow_dispatch" and requested and requested != "all": + wanted = [requested] + else: + wanted = allowed + return [r for r in allowed if r in wanted] + + +def cmd_regions() -> None: + payload = None + event_path = os.environ.get("GITHUB_EVENT_PATH") + if event_path and Path(event_path).is_file(): + regions = (json.loads(Path(event_path).read_text(encoding="utf-8")).get("client_payload") or {}).get("regions") + payload = [r for r in regions if isinstance(r, str)] if isinstance(regions, list) else None + regions = select_regions(env("STORY_REGIONS", "hk-tw-mo"), env("GITHUB_EVENT_NAME", ""), payload, + os.environ.get("REQUESTED_REGION", "")) + output("regions", json.dumps(regions)) + output("count", str(len(regions))) + + +def configure_region() -> None: + """Keep build inputs from one release; JP client/API/CDN are read from the public snapshot.""" + region = env("MASTERDATA_REGION") + names = {"hk-tw-mo": "tw", "en": "en", "kr": "kr", "jp": "jp"} + if region not in names or env("NNNOTES_CATALOG_REGION") != names[region]: + sys.exit("story_site: master region and catalog region differ") + if region != "jp": + return + _, snapshot = master_index() + entry = snapshot.get("entry") or {} + assets = entry.get("assets") or {} + upstream = entry.get("upstream") or {} + api = assets.get("api_root") or upstream.get("api_root") + bundle = assets.get("bundle_root") + from urllib.parse import urlsplit + cdn = upstream.get("cdn_root") + if bundle: + u = urlsplit(bundle) + cdn = f"{u.scheme}://{u.netloc}" + client = entry.get("client_version") + if not all(isinstance(v, str) and v for v in (api, cdn, client)): + sys.exit("story_site: JP snapshot lacks API/CDN/client metadata") + os.environ.update(NNNOTES_SERVERS_JP_API=api, NNNOTES_SERVERS_JP_CDN=cdn, + NNNOTES_SERVERS_JP_CLIENT_VERSION=client, NNNOTES_SERVERS_JP_PROVIDER="jp") + + +def fetch_table(url: str, name: str, digest: str, dest: Path) -> None: + if not re.fullmatch(r"[A-Za-z0-9_-]+\.json", name): + sys.exit(f"story_site: unexpected master file name {name!r}") + for _ in range(4): + data = get(url + name) + if hashlib.sha256(data).hexdigest() == digest: + dest.write_bytes(data) + return + sys.exit(f"story_site: {name}: SHA-256 differs from index.json (the service changed snapshots? run again)") + + +def cmd_master(out: str) -> None: + url, region = master_index() + out_dir = Path(out) + out_dir.mkdir(parents=True, exist_ok=True) + parallel(lambda item: fetch_table(url, item[0], item[1], out_dir / item[0]), region["files"].items(), 8) + print(f"master data: {len(region['files'])} tables into {out_dir}") + + +# ---------------------------------------------------------------- plan +def parse_ids(text: str) -> list[int]: + ids = [] + for token in re.split(r"[\s,]+", text.strip()): + if token: + if not token.isdigit(): + sys.exit(f"story_site: not a story id: {token!r}") + ids.append(int(token)) + return list(dict.fromkeys(ids)) + + +def cmd_plan() -> None: + url, region = master_index() + data = get(url + "MasterAdv.json") + if hashlib.sha256(data).hexdigest() != region["files"]["MasterAdv.json"]: + sys.exit("story_site: MasterAdv.json: SHA-256 differs from index.json (run again)") + stories = sorted(row["_id"] for row in json.loads(data)["_allData"]) # nnnotes storysite.all_stories + have = {int(m.group(1)) for k in Bucket().keys("stories/") if (m := STORY_MANIFEST.fullmatch(k))} + requested = parse_ids(os.environ.get("REQUESTED", "")) + if requested: + unknown = [i for i in requested if i not in set(stories)] + if unknown: + sys.exit(f"story_site: no MasterAdv row for {', '.join(map(str, unknown))}") + todo, missing = requested, [i for i in requested if i not in have] + else: + missing = [i for i in stories if i not in have] + todo = missing[:int(env("STORY_LIMIT", "40"))] + summary(f"### Story site\n\n- master: {region.get('entry', {}).get('version', '?')} " + f"({len(stories)} stories), the site has {len(have)}\n- missing: {len(missing)}; this run builds " + f"{len(todo)}" + (f": {' '.join(map(str, todo))}" if todo else "")) + output("stories", " ".join(map(str, todo))) + output("count", str(len(todo))) + + +# ---------------------------------------------------------------- fetch / build / publish +def fetched_file(site: Path) -> Path: + return site.parent / f"{site.name}.fetched.json" + + +def cmd_fetch(site_dir: str) -> None: + site, bucket = Path(site_dir), Bucket() + keys = [k for k in bucket.keys() if not k.startswith(ASSETS) and not k.endswith("/")] + endpoint = env("STORY_S3_ENDPOINT").rstrip("/") + + # The bucket serves public read + list, so the fetch is plain HTTP like the player's, not a signed S3 + # call: Cloudflare's edge intermittently returned SignatureDoesNotMatch on signed ranged downloads. + def fetch(key: str) -> None: + dest = site / key + dest.parent.mkdir(parents=True, exist_ok=True) + dest.write_bytes(get(f"{endpoint}/{env('STORY_S3_BUCKET')}/{bucket.prefix}{key}", timeout=300)) + + parallel(fetch, keys) + (site / "assets").mkdir(parents=True, exist_ok=True) + digests = {k: sha256(site / k) for k in keys} + fetched_file(site).write_text(json.dumps(digests, indent=1, sort_keys=True), encoding="utf-8") + print(f"fetched {len(keys)} files ({sum((site / k).stat().st_size for k in keys) / 1e6:.1f} MB) into {site}") + + +def cmd_build(site_dir: str, ids: list[str]) -> None: + configure_region() + site = Path(site_dir).resolve() + stories = parse_ids(" ".join(ids)) + tmp = site.parent / f"{site.name}.tmp" + nnnotes = [sys.executable, "-m", "nnnotes", "web", str(site), "--tmp", str(tmp)] + if env("MASTERDATA_REGION") == "jp": + nnnotes += ["--story-languages", "ja"] + status = 0 + if stories: + cmd = nnnotes + [a for i in stories for a in ("--story", str(i))] + cmd += ["--workers", env("STORY_WORKERS", "2")] + if os.environ.get("FORCE") == "true": + cmd.append("--force") + print("+ " + " ".join(cmd[1:]), flush=True) + status = subprocess.run(cmd, stdout=open(site.parent / f"{site.name}.build.json", "wb")).returncode + # The indexes and player files from every manifest present, whatever the build did. + index = subprocess.run(nnnotes + ["--player-only"], stdout=subprocess.PIPE).returncode + report = site.parent / f"{site.name}.build.json" + if report.is_file(): + try: + r = json.loads(report.read_text(encoding="utf-8")) + summary(f"- built {len(r.get('storiesBuilt', []))}, failed {len(r.get('storiesFailed', []))}, " + f"skipped {len(r.get('storiesSkipped', []))}; models built " + f"{len((r.get('storyModels') or {}).get('modelsBuilt') or [])} in {r.get('storySeconds', '?')} s") + except ValueError: + pass + if status or index: + sys.exit(f"story_site: nnnotes web exited with {status or index}") + + +def cmd_publish(site_dir: str, dry_run: bool = False) -> None: + site, bucket = Path(site_dir), Bucket() + if not bucket.writable and not dry_run: + sys.exit("story_site: publish needs STORY_S3_ACCESS_KEY and STORY_S3_SECRET_KEY") + before = json.loads(fetched_file(site).read_text(encoding="utf-8")) + local = sorted(p.relative_to(site).as_posix() for p in site.rglob("*") if p.is_file()) + assets = [k for k in local if k.startswith(ASSETS)] + have = bucket.keys(ASSETS) + new_assets = [k for k in assets if have.get(k) != (site / k).stat().st_size] + changed = [k for k in local if not k.startswith(ASSETS) and before.get(k) != sha256(site / k)] + last = [k for k in changed if k in INDEXES] + first = [k for k in changed if k not in INDEXES] + if dry_run: + for k in first + last: + print(f"would upload {k}") + else: + # manifests before the indexes that list them, assets before the manifests that name them + for group in (new_assets, first, last): + parallel(lambda k: bucket.upload(k, site / k), group) + summary(f"- {'would publish' if dry_run else 'published'} {len(new_assets)} assets ({sum((site / k).stat().st_size for k in new_assets) / 1e6:.1f} MB)" + f", {len(first)} files, indexes: {', '.join(last) or 'unchanged'}") + + +def main(argv: list[str]) -> None: + if not argv: + sys.exit(__doc__) + cmd, args = argv[0], argv[1:] + if cmd == "regions" and not args: + cmd_regions() + elif cmd == "plan" and not args: + cmd_plan() + elif cmd == "master" and len(args) == 1: + cmd_master(args[0]) + elif cmd == "fetch" and len(args) == 1: + cmd_fetch(args[0]) + elif cmd == "build" and args: + cmd_build(args[0], args[1:]) + elif cmd == "publish" and len(args) in (1, 2) and args[1:] in ([], ["--dry-run"]): + cmd_publish(args[0], dry_run=bool(args[1:])) + else: + sys.exit(__doc__) + + +if __name__ == "__main__": + main(sys.argv[1:]) diff --git a/.github/scripts/test_jp_workflow.py b/.github/scripts/test_jp_workflow.py new file mode 100644 index 0000000..e6eeccd --- /dev/null +++ b/.github/scripts/test_jp_workflow.py @@ -0,0 +1,85 @@ +"""The workflow selects one release and derives JP endpoints without copying credentials.""" +import os +from pathlib import Path +import sys + +import pytest + +sys.path.insert(0, str(Path(__file__).parent)) +import story_site +import music_data + + +def test_configure_jp_from_public_snapshot(monkeypatch): + monkeypatch.setenv('MASTERDATA_REGION', 'jp') + monkeypatch.setenv('NNNOTES_CATALOG_REGION', 'jp') + entry = {'client_version': '1.0.4', 'assets': {'api_root': 'https://api.example.test', + 'bundle_root': 'https://cdn.example.test/asset/1.0.0.300/Android/' + 'a' * 32}} + monkeypatch.setattr(story_site, 'master_index', lambda: ('unused', {'entry': entry})) + for key in ('API', 'CDN', 'CLIENT_VERSION', 'PROVIDER'): + monkeypatch.setenv('NNNOTES_SERVERS_JP_' + key, '') + story_site.configure_region() + assert os.environ['NNNOTES_SERVERS_JP_CDN'] == 'https://cdn.example.test' + assert os.environ['NNNOTES_SERVERS_JP_API'] == 'https://api.example.test' + assert os.environ['NNNOTES_SERVERS_JP_CLIENT_VERSION'] == '1.0.4' + + +def test_reject_mixed_master_and_catalog(monkeypatch): + monkeypatch.setenv('MASTERDATA_REGION', 'jp') + monkeypatch.setenv('NNNOTES_CATALOG_REGION', 'tw') + with pytest.raises(SystemExit, match='differ'): + story_site.configure_region() + + +def test_asset_hash_changes_build_inputs(monkeypatch): + monkeypatch.setenv('MASTERDATA_REGION', 'jp') + monkeypatch.setattr(music_data, 'deck_commit', lambda root: 'a' * 40) + monkeypatch.setattr(music_data, 'nnnotes_commit', lambda root: 'b' * 40) + a = music_data.inputs({'version': 'v', 'resource_hash': 'a' * 32}) + b = music_data.inputs({'version': 'v', 'resource_hash': 'b' * 32}) + assert a != b and a['resourceHash'] == 'a' * 32 + + +def test_jp_catalog_provenance_matches_snapshot(): + from test_music_data import sample + doc = sample() + doc['provenance']['catalog']['resourceVersion'] = '1.0.0.300' + doc['provenance']['catalog']['resourceHash'] = 'a' * 32 + context = music_data.Context(snapshot={'region': 'jp', 'entry': { + 'version': doc['provenance']['master']['version'], 'resource_version': '1.0.0.300', 'resource_hash': 'b' * 32}}) + gate = music_data.Gate() + music_data.gate_provenance(doc, context, gate) + assert any('resourceHash differs' in s for s in gate.failures) + + +def test_regions_follow_the_dispatch_payload(): + assert story_site.select_regions('hk-tw-mo jp', 'repository_dispatch', ['jp'], '') == ['jp'] + assert story_site.select_regions('hk-tw-mo jp', 'repository_dispatch', ['hk-tw-mo', 'en', 'kr'], '') == ['hk-tw-mo'] + # an older dispatch without regions: every enabled region + assert story_site.select_regions('hk-tw-mo,jp', 'repository_dispatch', None, '') == ['hk-tw-mo', 'jp'] + + +def test_regions_of_schedule_and_manual_runs(): + assert story_site.select_regions('hk-tw-mo jp', 'schedule', None, '') == ['hk-tw-mo', 'jp'] + assert story_site.select_regions('hk-tw-mo jp', 'workflow_dispatch', None, 'all') == ['hk-tw-mo', 'jp'] + assert story_site.select_regions('hk-tw-mo jp', 'workflow_dispatch', None, 'jp') == ['jp'] + # a region STORY_REGIONS leaves out is not built + assert story_site.select_regions('hk-tw-mo', 'workflow_dispatch', None, 'jp') == [] + + +def test_regions_reject_a_region_without_a_story_site(): + with pytest.raises(SystemExit, match='no story site'): + story_site.select_regions('hk-tw-mo kr', 'schedule', None, '') + + +def test_regions_command_reads_the_event(tmp_path, monkeypatch): + import json + event = tmp_path / 'event.json' + event.write_text(json.dumps({'client_payload': {'regions': ['jp']}}), encoding='utf-8') + out = tmp_path / 'out' + monkeypatch.setenv('GITHUB_EVENT_PATH', str(event)) + monkeypatch.setenv('GITHUB_EVENT_NAME', 'repository_dispatch') + monkeypatch.setenv('GITHUB_OUTPUT', str(out)) + monkeypatch.setenv('STORY_REGIONS', 'hk-tw-mo jp') + story_site.cmd_regions() + assert out.read_text(encoding='utf-8').splitlines() == ['regions=["jp"]', 'count=1'] diff --git a/.github/scripts/test_music_data.py b/.github/scripts/test_music_data.py new file mode 100644 index 0000000..3df390c --- /dev/null +++ b/.github/scripts/test_music_data.py @@ -0,0 +1,709 @@ +"""Self-test of the music data gates (music_data.py): a synthetic music-data.json passes every gate, and each gate +fails on the defect it is there for. Runs in seconds, without the network: + + python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py + +The JSON Schema gate uses docs/schema/music-data.schema.json (or $MUSIC_DATA_SCHEMA) when the checkout has it; the +page smoke test runs when $MUSIC_DATA_PAGE names ournotes-player's examples/songs (and Node.js is installed). Real +files, when named: $MUSIC_DATA_SAMPLE (a file with the play scenario fields: the content gates pass; one made before +the ranges' luckPoints and the Gekisou skill aptitude, the deck and aptitude gates stop it on those alone) and +$MUSIC_DATA_OLD_SAMPLE (one without the play scenario fields: the scenario gate stops it). +""" +import copy +import hashlib +import importlib.util +import io +import json +import os +import shutil +import sys +from pathlib import Path + +import pytest + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import music_data # noqa: E402 +from music_data import Context, gates # noqa: E402 + +ROOT = Path(__file__).resolve().parents[2] +LANGS = ["ja", "en", "zh-Hant", "zh-Hans", "ko"] +TABLES = music_data.SONG_TABLES + ("MasterLiveSkillEffect", "MasterLiveGekisouRankingScoreBonus") +BIN = {t: hashlib.sha256(t.encode()).hexdigest() for t in TABLES} # the manifest's hashes of the served files +DECK = "d4" * 20 +CONTENT = ("deck", "scenarios", "finite", "references", "bgm") + + +def schema_path(): + p = Path(os.environ.get("MUSIC_DATA_SCHEMA") or ROOT / music_data.SCHEMA) + if not p.is_file(): + return None + return p if importlib.util.find_spec("jsonschema") else None + + +def page_path(): + p = os.environ.get("MUSIC_DATA_PAGE") + return Path(p) if p and Path(p, "ranking.js").is_file() and shutil.which("node") else None + + +# ---------------------------------------------------------------- a synthetic file +def text(stem): + return {lang: f"{stem}-{lang}" for lang in LANGS} + + +def check_deck(exact): + return {"deck": [[0, 5000], None], "exact": exact, "predicted": exact + 0.25, "bound": 7.0} + + +LUCK_SEEDS = [11, 22] # a luck chart's seeds; else the one seed 0 +# the Gekisou skill aptitude's shapes: id, source, mission, band condition +SHAPES = [(0, "member", 1, False), (1, "support", 2, True), (2, "member", 3, False), (3, "support", 4, False)] +SEED_RULE = {"deterministicTest": 4, "batches": [32, 64, 128, 256, 512, 1024], "relative": 0.01, "baseline": 0.001, + "crossSeeds": 64} + + +def effect(band=False): + return {"effectType": 2000, "triggerType": 7010, "activationTimeSecond": 5.0, "effectValue": 1000, + "maxEffectValue": 0, "effectLimitCount": 0, "effectExecuteLimitCount": 0, "skillTargetIds": [], + "trigger": [[{"type": 7010, "values": [1], "positive": True, "targetIds": []}]], + "condition": [[{"type": 5000, "values": [1], "positive": True, "targetIds": None}]] if band else [], + "release": [], "reset": [], "cumulative": None} + + +def aptitude_head(): + return {"plainKind": 0, "host": "one performer with a synthetic empty Gekisou skill", "seedRule": SEED_RULE, + "shapes": [{"id": i, "source": src, "mission": m, "bandCondition": band, "effects": [effect(band)], + "skills": [{"id": 10 + i, "level": 5, "memberTargetIds": [41] if band else None, + "bandIds": [1] if band else None}]} for i, src, m, band in SHAPES]} + + +def variant(deck, shape, match, deterministic): + """One shape's aptitude on a chart: deterministic on a chart without a luck range, else on 32 seeds.""" + se = 0 if deterministic else 12.5 + + def p(m): + return [m, se] + ranges = [{"rangeScore": p(300), "rankBonus": p(750), "rangeScorePerfect": p(300), "maxCombo": p(0), + "justCount": p(0), "luckPoints": p(2)} for _ in deck["ranges"]] + n = 1 if deterministic else 32 + linear = deck["seeds"][0]["rangeWeights"] is not None + result = {"shape": shape, "bandMatch": match, "deterministic": deterministic, "seeds": n, "seTargetMet": True, + "crossSeeds": min(n, SEED_RULE["crossSeeds"]), "score": p(1500), "scorePerfect": p(1500), + "tail": p(1500 - 1050 * len(ranges)), "tailPerfect": p(1500 - 1050 * len(ranges)), + "converted": p(0), "ranges": ranges, + "weights": [p(0.01), p(0.02)], + "rangeWeights": [[p(0.001)] * len(ranges), [p(0.002)] * len(ranges)] if linear else None, + "check": {"seed": deck["seeds"][0]["seed"], "ranks": [2] * len(ranges) if linear else [1] * len(ranges), + "deck": [[0, 7000], None], "exact": 5000, "predicted": 5000.5, "bound": 3.0}} + base = deck["seeds"][0] + score = base["score"] + 1500 + weight = base["weights"][0][0] + 0.01 + if linear: + for r, info in zip(base["ranges"], deck["ranges"], strict=True): + percent = info["rankBonusPercents"][1] + score += int(r["rangeScore"] * percent / 100) - r["rankBonus"] + 300 * (percent - 250) / 100 + weight += (190 - 250) / 100 * (0.1 + 0.001) + predicted = 1000003 * (score / 300000 + 0.7 * weight) + result["check"].update(exact=round(predicted), predicted=predicted) + return result + + +def aptitude(deck, luck): + missions = {r["mission"] for r in deck["ranges"]} + variants = [variant(deck, i, match, deterministic=not luck) for i, _, m, band in SHAPES + if m == 4 or m in missions for match in ((True, False) if band else (None,))] + factors = [{"judgedNotes": 12, "justNotes": 0, "perfectNotes": 0, "tailNotes": 2, "comboAtStart": 5, + "lotteries": [4.0, 0.0] if r["mission"] == 2 else [0, 0]} for r in deck["ranges"]] + return {"factors": factors, "variants": variants} + + +def deck_chart(luck=False): + def one(n): + return {"seed": n, "score": 120000 + n, "ranges": [{"rangeScore": 4000, "rankBonus": 10000, "maxCombo": 10, + "justCount": 0, "luckPoints": 3 if luck else 0, + "lotResults": [1, 2, 0, 1] if luck else [0, 0, 0, 0], + "rangeScorePerfect": 4000}], + "weights": [[0.5, 0.25]], "check": check_deck(2000), "scorePerfect": 120000 + n, + "rangeWeights": [[[0.1], [0.05]]], "rankCheck": dict(check_deck(1900), ranks=[3])} + d = {"convertedNoteCount": 20, "skip": 0.01, "events": [[0, 1000], [1, 3000]], "positions": 2, + "ranges": [{"index": 0, "mission": 2 if luck else 1, "startMs": 1000, "endMs": 5000, + "rankBonusPercent": 250, "rankBonusPercents": [250, 190, 160, 100, 100]}], + "justNotes": 0, "seeds": [one(n) for n in (LUCK_SEEDS if luck else [0])], + "offSeeds": [{"seed": 0, "score": 90000, "weights": [[0.4, 0.2]], "check": check_deck(1500)}], + "unplayable": None} + d["gekisouAptitude"] = aptitude(d, luck) + return d + + +def chart(difficulty, score_id, last=60000, luck=False): + return {"difficulty": difficulty, "scoreId": score_id, "level": 10, "displayLevel": 10.5, "fullComboCount": 20, + "asset": {"key": f"Live/MusicScore/c/c_{score_id}", "sha256": "ab" * 32}, + "notes": {"judged": 20, "total": 22, "byOperateType": {"1": 20, "120": 2}}, + "bpm": {"main": 120.0, "min": 120.0, "max": 120.0, "changes": [{"timeMs": 0, "bpm": 120.0}]}, + "firstNoteMs": 1000, "lastJudgedNoteMs": last, "lastNoteMs": last, "musicLengthMs": last + 1000, + "skillEventsMs": [1000, 3000], "fevers": [[1000, 5000]], "deck": deck_chart(luck)} + + +def song(i, charts): + ranks = ["D", "C", "B", "A", "S", "SS"] + return {"id": i, "sortOrder": i, "startAt": "2026/01/01 0:00:00", "defaultUnlock": True, "title": text(f"t{i}"), + "ruby": None, "phonetic": text(f"p{i}"), "bandIds": [1], "bandName": None, "vocalCharacterIds": [1], + "lyricist": text("l"), "composer": text("c"), "arranger": None, "musicType": 1, "musicCategories": [1], + "bestMusicTagIds": [1], "jacket": f"jkt_{i}", "gekisouMissions": [1, 2, 3], + "bgm": {"soundId": i, "cueSheet": f"Bgm{i}", "cue": f"song{i}", + "length": {"lengthMs": 90000, "samples": 48000 * 90, "sampleRate": 48000, "durationMs": 90000}}, + "scoreRanks": [{"rank": r, "requiredScore": n * 1000, "battleRequiredScore": n * 2000} + for n, r in enumerate(ranks)], + "charts": charts, "master": {"MasterLiveMusic": {"_id": i, "_rate": 1.5}, "MasterLiveScoreRank": []}} + + +def sample() -> dict: + kind = {"id": 0, "effectType": 2000, "activationTimeSecond": 5.0, "durationMs": 5000, "skillTargetIds": [], + "skillConditionGroup": 0, "skillReleaseConditionGroup": 0, "effectLimitCount": 0, + "effectExecuteLimitCount": 0, "effectExecuteLimitResetConditionGroup": 0, "rows": [1], "values": [10000]} + return { + "format": "nnnotes.music-data/1", + "provenance": { + "region": "tw", "client": {"versionName": "1.0.1", "versionCode": 25}, + "catalog": {"resourceVersion": None, "sha256": "cd" * 32}, + "master": {"source": "api", "version": "v-test", "tables": {t: {"sha256": BIN[t]} for t in TABLES}}, + "exporter": {"name": "nnnotes", "version": "0.1.2", "chartFormat": "nnnotes.live-score/1"}, + "deck": {"name": "ournotes-deck", "version": "0.0.1", + "source": "https://github.com/empty-sekai/ournotes-deck", "commit": DECK, + "format": "ournotes-deck.chart-stats/2"}}, + "languages": LANGS, + "bands": [{"id": 1, "name": text("band"), "mainColor": "#3388BB", "subColor": "#FFFFFF"}], + "characters": [{"id": 1, "bandId": 1, "name": text("ch"), "shortName": text("c"), "mainColor": "#77BBDD"}], + "tags": [{"id": 1, "name": text("tag")}], + "categories": [{"id": 1, "musicCategories": [1], "name": text("cat")}], + "deck": {"model": {"power": 300000, "checkPower": 1000003, "gekisouAptitude": "one skill at a time"}, + "kinds": [kind], "gekisouAptitude": aptitude_head()}, + "songs": [song(100001, [chart("easy", 10), chart("expert", 30, luck=True)]), + song(100002, [chart("expert", 40, luck=True)])], + } + + +def context(tmp_path: Path, doc: dict, **kw) -> tuple[bytes, Context]: + """The file's bytes and what check reads next to it: the decoded master data and its snapshot, the jackets (the + page smoke test only with page=page_path(): it starts Node.js).""" + master = tmp_path / "master" + master.mkdir(exist_ok=True) + files = {} + for t in TABLES: + data = json.dumps({"_allData": []}).encode() + (master / f"{t}.json").write_bytes(data) + files[f"{t}.json"] = hashlib.sha256(data).hexdigest() + listed = [{"name": f"{t}.bin", "hash": BIN[t], "size": 1} for t in TABLES] + (master / music_data.MANIFEST).write_text(json.dumps({"version": "v-test", "files": listed}), encoding="utf-8") + jackets = tmp_path / "jackets" + jackets.mkdir(exist_ok=True) + for s in doc.get("songs") or []: + if s.get("jacket"): + (jackets / f"{s['jacket']}.webp").write_bytes(b"RIFF....WEBP") + # an infinity as nnnotes writes it (1e999: JSON that JavaScript reads too); a NaN stays NaN (nnnotes writes none) + raw = json.dumps(doc, ensure_ascii=False).replace("Infinity", "1e999").encode("utf-8") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(region="tw", language="zh-Hant", master=master, + snapshot={"entry": {"version": "v-test", "client_version": "1.0.1"}, "files": files}, + deck_commit=DECK, nnnotes_version="0.1.2", jackets=jackets, schema=schema_path(), published=None, + page=None, file=tmp_path / "music-data.json") + for k, v in kw.items(): + setattr(ctx, k, v) + return raw, ctx + + +def run(tmp_path, doc, **kw) -> dict: + return gates(*context(tmp_path, doc, **kw)) + + +def gate(report: dict, name: str) -> dict: + return next(g for g in report["gates"] if g["gate"] == name) + + +def failures(report: dict) -> list: + return [(g["gate"], g["failures"]) for g in report["gates"] if not g["passed"]] + + +def seed(doc, song=0, chart=0): + return doc["songs"][song]["charts"][chart]["deck"]["seeds"][0] + + +def apt(doc, song=0, chart=0): + return doc["songs"][song]["charts"][chart]["deck"]["gekisouAptitude"] + + +def var(doc, i=0, song=0, chart=0): + return apt(doc, song, chart)["variants"][i] + + +def shape(doc, i): + return doc["deck"]["gekisouAptitude"]["shapes"][i] + + +def nonlinear(doc, song=0, chart=0): + """A chart whose ranks are not linear (overlapping ranges): no rangeWeights, the checks at rank 1.""" + d = doc["songs"][song]["charts"][chart]["deck"] + for s in d["seeds"]: + s.update(rangeWeights=None, rankCheck=None) + for v in d["gekisouAptitude"]["variants"]: + v["rangeWeights"] = None + v["check"]["ranks"] = [1] * len(d["ranges"]) + predicted = 1000003 * ((d["seeds"][0]["score"] + v["score"][0]) / 300000 + 0.7 * 0.51) + v["check"].update(exact=round(predicted), predicted=predicted) + + +# ---------------------------------------------------------------- the gates +def test_the_sample_passes_every_gate(tmp_path): + r = run(tmp_path, sample(), page=page_path()) + assert r["passed"], failures(r) + assert [g["gate"] for g in r["gates"]] == [name for name, _ in music_data.GATES] + assert not [g for g in r["gates"] if g["warningCount"]] + assert r["sha256"] == hashlib.sha256((tmp_path / "music-data.json").read_bytes()).hexdigest() + + +def test_schema(tmp_path): + if schema_path() is None: + pytest.skip("no docs/schema/music-data.schema.json (a fork not synced with upstream) or no jsonschema") + doc = sample() + doc["songs"][0]["charts"][0]["level"] = "10" + g = gate(run(tmp_path, doc), "schema") + assert not g["passed"] and any("songs/0/charts/0/level" in f for f in g["failures"]) + + +def test_page_smoke(tmp_path): + if page_path() is None: + pytest.skip("set MUSIC_DATA_PAGE to ournotes-player's examples/songs (and install Node.js)") + assert gate(run(tmp_path, sample(), page=page_path()), "page")["passed"] + doc = sample() + for s in doc["songs"]: + for c in s["charts"]: + c["deck"].pop("offSeeds") + g = gate(run(tmp_path, doc, page=page_path()), "page") + assert not g["passed"] and any("free scenario" in f for f in g["failures"]) + + +def test_page_aptitude_check_reconstruction(tmp_path): + page = page_path() + if page is None or "aptitudeFigures" not in (page / "ranking.js").read_text(encoding="utf-8"): + pytest.skip("page does not yet expose the aptitude API") + doc = sample() + # Stochastic check values are not reconstructible from means; changing them must not trigger reconstruction. + var(doc, 0, 0, 1)["check"].update(exact=99999999, predicted=99999999) + assert gate(run(tmp_path, doc, page=page), "page")["passed"] + # Even a self-consistent exported check is independently rejected when deterministic deltas disagree. + var(doc)["check"].update(exact=99999999, predicted=99999999) + g = gate(run(tmp_path, doc, page=page), "page") + assert any("deterministic check reconstruction" in f for f in g["failures"]), g + + +def unplayable(doc): + doc["songs"][1]["charts"][0]["deck"]["unplayable"] = "more than three fevers" # its seeds kept + + +@pytest.mark.parametrize("change, name, match", [ + # the play scenario fields + (lambda d: d["songs"][0]["charts"][0]["deck"].pop("offSeeds"), "scenarios", "offSeeds missing"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["offSeeds"].append({}), "scenarios", "offSeeds has 2 entries"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].pop("rankBonusPercents"), "scenarios", + "rankBonusPercents missing"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[250, 190, 160, 100]), + "scenarios", "rankBonusPercents not five ints"), + (lambda d: d["songs"][0]["charts"][0]["deck"]["ranges"][0].update(rankBonusPercents=[251, 190, 160, 100, 100]), + "scenarios", "is not rankBonusPercent"), + (lambda d: seed(d).pop("scorePerfect"), "scenarios", "no scorePerfect"), + (lambda d: seed(d).pop("rangeWeights"), "scenarios", "no rangeWeights"), + (lambda d: seed(d).pop("rankCheck"), "scenarios", "no rankCheck"), + (lambda d: seed(d)["ranges"][0].pop("rangeScorePerfect"), "scenarios", "rangeScorePerfect missing"), + (lambda d: seed(d).update(rangeWeights=[[[0.1]]]), "scenarios", "rangeWeights are not"), + (lambda d: seed(d)["rankCheck"].update(exact=99999), "scenarios", "rank check deck is not within"), + # the Gekisou skill aptitude + (lambda d: d["deck"].pop("gekisouAptitude"), "aptitude", "deck.gekisouAptitude missing"), + (lambda d: d["deck"]["model"].pop("gekisouAptitude"), "aptitude", "deck.model.gekisouAptitude: no text"), + (lambda d: d["deck"]["gekisouAptitude"].update(plainKind=1), "aptitude", "the page's plain kind is 0"), + (lambda d: d["deck"]["kinds"][0].update(durationMs=6000), "aptitude", "plainKind 0, the page's plain kind is None"), + (lambda d: d["deck"]["kinds"][0].update(durationMs=6000), "aptitude", "weights or rangeWeights without a plain"), + (lambda d: d["deck"]["gekisouAptitude"].pop("host"), "aptitude", "deck.gekisouAptitude: no host"), + (lambda d: d["deck"]["gekisouAptitude"].update(seedRule=dict(SEED_RULE, batches=[64, 32])), "aptitude", + "deck.gekisouAptitude.seedRule"), + (lambda d: shape(d, 1).update(id=5), "aptitude", "ids are not 0, 1, 2, ... in order"), + (lambda d: shape(d, 0).update(source="card"), "aptitude", "shape 0: source 'card'"), + (lambda d: shape(d, 2).update(mission=5), "aptitude", "shape 2: mission 5"), + (lambda d: shape(d, 0).update(bandCondition=True), "aptitude", "bandCondition True (a support skill's alone)"), + (lambda d: shape(d, 3).update(bandCondition=True), "aptitude", + "shape 3: bandCondition True, its effects have 0 condition 5000"), + (lambda d: shape(d, 1)["effects"][0]["condition"][0][0].update(targetIds=[41]), "aptitude", + "condition 5000 targetIds not null"), + (lambda d: shape(d, 0)["effects"][0].pop("cumulative"), "aptitude", "effect 0: cumulative missing or malformed"), + (lambda d: shape(d, 0)["effects"][0].update(effectValue=1.5), "aptitude", "effectValue missing or malformed"), + (lambda d: shape(d, 0)["effects"][0].update(reset=None), "aptitude", "reset missing or malformed"), + (lambda d: shape(d, 1)["skills"][0].update(memberTargetIds=None), "aptitude", "is not an id, a level"), + (lambda d: shape(d, 0)["skills"][0].update(bandIds=[1]), "aptitude", "is not an id, a level"), + (lambda d: shape(d, 0).update(skills=[]), "aptitude", "shape 0: no skills"), + (lambda d: d["songs"][0]["charts"][0]["deck"].pop("gekisouAptitude"), "aptitude", "no deck.gekisouAptitude"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(gekisouAptitude=None), "aptitude", + "deck.gekisouAptitude is null, but the chart is playable"), + (unplayable, "aptitude", "a Gekisou skill aptitude on a chart unplayable with Gekisou on"), + (lambda d: apt(d)["factors"].pop(), "aptitude", "0 factors for 1 ranges"), + (lambda d: apt(d)["factors"][0].update(justNotes=3), "aptitude", "Just or Perfect notes in a range without"), + (lambda d: apt(d)["factors"][0].update(tailNotes=-1), "aptitude", "tailNotes missing or not counts"), + (lambda d: apt(d)["factors"][0].update(lotteries=[1, 0]), "aptitude", "in a range without the luck mission"), + (lambda d: apt(d, 0, 1)["factors"][0].update(lotteries=[3.0, 0]), "aptitude", "deck.seeds' lotResults give 4.0"), + (lambda d: apt(d)["variants"].pop(0), "aptitude", "variants lack [(0, None)]"), + (lambda d: apt(d)["variants"].reverse(), "aptitude", "not in shape order, true first"), + (lambda d: apt(d, 0, 1)["variants"].pop(1), "aptitude", "variants lack [(1, False)]"), + (lambda d: var(d).update(shape=9), "aptitude", "variants of shapes ['9'] not in deck.gekisouAptitude.shapes"), + (lambda d: var(d).update(shape=2), "aptitude", "variants of shapes [2] of a mission the chart does not play"), + (lambda d: var(d, 0, 0, 1).update(bandMatch=None), "aptitude", "bandMatch None for a shape with a band"), + (lambda d: var(d).update(bandMatch=True), "aptitude", "bandMatch True for a shape without a band"), + (lambda d: var(d).pop("tail"), "aptitude", "shape 0: no tail"), + (lambda d: var(d).update(score=[1500]), "aptitude", "score not [mean, se]"), + (lambda d: var(d).update(score=[1500, -1]), "aptitude", "score not [mean, se]"), + (lambda d: var(d).update(converted=[float("nan"), 0]), "aptitude", "converted not [mean, se]"), + (lambda d: var(d)["ranges"][0].update(luckPoints=None), "aptitude", "ranges[0].luckPoints not [mean, se]"), + (lambda d: var(d).update(score=[1500, 1]), "aptitude", "deterministic, but a standard error is not 0: score"), + (lambda d: var(d).update(seeds=2), "aptitude", "deterministic, but 2 seeds"), + (lambda d: var(d, 0, 0, 1).update(seeds=33), "aptitude", "33 seeds (seTargetMet True), not a batch"), + (lambda d: var(d, 0, 0, 1).update(seTargetMet=False), "aptitude", "32 seeds (seTargetMet False), not a batch"), + (lambda d: var(d, 0, 0, 1).update(crossSeeds=64), "aptitude", "crossSeeds 64, expected min(32, 64)"), + (lambda d: var(d)["ranges"].append(dict(var(d)["ranges"][0])), "aptitude", "2 range results for 1 ranges"), + (lambda d: var(d).update(tail=[451, 0]), "aptitude", "tail 451 is not score 1500 less the ranges' rangeScore"), + (lambda d: var(d)["weights"].append([0, 0]), "aptitude", "weights are not one [mean, se] per position"), + (lambda d: seed(d).update(rangeWeights=None, rankCheck=None), "aptitude", + "rangeWeights, but deck.seeds[0].rangeWeights is null"), + (lambda d: var(d).update(rangeWeights=[[[0, 0]]]), "aptitude", "rangeWeights are not [position][range]"), + (lambda d: (nonlinear(d), var(d)["check"].update(ranks=[2])), "aptitude", "ranks are not linear (rank 1 alone)"), + (lambda d: var(d)["check"].update(seed=5), "aptitude", "the check's seed 5 is not deck.seeds[0]'s 0"), + (lambda d: var(d, 0, 0, 1)["check"].update(seed=22), "aptitude", "the check's seed 22 is not deck.seeds[0]'s 11"), + (lambda d: var(d)["check"].update(ranks=[6]), "aptitude", "the check's ranks [6] are not one rank per range"), + (lambda d: var(d)["check"].update(deck=[[1, 7000], None]), "aptitude", "not a [plain kind, value] or null"), + (lambda d: var(d)["check"].update(exact=99999999), "aptitude", "the check is not within its bound"), + (lambda d: var(d)["check"].pop("bound"), "aptitude", "the check has no seed, ranks"), + (lambda d: var(d).update(tailPerfect=[449, 0]), "aptitude", "tailPerfect is not scorePerfect"), + (lambda d: var(d).update(converted=[0.5, 0]), "aptitude", "a point delta is not an integer"), + # the deck statistics + (lambda d: d.update(deck=None), "deck", "deck is null"), + (lambda d: d["songs"][0]["charts"][0].update(deck=None), "deck", "no deck statistics"), + (lambda d: seed(d)["check"].update(exact=99999), "deck", "check deck is not within"), + (lambda d: seed(d).update(weights=[[0.5]]), "deck", "weights are not"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(seeds=[]), "deck", "no seeds"), + (unplayable, "deck", "unplayable, but has Gekisou on seeds"), + (lambda d: d["songs"][0]["charts"][0]["deck"].update(events=[[0, 1000]]), "deck", "1 skill events"), + (lambda d: seed(d).update(seed=5), "deck", "seeds [5] without a luck range, expected the one seed 0"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["seeds"].pop(), "deck", "1 seeds on a luck chart"), + (lambda d: d["songs"][0]["charts"][1]["deck"]["seeds"][1].update(seed=11), "deck", "2 seeds on a luck chart"), + (lambda d: d["songs"][1]["charts"][0]["deck"]["seeds"][1].update(seed=33), "deck", + "its 2 seeds are not the 2 of the first luck chart"), + (lambda d: seed(d)["ranges"][0].update(rankBonus=9999), "deck", "rankBonus 9999 is not trunc(4000 * 250 / 100)"), + (lambda d: seed(d)["ranges"][0].pop("luckPoints"), "deck", "seed 0 range 0: luckPoints missing"), + (lambda d: seed(d, 0, 1)["ranges"][0].update(luckPoints=2.5), "deck", "seed 11 range 0: luckPoints not an int"), + # numbers + (lambda d: seed(d)["weights"][0].__setitem__(1, float("nan")), "finite", "weights[0][1]: not finite"), + (lambda d: d["songs"][1]["charts"][0]["bpm"].update(main=float("inf")), "finite", "bpm.main: not finite"), + # references and texts + (lambda d: d["songs"][1]["bandIds"].append(9), "references", "band 9 not in"), + (lambda d: d["songs"][1]["vocalCharacterIds"].append(7), "references", "character 7 not in"), + (lambda d: d["songs"][0].update(title=None), "references", "title: no text"), + (lambda d: d["songs"][0].update(title={lang: "" for lang in LANGS}), "references", "empty in every language"), + (lambda d: d["songs"][0]["composer"].pop("ko"), "references", "composer: not a text"), + (lambda d: d["songs"][0].update(jacket=None), "references", "no jacket"), + (lambda d: d["songs"].reverse(), "references", "not sorted"), + (lambda d: d["songs"][1]["charts"][0].update(scoreId=10), "references", "occurs twice"), + (lambda d: d["songs"][0]["charts"].reverse(), "references", "charts ['expert', 'easy']"), + (lambda d: d["songs"][0].update(scoreRanks=[]), "references", "score ranks"), + (lambda d: d["songs"][0].update(bandIds=[]), "references", "neither a band nor a band name"), + # BGM + (lambda d: d["songs"][0]["bgm"].update(length=None), "bgm", "no BGM length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=50000, samples=48000 * 50), "bgm", + "ends before the last note"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=90500), "bgm", "is not samples"), + (lambda d: d["songs"][0]["bgm"]["length"].update(lengthMs=95000), "bgm", "cue length"), + (lambda d: d["songs"][0]["bgm"]["length"].update(durationMs=1200000, samples=48000 * 1200), "bgm", + "bounds"), + # provenance + (lambda d: d.update(format="nnnotes.music-data/2"), "provenance", "format"), + (lambda d: d["provenance"].update(region="en"), "provenance", "region 'en'"), + (lambda d: d["provenance"]["master"].update(version="other"), "provenance", "the snapshot's 'v-test'"), + (lambda d: d["provenance"]["master"].update(source="embedded"), "provenance", "master.source"), + (lambda d: d["provenance"]["master"]["tables"]["MasterText"].update(sha256="00" * 32), "provenance", + "MasterText.sha256 is not"), + (lambda d: d["provenance"]["master"]["tables"].pop("MasterBand"), "provenance", "lacks MasterBand"), + (lambda d: d["provenance"]["deck"].update(commit="ee" * 20), "provenance", "rust/Cargo.lock pins"), + (lambda d: d["provenance"]["exporter"].update(version="0.0.9"), "provenance", "installed nnnotes"), + (lambda d: d["provenance"]["client"].update(versionName=None), "provenance", "no APK version"), +]) +def test_a_gate_fails(tmp_path, change, name, match): + doc = sample() + change(doc) + r = run(tmp_path, doc) + g = gate(r, name) + assert not r["passed"] and not g["passed"] and any(match in f for f in g["failures"]), g + + +def test_a_table_read_is_not_the_listed_one(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.master / "MasterText.json").write_text('{"_allData": [{}]}', encoding="utf-8") + g = gate(gates(raw, ctx), "provenance") + assert not g["passed"] and g["failures"] == ["MasterText.json read is not the one index.json lists"] + + +def test_a_jacket_file_is_missing(tmp_path): + raw, ctx = context(tmp_path, sample()) + (ctx.jackets / "jkt_100002.webp").unlink() + g = gate(gates(raw, ctx), "references") + assert not g["passed"] and g["failures"] == ["song 100002: no jacket file jackets/jkt_100002.webp"] + + +def test_warnings_do_not_fail(tmp_path): + doc = sample() + nonlinear(doc) # overlapping ranges + seed(doc, 0, 1)["rangeWeights"][0] = None # a kind reading the confirmed rank + doc["songs"][1]["charts"][0]["deck"]["offSeeds"][0]["weights"][0] = None + doc["songs"][0]["master"]["MasterLiveMusic"]["_rate"] = float("inf") # 1e999 in master data as served + doc["songs"][1]["title"]["zh-Hant"] = "" + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + assert gate(r, "scenarios")["warningCount"] == 3 and gate(r, "finite")["warningCount"] == 1 + assert gate(r, "references")["warnings"] == ["songs without a zh-Hant title: 100002"] + + +def test_aptitude_warnings(tmp_path): + doc = sample() + var(doc, 0, 0, 1).update(seeds=1024, seTargetMet=False, crossSeeds=64) # the last batch, not the target + r = run(tmp_path, doc) + assert r["passed"], failures(r) + assert gate(r, "aptitude")["warnings"] == [ + "1 variants missed the seed rule's standard error target: 30 shape 1"] + assert gate(r, "aptitude")["note"] == ("4 shapes; 3 charts with an aptitude, 0 null; 8 variants, " + "2 deterministic; plain kind 0") + + +def test_no_gekisou_skill_shapes(tmp_path): + doc = sample() + doc["deck"]["gekisouAptitude"]["shapes"] = [] + kept = apt(doc) + for _, c in music_data.charts_of(doc): + c["deck"]["gekisouAptitude"] = None + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + doc["songs"][0]["charts"][0]["deck"]["gekisouAptitude"] = kept + g = gate(run(tmp_path, doc), "aptitude") + assert g["failures"] == ["chart 10 (100001 easy): a Gekisou skill aptitude on a chart without a Gekisou skill " + "shape"] + + +def test_an_unplayable_chart_has_no_aptitude(tmp_path): + doc = sample() + doc["songs"][1]["charts"][0]["deck"].update(unplayable="more than three fevers", seeds=[], gekisouAptitude=None) + r = run(tmp_path, doc, page=page_path()) + assert r["passed"], failures(r) + assert gate(r, "aptitude")["note"].startswith("4 shapes; 2 charts with an aptitude, 1 null") + + +def test_gzip_caps(tmp_path, monkeypatch): + raw, ctx = context(tmp_path, sample()) + g = gate(gates(raw, ctx, only=("gzip",)), "gzip") + assert g["passed"] and g["note"].startswith(f"{len(music_data.gzip.compress(raw, 6))} bytes gzipped") + monkeypatch.setattr(music_data, "APTITUDE_GZIP_MAX", 100) + g = gate(gates(raw, ctx, only=("gzip",)), "gzip") + assert not g["passed"] and g["failures"][0].startswith("the Gekisou skill aptitude: ") + monkeypatch.setattr(music_data, "FILE_GZIP_MAX", 100) + assert len(gate(gates(raw, ctx, only=("gzip",)), "gzip")["failures"]) == 2 + + +def test_against_the_published_file(tmp_path): + doc = sample() + raw, ctx = context(tmp_path, doc) + more = copy.deepcopy(doc) + more["songs"].append(song(100003, [chart("expert", 50)])) + ctx.published = json.dumps(more).encode() + r = gates(raw, ctx) + assert not gate(r, "counts")["passed"] and gate(r, "counts")["failures"] == [ + "2 songs, the published file has 3", "3 charts, the published file has 4"] + assert gate(r, "counts")["warnings"] == ["songs no longer in the file: 100003", "charts no longer in the file: 50"] + ctx.published = raw + b" " * (len(raw) * 2) # the same songs in a file three times as big + r = gates(raw, ctx) + assert gate(r, "counts")["passed"] and not gate(r, "size")["passed"] + ctx.published = raw[:len(raw) // 3] # a third: not JSON, and twice exceeded + r = gates(raw, ctx) + assert gate(r, "counts")["warnings"] == ["the published music-data.json is not JSON: skipped"] + assert not gate(r, "size")["passed"] + ctx.published = raw + assert gates(raw, ctx)["passed"] + + +def test_not_json(tmp_path): + r = gates(b"{", Context()) + assert not r["passed"] and r["gates"][0]["gate"] == "json" + + +def test_report_markdown(tmp_path): + doc = sample() + doc["songs"][0]["charts"][0]["deck"].pop("offSeeds") + text_ = music_data.report_markdown(run(tmp_path, doc)) + assert "| scenarios | **failed** (1) |" in text_ + assert "- scenarios failure: chart 10 (100001 easy): offSeeds missing, expected exactly one" in text_ + + +def test_deck_commit_from_cargo_lock(tmp_path): + lock = tmp_path / "rust" / "Cargo.lock" + lock.parent.mkdir() + lock.write_text('[[package]]\nname = "pyo3"\nversion = "0.29.0"\n\n[[package]]\nname = "ournotes-deck"\n' + 'version = "0.0.1"\nsource = "git+https://github.com/empty-sekai/ournotes-deck?rev=' + "a" * 40 + + "#" + "b" * 40 + '"\n', encoding="utf-8") + assert music_data.deck_commit(tmp_path) == "b" * 40 + with pytest.raises(SystemExit, match="sync the fork"): + music_data.deck_commit(tmp_path / "nowhere") + + +# ---------------------------------------------------------------- real files (when named) +def real(name): + p = os.environ.get(name) + if not p or not Path(p).is_file(): + pytest.skip(f"set {name} to a music-data.json") + return Path(p).read_bytes() + + +def luck_points(raw: bytes) -> bool: + """Whether a file's deck.seeds ranges have luckPoints (a file made before them has none).""" + return any("luckPoints" in r for _, c in music_data.charts_of(json.loads(raw)) + for s in (c.get("deck") or {}).get("seeds") or [] for r in s.get("ranges") or []) + + +def before_luck_points(r: dict) -> bool: + """The deck gate stopped only on the ranges' missing luckPoints.""" + g = gate(r, "deck") + return not g["passed"] and all(f.endswith("luckPoints missing") for f in g["failures"]) + + +def test_a_real_file_without_the_scenario_fields(): + raw = real("MUSIC_DATA_OLD_SAMPLE") + r = gates(raw, Context(language="zh-Hant"), only=CONTENT) + assert [n for n, _ in failures(r)] == (["scenarios"] if luck_points(raw) else ["deck", "scenarios"]) + assert luck_points(raw) or before_luck_points(r) + g = gate(r, "scenarios") + charts = sum(len(s["charts"]) for s in json.loads(raw)["songs"]) + assert g["failureCount"] >= charts and all("offSeeds missing" in f or "rankBonusPercents missing" in f + or "no scorePerfect" in f for f in g["failures"]) + + +def test_a_real_file_with_the_scenario_fields(tmp_path): + raw = real("MUSIC_DATA_SAMPLE") + (tmp_path / "music-data.json").write_bytes(raw) + ctx = Context(language="zh-Hant", page=page_path(), file=tmp_path / "music-data.json", published=raw) + r = gates(raw, ctx, only=CONTENT + ("aptitude", "counts", "size", "gzip", "page")) + new = [n for n, ok in (("deck", luck_points(raw)), ("aptitude", "gekisouAptitude" in json.loads(raw)["deck"])) + if not ok] # the gates of fields made after the file + assert [n for n, _ in failures(r)] == new, failures(r) + assert "deck" not in new or before_luck_points(r) + g = gate(r, "aptitude") + assert "aptitude" not in new or all(f.endswith(("gekisouAptitude missing", "gekisouAptitude: no text", + ": no deck.gekisouAptitude")) for f in g["failures"]), g + assert gate(r, "scenarios")["warningCount"] == 0 + print(f"aptitude: {g['note']}; warnings: {g['warnings']}; gzip: {gate(r, 'gzip')['note']}") + + +def test_real_aptitude_deck_sample(): + """Optional real chart-stats output, wrapped without modifying its statistics or touching the source file.""" + stats = json.loads(real("MUSIC_DATA_APTITUDE_DECK_SAMPLE")) + doc = {"deck": {k: stats[k] for k in ("model", "kinds", "gekisouAptitude")}, + "songs": [{"id": c["musicId"], "charts": [{"scoreId": c["scoreId"], "difficulty": c["difficulty"], + "skillEventsMs": [e[1] for e in c["events"]], "deck": c}]} for c in stats["charts"]]} + report = gates(json.dumps(doc).encode(), Context(), only=("deck", "aptitude", "gzip")) + assert report["passed"], failures(report) + + +# ---------------------------------------------------------------- publish (a stand-in bucket) +class FakeS3: + def __init__(self, corrupt=None): + self.store, self.log, self.corrupt = {}, [], corrupt + + def upload_file(self, src, bucket, key, ExtraArgs): + self.store[key] = b"not it" if key == self.corrupt else Path(src).read_bytes() + self.log.append((key, ExtraArgs["CacheControl"], ExtraArgs["ContentType"])) + + def get_object(self, Bucket, Key): + return {"Body": io.BytesIO(self.store[Key])} + + +class FakeBucket: + name, prefix, writable = "moenotes", "music-data/", True + + def __init__(self, s3): + self.s3 = s3 + + def keys(self, sub=""): + return {k[len(self.prefix):]: len(v) for k, v in self.s3.store.items() if k.startswith(self.prefix + sub)} + + +def published_out(tmp_path, monkeypatch, s3): + out = tmp_path / "out" + out.mkdir() + raw, ctx = context(out, sample()) + report = gates(raw, ctx) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "build.json").write_text(json.dumps({"sha256": report["sha256"], + "archive": f"archive/v-test/{report['sha256']}.json"}), + encoding="utf-8") + monkeypatch.setattr(music_data, "bucket", lambda: FakeBucket(s3)) + monkeypatch.setattr(music_data.time, "sleep", lambda s: None) + monkeypatch.setenv("STORY_S3_ENDPOINT", "https://storage.example") + monkeypatch.setenv("STORY_S3_BUCKET", "moenotes") + monkeypatch.delenv("FORCE", raising=False) + monkeypatch.setenv("MUSIC_DATA_PUBLISH", "true") + return out, report + + +def test_publish_order_and_read_back(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + music_data.cmd_publish(str(out)) + keys = [k for k, _, _ in s3.log] + archive = f"music-data/archive/v-test/{report['sha256']}.json" + assert sorted(keys[:2]) == ["music-data/jackets/jkt_100001.webp", "music-data/jackets/jkt_100002.webp"] + assert keys[2:] == [archive, "music-data/music-data.json", "music-data/build.json"] + caches = {k: c for k, c, _ in s3.log} + assert caches[archive].endswith("immutable") and caches["music-data/music-data.json"] == "no-cache" + assert s3.store["music-data/music-data.json"] == (out / "music-data.json").read_bytes() + s3.log.clear() + music_data.cmd_publish(str(out)) # again: the jackets and the archive copy are there + assert [k for k, _, _ in s3.log] == ["music-data/music-data.json", "music-data/build.json"] + monkeypatch.setenv("FORCE", "true") + s3.log.clear() + music_data.cmd_publish(str(out), dry_run=True) + assert s3.log == [] + + +def test_publish_stops_before_the_marker_when_the_file_does_not_read_back(tmp_path, monkeypatch): + s3 = FakeS3(corrupt="music-data/music-data.json") + out, _ = published_out(tmp_path, monkeypatch, s3) + + def offline(*a, **kw): + raise music_data.urllib.error.URLError("offline") + monkeypatch.setattr(music_data, "get", offline) + with pytest.raises(SystemExit, match="music-data.json: the bucket does not serve what was uploaded"): + music_data.cmd_publish(str(out)) + assert "music-data/build.json" not in s3.store + + +@pytest.mark.parametrize("switch", [None, "", "false", "True", "1"]) +def test_publishing_is_off_unless_the_switch_is_true(tmp_path, monkeypatch, switch): + s3 = FakeS3() + out, _ = published_out(tmp_path, monkeypatch, s3) + if switch is None: + monkeypatch.delenv("MUSIC_DATA_PUBLISH") + else: + monkeypatch.setenv("MUSIC_DATA_PUBLISH", switch) + music_data.cmd_publish(str(out)) # a dry run, whatever the command line says + assert s3.log == [] and s3.store == {} + + +def test_publish_needs_passed_gates(tmp_path, monkeypatch): + s3 = FakeS3() + out, report = published_out(tmp_path, monkeypatch, s3) + (out / "check.json").write_text(json.dumps(dict(report, passed=False)), encoding="utf-8") + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + (out / "check.json").write_text(json.dumps(report), encoding="utf-8") + (out / "music-data.json").write_bytes(b"{}") # not the file that was checked + with pytest.raises(SystemExit, match="has not passed the gates"): + music_data.cmd_publish(str(out)) + assert s3.store == {} diff --git a/.github/scripts/tools.sh b/.github/scripts/tools.sh new file mode 100755 index 0000000..5d925d4 --- /dev/null +++ b/.github/scripts/tools.sh @@ -0,0 +1,16 @@ +#!/usr/bin/env bash +# vgmstream-cli (the CRI HCA voices and music), pinned by version and SHA-256, into $1. ffmpeg comes from apt. +set -euo pipefail +dir="$1" +version=r2117 +digest=2f98c77f756079f63fbd119939067f1ed461d77e70993bc4cc372736d859c84a +mkdir -p "$dir" +if [ ! -x "$dir/vgmstream-cli" ]; then + curl -fsSL -o "$dir/vgmstream.zip" "https://github.com/vgmstream/vgmstream/releases/download/$version/vgmstream-linux.zip" + echo "$digest $dir/vgmstream.zip" | sha256sum -c - + unzip -q -o "$dir/vgmstream.zip" vgmstream-cli -d "$dir" + rm "$dir/vgmstream.zip" + chmod 755 "$dir/vgmstream-cli" +fi +"$dir/vgmstream-cli" -V 2>/dev/null | head -1 || true +ffmpeg -version | head -1 diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 4889741..fc8d103 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -29,10 +29,11 @@ jobs: - uses: Swatinem/rust-cache@v2 with: workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true - name: Install (builds the extension module nnnotes._deck) - run: | - python -m pip install --upgrade pip - python -m pip install -e ".[test]" + run: uv pip install --system -e ".[test]" - name: Lint run: python -m pyflakes src tests - name: Command line @@ -56,10 +57,11 @@ jobs: - uses: Swatinem/rust-cache@v2 with: workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true - name: Install (pinned) - run: | - python -m pip install --upgrade pip - python -m pip install -c .github/goldens-constraints.txt -e ".[test]" + run: uv pip install --system -c .github/goldens-constraints.txt -e ".[test]" - name: Golden records env: GOLDENS_STRICT: "1" diff --git a/.github/workflows/music-data.yml b/.github/workflows/music-data.yml new file mode 100644 index 0000000..7594233 --- /dev/null +++ b/.github/workflows/music-data.yml @@ -0,0 +1,152 @@ +name: Music data + +# Publishes the music data file of the current master data for the chart data page (ournotes-player examples/songs): +# `nnnotes music-data` (every song and chart with the deck model's statistics, docs/music-data.md) of +# moenotes-masterdata-sync's decoded master data, into the story site's bucket under music-data/: music-data.json, +# jackets/, archive//.json and the build marker build.json. A run first compares what it would +# build (the master data snapshot, the deck commit, the nnnotes commit) with the published build.json and ends there +# when nothing changed; else it builds, checks the file (quality gates and a smoke test with the page's own modules) +# and uploads it: the jackets and the archive copy first, then music-data.json, the build marker last. A file that +# fails a gate is not published. Nothing is deleted from the bucket. +# +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` when a region starts serving a new +# master snapshot; a daily schedule catches a missed dispatch; workflow_dispatch builds although nothing changed +# (`force`) or does a dry run. Publishing switch: unless the repository variable MUSIC_DATA_PUBLISH is `true`, every +# run, whatever its trigger, is a dry run (it builds and checks, and uploads nothing). docs: .github/MUSIC_DATA.md + +on: + repository_dispatch: + types: [masterdata-updated] + workflow_dispatch: + inputs: + force: + description: "build and publish although the published file was made from the same inputs" + type: boolean + default: false + dry_run: + description: "build and check, but upload nothing" + type: boolean + default: false + schedule: + - cron: "41 3 * * *" + +permissions: + contents: read + +# One run at a time: a run replaces music-data.json and build.json. +concurrency: + group: music-data-${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} + cancel-in-progress: false + +env: + # The story site's bucket (its repository variables) and the key prefix of the music data. + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + MUSIC_DATA_S3_PREFIX: ${{ vars.MUSIC_DATA_S3_PREFIX || (vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'jp/music-data' || 'music-data') }} + # The publishing switch: anything but 'true' makes every run a dry run. + MUSIC_DATA_PUBLISH: ${{ vars.MUSIC_DATA_PUBLISH || 'false' }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every file with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION || 'hk-tw-mo' }} + # The ournotes-player whose chart data page reads the file: its modules smoke-test every build. No default yet: the + # page with the play scenarios is not merged; a run stops before building until MUSIC_DATA_PLAYER_REF is set. + PLAYER_REPOSITORY: ${{ vars.MUSIC_DATA_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.MUSIC_DATA_PLAYER_REF || '' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'com.bushiroad.sirius' || vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the bundle key and the CDN come from secrets in the build step. + NNNOTES_CATALOG_REGION: ${{ vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'jp' || 'tw' }} + NNNOTES_CATALOG_LANGUAGE: ${{ vars.MUSIC_DATA_MASTERDATA_REGION == 'jp' && 'ja' || 'zh-Hant' }} + NNNOTES_SERVERS_JP_NAME: JP + NNNOTES_SERVERS_JP_LANGUAGES: ja + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 10 + outputs: + build: ${{ steps.plan.outputs.build }} + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 # the nnnotes commit: the last one that changed src/, rust/ or pyproject.toml + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - name: Inputs against the published build + id: plan + env: + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py plan + + build: + needs: plan + if: needs.plan.outputs.build == 'true' + runs-on: ubuntu-latest + timeout-minutes: 90 + steps: + - uses: actions/checkout@v7 + with: + fetch-depth: 0 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: The chart data page + run: .github/scripts/songs_page.sh "$WORK/page" + + - uses: Swatinem/rust-cache@v2 + with: + workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - name: Install nnnotes + # builds the extension module nnnotes._deck: the deck model at the commit rust/Cargo.toml pins + run: uv pip install --system -e ".[test]" boto3 + + - name: Gate self-test + env: + MUSIC_DATA_PAGE: ${{ env.WORK }}/page/examples/songs + run: python -m pytest -q -p no:cacheprovider .github/scripts/test_music_data.py .github/scripts/test_jp_workflow.py + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/music_data.py master "$WORK/master" + + - name: Build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + # A new cache on every run, never actions/cache: nnnotes keeps a downloaded catalog file for good (a cached + # one would never see a new resource version), and the cache holds decrypted game files. + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + run: python .github/scripts/music_data.py build "$WORK/out" + + - name: Gates + run: python .github/scripts/music_data.py check "$WORK/out" "$WORK/master" "$WORK/page/examples/songs" + + - name: Publish + # Only after every gate passed (a failed step skips it). A dry run (the dry_run input, or the publishing switch + # MUSIC_DATA_PUBLISH off) lists what it would upload; it gets no bucket key: anonymous, it can only read. + env: + STORY_S3_ACCESS_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_ACCESS_KEY || '' }} + STORY_S3_SECRET_KEY: ${{ vars.MUSIC_DATA_PUBLISH == 'true' && secrets.STORY_S3_SECRET_KEY || '' }} + FORCE: ${{ github.event.inputs.force }} + run: python .github/scripts/music_data.py publish "$WORK/out" ${{ (github.event.inputs.dry_run == 'true' || vars.MUSIC_DATA_PUBLISH != 'true') && '--dry-run' || '' }} diff --git a/.github/workflows/story-site-region.yml b/.github/workflows/story-site-region.yml new file mode 100644 index 0000000..aa9203d --- /dev/null +++ b/.github/workflows/story-site-region.yml @@ -0,0 +1,174 @@ +name: Story site (one region) + +# One region's story site: plan the stories its site lacks, build them, publish. Called by story-site.yml once per +# region (a matrix); every region's site is separate (its own bucket prefix, master data, APK and catalog). + +on: + workflow_call: + inputs: + region: + description: "moenotes-masterdata-sync region: hk-tw-mo or jp" + type: string + required: true + stories: + type: string + default: "" + force: + type: boolean + default: false + dry_run: + type: boolean + default: false + +permissions: + contents: read + +env: + # The bucket of the site (repository variables; the defaults are the StarMoe site). JP is published under jp/: + # its Live2D model ids overlap the international ones. + STORY_S3_ENDPOINT: ${{ vars.STORY_S3_ENDPOINT || 'https://storage.bdon.moe' }} + STORY_S3_BUCKET: ${{ vars.STORY_S3_BUCKET || 'moenotes' }} + STORY_S3_PREFIX: ${{ inputs.region == 'jp' && (vars.STORY_S3_PREFIX_JP || 'jp') || (vars.STORY_S3_PREFIX || '') }} + # moenotes-masterdata-sync: decoded master data by region (index.json lists every table with its SHA-256). + MASTERDATA_BASE_URL: ${{ vars.MASTERDATA_BASE_URL || 'https://metadata.bdon.moe' }} + MASTERDATA_REGION: ${{ inputs.region }} + # At most this many stories per run (a run must end within the job limit); the next run builds the rest. + STORY_LIMIT: ${{ vars.STORY_LIMIT || '40' }} + # The ournotes-player the site is written for (its page files are part of the site). + PLAYER_REPOSITORY: ${{ vars.STORY_PLAYER_REPOSITORY || 'empty-sekai/ournotes-player' }} + PLAYER_REF: ${{ vars.STORY_PLAYER_REF || '3774d8ac3987' }} + PLAYFETCH_VERSION: ${{ vars.PLAYFETCH_VERSION || 'v0.92' }} + APK_PACKAGE: ${{ inputs.region == 'jp' && 'com.bushiroad.sirius' || vars.STORY_APK_PACKAGE || 'com.bilibili.sirius' }} + PYTHON_VERSION: "3.13" + WORK: ${{ github.workspace }}/../work + # nnnotes settings (docs/configuration.md); the keys and the CDN come from secrets in the build step, JP's API/CDN/ + # client version from its public master snapshot (story_site.py configure_region). + NNNOTES_CATALOG_REGION: ${{ inputs.region == 'jp' && 'jp' || 'tw' }} + NNNOTES_CATALOG_LANGUAGE: ${{ inputs.region == 'jp' && 'ja' || 'zh-Hant' }} + NNNOTES_SERVERS_JP_NAME: JP + NNNOTES_SERVERS_JP_LANGUAGES: ja + NNNOTES_SERVERS_TW_NAME: TW/HK/MO + NNNOTES_SERVERS_TW_LANGUAGES: zh-Hans,zh-Hant,en,ko,ja + +jobs: + plan: + runs-on: ubuntu-latest + timeout-minutes: 15 + outputs: + stories: ${{ steps.plan.outputs.stories }} + count: ${{ steps.plan.outputs.count }} + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - run: uv pip install --system --quiet boto3 + - name: Stories to build + id: plan + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + REQUESTED: ${{ inputs.stories }} + run: python .github/scripts/story_site.py plan + + build: + needs: plan + if: needs.plan.outputs.count != '0' + runs-on: ubuntu-latest + timeout-minutes: 340 + steps: + - uses: actions/checkout@v7 + - uses: actions/setup-python@v7 + with: + python-version: ${{ env.PYTHON_VERSION }} + cache: pip + cache-dependency-path: pyproject.toml + - uses: actions/setup-node@v7 + with: + node-version: "22" + + - name: Free disk space + # The runner's preinstalled SDKs take most of its disk; a story build keeps its bundles and work files. + run: sudo rm -rf /usr/local/lib/android /usr/share/dotnet /opt/ghc /usr/local/.ghcup /opt/hostedtoolcache/CodeQL && df -h "$GITHUB_WORKSPACE" + + - uses: Swatinem/rust-cache@v2 # the install builds the extension module nnnotes._deck (rust/) + with: + workspaces: rust + - uses: astral-sh/setup-uv@v7 + with: + enable-cache: true + - name: Install nnnotes + # UnityPy 1.25.x requires etcpak, which has no manylinux x86_64 wheel on PyPI (only PyPy / cp37), so the + # runner's cp313 environment cannot resolve it. pyproject.toml already pins it to the 1.24 line. + run: uv pip install --system -e ".[fonts]" boto3 + + - name: Tools + run: | + sudo apt-get update -qq && sudo apt-get install -y -qq ffmpeg > /dev/null + .github/scripts/tools.sh "$WORK/tools" + + - name: Fonts + uses: actions/cache@v6 + with: + path: ${{ env.WORK }}/fonts + key: story-fonts-${{ hashFiles('.github/scripts/fonts.sh') }} + - run: .github/scripts/fonts.sh "$WORK/fonts" + + - name: ournotes-player + uses: actions/cache@v6 + with: + path: ${{ env.WORK }}/player + key: story-player-${{ env.PLAYER_REPOSITORY }}-${{ env.PLAYER_REF }} + - run: .github/scripts/player.sh "$WORK/player" + + - name: APK + env: + PLAYFETCH_CREDENTIALS_JSON: ${{ secrets.PLAYFETCH_CREDENTIALS }} + run: .github/scripts/apk.sh "$WORK/apk" + + - name: Master data + run: python .github/scripts/story_site.py master "$WORK/master" + + - name: The published site + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + run: python .github/scripts/story_site.py fetch "$WORK/site" + + - name: Build + id: build + env: + NNNOTES_BUNDLE_KEY: ${{ secrets.NNNOTES_BUNDLE_KEY }} + NNNOTES_BUNDLE_NONCE_SEED: ${{ secrets.NNNOTES_BUNDLE_NONCE_SEED }} + NNNOTES_SERVERS_TW_CDN: ${{ secrets.NNNOTES_SERVERS_TW_CDN }} + NNNOTES_PATHS_CACHE: ${{ env.WORK }}/cache + NNNOTES_PATHS_MASTER: ${{ env.WORK }}/master + NNNOTES_PATHS_APK: ${{ env.WORK }}/apk/base.apk + NNNOTES_PATHS_PLAYER: ${{ env.WORK }}/player + NNNOTES_PATHS_VGMSTREAM: ${{ env.WORK }}/tools/vgmstream-cli + NNNOTES_PATHS_FONTS_JA: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf + NNNOTES_PATHS_FONTS_EN: ${{ env.WORK }}/fonts/NotoSansCJKjp-Regular.otf + NNNOTES_PATHS_FONTS_ZH_HANT: ${{ env.WORK }}/fonts/NotoSansCJKtc-Regular.otf + NNNOTES_PATHS_FONTS_ZH_HANS: ${{ env.WORK }}/fonts/NotoSansCJKsc-Regular.otf + NNNOTES_PATHS_FONTS_KO: ${{ env.WORK }}/fonts/Pretendard-SemiBold.otf + NNNOTES_PATHS_FONTS_EMOJI: ${{ env.WORK }}/fonts/NotoColorEmoji.ttf + FORCE: ${{ inputs.force }} + run: python .github/scripts/story_site.py build "$WORK/site" ${{ needs.plan.outputs.stories }} + + - name: Publish + # Also after a partial build: the stories that were built are published, the failed ones are built again by + # the next run (they have no manifest). A dry run lists what it would upload. + if: always() && steps.build.outcome != 'skipped' && steps.build.outcome != 'cancelled' + env: + STORY_S3_ACCESS_KEY: ${{ secrets.STORY_S3_ACCESS_KEY }} + STORY_S3_SECRET_KEY: ${{ secrets.STORY_S3_SECRET_KEY }} + run: python .github/scripts/story_site.py publish "$WORK/site" ${{ inputs.dry_run && '--dry-run' || '' }} + + - name: Failures + if: always() + run: | + f="$WORK/site.story-failures.json" + if [ -f "$f" ]; then echo "::error::stories failed (${{ inputs.region }}), see the job summary"; { echo '### Failed stories (${{ inputs.region }})'; echo '```json'; cat "$f"; echo '```'; } >> "$GITHUB_STEP_SUMMARY"; fi diff --git a/.github/workflows/story-site.yml b/.github/workflows/story-site.yml new file mode 100644 index 0000000..05dfe07 --- /dev/null +++ b/.github/workflows/story-site.yml @@ -0,0 +1,75 @@ +name: Story site + +# Adds the stories that the published story sites do not have yet to them, one site per game region (hk-tw-mo at the +# bucket root, jp under jp/). A site (stories.json, stories/, models/, assets/, story/: the layout `nnnotes web` +# writes) lives in an S3 bucket; a region's run (story-site-region.yml) fetches its manifests (not its assets), builds +# the missing stories with `nnnotes web --story` on top of them, and uploads what the build added or changed: new +# assets first, then the story and model manifests, the indexes last. Nothing is deleted from the bucket. +# +# Triggers: moenotes-masterdata-sync sends repository_dispatch `masterdata-updated` (client_payload.regions) when a +# region starts serving a new master snapshot: those regions build; a daily schedule catches a missed dispatch (every +# region); workflow_dispatch builds one region or all, given stories (with `force`: again) or does a dry run. +# The regions: repository variable STORY_REGIONS (default "hk-tw-mo jp"). docs: .github/STORY_SITE.md + +on: + repository_dispatch: + types: [masterdata-updated] + workflow_dispatch: + inputs: + region: + description: "the region to build (all: every region of STORY_REGIONS)" + type: choice + options: [all, hk-tw-mo, jp] + default: all + stories: + description: "MasterAdv ids to build, separated by spaces or commas (empty: every story the site lacks)" + required: false + default: "" + force: + description: "rebuild the given stories (and their Live2D models) even though their manifests exist" + type: boolean + default: false + dry_run: + description: "build, but upload nothing" + type: boolean + default: false + schedule: + - cron: "23 3 * * *" + +permissions: + contents: read + +jobs: + regions: + runs-on: ubuntu-latest + timeout-minutes: 5 + outputs: + regions: ${{ steps.regions.outputs.regions }} + count: ${{ steps.regions.outputs.count }} + steps: + - uses: actions/checkout@v7 + - name: Regions to build + id: regions + env: + STORY_REGIONS: ${{ vars.STORY_REGIONS || 'hk-tw-mo jp' }} + REQUESTED_REGION: ${{ github.event.inputs.region }} + run: python3 .github/scripts/story_site.py regions + + site: + needs: regions + if: needs.regions.outputs.count != '0' + strategy: + fail-fast: false + matrix: + region: ${{ fromJSON(needs.regions.outputs.regions) }} + # One run per region at a time: every run rewrites its site's indexes from the manifests it fetched. + concurrency: + group: story-site-${{ matrix.region }} + cancel-in-progress: false + uses: ./.github/workflows/story-site-region.yml + with: + region: ${{ matrix.region }} + stories: ${{ github.event.inputs.stories || '' }} + force: ${{ github.event.inputs.force == 'true' }} + dry_run: ${{ github.event.inputs.dry_run == 'true' }} + secrets: inherit diff --git a/README.en.md b/README.en.md index e6c39ee..7bb923e 100644 --- a/README.en.md +++ b/README.en.md @@ -1,8 +1,6 @@ # nnnotes?! Japanese-release support includes anonymous version discovery, authenticated CDN downloads, gzip catalogs, -snapshot-isolated caches and split APKs. See the [JP guide](https://github.com/MetaSekaiLab/nnnotes/blob/main/docs/jp.md) -for configuration and validation limits. [简体中文](https://github.com/MetaSekaiLab/nnnotes/blob/main/README.md) | [English](https://github.com/MetaSekaiLab/nnnotes/blob/main/README.en.md) diff --git a/docs/jp.md b/docs/jp.md index 43a95d5..be4110a 100644 --- a/docs/jp.md +++ b/docs/jp.md @@ -93,6 +93,18 @@ can be read offline; downloading a missing historical file fails if Version now Refresh the catalog to use the new snapshot. Historical local bundles require the APK set matching the imported APK catalog. A JP master download likewise rejects a different current master snapshot. +## Workflow integration + +For the StarMoe workflows, set `MUSIC_DATA_MASTERDATA_REGION=jp`; the story site builds JP beside hk-tw-mo (`STORY_REGIONS`, +default `hk-tw-mo jp`, one run per region). This selects the +JP package and Japanese catalog language; the build reads API/CDN/client-version from the JP entry in the public +masterdata index. The default output prefixes become `jp/music-data` and `jp`, with separate concurrency groups. +Custom output prefixes must also be separate from the international outputs. Publishing remains controlled by +the existing workflow switches. No GitHub variable, secret, deployment or bucket is changed by installing nnnotes. + +The music build marker includes resource hash, so hash-only updates rebuild. Its provenance gate checks that the +catalog version/hash matches the master snapshot. A mixed master/catalog region is rejected before building. + ## Validation and current limits The synthetic tests use local gRPC and HTTP servers, including gzip catalogs, split packages, authentication diff --git a/pyproject.toml b/pyproject.toml index 2b91240..ef4dedc 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -11,6 +11,8 @@ license-files = ["LICENSE"] requires-python = ">=3.11" authors = [{ name = "MetaMiku" }, { name = "emptysekai" }] dependencies = [ + # 1.25.3's PyPI metadata briefly required etcpak (no manylinux x86_64 wheel); the same version now declares + # tpk-ar instead. pip 26's resolver keeps the cached old metadata and fails, so we install with uv. "UnityPy>=1.25", "numpy>=2.0", "Pillow>=10.0",