Skip to content

promote: a stale _admonition/ copy refuses the refresh of every lecture that includes it, while drift-check calls it info #22

Description

@mmcky

Follow-up from the review of #14. The behaviour predates #14 (main refuses the same entries). Nothing has triggered it yet, but the next upstream edit of the GPU admonition will. Line numbers are at #14's head as it merges, with #17 included.

What happens

check_admonition_identity (tools/promote:970-984) runs for each _admonition/ include referenced by a lecture in the batch (:1379-1380). It hashes the include in every considered, canonical-eligible series and refuses the lecture unless all the copies are byte-identical. It has no stale rule. Same-name lecture copies have had one since 7128fb1 (resolve_sources, tools/promote:718-726): a lecture copy still at a version the pool already took stays a source. A copy of an include that has merely not caught up, though, counts as a conflict. drift-check judges that same copy as stale: check_asset reports stale-copy under "INFO — nothing to act on now" (tools/drift-check:510-515).

The estate has one include, _admonition/gpu.md. Its pool copy, _static/_shared/_admonition/gpu.md, has 13 users. Eleven are canonical in intermediate: career, ifp_advanced, ifp_discrete, ifp_egm, ifp_egm_transient_shocks, ifp_opi, jv, mccall_fitted_vfi, mccall_model, mccall_model_with_separation and mccall_persist_trans. Two are canonical in dp-test: mccall_model_with_sep_markov and os_egm_jax. In CI the guard compares programming's, intermediate's and dp-test's copies, because dp is a consumer and CI excludes jax. The live series edit the file together: programming, intermediate and jax all changed it on 2025-11-30 and again on 2026-01-13. dp-test, the sandbox copy of lecture-dp in QuantEcon/lecture-dp.monorepo, has not changed it since it was seeded on 2026-07-09, and nothing keeps it in step.

On a copy of the real data, with the same edit applied to the gpu.md of programming, intermediate and jax and their pins moved, drift-check exits 1. The asset's recorded canonical is programming, the first series in manifest order that holds the file:

REFRESH — the canonical copy moved; `tools/promote --refresh` (CI adds --exclude $PROMOTE_EXCLUDE) brings the pool up to date (1)
  [asset-moved] lectures/_static/_shared/_admonition/gpu.md: canonical asset in programming changed upstream

INFO — nothing to act on now (1)
  [stale-copy] lectures/_static/_shared/_admonition/gpu.md: the dp-test copy is unchanged while programming moved

The refresh job's tools/promote --refresh --exclude dp jax then refuses all 11 intermediate-canonical users, although their own text has not changed (output abridged):

[career] FAILED: _admonition/gpu.md differs across series (programming:a73e4512bcf5, intermediate:a73e4512bcf5, dp-test:fb5f6f926043) — cannot merge into _static/_shared/; reconcile upstream first
summary [refresh]: 0 promoted, 30 unchanged, 0 skipped, 11 failed
  FAIL: career, ifp_advanced, ifp_discrete, ifp_egm, ifp_egm_transient_shocks, ifp_opi, jv, mccall_fitted_vfi, mccall_model, mccall_model_with_separation, mccall_persist_trans

The other 30 targets refresh with no change to their text, so the refresh PR carries only their promoted_at bumps, and the run is red. The next night drift-check reports asset-moved again, and the refresh refuses the same 11. This repeats every night.

Nothing points the way out:

  • promote says "reconcile upstream first", but the copies that moved upstream already agree.
  • drift-check files the one differing copy under "nothing to act on now".
  • A sync/canonical.yml line cannot settle an asset; 8 of the 11 already have one.

Exempting the stale copy alone is not enough

If dp-test's copy is left out of the comparison, the 11 are still refused, now by the outside-batch guard (tools/promote:1414-1435). That guard stops a pool asset from changing while lectures outside the batch use it. The two dp-test-canonical users are outside the batch, because they are not refresh targets while dp-test's pin stands still. With a stale rule patched into the guard, the same run gives:

[career] FAILED: pool asset(s) exist with different bytes and may not be overwritten — reconcile deliberately: _static/_shared/_admonition/gpu.md (referenced by lectures outside this batch: mccall_model_with_sep_markov, os_egm_jax)

--exclude dp-test gives the same line. Promoting all 13 by name fails too, because the two dp-test lectures bring dp-test's old bytes to the same pool path. Without the patch, all 13 fail the identity guard. With it, the two dp-test lectures fail with an asset destination collision, and that leaves the 11 blocked as above. The only ways through are moving dp-test's copy, or removing dp-test.

Reproduction (with the tools/tests harness)

INCLUDE = "```{include} _admonition/gpu.md\n```"

repo = make_repo(["intermediate", "dp-test"])
for s in ("intermediate", "dp-test"):
    repo.mirror_file(s, "_admonition/gpu.md", "GPU note, version 1\n")
repo.lecture("intermediate", "mccall_model", lecture_text("McCall", INCLUDE))
repo.lecture("dp-test", "os_egm_jax", lecture_text("EGM", INCLUDE))
assert repo.promote("mccall_model", "os_egm_jax").returncode == 0
# The admonition is edited upstream; the sandbox's copy stays at its pin.
repo.mirror_file("intermediate", "_admonition/gpu.md", "GPU note, version 2\n")
repo.move_pin("intermediate")

dc = repo.drift_check()  # exit 1: asset-moved (refresh) and stale-copy (info)
assert "[stale-copy] lectures/_static/_shared/_admonition/gpu.md: the dp-test copy is unchanged" in dc.stdout
res = repo.promote("--refresh")
assert res.returncode == 1
assert "[mccall_model] FAILED: _admonition/gpu.md differs across series" in res.stderr
# Leaving dp-test's copy out, as a stale rule would, still refuses it:
res = repo.promote("--refresh", "--exclude", "dp-test")
assert res.returncode == 1
assert "referenced by lectures outside this batch: os_egm_jax" in res.stderr

The two refusals:

[mccall_model] FAILED: _admonition/gpu.md differs across series (intermediate:cabee308aca6, dp-test:3930471fa882) — cannot merge into _static/_shared/; reconcile upstream first
[mccall_model] FAILED: pool asset(s) exist with different bytes and may not be overwritten — reconcile deliberately: _static/_shared/_admonition/gpu.md (referenced by lectures outside this batch: os_egm_jax)

The same checks pass against main's tools, run with #14's harness.

When it bites

Not yet. At the current pins every series' copy of gpu.md is identical, and drift-check is clean. It bites at the next edit of lectures/_admonition/gpu.md in programming or intermediate. It keeps biting every night until that edit is copied into the sandbox or QuantEcon/project-monorepo#24 step 1 retires dp-test by removing its stanza from sync/manifest.yml.

It predates #14. On main, the same guard (tools/promote:627-640 there) refuses the same 11 entries, and the refresh step then fails without opening a PR. With #14, the nightly opens a refresh PR holding only the other 30 entries' promoted_at bumps, and the run stays red.

A milder form of the problem will remain after dp-test is gone. Suppose one series merges an admonition edit a day after another. The lagging copy is then stale in drift-check's terms, and the guard refuses the intermediate users at each refresh until that copy catches up (checked in the harness, with programming ahead of intermediate).

Until then

Whenever lectures/_admonition/gpu.md changes in QuantEcon/lecture-python-programming, QuantEcon/lecture-python.myst or QuantEcon/lecture-jax, copy the new file into QuantEcon/lecture-dp.monorepo, ideally before the next nightly sync. dp-test's pin then moves, its two users join the refresh batch, and all 13 refresh together. On the same copy of the real data, with dp-test's copy also updated, the refresh gives summary [refresh]: 0 promoted, 51 unchanged, 0 skipped, 0 failed, and drift-check is clean afterwards. "Unchanged" refers to the lecture texts; the pool's gpu.md takes the new version.

Suggested fix

  1. Give check_admonition_identity a stale rule. Pass it the pool asset's ledger record, and leave out any copy still at the recorded digest: the version the pool already holds, which drift-check calls stale-copy. Refuse only when the copies that moved disagree. This makes the two tools agree and clears the lagging-series case. For dp-test, the refusal moves to the outside-batch guard.
  2. Have that refusal name the real remedy. Sometimes the outside-batch guard blocks a shared asset because its users outside the batch are canonical in a series whose copy is still the recorded one. In that case the message should name that series and say its copy has to move. For example: "_static/_shared/_admonition/gpu.md is still used by mccall_model_with_sep_markov, os_egm_jax (canonical in dp-test, whose copy is the version in the pool): update dp-test's _admonition/gpu.md upstream so they refresh in the same batch". In that case drift-check's asset stale-copy finding could say the same, instead of "nothing to act on now".

A test alongside test_a_stale_copy_stays_a_source_and_the_refresh_goes_through could cover both: a lagging series' copy no longer blocks the refresh, and the dp-test case is refused with a message that names dp-test.

Activity

  1. added
    bugSomething isn't working
    on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions