Skip to content

Release 0.77.11 - #3833

Merged
odlbot merged 22 commits into
releasefrom
release-candidate
Aug 25, 2026
Merged

odlbot merged 22 commits into
releasefrom
release-candidate

Conversation

@odlbot

@odlbot odlbot commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Ahtesham Quraish

Shankar Ambady

Tobias Macey

Anastasia Beglova

renovate[bot]

Matt Bertrand

Danielle Frappier

Sar

Chris Chudzicki

blarghmatey and others added 22 commits August 20, 2026 12:26
* chore(deps): mitol-django-observability 2026.3.11 -> 2026.8.19

Raises the floor rather than only relocking, so the OTLP endpoint work in
ol-infrastructure has a version it can rely on. `>=2026.1.0` already admitted
this release; nothing guaranteed it.

What the app gains, in dependency terms only -- no application code changes and
no new environment variables are set here:

- OTLP endpoint resolution is fixed. Previously the library read
  OTEL_EXPORTER_OTLP_ENDPOINT itself and passed it verbatim to the exporter, so
  the spec-correct base URL produced a POST at the collector root and a 404 per
  batch. OTEL_EXPORTER_OTLP_TRACES_ENDPOINT was ignored outright.
- A MeterProvider is installed, which is what lets DjangoInstrumentor build
  http.server.duration and http.server.active_requests from every request rather
  than the fraction of traces that survive tail sampling. This app already
  depends on opentelemetry-instrumentation-django directly, so the provider has
  something feeding it.
- trace_id/span_id stay on logs for spans the head sampler declined.
- Exception tracebacks render properly in production JSON logs.

Metrics stay off until the environment sets OTEL_EXPORTER_OTLP_ENDPOINT (a base
URL) or OTEL_EXPORTER_OTLP_METRICS_ENDPOINT, so this release changes no runtime
behaviour on its own.

2026.8.19 drops Python 3.10 and declares requires-python >=3.11. This app is
pinned to >=3.12,<3.13, so that is not a constraint here.

The `<2027` upper bound is preserved. `uv lock --upgrade-package` kept the change
surgical: the only version line that moved in uv.lock is this package's, with no
packages added or removed. Verified after `uv sync --frozen` that 2026.8.19 is
what installs and that both new entry points are present.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016KwQ8MzvnJiP5QJq6dUn4j

* docs(observability): correct the endpoint comment for 2026.8.19

The comment said OTEL_EXPORTER_OTLP_TRACES_ENDPOINT "is not consulted by the
released library" and that only OPENTELEMETRY_ENDPOINT and
OTEL_EXPORTER_OTLP_ENDPOINT enable tracing. That was true of 2026.3.11 and is
the exact bug 2026.8.19 fixes, so the bump in this PR turns the guidance into a
trap: an operator following it would avoid a configuration that now works.

Current behaviour, from mitol/observability/telemetry.py: _endpoint_from_env
reads OTEL_EXPORTER_OTLP_<SIGNAL>_ENDPOINT then OTEL_EXPORTER_OTLP_ENDPOINT,
and the enable guard considers the traces endpoint, the metrics endpoint and
the OPENTELEMETRY_ENDPOINT setting. Confirmed against the installed 2026.8.19
by booting an app with only OTEL_EXPORTER_OTLP_ENDPOINT set and seeing the
exporter configured from the environment.

Raised by copilot-pull-request-reviewer on this PR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016KwQ8MzvnJiP5QJq6dUn4j

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#2874 set the browser staleTime to 30 minutes to match the CDN s-maxage, but
#3280 subsequently moved s-maxage to an env var and did not update
getQueryClient. Production sets 7200, so CDN-cached HTML older than 30 minutes
once again hydrates into stale queries that refetch immediately -- the
featured-carousel flicker #2874 fixed. QA sets 900, below the hardcoded 30
minutes, which is why it did not resurface there.

Both values now come from getCacheSMaxageSeconds() in src/common/config.ts, so
the Cache-Control header proxy.ts sends and the staleTime the browser hydrates
with cannot diverge again.

The var is renamed NEXT_PUBLIC_CACHE_S_MAXAGE_SECONDS so the browser can read
it: NEXT_PUBLIC_* values are delivered at runtime via the x-public-env <meta>
(see src/env.ts), which is the only mechanism that gets a per-environment
Kubernetes value into the client -- process.env is inlined at build time and
empty in the standalone image.

The matching ol-infrastructure rename can merge in either order. Whichever side
is ahead, the app finds no value and falls back to the 1800-second default
until the other lands -- the wrong TTL for a window, not a broken one -- so the
rename needs no transitional both-names step.

Reading through env() rather than process.env is also what the
no-restricted-syntax lint rule requires for NEXT_PUBLIC_* keys.

An explicit 0 is still honored, matching the previous `|| "1800"` behavior,
since "" and unset are the only values that should fall back to the default.
RELEASE.rst records a past deployment that set s-maxage=0, so this is a value
someone has reached for before.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
* Revert "Update dependency tiktoken to >=0.13,<0.14 (#3816)"

This reverts commit 7d90b07.

* Revert "Update dependency ruff to v0.16.2 (#3815)"

This reverts commit cdb6619.

* Revert "Update dependency llama-index-llms-openai to >=0.7.10,<0.8 (#3814)"

This reverts commit 690f08d.

* Revert "Update dependency litellm to v1.95.0 (#3813)"

This reverts commit c9c3139.

* Revert "Update dependency drf-spectacular to >=0.30,<0.31 (#3812)"

This reverts commit ab6ab34.
* Update dependency Django to v5 [SECURITY]

* Fix Django 5.2 incompatibilities

Django 5 gives RemoteUserMiddleware its own __call__ that discards the
return value of process_request, so ApisixUserMiddleware returning
self.get_response(request) made every request run its downstream chain
twice — doubled writes, doubled side effects, throttle limits consumed at
2x. Returning None and letting __call__ drive the chain fixes it, and
behaves identically on 4.2. Same fix as learn-ai 2fe143b.

The rest:

- index_together was removed in 5.1; Meta.indexes replaces it and the
  migration renames the existing index in place rather than rebuilding it
- USE_L10N was removed in 5.0
- django-filter 2.4.0 calls ChoiceField._set_choices, which 5.0 removed.
  The >=2.4.0,<3 pin could never reach a fixed release because upstream
  switched to CalVer (21.1 ... 26.1), so the upper bound comes off
- djangorestframework 3.17.2 fixes RawPostDataException when DRF reads
  request.data under 5.2
- django-safedelete 1.4.1 still calls the deprecated log_action(); ignore
  the warning until upstream moves to log_actions()
- renovate was capped at Django <5, which is why 4.2 sat here through its
  EOL; raise it to <6 so 5.2.x patches come through

Full suite passes apart from the ordering flakes tracked in mitodl/hq#12841.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Bump CI postgres to 16

Django 5.2 requires PostgreSQL 14+. CI was still on 12.22; local dev
already runs 16 via docker-compose.services.yml.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Give the ChannelGroupRole index an explicit name

Matches taskbatch_job_status_idx and content_fb_course_time_idx rather
than Django's auto-generated hash name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Exercise the same-user branch through the middleware call

The same_user case called process_request() directly, so it couldn't catch
the duplicate get_response() run this branch fixes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Scope the log_action() deprecation filter to safedelete.admin

Matching on message alone would also swallow the warning for any
first-party log_action() call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Use postgres 15 in CI to match deployed RDS (15.17 on CI/RC/production)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Drop pytest-freezegun and finish the CheckConstraint.condition rename

pytest-freezegun (unreleased since 2019) calls distutils LooseVersion from
its marker helper, which runs once per collected item, so it accounted for
29376 of the 29399 warnings in CI. The repo used its marker exactly once;
everything else already imports freeze_time from freezegun directly.

Also renames check= to condition= in the two migrations that still had it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
add pypdf explicitly

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
* feat(cohort-1): StarRocks warehouse-pull ETL machinery

Generic Celery-task machinery for pulling Cohort-1 catalog data out of
the data platform's integrations__learn__* warehouse views: a
DB-API-backed BaseWarehouseETLTask, backend-agnostic iter_rows, and a
per-source cutover switch (WAREHOUSE_ETL_CUTOVER_SOURCES). Targets
StarRocks (MySQL wire protocol via pymysql) rather than Trino/Starburst
per the org's decision to draw down Trino usage — see mit-learn#3566,
closed in favor of this stacked-PR series.

Also carries the loaders.py None-vs-[] sentinel fix from #3566:
load_instructors/load_prices now distinguish "not provided by this
source" (None, leave existing values alone) from "explicitly empty"
([], clear them) — needed by any warehouse-pull transform that lacks
pricing/instructor data outright.

Adds a local/CI StarRocks instance (starrocks/allin1-ubuntu, pinned to
the same version running in production) plus a synthetic warehouse-row
factory, so BaseWarehouseETLTask/iter_rows are exercised against a real
StarRocks server, not only mocks — see warehouse_integration_test.py.

No sources are wired up yet (zero beat entries, zero cutover-eligible
sources); those land as separate stacked PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014XBMU19LUWbQjqsXjWWHHH

* fix: restore pluggable warehouse backend, fix StarRocks since-filter syntax

- Restore the WAREHOUSE_BACKEND / _CONNECTORS registry design this PR had
  dropped in favor of hard-locking to StarRocks. The backend must stay
  pluggable so a future engine (e.g. DuckDB for a local/offline mode) is
  a drop-in _connect_* addition, not another rewrite.
- Fix iter_rows' incremental (since=) query: `WHERE last_modified >
  TIMESTAMP '...'` is a syntax error on StarRocks (confirmed against a
  live starrocks/allin1-ubuntu instance — StarRocks implicitly casts a
  bare string literal compared against a DATETIME column instead, no
  TIMESTAMP keyword). Every full_refresh=False task would have failed
  outright the first time it ran against real StarRocks. Verified the
  fix end to end (full pull and incremental pull, both correct row
  counts) against a live container via this module's actual
  connect_to_warehouse()/iter_rows() code path, not just mocks.

* test: cover the warehouse-pull cutover-switch mechanism

Landed with zero test coverage — WAREHOUSE_ETL_CUTOVER_SOURCES defaults
to empty, and any non-empty value is currently "unknown" since
_API_ETL_BEAT_ENTRIES_BY_SOURCE starts empty (no source has landed a
task yet). Covers both.

* fix: drop broken STARROCKS_CATALOG connect param, isolate integration tests per xdist worker

- Remove STARROCKS_CATALOG entirely rather than fix its usage: pymysql's
  database= becomes MySQL COM_INIT_DB, which StarRocks resolves as a
  database name or catalog.database pair, never a bare catalog name
  alone. Confirmed against a live StarRocks instance —
  database="default_catalog" (analogous to the documented
  STARROCKS_CATALOG value "ol_data_lake_production") fails outright with
  "Unknown database". Every view_name here is already fully
  catalog-qualified, so no connect-time database is needed at all.
- Scope warehouse_integration_test.py's scratch database name by
  PYTEST_XDIST_WORKER (falls back to "master" outside xdist). CI runs
  `pytest -n logical`; these tests share one real StarRocks
  CREATE/INSERT/DROP lifecycle, so two workers running this module
  concurrently would race on the same table.

Addresses copilot-pull-request-reviewer feedback on #3807. Verified
both fixes against a live starrocks/allin1-ubuntu container — the
existing 3 integration tests pass with STARROCKS_CATALOG gone, and the
worker-scoped naming produces distinct database names per worker.

* fix: add a lookback window to the incremental watermark

The stored watermark was the consumer's (MIT Learn's) wall-clock time
at fetch start, with no allowance for warehouse build/replication lag.
A row modified just before that timestamp but not yet visible in the
queried view at query time would be permanently skipped by every
future incremental run — only a later full_refresh would repair it.

Subtract a fixed 10-minute _WATERMARK_LOOKBACK before storing. Trades
a few harmlessly-reprocessed rows each incremental run (upsert is
idempotent) for closing that skip window, as long as warehouse lag
never exceeds 10 minutes.

Currently dormant in production — every wired beat entry uses
full_refresh=True today, so _get_watermark/_set_watermark are never
exercised yet — but this is the machinery every future incremental
schedule tier builds on, so it should be correct before that tier
lands, not after.

Addresses copilot-pull-request-reviewer feedback on #3807.

* fix: reject non-int fetch_and_upsert return before advancing the watermark

A subclass that forgets `return count` implicitly returns None. Without
a check, that would crash the %d-formatted log line — but only after
_set_watermark had already run, permanently skipping whatever that run
should have picked up (the watermark update doesn't roll back on the
crash). Check the return type before advancing the watermark, not just
before the log call, so a broken fetch_and_upsert fails loudly instead
of silently corrupting future incremental pulls.

Addresses sentry[bot] feedback on #3807.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* adding initial changes

* updating views tests

* refactor docstring

* modify penalty score to properly surface documents on exact match

* fixed prefetch limit with offsets
Co-authored-by: Ahtesham Quraish <ahtesham.quraish@192.168.1.111>
@odlbot
odlbot requested a review from a team as a code owner August 25, 2026 09:34
@github-actions

Copy link
Copy Markdown

OpenAPI Changes

No changes detected

View full changelog

Unexpected changes? Ensure your branch is up-to-date with main (consider rebasing).

# that don't have instructor/price data (e.g. the warehouse-pull
# transforms, which lack pricing entirely) pass `None` explicitly
# rather than `[]`, so a sync doesn't wipe data another pipeline wrote.
instructors_data = run_data.pop("instructors", [])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: When an ETL source omits instructors data, it defaults to [] instead of None, causing load_instructors to delete existing instructors instead of preserving them as intended.
Severity: HIGH

Suggested Fix

Change the default value in run_data.pop("instructors", []) to None. This will align the implementation with the documented convention that None is the sentinel value to indicate that existing instructor data should be preserved.

Prompt for AI Agent
Review the code at the location below. A potential bug has been identified by an AI
agent. Verify if this is a real issue. If it is, propose a fix; if not, explain why it's
not valid.

Location: learning_resources/etl/loaders.py#L370

Potential issue: In `learning_resources/etl/loaders.py`, the `load_run` function
processes instructor data. A comment explicitly states that if an ETL source does not
provide instructor data, it should pass `None` to leave existing data untouched.
However, the implementation at line 370, `instructors_data = run_data.pop("instructors",
[])`, defaults to an empty list `[]` if the `instructors` key is missing. The subsequent
call to `load_instructors` interprets this empty list as a signal to delete all existing
instructors for the run. This leads to silent data loss for any ETL source that omits
the `instructors` key, contrary to the documented intention.

Did we get this right? 👍 / 👎 to inform future reviews.

@odlbot
odlbot merged commit 5e5afdb into release Aug 25, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants