Skip to content

Release 0.77.6 - #3782

Merged
odlbot merged 2 commits into
releasefrom
release-candidate
Aug 17, 2026
Merged

odlbot merged 2 commits into
releasefrom
release-candidate

Conversation

@odlbot

@odlbot odlbot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Tobias Macey

blarghmatey and others added 2 commits August 17, 2026 16:43
…#3780)

* Stream PostHog parquet in batches instead of exploding it into Series

get_learning_resource_views has been OOM-killing mitlearn-default-celery-worker
in production continuously since 0.77.3 shipped on 2026-08-13, at 15-24 container
restarts an hour.

The extractor read every column of each parquet export into a DataFrame and then
wrapped iterrows() in list(), which holds a pandas Series object for every row in
the file simultaneously. The list() was already redundant -- iterrows() is a
generator and the next line consumes it exactly once.

This was not new in 0.77.3. #3714 removed the MultipleObjectsReturned crash that
had been killing this task within seconds of every run, which is what had been
keeping the extractor from ever reaching the backlog. With the crash gone the
task walks the full set of S3 objects for the first time and exhausts the cgroup
limit. Because it is acks_late with reject_on_worker_lost, an OOM kill is a
redelivery rather than a failure, so it retries forever; and because it never
succeeds, last_event never advances, so each retry re-reads a larger backlog than
the last. Pending deliveries have climbed from 12 to 280 per 12h and the VPA is
already pegged at its ceiling, so there is no sizing fix left.

Decode a batch at a time via pyarrow, projecting only the three columns the
transform actually reads. The export also carries person_properties,
elements_chain, distinct_id, person_id, created_at, _inserted_at and event;
person_properties is a JSON blob comparable in size to properties, and all of
them were being decoded per row and discarded.

pyarrow is already a direct dependency. pandas was not declared at all -- only
gensim references it, and only under a docs extra -- so this also stops the ETL
importing an undeclared package. posthog.py was the last pandas import in the
repository.

Verified field-for-field equivalence against the checked-in fixture: same row
count, and uuid/timestamp/properties identical for every row. event_date is now
a stdlib datetime rather than a pandas Timestamp, which is what
LearningResourceViewEvent.event_date wanted anyway; the test that unwrapped it
with to_pydatetime() no longer needs to.

Deliberately not filtering on event == 'lrd_view' at scan time: the fixture's
events are all named lrd_open, so that predicate would silently discard
everything. Narrowing the scan is worth doing, but needs the real event taxonomy
confirmed first.

This bounds memory, not runtime. The loader still issues ~3 queries per event
over a multi-million row backlog, and a long task under acks_late stays exposed
to redelivery on any eviction. Batching the loader is the natural follow-up,
ahead of moving this ETL to the data platform entirely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011yGLhunNHrh4DLyujeH7QR

* Stop retaining a model instance per event in the loader

Batching the parquet decode bounded the extract side, but the load side was
still unbounded: load_posthog_lrd_view_events built a list comprehension over
the whole generator and returned it, so the full set of loaded rows stayed live
for the duration of the ETL. load_posthog_lrd_view_event returns a row on every
path -- including the early return for an event that is already stored -- so on
a backlog re-read every event contributed an instance. Millions of model
instances can exhaust the worker's limit on their own, which would have left the
OOM in place after the extract fix.

Consume the generator one event at a time and keep only the learning resource
ids that need recounting. That set was already being computed for the recount
loop, so this costs nothing, and it is bounded by distinct resources rather than
by events.

Return that set instead of the rows. The per-resource loaders in loaders.py
return what they loaded, and that convention is safe at their scale -- they take
one course's instructors or one page of courses. This loader is called once with
a generator over the entire S3 backlog, where the same convention is the bug.
Nothing consumed the return value: get_learning_resource_views calls
pipelines.posthog_etl() and discards it, and pipelines_test asserts on the
composed result with the loader mocked, so it is unaffected.

Add a log line with attempted/loaded/recount counts. The task processes millions
of events and previously reported nothing until it finished, which it has not
done since 2026-08-13.

The dropped `len(loaded_events) == 4` assertion held in both branches -- a
missing resource still appended None -- so it measured attempts, not loads. The
LearningResourceViewEvent.objects.count() assertion above it already covers the
real outcome; the loader assertion now checks which resources were recounted.

Raised by Copilot on #3780.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011yGLhunNHrh4DLyujeH7QR

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@odlbot
odlbot requested a review from a team as a code owner August 17, 2026 20:45
@github-actions

Copy link
Copy Markdown

OpenAPI Changes

No changes detected

View full changelog

Unexpected changes? Ensure your branch is up-to-date with main (consider rebasing).

@odlbot
odlbot merged commit 36757a4 into release Aug 17, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants