Release 0.77.6 - #3782
Merged
Merged
Release 0.77.6#3782
Conversation
…#3780) * Stream PostHog parquet in batches instead of exploding it into Series get_learning_resource_views has been OOM-killing mitlearn-default-celery-worker in production continuously since 0.77.3 shipped on 2026-08-13, at 15-24 container restarts an hour. The extractor read every column of each parquet export into a DataFrame and then wrapped iterrows() in list(), which holds a pandas Series object for every row in the file simultaneously. The list() was already redundant -- iterrows() is a generator and the next line consumes it exactly once. This was not new in 0.77.3. #3714 removed the MultipleObjectsReturned crash that had been killing this task within seconds of every run, which is what had been keeping the extractor from ever reaching the backlog. With the crash gone the task walks the full set of S3 objects for the first time and exhausts the cgroup limit. Because it is acks_late with reject_on_worker_lost, an OOM kill is a redelivery rather than a failure, so it retries forever; and because it never succeeds, last_event never advances, so each retry re-reads a larger backlog than the last. Pending deliveries have climbed from 12 to 280 per 12h and the VPA is already pegged at its ceiling, so there is no sizing fix left. Decode a batch at a time via pyarrow, projecting only the three columns the transform actually reads. The export also carries person_properties, elements_chain, distinct_id, person_id, created_at, _inserted_at and event; person_properties is a JSON blob comparable in size to properties, and all of them were being decoded per row and discarded. pyarrow is already a direct dependency. pandas was not declared at all -- only gensim references it, and only under a docs extra -- so this also stops the ETL importing an undeclared package. posthog.py was the last pandas import in the repository. Verified field-for-field equivalence against the checked-in fixture: same row count, and uuid/timestamp/properties identical for every row. event_date is now a stdlib datetime rather than a pandas Timestamp, which is what LearningResourceViewEvent.event_date wanted anyway; the test that unwrapped it with to_pydatetime() no longer needs to. Deliberately not filtering on event == 'lrd_view' at scan time: the fixture's events are all named lrd_open, so that predicate would silently discard everything. Narrowing the scan is worth doing, but needs the real event taxonomy confirmed first. This bounds memory, not runtime. The loader still issues ~3 queries per event over a multi-million row backlog, and a long task under acks_late stays exposed to redelivery on any eviction. Batching the loader is the natural follow-up, ahead of moving this ETL to the data platform entirely. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011yGLhunNHrh4DLyujeH7QR * Stop retaining a model instance per event in the loader Batching the parquet decode bounded the extract side, but the load side was still unbounded: load_posthog_lrd_view_events built a list comprehension over the whole generator and returned it, so the full set of loaded rows stayed live for the duration of the ETL. load_posthog_lrd_view_event returns a row on every path -- including the early return for an event that is already stored -- so on a backlog re-read every event contributed an instance. Millions of model instances can exhaust the worker's limit on their own, which would have left the OOM in place after the extract fix. Consume the generator one event at a time and keep only the learning resource ids that need recounting. That set was already being computed for the recount loop, so this costs nothing, and it is bounded by distinct resources rather than by events. Return that set instead of the rows. The per-resource loaders in loaders.py return what they loaded, and that convention is safe at their scale -- they take one course's instructors or one page of courses. This loader is called once with a generator over the entire S3 backlog, where the same convention is the bug. Nothing consumed the return value: get_learning_resource_views calls pipelines.posthog_etl() and discards it, and pipelines_test asserts on the composed result with the loader mocked, so it is unaffected. Add a log line with attempted/loaded/recount counts. The task processes millions of events and previously reported nothing until it finished, which it has not done since 2026-08-13. The dropped `len(loaded_events) == 4` assertion held in both branches -- a missing resource still appended None -- so it measured attempts, not loads. The LearningResourceViewEvent.objects.count() assertion above it already covers the real outcome; the loader assertion now checks which resources were recounted. Raised by Copilot on #3780. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011yGLhunNHrh4DLyujeH7QR --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
OpenAPI ChangesNo changes detected Unexpected changes? Ensure your branch is up-to-date with |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tobias Macey