Skip to content

fix: flag cache deletion job once all tasks are created (MAPCO-11779) - #122

Open
almog8k wants to merge 2 commits into
masterfrom
fix/MAPCO-11779
Open

almog8k wants to merge 2 commits into
masterfrom
fix/MAPCO-11779

Conversation

@almog8k

@almog8k almog8k commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator
Question Answer
Bug fix ✔
New feature ✖
Breaking change ✖
Deprecations ✖
Documentation ✖
Tests added ✔
Chore ✖

Related issues: MAPCO-11779

Further information:

The cache deletion job is created with its first batch of tiles-deletion tasks (status IN_PROGRESS) and later batches are appended via POST /jobs/:jobId/tasks. Workers start immediately, job-tracker completed the job once the enqueued tasks were done, and the next batch was rejected with 409 (Cannot perform the requested operation on job with final state status: Completed) — the cache was left partially deleted.

Overseer now tells job-tracker when it has finished creating tasks, via raster-shared's CacheDeletionJobParams:

  • The job is created with parameters: { ingestionJobId, tasksCreationCompleted }. A job whose tasks all fit in the first batch (always the case for swap-update's single wipe task) is created with the flag already true — no follow-up update.
  • Otherwise, after the last batch is enqueued, overseer sets tasksCreationCompleted: true, sending the full parameters object since job-manager's updateJob overwrites parameters rather than merging.
  • Completing the job stays job-tracker's responsibility; overseer never completes it.
  • If an enqueue or the flag update fails, failPartialJob still marks the job Failed.
  • ingestionJobType is no longer sent (not part of the shared schema).
  • The creator's input interface is renamed CacheDeletionJobParams → CreateCacheDeletionJobParams to avoid clashing with the shared type.
  • Bumps @map-colonies/raster-shared to 9.0.0-alpha.2.

Deployment: pairs with MapColonies/job-tracker#78, which completes the job only once the flag is set. Deploy this first/together.

Known, accepted gap: if every task completes before the flag is written (flag update slow/retried, or overseer crashing in between), the job stays In Progress at 100%. Very unlikely under normal load.

Workers start on the first batch of tiles-deletion tasks while later batches
are still being enqueued, and job-tracker completed the job as soon as the
enqueued tasks were done, so later batches were rejected with 409 and the
cache was left partially deleted.

- create the job with raster-shared's CacheDeletionJobParams
  ({ ingestionJobId, tasksCreationCompleted }); a job whose tasks all fit in
  the first batch (always the case for swap-update) is created with the flag
  already set
- after the last batch is enqueued, set tasksCreationCompleted to true with
  the full parameters object (job-manager overwrites parameters on update);
  job-tracker completes the job
- drop ingestionJobType from the job parameters
- rename the creator input interface to CreateCacheDeletionJobParams to avoid
  clashing with the shared type
- bump @map-colonies/raster-shared to 9.0.0-alpha.2
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant