Skip to content

Translate beat schedules with explicit conditions and differences - #121

Merged
PhiLily merged 1 commit into
mainfrom
feature/beat-importer
Oct 4, 2026
Merged

PhiLily merged 1 commit into
mainfrom
feature/beat-importer

Conversation

@PhiLily

@PhiLily PhiLily commented Oct 4, 2026

Copy link
Copy Markdown
Member

Summary

Make ox_import_beat_schedules print explicit translation conditions,
a SCHEDULABLE_TASKS fragment containing only entries used by printed
rows, and one all-or-nothing created = create_schedules([...]) call.
The command writes nothing itself.

Who is affected

Anyone importing django-celery-beat schedules should check the named
differences and resolve omitted rows before switching schedulers.

If you applied importer output from 1.2.0 through 1.7.0, check schedules
restricting both day fields and crontabs from a zone other than
TIME_ZONE. Those versions passed fields through as stored and judged
zones by two sample offsets. Also check beat rows with microsecond
intervals, which were listed as below one second. Regenerate output
with this version to see these rows by name.

Schedule translation

  • Read crontabs with Celery's grammar, including wrap-around ranges,
    names, steps and lists. Use * for whole ranges and compact stepped
    ranges such as a-b/n. Use */n only for three or more values over
    the whole field.
  • Use compact expressions such as 1-59/2 1-23/2 1-31/2 * *.
    Long crontabs that 1.7.0 printed can print when a verified expression
    fits the 128-character cron column. Try a covering form when the
    canonical form is too long. List rows for which the command finds no
    verified expression within the column limit.
  • When weekdays are restricted, write a day-of-month field covering
    every date of its selected months as *. Omit rows that still
    restrict both day fields. Celery requires both to match; stored
    schedules run when either matches.
  • Compute intervals as beat does, with timedelta(**{period: every}),
    including weeks and milliseconds. For example, 2,000,000 microseconds
    becomes 2 seconds. Only SQLite can store a fractional every;
    0.1 days becomes 8,640 seconds. Apply Python's microsecond rounding
    before the whole-seconds check.
  • List intervals with fractional seconds, durations below one second,
    unknown periods or non-numeric values instead of translating them.
  • Compare crontab zone files with TIME_ZONE's data. Identical aliases
    qualify; matching current offsets alone do not. List incompatible
    or empty zones.
  • Decode arguments as beat decodes them. List rows whose arguments
    cannot be translated.

Differences and omitted rows

The "Translated, with a difference from beat" list names every interval.
Stored intervals count from a fixed instant; Celery counts from the
last run.

Clock-change notices follow actual tick patterns, including
spring-forward gaps and django-celery-beat 2.9.0's hour filter.
With USE_TZ on, stored schedules run matching repeated clock times
on both passes, while beat runs them on the first only. Notices give the
next relevant date within a ten-year scan from import.

A header line names rows that can stop beat's whole-table scheduling
pass. With USE_TZ and DJANGO_CELERY_BEAT_TZ_AWARE both off, rows
with a start time or expiry stored with a UTC offset are listed with
an explanation. Their due checks raise TypeError in
django-celery-beat 2.9.0. This can stop scheduling for the whole table.

Past starts are dropped from output. Rows with future starts are
omitted: beat runs once when the start arrives, while a stored
schedule waits for its next tick. Arrange those starts separately.

Destination checks cover NUL characters, surrogates, excessively deep
JSON, oversized integers, names equal under the name column's
comparison, and names already taken. Diagnostics quote names and
values with bounded output.

Applying the output

  • Run the command with beat's Python environment, timezone data and
    Django settings. The command cannot verify that they match.
  • Supply --beat-timezone ZONE when the crontab table has no timezone
    column or DJANGO_CELERY_BEAT_TZ_AWARE=False. If the option was given
    but was not needed, a comment directly below the header explains why.
    It appears only when at least one crontab row was read.
  • Where settings require TIME_ZONE to have UTC-identical data,
    distinguish a zone Python cannot load from missing files that prevent
    comparison. Neither failure establishes that the data differs
    from UTC's.
  • Check the output before applying the batch. Keep SCHEDULE_SOURCE
    in the same OPTIONS as SCHEDULABLE_TASKS; without it, no stored
    schedule ever runs.
  • Creation is all-or-nothing. Validation, permissions, concurrent
    changes and database errors can still prevent creation.

The translation covers run times while the scheduler is running.
It excludes downtime, late loading and a never-run row's first run.
The beat behavior behind the rules was measured with
django-celery-beat 2.9.0 and Celery 5.6.3.

No migration is needed. Stored-schedule dispatch is unchanged.
Read or stored-value conversion errors stop the command before
translation output. An unreadable database produces a non-zero exit
and a one-line error.

Validation

Execution covered 104,719 generated beat rows across SQLite,
PostgreSQL and MySQL. Every printed crontab and interval matched
Celery's and beat's own reading, with zero wrong schedules.

Every printed batch applied whole on three databases and six routed
pairs. Against real beat 2.9.0, there were zero missing clock-change
notices in the measured cases.

Tests cover each behavior the change adds or alters.

Make ox_import_beat_schedules print translation conditions, a registry
fragment for printed rows and one all-or-nothing batch:
created = create_schedules([...]). The command writes nothing itself.

Read crontabs with Celery's grammar. Use compact stepped ranges.
Long crontabs that 1.7.0 printed can print when a verified expression
fits the 128-character cron column. Try a covering form when the
canonical form is too long. List rows for which the command finds no
verified expression within the column limit. When weekdays are
restricted, write a day-of-month field covering every date of the
selected months as *. List rows that still restrict both day fields
because the schedulers combine them differently.

Compute intervals with timedelta(**{period: every}), including weeks
and milliseconds. Fractional every values can be stored only on SQLite.
Apply timedelta's microsecond rounding before checking whole seconds.
Decode arguments as beat does and compare crontab timezone files with
TIME_ZONE's data, rather than comparing current offsets.

List rows that cannot be translated or would fail destination checks.
Check names under the name column's comparison, including existing
names. Bound diagnostic quoting. Drop past starts and list rows with
future starts instead of translating them.

Name rows that can stop beat's whole-table scheduling pass in a header
line. With USE_TZ and DJANGO_CELERY_BEAT_TZ_AWARE both off, offset-bearing
start or expiry values are listed with an explanation. Their due checks
raise TypeError in django-celery-beat 2.9.0.

List every interval under "Translated, with a difference from beat":
stored intervals count from a fixed instant rather than the last run.
Derive clock-change notices from actual ticks, including spring-forward
gaps and beat 2.9.0's hour filter. Give the next relevant date within a
ten-year scan. With USE_TZ on, stored schedules run matching repeated
clock times on both passes, while beat runs them on the first only.

Add --beat-timezone for crontabs without a timezone column or with
DJANGO_CELERY_BEAT_TZ_AWARE=False. Explain an unused option below the
header when crontab rows were read. Distinguish an unloadable TIME_ZONE
from missing files when UTC-identical data is required.

Run the command with beat's Python environment, timezone data and
Django settings. Check differences and resolve omitted rows before
switching. Keep SCHEDULE_SOURCE in the same OPTIONS. Applying output
can still fail through validation, permissions, concurrent changes or
database errors.

Users of output from 1.2.0 through 1.7.0 should check both-day-field
restrictions, crontabs outside TIME_ZONE and microsecond intervals.
Regenerate output to see these rows by name.

Execution covered 104,719 generated beat rows on SQLite, PostgreSQL
and MySQL. Every printed crontab and interval matched Celery's and
beat's own reading, with zero wrong schedules. Every printed batch
applied whole on three databases and six routed pairs. No clock-change
notices were missing against real beat 2.9.0 in the measured cases.

No migration is needed. Stored-schedule dispatch is unchanged.

Tests cover each behavior the change adds or alters.
@PhiLily
PhiLily merged commit 31f1679 into main Oct 4, 2026
35 checks passed
@PhiLily
PhiLily deleted the feature/beat-importer branch October 4, 2026 15:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant