Repository navigation
docs(interruptible-tasks-and-queues): add on-demand fallback example - #1680
Merged
Merged
Conversation
LeonKolyang
requested review from
EngHabu,
cosmicbboy,
kumare3,
ppiegaze and
samhita-alla
as code owners
October 5, 2026 11:33
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The caller also runs on spot compute, preventing the fallback from executing when spot capacity is unavailable.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Adds guidance for falling back from unavailable spot capacity to on-demand compute.
Changes:
- Documents spot-capacity queue behavior and
max_queued_time. - Adds an asynchronous fallback example.
- Links to the equivalent GPU-capacity pattern.
| File | Description |
|---|---|
content/user-guide/tasks/task-configuration/interruptible-tasks-and-queues.md |
Adds spot-capacity fallback documentation and example. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| return {"accuracy": 0.95} | ||
|
|
||
|
|
||
| @env.task |
Collaborator
There was a problem hiding this comment.
GHA build & deploy previewBuilt by
Updated automatically on every push. |
Signed-off-by: LeonKolyang <leon.menkreo@googlemail.com>
Signed-off-by: LeonKolyang <leon.menkreo@googlemail.com>
LeonKolyang
force-pushed
the
docs/spot-ondemand-fallback
branch
from
October 5, 2026 11:47
4a4a990 to
cfc0378
Compare
fiedlerNr9
approved these changes
Oct 5, 2026
…ound (#1682) docs(interruptible-tasks): run the fallback parent on-demand, clear its queue bound - The parent was in the interruptible environment, so with no spot capacity it could not start and never reached the fallback. It now runs with interruptible=False. - override() keeps the task's timeout, so the on-demand retry inherited the 10-minute max_queued_time. Pass timeout=flyte.Timeout() to clear it (timeout=0 does not, since override uses `timeout or self.timeout`). - Note that max_queued_time is per attempt, so retries multiply the wait before the fallback. - Fix the existing invocation-time snippet: an async parent calling sync tasks without .aio() raises SyncTaskCallInAsyncContextError. - Task-shaped heading, shorter sentences, "the parent", "can sit", and typed list inputs to avoid the pickle fallback. Both snippets verified with a local run on flyte 2.10.7. Signed-off-by: Peeter Piegaze <1153481+ppiegaze@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Adds a "Behavior if spot instances are not available" section to the interruptible
tasks page.
The page covered what happens when a spot instance is reclaimed, but not what
happens when the spot pool has no capacity to begin with — the task sits in
Queued / Waiting for resources indefinitely. The new section shows bounding that
wait with
max_queued_timeand catchingMaxQueuedTimeExceededErrorto re-runthe task on on-demand compute via
override(interruptible=False), and links tothe equivalent GPU-capacity pattern in the error-handling guide.
The example uses
async defrather than thedefin the surrounding snippets:calling a sync task from inside an async task raises
SyncTaskCallInAsyncContextError.Verified against a live cluster (flyte 2.10.8.dev2): the spot action ends
TIMED_OUT with code
QUEUED_TIMEOUT_EXCEEDEDand the on-demand retry succeeds.