Repository navigation
Fix graceful engine shutdown on SIGTERM and SIGINT - #416
Merged
Merged
Conversation
|
Sorry @cooktheryan, you've used your own review budget of 250,000 diff characters for the last 7 days. You can request another review in 5 days and 22 hours by commenting |
Signed-off-by: Ryan Cook <rcook@redhat.com>
cooktheryan
force-pushed
the
fix/287-graceful-shutdown
branch
from
October 9, 2026 15:18
1928041 to
167a1e6
Compare
Reviewer's GuideThe PR changes FetchIt from a signal-isolated Bash child into a directly signaled process, then propagates a cancelable engine lifetime through initialization, scheduling, and reloads. Shutdown cancels new and active work, waits for the scheduler without blocking reload coordination, exits successfully after cooperative completion or a bounded five-second grace period, and is covered by unit, race, and built-image integration checks. Sequence diagram for graceful FetchIt shutdownsequenceDiagram
participant Podman
participant Entry as entry.sh
participant FetchIt
participant Scheduler
participant Jobs
Podman->>Entry: SIGTERM or SIGINT
Entry->>FetchIt: forward signal via exec
FetchIt->>FetchIt: signal.NotifyContext()
FetchIt->>FetchIt: cancel lifetime context
FetchIt->>Scheduler: Stop()
Scheduler->>Jobs: cancel and stop scheduling
alt jobs finish within 5 seconds
Jobs-->>Scheduler: completed
Scheduler-->>FetchIt: shutdown complete
FetchIt-->>Podman: exit status 0
else work remains after grace period
FetchIt-->>Podman: log grace-period expiry
FetchIt-->>Podman: exit status 0
end
Flow diagram for cancellation-aware engine startup and reloadflowchart TD
Signal["SIGTERM or SIGINT"] --> Cancel["Cancel engine lifetime context"]
Cancel --> BlockNew["Prevent queued methods and reloads from starting"]
BlockNew --> StopScheduler["scheduler.Stop()"]
StopScheduler --> Wait["Wait without holding configRestartMu"]
Wait --> Complete["Exit 0 after cooperative completion"]
Wait --> Expire["After five-second grace period"]
Expire --> Exit["Log expiry and exit 0"]
Cancel --> Startup["Cancel initialization and startup Git operations"]
Startup --> Exit
File-Level Changes
Assessment against linked issues
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
podman container stop fetchitcurrently waits for the stop timeout and escalates to SIGKILL because Bash runs the engine as a child without forwarding signals. FetchIt also has no termination handler. This change makes the entry script exec FetchIt and handles SIGTERM/SIGINT before engine initialization, so normal shutdown exits promptly with status 0.An engine lifetime context survives configuration reloads, cancels jobs and the Podman client, and prevents canceled reloads or queued methods from starting new work. Shutdown waits for the scheduler without holding the reload mutex and allows at most five seconds for startup or running jobs to finish. If work does not honor cancellation, the engine logs grace-period expiry and exits; this does not guarantee completion or rollback of arbitrary host operations. Stopping the engine leaves deployed workloads running.
Adds unit coverage for running jobs, uncooperative jobs, canceled methods/reloads, initialization cancellation, and a blocked startup/reload lock. The image workflow now checks SIGTERM/SIGINT against the built image, and the unit workflow includes shutdown tests in its race checks. Running guidance and unreleased notes describe the behavior.
Validation:
git diff --check, and Sphinx HTML build with warnings as errors passed.cleanupOnRemovalenabled.The original shutdown hang reproduces on both published latest and current main. On Podman 5.8.7 the stop command itself already returns 0; the forced container exit was 137. Podman sends SIGTERM first and escalates to SIGKILL at the timeout.
Fixes #287.
Summary by Sourcery
Enable FetchIt to terminate gracefully on SIGTERM and SIGINT while bounding shutdown time and preventing new work from starting.
New Features:
Bug Fixes:
Enhancements:
CI:
Documentation:
Tests: