Skip to content

RDKEMW-25196: Harden force-stop to SIGKILL the whole container cgroup - #474

Merged
B-Larsen merged 1 commit into
release/v3.19from
topic/RDKEMW-25196
Oct 1, 2026
Merged

B-Larsen merged 1 commit into
release/v3.19from
topic/RDKEMW-25196

Conversation

@ks734

@ks734 ks734 commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

DobbyManager::stopContainer(cd, withPrejudice=true) only sent SIGKILL to the container's tracked init process, not the whole cgroup, and trusted that the runc CLI call succeeding meant the container was actually dead. If a crashed/hung process left orphaned descendants that init can't reap, or a process is stuck in an uninterruptible sleep, the container could be left running indefinitely with no indication the force-stop failed.

Add DobbyManager::forceKillContainerAndVerify(), used by both force-kill paths in stopContainer(), which sends SIGKILL via killCont(..., all=true) and polls DobbyRunC::state() for up to ~500ms to confirm the container actually stopped, logging clearly and returning false if it didn't.

Add L1 tests covering both the all=true behaviour and the case where the container stays Running after SIGKILL. Update the daemon-core spec to document the new behaviour.

Description

What does this PR change/fix and why?

If there is a corresponding JIRA ticket, please ensure it is in the title of the PR.

Test Procedure

How to test this PR (if applicable)

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Other (doesn't fit into the above categories - e.g. documentation updates)

Requires Bitbake Recipe changes?

  • The base Bitbake recipe (meta-rdk-ext/recipes-containers/dobby/dobby.bb) must be modified to support the changes in this PR (beyond updating SRC_REV)

DobbyManager::stopContainer(cd, withPrejudice=true) only sent SIGKILL to
the container's tracked init process, not the whole cgroup, and trusted
that the runc CLI call succeeding meant the container was actually dead.
If a crashed/hung process left orphaned descendants that init can't reap,
or a process is stuck in an uninterruptible sleep, the container could be
left running indefinitely with no indication the force-stop failed.

Add DobbyManager::forceKillContainerAndVerify(), used by both force-kill
paths in stopContainer(), which sends SIGKILL via killCont(..., all=true)
and polls DobbyRunC::state() for up to ~500ms to confirm the container
actually stopped, logging clearly and returning false if it didn't.

Add L1 tests covering both the all=true behaviour and the case where the
container stays Running after SIGKILL. Update the daemon-core spec to
document the new behaviour.
@ks734 ks734 changed the title RDKEMW-25126: Harden force-stop to SIGKILL the whole container cgroup RDKEMW-25196: Harden force-stop to SIGKILL the whole container cgroup Sep 21, 2026
@B-Larsen
B-Larsen merged commit 35efc62 into release/v3.19 Oct 1, 2026
39 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 1, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants