Skip to content

GUI training hangs indefinitely at Epoch 1 (macOS) — deadlock in TF Eigen thread pool #995

Description

@Wallefer94

Describe the bug

Training via the CLI (python -m birdnet_analyzer.train ...) completes normally.
Training the exact same dataset via the GUI's Train tab hangs indefinitely.
The process prints:

Training model...
Training on 1694 samples, validating on 1 samples.
Epoch 1/50

...and then never progresses. No further output, no error, no crash. I waited
tens of minutes (vs. seconds for the equivalent CLI run) before concluding it
was genuinely stuck rather than just slow.

This may be related to #626, which reports a similar GUI training stall on Linux.

Diagnostic evidence

I captured a macOS sample of the stuck process (attached). It shows:

  • The main thread is blocked in every one of the 1635 captured samples at the
    same point: inside TensorFlow's eager execution path
    (TFE_Py_ExecuteCancelable → ProcessFunctionLibraryRuntime::RunSync),
    waiting on absl::Notification::WaitForNotification() →
    absl::Mutex::Block.
  • Meanwhile, the Eigen thread-pool worker threads (TensorFlow's internal CPU
    compute pool) are idle, sitting in Eigen::EventCount::CommitWait —
    i.e. waiting for work, not doing any.

This is the signature of a deadlock: the main thread is waiting for a
completion signal that the worker pool never sends, while the worker pool is
waiting for work it never receives.

To Reproduce

  1. Launch the GUI (cd BirdNET-Analyzer
    source venv-birdnet/bin/activate
    python -m birdnet_analyzer.gui)

  2. Go to the Train tab, point it at a training/testing dataset, leave default
    training settings

  3. Start training

  4. Observe it hang after the "Epoch 1/50" line

Expected behavior

Training should proceed through epochs the same way it does from the CLI.

Environment

  • OS: macOS 26.6.2 (25G83), Apple Silicon (ARM64)
  • BirdNET-Analyzer: commit 3286d78924ccb4413449f28f407b9fef842a17a1
    (birdnet_analyzer 2.4.0, birdnet 0.2.15)
  • TensorFlow: 2.21.0
  • Python: 3.11.14
  • tensorflow-metal: not installed
  • Installed in a venv (venv-birdnet), same venv used for both CLI and GUI

Additional context

Confirmed this is not a tensorflow-metal GPU-deadlock issue (a known,
separate bug pattern) since that package isn't installed here — training on
this machine runs CPU-only.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions