Skip to content

Streaming + autoreleasepool drain + memcheck harness - #17

Merged
Jud merged 17 commits into
mainfrom
memory-wins
May 14, 2026
Merged

Jud merged 17 commits into
mainfrom
memory-wins

Conversation

@Jud

@Jud Jud commented May 14, 2026

Copy link
Copy Markdown
Owner

Summary

Bundles three previously-merged-elsewhere streams of memory work into a single PR against main. Targets #11 (iPhone SE jetsam) by removing per-utterance buffer accumulation and giving us a CLI memcheck harness for future regressions. No model re-export required — pure Swift + scripts.

What's in the box (17 commits)

Streaming speak() — Per-chunk peak normalization replaced with a fixed gain locked from the first chunk (allows attenuation, peak-rescales overshooting chunks instead of hard-clipping). 100 ms inter-chunk silence buffers allocated fresh per gap (no shared mutable singleton on the public API). Producer paces to within 2 s of a real-time consumer using CFAbsoluteTimeGetCurrent, sleeping in 50 ms increments so onTermination tears the worker down promptly. New paceToRealtime: Bool = true parameter on speak() lets batch consumers (file export, tests) opt out.

KokoroApp playback — AudioEngine rewritten around an AsyncStream<SpeakEvent> with a generation counter for race-safe start()/stop(). Tap installed at 24 kHz on the player node (was incorrectly at 48 kHz via mainMixerNode).

autoreleasepool drain — Each chunk iteration in streamSpeakLoop/synthesizeTokens/synthesizeTokensWithEmbedding wraps in autoreleasepool so CoreML's autoreleased MLMultiArray/MLFeatureProvider/NSError instances don't accumulate for the duration of the utterance. Cuts simulator per-suite duration ~17%.

CLI memcheck harness — scripts/memcheck.sh boots an iPhone SE (3rd gen) simulator, builds and installs the app, runs a deterministic 3-voice × 3-length suite while sampling phys_footprint at 10 Hz, parses the fenced JSON dump, and asserts peak/delta/residual budgets and case-completion counts. Catches both leaks (post-teardown residual) and regressions (peak run-over-run). Documented that simulator memory is NOT a faithful jetsam proxy — use for relative comparison, not absolute thresholds.

Review chain

Every unit went through audit → optional codex → comment cleanup → chunk-close simplify + codex. Notable findings caught and fixed:

  • Stale .dataPlayedBack callbacks decrementing a generation-stale counter (audit).
  • Consumer-side player-queue cap caused underruns when a single chunk exceeded the 2 s cap — removed in favor of producer-only pacing (chunk-close codex).
  • Thread.sleep ignored cancellation — replaced with 50 ms-incremented polling (chunk-close codex).
  • Magic numbers (0.95, 0.001, 0.99) extracted to named constants so streaming and batch paths share one source of truth (chunk-close simplify).

What this PR does not do

  • No iOS Simulator memory drop (peak is structural MPSGraph working set, ~4 GB on sim). Real-device numbers are typically 10-20× lower; this PR's leak-and-residual checks pass on the simulator, but the absolute peak budget validation needs real hardware.
  • No model export changes. Hybrid 4-bit/8-bit palettization and lazy voice packs (~24 MB resident win) remain queued for a separate PR.
  • No forceCPU: true improvements — that path needs profiling on real hardware to identify the dominant memory consumer before any code or export changes.

Test plan

  • swift build — green locally.
  • scripts/memcheck.sh on iPhone SE (3rd gen) simulator — peak/delta/residual within budgets, all 9 cases complete with audio.
  • CLI streaming: kokoro say --stream "<long text>" --voice af_heart — listen for chunk-boundary stutter (should be unchanged), no producer-side runaway.
  • KokoroApp on real device — start/stop several times mid-utterance, verify clean teardown and no leaked audio session.
  • Outstanding: real-device iPhone SE memory measurement to validate the issue it gets killed by the system due to memory issues #11 closeout.

🤖 Generated with Claude Code

Jud and others added 17 commits May 14, 2026 12:02
Per-chunk peak normalization in speak() caused chunk-to-chunk loudness
jumps. Compute the gain once from the first chunk (clamped 1.0-2.0) and
apply it to every subsequent chunk so amplitude stays consistent across
the utterance. Emit a 100ms silence buffer between chunks to match the
cadence of the non-streaming synthesize() path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replace the AVAudioPlayer + WAV-buffer flow with AVAudioEngine +
AVAudioPlayerNode driven by engine.speak()'s AsyncStream. PCM buffers
get scheduled directly on the player node as they arrive. A tap on the
main mixer drives the spectrum analyzer at the actual playback rate, so
the waveform animation no longer needs to track elapsed time against a
held [Float] of the whole utterance.

Drops three full waveform copies (raw samples, encoded WAV, AVAudioPlayer
internal buffer) — peak audio memory goes from O(utterance) to O(chunk).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Reset isSynthesizing in the synthesis-throw error path so the UI
  doesn't get stuck.
- Tap on the player node at 24kHz instead of the main mixer's
  output bus, so the spectrum analyzer (constructed at 24kHz) sees
  buffers at its expected sample rate. The mixer-bus tap was
  delivering at the device's mix rate (e.g. 48kHz), which compressed
  the band-frequency mapping toward the low end.
- Guard run() and the sentinel completion against a stop()-then-speak()
  race by checking self.playerNode === player. A stale sentinel
  callback firing after a new playback started would otherwise tear
  down the new engine.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Stale stream tasks resuming after stop()/speak() race were stomping
isSynthesizing and error state on the new utterance, leaving the UI
stuck mid-state. Gate every shared-state mutation in run() on
playerNode === player so a cancelled task only writes its own state.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Dedup the two speak() overloads into a shared streamSpeakLoop helper
  parameterized by a style-vector closure. Removes ~45 lines of
  byte-identical copy-paste plus the streamGain force-unwrap.
- Promote the inter-chunk silence buffer from per-call allocation to
  a static singleton — it's read-only after init, so AVAudioPlayerNode
  can enqueue the same instance for every gap.
- Replace AudioEngine's five playerNode === player race guards with a
  single generation counter. The identity check leaked implementation
  detail across method boundaries; the counter is one named predicate.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Move engine.speak() onto a detached task. Phonemization runs
  synchronously inside speak() before the worker thread starts; on
  long text that was happening on MainActor and could freeze the UI
  for a few hundred ms before the first audio buffer arrived.
- Flip isSynthesizing to false on the first audio buffer instead of
  after the stream finishes. ContentView only transitions to
  .speaking when isSynthesizing turns false, so multi-chunk streams
  were showing the synthesizing UI while audio was already playing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
scripts/memcheck.sh boots an iPhone SE (3rd gen) simulator (creates
one if needed), builds and installs the app, launches with
--memory-test, parses the JSON dump, and asserts peak phys_footprint
stays within budget. Optional XCTRACE=1 also records an Allocations
trace.

App side: --memory-test launch arg routes KokoroAppApp to a minimal
test view that loads the engine, runs a deterministic 3-voice ×
3-length synthesis suite while sampling phys_footprint at 10Hz, prints
a JSON summary fenced with MEMORY_TEST_RESULT_START/_END markers, and
exits.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wrap each iteration of streamSpeakLoop, synthesizeTokens, and
synthesizeTokensWithEmbedding in autoreleasepool. CoreML allocates
MLMultiArrays, MLFeatureProviders, NSError instances etc. that get
autoreleased rather than ARC-released. Without an explicit pool the
producer thread's pool only drains at thread exit, so intermediate
state accumulates for the entire utterance.

In the iPhone SE simulator memcheck this cuts per-suite duration ~17%
(53s → 44s) by reducing memory pressure between chunks. Doesn't drop
the absolute peak — that's CoreML's GPU/MPS working set on the
simulator, which counts in phys_footprint differently than ANE on
device.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Document that simulator memory is not a faithful proxy for device:
  CoreML's GPU/MPS working set on simulator is 10-20x ANE on hardware.
- Default budgets calibrated to simulator-realistic numbers so the
  script flags regressions, not absolute jetsam risk.
- Add LEAK_MB residual check: post-teardown phys_footprint should
  return near baseline. Default 150 MB. Catches accumulated state
  that doesn't release between runs.
- Document exit codes.

First passing run: peak 4 GB (simulator artifact), residual +36 MB,
streaming drains between cases. Use real hardware for jetsam
validation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Auto-generate the xcodeproj via xcodegen if missing. The .xcodeproj
  is .gitignored, so a clean checkout doesn't have one — without this
  the script bails before running.
- Fail the gate if any error or chunk_failed event was logged, or if
  case_end/first_buffer counts don't match case_start. A clean sim
  without bundled models would emit error events and otherwise pass
  on low memory; that silent pass is now caught.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The streaming path had no upstream bound: the producer ran synthesis
as fast as the CoreML stack allowed, the AsyncStream buffered every
yield unbounded, and the consumer scheduled each buffer on
AVAudioPlayerNode immediately. On a fast-synthesis-slow-playback
device (long utterance, real-time playback) all three queues could
hold the full utterance simultaneously.

Two fixes, one per side:

- Producer paces in streamSpeakLoop. After each yield, track total
  audio produced vs. wall-clock elapsed; if we're more than
  producerLeadSeconds ahead of a hypothetical real-time consumer,
  sleep the difference before the next chunk. Bounds the AsyncStream
  buffer regardless of consumer behavior.

- Consumer caps the player queue. AudioEngine tracks queuedFrames,
  awaits a CheckedContinuation when the queue exceeds maxQueuedSeconds,
  resumes from the .dataPlayedBack callback. teardown() resumes any
  pending continuation so stop() never strands the producer task.

The memcheck suite consumes the stream without playback, so its
wall time now ≈ total synthesized audio (~4 min for the 3×3 suite).
Bump TIMEOUT_S default 180 → 480.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Generation-guard the .dataPlayedBack callback's Task. Without this,
stale callbacks from a previous speak() generation continue firing
after teardown() resets queuedFrames to 0, driving the counter
negative. The next speak()'s enqueue then schedules over the cap
until the deficit is paid back.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The fixed-gain streaming path had two failure modes a quiet first
chunk could trigger:

- Gain floor of 1.0 prevented ever attenuating, so a naturally loud
  first chunk could push the whole utterance toward clipping.
- After-gain hard clip at ±0.95 caused audible distortion whenever
  a later chunk's peak exceeded the budget locked in from chunk one.

Allow attenuation by dropping the floor to a numerical safety value
(0.01) and replace the per-sample hard clip with a per-chunk peak
rescale: if the chunk peaks above 1.0 after gain, scale it down to
±0.99. This keeps the chunk-to-chunk loudness steady for normal
content (every TTS chunk lands inside the budget) and gives the rare
overshooting chunk a smooth attenuation instead of a clipped
waveform.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Trim doc comments that narrated the prior hard-clip behavior — kept
the WHY (rescale for overshoot, attenuation headroom) and removed
the historical "before this change" framing that rots.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The shared static interChunkSilenceBuffer let SpeakEvent.audio yield
the same AVAudioPCMBuffer instance to every consumer for every gap.
For an internal helper that was fine, but the buffer crosses a
public API surface — any caller retaining a yielded `.audio` payload
shares state with future calls.

Replace the static let with a makeSilenceBuffer() factory. One
allocation per gap (~9.6 KB at 100 ms / 24 kHz) — negligible next to
chunk-sized PCM buffers — and SpeakEvent payloads are now unaliased.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Name the peak-related magic numbers (targetPeakAmplitude 0.95,
  silencePeakThreshold 0.001, overshootRescalePeak 0.99) so the
  streaming and batch paths share a single source of truth instead
  of repeating the literal in each call site.
- makeSilenceBuffer now flows through makePCMBuffer(from:format:)
  instead of re-rolling AVAudioPCMBuffer alloc + memset.
- Switch producer pacing to CFAbsoluteTimeGetCurrent (monotonic),
  matching how KokoroSay times its streamPlayback path. Date() is
  wall-clock and vulnerable to NTP skew.
- Drop the redundant `audioProduced > 0` guard — when nothing has
  been produced, `lead = -elapsed` is already < producerLeadSeconds
  and the inner branch is skipped naturally.
- Add a `paceToRealtime: Bool = true` parameter to both speak()
  overloads. Real-time consumers (AudioEngine, KokoroSay
  --stream) keep the default; batch consumers (MemoryTestRunner,
  future file-export paths) opt out and run synthesis at full
  throughput. Lets the memcheck suite finish in ~real wall time
  again (revert TIMEOUT_S 480 → 240).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
[P2] Drop the consumer-side player-queue cap in AudioEngine. Single
chunks can run up to ~13 s (maxTokens=512), well over the 2 s cap;
the cap caused the consumer to await the entire chunk's
.dataPlayedBack before scheduling the next, leaving the player
silent at chunk boundaries. The producer-side wall-time pacing
already bounds in-flight buffers under the steady-state assumption
(real-time consumer), and that's the AudioEngine path's
assumption — relying on it removes the underrun.

[P2] Make the producer pacing sleep cancellation-aware. The previous
Thread.sleep(forTimeInterval:) blocked until expiry no matter what
onTermination's Thread.cancel() did; for long-chunk live playback a
user stop() could leave the worker thread alive for seconds. Sleep
in 50 ms increments and check isCancelled between them.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Jud
Jud merged commit 0d5a21e into main May 14, 2026
2 checks passed
@Jud
Jud deleted the memory-wins branch May 14, 2026 18:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant