Repository navigation
Streaming + autoreleasepool drain + memcheck harness - #17
Merged
Merged
Conversation
Per-chunk peak normalization in speak() caused chunk-to-chunk loudness jumps. Compute the gain once from the first chunk (clamped 1.0-2.0) and apply it to every subsequent chunk so amplitude stays consistent across the utterance. Emit a 100ms silence buffer between chunks to match the cadence of the non-streaming synthesize() path. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replace the AVAudioPlayer + WAV-buffer flow with AVAudioEngine + AVAudioPlayerNode driven by engine.speak()'s AsyncStream. PCM buffers get scheduled directly on the player node as they arrive. A tap on the main mixer drives the spectrum analyzer at the actual playback rate, so the waveform animation no longer needs to track elapsed time against a held [Float] of the whole utterance. Drops three full waveform copies (raw samples, encoded WAV, AVAudioPlayer internal buffer) — peak audio memory goes from O(utterance) to O(chunk). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Reset isSynthesizing in the synthesis-throw error path so the UI doesn't get stuck. - Tap on the player node at 24kHz instead of the main mixer's output bus, so the spectrum analyzer (constructed at 24kHz) sees buffers at its expected sample rate. The mixer-bus tap was delivering at the device's mix rate (e.g. 48kHz), which compressed the band-frequency mapping toward the low end. - Guard run() and the sentinel completion against a stop()-then-speak() race by checking self.playerNode === player. A stale sentinel callback firing after a new playback started would otherwise tear down the new engine. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Stale stream tasks resuming after stop()/speak() race were stomping isSynthesizing and error state on the new utterance, leaving the UI stuck mid-state. Gate every shared-state mutation in run() on playerNode === player so a cancelled task only writes its own state. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Dedup the two speak() overloads into a shared streamSpeakLoop helper parameterized by a style-vector closure. Removes ~45 lines of byte-identical copy-paste plus the streamGain force-unwrap. - Promote the inter-chunk silence buffer from per-call allocation to a static singleton — it's read-only after init, so AVAudioPlayerNode can enqueue the same instance for every gap. - Replace AudioEngine's five playerNode === player race guards with a single generation counter. The identity check leaked implementation detail across method boundaries; the counter is one named predicate. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Move engine.speak() onto a detached task. Phonemization runs synchronously inside speak() before the worker thread starts; on long text that was happening on MainActor and could freeze the UI for a few hundred ms before the first audio buffer arrived. - Flip isSynthesizing to false on the first audio buffer instead of after the stream finishes. ContentView only transitions to .speaking when isSynthesizing turns false, so multi-chunk streams were showing the synthesizing UI while audio was already playing. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
scripts/memcheck.sh boots an iPhone SE (3rd gen) simulator (creates one if needed), builds and installs the app, launches with --memory-test, parses the JSON dump, and asserts peak phys_footprint stays within budget. Optional XCTRACE=1 also records an Allocations trace. App side: --memory-test launch arg routes KokoroAppApp to a minimal test view that loads the engine, runs a deterministic 3-voice × 3-length synthesis suite while sampling phys_footprint at 10Hz, prints a JSON summary fenced with MEMORY_TEST_RESULT_START/_END markers, and exits. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wrap each iteration of streamSpeakLoop, synthesizeTokens, and synthesizeTokensWithEmbedding in autoreleasepool. CoreML allocates MLMultiArrays, MLFeatureProviders, NSError instances etc. that get autoreleased rather than ARC-released. Without an explicit pool the producer thread's pool only drains at thread exit, so intermediate state accumulates for the entire utterance. In the iPhone SE simulator memcheck this cuts per-suite duration ~17% (53s → 44s) by reducing memory pressure between chunks. Doesn't drop the absolute peak — that's CoreML's GPU/MPS working set on the simulator, which counts in phys_footprint differently than ANE on device. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Document that simulator memory is not a faithful proxy for device: CoreML's GPU/MPS working set on simulator is 10-20x ANE on hardware. - Default budgets calibrated to simulator-realistic numbers so the script flags regressions, not absolute jetsam risk. - Add LEAK_MB residual check: post-teardown phys_footprint should return near baseline. Default 150 MB. Catches accumulated state that doesn't release between runs. - Document exit codes. First passing run: peak 4 GB (simulator artifact), residual +36 MB, streaming drains between cases. Use real hardware for jetsam validation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Auto-generate the xcodeproj via xcodegen if missing. The .xcodeproj is .gitignored, so a clean checkout doesn't have one — without this the script bails before running. - Fail the gate if any error or chunk_failed event was logged, or if case_end/first_buffer counts don't match case_start. A clean sim without bundled models would emit error events and otherwise pass on low memory; that silent pass is now caught. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The streaming path had no upstream bound: the producer ran synthesis as fast as the CoreML stack allowed, the AsyncStream buffered every yield unbounded, and the consumer scheduled each buffer on AVAudioPlayerNode immediately. On a fast-synthesis-slow-playback device (long utterance, real-time playback) all three queues could hold the full utterance simultaneously. Two fixes, one per side: - Producer paces in streamSpeakLoop. After each yield, track total audio produced vs. wall-clock elapsed; if we're more than producerLeadSeconds ahead of a hypothetical real-time consumer, sleep the difference before the next chunk. Bounds the AsyncStream buffer regardless of consumer behavior. - Consumer caps the player queue. AudioEngine tracks queuedFrames, awaits a CheckedContinuation when the queue exceeds maxQueuedSeconds, resumes from the .dataPlayedBack callback. teardown() resumes any pending continuation so stop() never strands the producer task. The memcheck suite consumes the stream without playback, so its wall time now ≈ total synthesized audio (~4 min for the 3×3 suite). Bump TIMEOUT_S default 180 → 480. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Generation-guard the .dataPlayedBack callback's Task. Without this, stale callbacks from a previous speak() generation continue firing after teardown() resets queuedFrames to 0, driving the counter negative. The next speak()'s enqueue then schedules over the cap until the deficit is paid back. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The fixed-gain streaming path had two failure modes a quiet first chunk could trigger: - Gain floor of 1.0 prevented ever attenuating, so a naturally loud first chunk could push the whole utterance toward clipping. - After-gain hard clip at ±0.95 caused audible distortion whenever a later chunk's peak exceeded the budget locked in from chunk one. Allow attenuation by dropping the floor to a numerical safety value (0.01) and replace the per-sample hard clip with a per-chunk peak rescale: if the chunk peaks above 1.0 after gain, scale it down to ±0.99. This keeps the chunk-to-chunk loudness steady for normal content (every TTS chunk lands inside the budget) and gives the rare overshooting chunk a smooth attenuation instead of a clipped waveform. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Trim doc comments that narrated the prior hard-clip behavior — kept the WHY (rescale for overshoot, attenuation headroom) and removed the historical "before this change" framing that rots. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The shared static interChunkSilenceBuffer let SpeakEvent.audio yield the same AVAudioPCMBuffer instance to every consumer for every gap. For an internal helper that was fine, but the buffer crosses a public API surface — any caller retaining a yielded `.audio` payload shares state with future calls. Replace the static let with a makeSilenceBuffer() factory. One allocation per gap (~9.6 KB at 100 ms / 24 kHz) — negligible next to chunk-sized PCM buffers — and SpeakEvent payloads are now unaliased. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Name the peak-related magic numbers (targetPeakAmplitude 0.95, silencePeakThreshold 0.001, overshootRescalePeak 0.99) so the streaming and batch paths share a single source of truth instead of repeating the literal in each call site. - makeSilenceBuffer now flows through makePCMBuffer(from:format:) instead of re-rolling AVAudioPCMBuffer alloc + memset. - Switch producer pacing to CFAbsoluteTimeGetCurrent (monotonic), matching how KokoroSay times its streamPlayback path. Date() is wall-clock and vulnerable to NTP skew. - Drop the redundant `audioProduced > 0` guard — when nothing has been produced, `lead = -elapsed` is already < producerLeadSeconds and the inner branch is skipped naturally. - Add a `paceToRealtime: Bool = true` parameter to both speak() overloads. Real-time consumers (AudioEngine, KokoroSay --stream) keep the default; batch consumers (MemoryTestRunner, future file-export paths) opt out and run synthesis at full throughput. Lets the memcheck suite finish in ~real wall time again (revert TIMEOUT_S 480 → 240). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
[P2] Drop the consumer-side player-queue cap in AudioEngine. Single chunks can run up to ~13 s (maxTokens=512), well over the 2 s cap; the cap caused the consumer to await the entire chunk's .dataPlayedBack before scheduling the next, leaving the player silent at chunk boundaries. The producer-side wall-time pacing already bounds in-flight buffers under the steady-state assumption (real-time consumer), and that's the AudioEngine path's assumption — relying on it removes the underrun. [P2] Make the producer pacing sleep cancellation-aware. The previous Thread.sleep(forTimeInterval:) blocked until expiry no matter what onTermination's Thread.cancel() did; for long-chunk live playback a user stop() could leave the worker thread alive for seconds. Sleep in 50 ms increments and check isCancelled between them. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Bundles three previously-merged-elsewhere streams of memory work into a single PR against
main. Targets #11 (iPhone SE jetsam) by removing per-utterance buffer accumulation and giving us a CLI memcheck harness for future regressions. No model re-export required — pure Swift + scripts.What's in the box (17 commits)
Streaming
speak()— Per-chunk peak normalization replaced with a fixed gain locked from the first chunk (allows attenuation, peak-rescales overshooting chunks instead of hard-clipping). 100 ms inter-chunk silence buffers allocated fresh per gap (no shared mutable singleton on the public API). Producer paces to within 2 s of a real-time consumer usingCFAbsoluteTimeGetCurrent, sleeping in 50 ms increments soonTerminationtears the worker down promptly. NewpaceToRealtime: Bool = trueparameter onspeak()lets batch consumers (file export, tests) opt out.KokoroApp playback —
AudioEnginerewritten around anAsyncStream<SpeakEvent>with a generation counter for race-safestart()/stop(). Tap installed at 24 kHz on the player node (was incorrectly at 48 kHz viamainMixerNode).autoreleasepool drain — Each chunk iteration in
streamSpeakLoop/synthesizeTokens/synthesizeTokensWithEmbeddingwraps inautoreleasepoolso CoreML's autoreleasedMLMultiArray/MLFeatureProvider/NSErrorinstances don't accumulate for the duration of the utterance. Cuts simulator per-suite duration ~17%.CLI memcheck harness —
scripts/memcheck.shboots an iPhone SE (3rd gen) simulator, builds and installs the app, runs a deterministic 3-voice × 3-length suite while samplingphys_footprintat 10 Hz, parses the fenced JSON dump, and asserts peak/delta/residual budgets and case-completion counts. Catches both leaks (post-teardown residual) and regressions (peak run-over-run). Documented that simulator memory is NOT a faithful jetsam proxy — use for relative comparison, not absolute thresholds.Review chain
Every unit went through audit → optional codex → comment cleanup → chunk-close simplify + codex. Notable findings caught and fixed:
.dataPlayedBackcallbacks decrementing a generation-stale counter (audit).Thread.sleepignored cancellation — replaced with 50 ms-incremented polling (chunk-close codex).What this PR does not do
forceCPU: trueimprovements — that path needs profiling on real hardware to identify the dominant memory consumer before any code or export changes.Test plan
swift build— green locally.scripts/memcheck.shon iPhone SE (3rd gen) simulator — peak/delta/residual within budgets, all 9 cases complete with audio.kokoro say --stream "<long text>" --voice af_heart— listen for chunk-boundary stutter (should be unchanged), no producer-side runaway.🤖 Generated with Claude Code