fix(tts/kokoro-ane): adopt COLA-corrected KokoroTail_v2 + native output level for all variants - #868
Conversation
…ut level for all variants Point all three variants (ANE/, ANE-zh/, ANE-ja/) at KokoroTail_v2.mlmodelc and drop peak normalization for English/Mandarin, closing the follow-up scoped in #699 and reported in #852. The v1 tail omits the overlap-add/COLA normalization torch.istft applies (interior envelope sum(w^2) = 1.5 for periodic Hann, win 20 / hop 5), leaving raw output exactly 1.5x the PyTorch reference. #699 fixed the conversion and moved Japanese to native output level, but the corrected tail was never published — .japanese has been shipping unmasked 1.5x-hot audio. The corrected tails are now live as KokoroTail_v2.mlmodelc alongside the originals (kokoro-82m-coreml commit acac8811; conversion fix recommitted in mobius #83, reconstruction matches the #699-measured artifact to 1 ulp). The _v2 rename (same pattern as KokoroNoise_v2) doubles as cache invalidation: DownloadUtils skips files that already exist and the weight.bin size is unchanged, so a same-path replacement would never reach existing caches. With the rename, existing caches fetch just the missing file on next ensureModels. Verified e2e on M5 Pro: existing cache self-healed (only KokoroTail_v2 downloaded per variant); native render levels en peak 0.343 / zh 0.400 / ja 0.354 — reference territory (#699 measured PyTorch jf_alpha at 0.299), vs 1.0 forced by the old 0 dBFS normalization. CoreML A/B of v1 vs v2 tails: interior per-sample ratio 1.500000, corr 1.0 (pure scalar, no spectral change), for both the en/ja and zh tails. Fixes #852.
PocketTTS Smoke Test ✅
Runtime: 0m32s Note: PocketTTS uses CoreML MLState (macOS 15) KV cache + Mimi streaming state. CI VM lacks physical GPU — audio quality and performance may differ from Apple Silicon. |
Parakeet EOU Benchmark Results ✅Status: Benchmark passed Performance Metrics
Streaming Metrics
Test runtime: 2m21s • 08/19/2026, 10:53 AM EST RTFx = Real-Time Factor (higher is better) • Processing includes: Model inference, audio preprocessing, state management, and file I/O |
Offline VBx Pipeline ResultsSpeaker Diarization Performance (VBx Batch Mode)Optimal clustering with Hungarian algorithm for maximum accuracy
Offline VBx Pipeline Timing BreakdownTime spent in each stage of batch diarization
Speaker Diarization Research ComparisonOffline VBx achieves competitive accuracy with batch processing
Pipeline Details:
🎯 Offline VBx Test • AMI Corpus ES2004a • 1049.0s meeting audio • 139.4s processing • Test runtime: 2m 43s • 08/19/2026, 10:55 AM EST |
Speaker Diarization Benchmark ResultsSpeaker Diarization PerformanceEvaluating "who spoke when" detection accuracy
Diarization Pipeline Timing BreakdownTime spent in each stage of speaker diarization
Speaker Diarization Research ComparisonResearch baselines typically achieve 18-30% DER on standard datasets
Note: RTFx shown above is from GitHub Actions runner. On Apple Silicon with ANE:
🎯 Speaker Diarization Test • AMI Corpus ES2004a • 1049.0s meeting audio • 68.3s diarization time • Test runtime: 4m 50s • 08/19/2026, 10:56 AM EST |
Supertonic3 Smoke Test ✅
Runtime: 0m37s Note: CI VMs lack a physical Neural Engine; the ANE-bucketed VectorEstimator falls back to CPU here. This validates download + variant resolution + synthesis, not ANE residency/perf. |
VAD Benchmark ResultsPerformance Comparison
Dataset Details
✅: Average F1-Score above 70% |
Sortformer High-Latency Benchmark ResultsES2004a Performance (30.4s latency config)
Sortformer High-Latency • ES2004a • Runtime: 4m 30s • 2026-08-19T15:06:17.316Z |
ASR Benchmark Results ✅Status: All benchmarks passed Parakeet v3 (multilingual)
Parakeet v2 (English-optimized)
Streaming (v3)
Streaming (v2)
Streaming tests use 5 files with 0.5s chunks to simulate real-time audio streaming 25 files per dataset • Test runtime: 13m59s • 08/19/2026, 11:06 AM EST RTFx = Real-Time Factor (higher is better) • Calculated as: Total audio duration ÷ Total processing time Expected RTFx Performance on Physical M1 Hardware:• M1 Mac: ~28x (clean), ~25x (other) Testing methodology follows HuggingFace Open ASR Leaderboard |
…-benchmark (#870) Follow-up to #868 (issue #852): two CLI artifact writers still peak-normalized KokoroAne output to 0 dBFS via the `AudioWAV.data` default, contradicting the native-level behavior #868 shipped. - **`TTSAsrVerifyCommand`** (KokoroAne-only): persisted `--audio-dir` WAVs now write at native level via `normalize: false`. - **`TtsBenchmarkCommand`**: the shared phrase loop takes an explicit `normalizeWavs` per-backend policy — `kokoro-ane` writes native level; `pocket-tts` / `styletts2` / `supertonic3` keep peak normalization, matching each backend's own shipping output path. ASR scoring is level-invariant, so benchmark WER/CER numbers are unaffected; only the persisted WAV levels change. `swift build` + `swift format lint` clean.
Closes #852. Completes the follow-up scoped in #699.
Background
#699 found that
CoreMLCustomSTFT(KokoroTail) omits the overlap-add/COLA normalizationtorch.istftapplies — the interior envelope is a constantsum(w²) = 1.5(periodic Hann, win 20 / hop 5), so raw output is exactly 1.5× the PyTorch reference. The PR fixed the conversion and moved.japaneseto native output level, but as #852 correctly reports, the corrected tail never made it to HuggingFace (and the script fix never made it into mobius):.japanesehas been shipping unmasked 1.5×-hot audio, while English/Mandarin were protected only by peak normalization.The fix is a pure weight rescale (iSTFT deconv kernels ÷ 1.5) —
model.milis unchanged, which is why the issue's mil-graph analysis couldn't see it; it shows only in theweight.binLFS oid.What landed where
FluidInference/kokoro-82m-coremlcommitacac8811— corrected tails published asKokoroTail_v2.mlmodelcalongside the originals inANE/,ANE-ja/,ANE-zh/(originals kept; consumers opt in by pointer change). The en/ja tail is the artifact feat(tts/kokoro-ane): Japanese variant + PyTorch-matched output level #699 measured at 1.02× PyTorch; the zh tail carries the same corrected kernel spans (the iSTFT deconv constants are deterministic DFT×window kernels, byte-identical between en and zh).KokoroTail_v2.mlmodelcand drops peak normalization for English/Mandarin (normalize: falseinKokoroAneManager.wavDataand the CLI), so all variants write at reference-accurate native level.The
_v2rename (same pattern asKokoroNoise_v2) doubles as cache invalidation:DownloadUtilsskips files that already exist and the weight.bin size is unchanged, so a same-path replacement would never reach existing caches. With the rename, existing caches fetch just the missing file on the nextensureModels.Validation
KokoroTail_v2.mlmodelcper variant; native render peaks en 0.343 / zh 0.400 / ja 0.354 (reference territory — feat(tts/kokoro-ane): Japanese variant + PyTorch-matched output level #699 measured PyTorchjf_alphaat 0.299), vs the 1.0 previously forced by 0 dBFS normalization.swift build+swift format lintclean.Behavior change
KokoroAne WAV output is no longer slammed to 0 dBFS for English/Mandarin — output now sits at the model's native (PyTorch-matched) level, consistent with
.japanese, LuxTTS, and NeuTTS. Downstream code that assumed peak-normalized output will hear quieter (correct) levels.🤖 Generated with Claude Code