Repository navigation
feat(tts): Paradee-8M CoreML backend [beta] - #998
Conversation
Parakeet EOU Benchmark Results ✅Status: Benchmark passed Performance Metrics
Streaming Metrics
Test runtime: 1m9s • 10/07/2026, 10:01 PM EST RTFx = Real-Time Factor (higher is better) • Processing includes: Model inference, audio preprocessing, state management, and file I/O |
ASR Benchmark Results ✅Status: All benchmarks passed Parakeet v3 (multilingual)
Parakeet v2 (English-optimized)
Streaming (v3)
Streaming (v2)
Streaming tests use 5 files with 0.5s chunks to simulate real-time audio streaming 25 files per dataset • Test runtime: 9m2s • 10/07/2026, 09:59 PM EST RTFx = Real-Time Factor (higher is better) • Calculated as: Total audio duration ÷ Total processing time Expected RTFx Performance on Physical M1 Hardware:• M1 Mac: ~28x (clean), ~25x (other) Testing methodology follows HuggingFace Open ASR Leaderboard |
Supertonic3 Smoke Test ✅
Runtime: 0m29s Note: CI VMs lack a physical Neural Engine; the ANE-bucketed VectorEstimator falls back to CPU here. This validates download + variant resolution + synthesis, not ANE residency/perf. |
VAD Benchmark ResultsPerformance Comparison
Dataset Details
✅: Average F1-Score above 70% |
Offline VBx Pipeline ResultsSpeaker Diarization Performance (VBx Batch Mode)Optimal clustering with Hungarian algorithm for maximum accuracy
Offline VBx Pipeline Timing BreakdownTime spent in each stage of batch diarization
Speaker Diarization Research ComparisonOffline VBx achieves competitive accuracy with batch processing
Pipeline Details:
🎯 Offline VBx Test • AMI Corpus ES2004a • 1049.0s meeting audio • 158.3s processing • Test runtime: 2m 42s • 10/07/2026, 09:52 PM EST |
Speaker Diarization Benchmark ResultsSpeaker Diarization PerformanceEvaluating "who spoke when" detection accuracy
Diarization Pipeline Timing BreakdownTime spent in each stage of speaker diarization
Speaker Diarization Research ComparisonResearch baselines typically achieve 18-30% DER on standard datasets
Note: RTFx shown above is from GitHub Actions runner. On Apple Silicon with ANE:
🎯 Speaker Diarization Test • AMI Corpus ES2004a • 1049.0s meeting audio • 36.6s diarization time • Test runtime: 2m 5s • 10/07/2026, 10:09 PM EST |
Sortformer High-Latency Benchmark ResultsES2004a Performance (30.4s latency config)
Sortformer High-Latency • ES2004a • Runtime: 3m 30s • 2026-10-08T02:00:17.178Z |
PocketTTS Smoke Test ✅
Runtime: 0m9s Note: PocketTTS uses CoreML MLState (macOS 15) KV cache + Mimi streaming state. CI VM lacks physical GPU — audio quality and performance may differ from Apple Silicon. |
Paradee-8M v1.0 (Kokoro-82M distilled to 8.07M params, af_heart, 24 kHz
English) via FluidInference/paradee-8m-coreml (int8 default 12 MB, fp32 34 MB).
Two CoreML graphs with host-side steps in Swift:
- ParadeeText: ids -> durations, prosody features d, text features asr_tok
- host: max(1, round(dur / speed)) with half-to-even rounding (torch.round),
column expansion, seeded N(0,1) harmonic-source noise (InflectNoise)
- ParadeeAcoustic: F0/N predictor + iSTFTNet decoder + phase-lock filter
Text path reuses the KokoroAne English frontend (NeMo TN, Misaki lexicon,
BART G2P) and KokoroAneVocab (Paradee's vocab is byte-identical to Kokoro's).
Sentences are synthesized separately, as upstream does. The frontend's
lexicon keeps misaki's raw flap/glottal stop (ɾ, ʔ); misaki rewrites them to
T/t for Kokoro v1.0 and Paradee was trained only on that output, so the
Paradee text path applies the same rewrite. MiniMax English WER 1.76% ->
1.20%, CER 0.33% -> 0.14% ("kittens", "satellite", "patterns" were garbled).
.all/.cpuAndGPU are rejected: the LSTMs abort in MPSGraph (GPURNNOps JIT
not supported). cpuOnly default; cpuAndNeuralEngine also supported.
CLI: tts --backend paradee [--variant int8|fp32] [--speed] [--seed]
[--phonemes]; tts-benchmark --backend paradee.
tts-benchmark minimax-english (100 phrases, Parakeet round trip, M5 Pro,
macOS 27): int8 and fp32 both WER 1.20% / CER 0.14%, RTFx ~100 (text
frontend + both models), synth p50 77 ms; all-ane 95.6x, same WER.
335dddb to
42be1df
Compare
…es (#999) The English frontend reads the misaki lexicon raw, so it keeps the flap `ɾ`. misaki's own `G2P.__call__` rewrites it to `T` before Kokoro v1.0 sees it. Baseline Kokoro read "metal" as "mattle". - Apply `ɾ → T` in the English phonemizer path (also used by `englishPhonemes(for:)`) - Keep the glottal stop `ʔ`: misaki also rewrites it to `t`, but that made Kokoro read "button" as "butt" and "mitten" as "mit" **Results** (`tts-benchmark --backend kokoro-ane`, M5 Pro, macOS 27): | | 12 flap/glottal-dense sentences | minimax-english (100) | |---|---:|---:| | main | WER 3.67% / CER 0.56% | 0.68% / 0.17% | | `ɾ→T` + `ʔ→t` | 1.67% / 1.04% | 0.68% / 0.17% | | **this PR (`ɾ→T`)** | **0.00% / 0.00%** | 0.68% / 0.17% | **Reviewer notes:** the gain is small (one word in a targeted set; MiniMax unchanged). The Mandarin variant's English runs (`englishPhonemes(for:)`) also get the rewrite; that path is untested. Unit test added, not run locally (no XCTest here). Found while adding Paradee (#998), which needs both rewrites. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Note
Beta — API and hosted model assets may change.
Adds Paradee-8M v1.0 (Kokoro-82M distilled to 8.07M params,
af_heart, 24 kHz English) as a TTS backend. Models: FluidInference/paradee-8m-coreml (int8/12 MB default,fp32/34 MB). Conversion code lives in our private model-lab repo.ParadeeManager:ParadeeText→ host (round-half-even durations, column expansion, seeded source noise) →ParadeeAcousticKokoroAneVocab(Paradee's vocab is identical); applies misaki'sɾ→T,ʔ→toutput step, which Paradee was trained on (MiniMax WER 1.76% → 1.20%).all/.cpuAndGPU: the LSTMs abort in MPSGraph (GPURNNOps JIT not supported)tts --backend paradee [--variant int8|fp32] [--speed] [--seed] [--phonemes],tts-benchmark --backend paradeeResults (
tts-benchmark --corpus minimax-english, M5 Pro, macOS 27): int8 and fp32 both WER 1.20% / CER 0.14%, RTFx ~100 including G2P, p50 77 ms per phrase. Model-only, Core ML runs ~152× real time vs upstream ONNX 20× (1 thread) / 47× (15 threads).Testing: release build + swift-format clean; unit tests added (
ParadeeTests) but not run locally (no XCTest here), so CI is their first run.🤖 Generated with Claude Code