Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 56 additions & 2 deletions Documentation/ANE_Profiler.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,9 @@
# ANE Profiler

The header below describes the original June audit. Later model sections state
their own dates, configurations and metrics; see [Kokoro ANE v3](#kokoro-ane-v3)
for the October M5 Pro routing, latency and telemetry checks.

| | |
|---|---|
| **Measured** | 2026-06-05 |
Expand Down Expand Up @@ -237,14 +241,64 @@ Latency **measured on real synthesis**, warm (one short sentence; `tts --backend

| Model | Type | ANE | GPU | CPU | ops | Size | Heavy graph → device |
|-------|------|----:|----:|----:|----:|-----:|----------------------|
| Kokoro ANE (7-stage) | batch (per utterance) | 75% | 0% | 25% | 1472 | 83 MB | Vocoder → ANE |
| Kokoro ANE (legacy capability profile) | batch (per utterance) | 75% | 0% | 25% | 1472 | 83 MB | Vocoder → ANE |
| Supertonic (`--ve-variant fp16`, legacy) | batch (8-step diffusion) | 30% | 0% | 70% | 1365 | 192 MB | VectorEstimator → **CPU** (dynamic shapes can't use ANE) |
| Supertonic (default, int4 L-bucketed) | batch (8-step diffusion) | ~90% | 0% | ~10% | 1289 | 102 MB | VectorEstimator → **ANE** (fixed L-buckets) |
| PocketTTS (v2.1) | streaming (autoregressive) | ~9% | ~31% | ~60% | 2629 | ~330 MB | flow_decoder_fused → **ANE**; flowlm/cond → GPU; mimi → CPU |

**Component detail**

### Kokoro ANE (7-stage)
### Kokoro ANE v3

**Separate check: M5 Pro, 24 GiB, macOS 27.0 (26A428), October 7, 2026.**
The historical seven-stage op-count percentages below do not describe v3.

| Stage | Requested compute policy in v0.17.6 |
| --- | --- |
| Static/dynamic Albert, PostAlbert, Alignment, Prosody, masked decoder | `.cpuAndNeuralEngine` |
| Fast waveform generator | `.cpuAndGPU` |
| Native source/STFT | CPU / Accelerate (outside Core ML) |
| Long-input fallback Vocoder | `.cpuAndNeuralEngine` |
| Long-input fallback Noise / Tail | `.cpuAndGPU` |

This is hybrid routing, not `.all`. Allowed devices are not a guarantee that
all operations execute on one device; CPU fallback is permitted. There is no
strict ANE-only Core ML compute policy. The public v3 API retains this fixed
routing; the policy overrides below were local diagnostic changes.

Across the three language variants, `MLComputePlan` preferred ANE for about
**98.8% of estimated operation cost in static Albert-64** and approximately
**100% in DecoderPre-200**; the source/generator plan preferred GPU for about
**100%**. PostAlbert, Alignment and fp32 Prosody preferred CPU in this audit.
These percentages are **estimated cost weights within individual graphs**, not
op counts, measured utilization, wall-time shares or a runtime execution trace.
The 32-token and 120-frame buckets are not covered by those percentages.

A separate real-input check measured hybrid medians of **49.45 / 35.85 / 53.97 ms**
for English / Japanese / Mandarin. Forcing `.cpuAndNeuralEngine` throughout
measured **511.40 / 474.45 / 584.19 ms**. Both `.all` and `.cpuAndGPU` hit
`GPURNNOps.mm: JIT not supported` during Prosody on the second warmup in all
three cases. See the [timing protocol and raw records](TTS/Benchmarks.md#compute-policy-comparison).

**Interpreting asitop:** the live demo performs short bursts, while the observed
powermetrics collector sampled roughly once per second. Samples near three demo
clicks reported about **30.0 / 9.7 / 11.7 mW** of system-wide ANE power for
English / Japanese / Mandarin. Those observations are not synchronized per-process
energy measurements and are separate from the compute-policy timing run.
The installed asitop display rounded low power to `0.0 W` / `0%`; a local viewer
showed mW and the maximum observed sample over 60 seconds. That maximum is still
a sampled average, not an instantaneous utilization peak. Neither those readings
nor a compute plan establish which Kokoro operations actually ran on ANE;
process-attributed runtime tracing remains outstanding.
[Compute-plan records](TTS/Measurements/KokoroV3/compute-placement.json) and
[system-wide power samples](TTS/Measurements/KokoroV3/ane-latest-clicks.json)
preserve the separate sources for those observations.

### Kokoro ANE (legacy seven-stage capability profile)

Historical June profile under `.cpuAndNeuralEngine` throughout. Percentages are
preferred-device operation counts, not production utilization or v3 routing.

| Component | ANE | GPU | CPU | ops | Size | Lat ms |
|-----------|----:|----:|----:|----:|-----:|-------:|
| Albert | 94% | 0% | 6% | 310 | 6 MB | 6.5 |
Expand Down
15 changes: 15 additions & 0 deletions Documentation/Benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -248,6 +248,21 @@ Derived metrics:

## Text-to-Speech

### Kokoro ANE v3 (FluidAudio v0.17.6)

Selected M5 Pro / macOS 27.0 checks measured **49.45 / 35.85 / 53.97 ms**
median inference for English / Japanese / Mandarin, producing 3.350 / 3.050 /
3.675 seconds of audio. Each median uses five calls after two warmups; input
phonemes are prepared in advance. Current v3 uses per-stage CPU/ANE and CPU/GPU
policies plus native CPU DSP, rather than `.all` or an ANE-only policy.

Applying CPU+ANE to every graph measured 511.40 / 474.45 / 584.19 ms; `.all`
and CPU+GPU aborted in Apple's GPU RNN path before completing warmups.
These are bounded local checks, not full-corpus results. See the
[complete protocol, inputs, ranges, raw records and separate live-demo timings](TTS/Benchmarks.md#kokoro-ane-v3-selected-m5-pro-measurements).

### Historical M4 Pro comparison

We generated the same strings with to generate audio between 1s to ~300s in order to test the speed across a range of varying inputs on Pytorch CPU, MPS, and MLX pipeline, and compared it against the native Swift version with Core ML models.

Each pipeline warmed up the models by running through it once with pesudo inputs, and then comparing the raw inference time with the model already loaded. You can see that for the Core ML model, we traded lower memory and very slightly faster inference for longer initial warm-up.
Expand Down
16 changes: 14 additions & 2 deletions Documentation/CLI.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,14 +7,22 @@ This guide collects commonly used `fluidaudio` CLI commands for ASR, diarization
TTS is built into the CLI. Run it directly:

```bash
# Default Kokoro (CPU+GPU, multi-voice, chunker, custom lexicon)
# Default TTS backend (use --backend to select explicitly)
swift run fluidaudiocli tts "Hello from FluidAudio" --output out.wav

# Kokoro ANE (7-stage, ANE-resident, 3-11× RTFx, single voice af_heart)
# Kokoro legacy seven-stage runtime (variant voices; OS-dependent routing)
swift run fluidaudiocli tts "Hello from FluidAudio" \
--backend kokoro-ane \
--output out-ane.wav

# Kokoro ANE v3: fixed hybrid CPU/ANE/GPU routing (v0.17.6+)
swift run -c release fluidaudiocli tts "Hello from FluidAudio" \
--backend kokoro-ane --kokoro-version v3 --voice af_heart --output en.wav
swift run -c release fluidaudiocli tts "こんにちは、世界。" \
--backend kokoro-ane --kokoro-version v3 --variant ja --voice jf_alpha --output ja.wav
swift run -c release fluidaudiocli tts "你好世界。" \
--backend kokoro-ane --kokoro-version v3 --variant zh --voice zf_001 --output zh.wav

# PocketTTS (streaming, voice cloning)
swift run fluidaudiocli tts "Hello from FluidAudio" \
--backend pocket \
Expand All @@ -24,6 +32,10 @@ swift run fluidaudiocli tts "Hello from FluidAudio" \
swift run fluidaudiocli g2p-benchmark
```

V3 requires macOS 15 / iOS 18. It supports the fixed hybrid route, not arbitrary
global compute policies. See [Kokoro usage and limits](TTS/KokoroAne.md) and
[latency measurement boundaries](TTS/Benchmarks.md#kokoro-ane-v3-selected-m5-pro-measurements).

## ASR

```bash
Expand Down
4 changes: 2 additions & 2 deletions Documentation/Models.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ maps the supported artifacts to their upstream sources and documents the limits

| Model | Description | Context |
|-------|-------------|---------|
| **Kokoro ANE (7-stage)** | Kokoro 82M weights split into 7 CoreML stages so the ANE-friendly layers (Albert / Prosody / Vocoder) stay resident on the Neural Engine while PostAlbert / Alignment / Noise / Tail run on CPU. 3-11× RTFx. English (`af_heart`) and Mandarin (`ANE-zh`) variants. ≤510 IPA phonemes per call, no chunker / SSML / custom lexicon. Managed by `KokoroAneManager`. | ANE-optimized variant derived (with permission) from [laishere/kokoro-coreml](https://github.com/laishere/kokoro-coreml). The original single-graph (mono) Kokoro backend was removed in favor of this ANE pipeline; the kokoro repo root is now retained only for shared G2P assets. |
| **Kokoro ANE (v3 / legacy)** | `KokoroAneManager` provides native English, Japanese and Mandarin text synthesis with opt-in `version: .v3` in v0.17.6. V3 uses static CPU/ANE stages, a CPU/GPU generator and native CPU DSP, with legacy fallback beyond the fast buckets. Existing callers keep the legacy runtime; Spanish/French remain legacy-only. See [KokoroAne.md](TTS/KokoroAne.md) for limits and [measured latency](TTS/Benchmarks.md#kokoro-ane-v3-selected-m5-pro-measurements). | Base weights: [hexgrad/Kokoro-82M](https://huggingface.co/hexgrad/Kokoro-82M) (en/ja), [hexgrad/Kokoro-82M-v1.1-zh](https://huggingface.co/hexgrad/Kokoro-82M-v1.1-zh) (zh). Legacy conversion derives from laishere; v3 adapts Matt Mireles’s optimized GPU graph. |
| **PocketTTS** | TTS backend (~155M params). Autoregressive frame-by-frame generation with dynamic audio chunking. No phoneme stage, works directly on text tokens. Managed by `PocketTtsManager`. | Supports streaming, minimal RAM usage, excellent quality |
| **Supertonic-3** | Multilingual TTS, 31 languages (`en`, `ko`, `ja`, `ar`, `bg`, `cs`, `da`, `de`, `el`, `es`, `et`, `fi`, `fr`, `hi`, `hr`, `hu`, `id`, `it`, `lt`, `lv`, `nl`, `pl`, `pt`, `ro`, `ru`, `sk`, `sl`, `sv`, `tr`, `uk`, `vi`, plus `na` for numeric/language-agnostic input). 4-stage CoreML pipeline (text_encoder → duration_predictor → vector_estimator → vocoder, ~398 MB). Caller-supplied voice styles loaded from Supertonic preset JSON. 44.1 kHz mono fp32 output. Managed by `Supertonic3Manager`. | CoreML conversion of upstream `Supertone/supertonic-3`; see `Scripts/convert_supertonic3_to_coreml.py`. |
| **StyleTTS2 (LibriTTS, iteration_3)** | Reference-audio–driven zero-shot English TTS. 8-stage CoreML pipeline (`text_encoder → bert → ref_encoder → fused_diffusion_sampler → duration_predictor → fused_f0n_har_source → decoder_pre → decoder_upsample`) with 3 lazily-loaded T = 64 / 128 / 256 bucket variants of `bert` / `fused_diffusion_sampler`. 5-step ADPM2 Karras-σ diffusion sampler with α/β style blending against a speaker reference clip. 24 kHz mono fp32 output. Phonemizer reuses Kokoro's Misaki lexicon cache + BART G2P CoreML model with Misaki uppercase diphthong shorthand (`A O I Y W` → `eɪ oʊ aɪ ɔɪ aʊ`) expanded before encoding so the output matches the espeak IPA the model was trained on. Callers with a higher-quality phonemizer can bypass the stack via `StyleTTS2Manager.synthesize(ipa:...)`. See [StyleTTS2.md](TTS/StyleTTS2.md). | Zero-shot voice cloning from a single reference WAV; English only |
Expand Down Expand Up @@ -106,7 +106,7 @@ Models we converted and tested but are not supported: too large for on-device de
| Diarization (Pyannote) | [FluidInference/speaker-diarization-coreml](https://huggingface.co/FluidInference/speaker-diarization-coreml) |
| LS-EEND | [FluidInference/ls-eend-coreml](https://huggingface.co/FluidInference/ls-eend-coreml) (per-dataset optimized variants: `/optimized/ami`, `/optimized/ch`, `/optimized/dih2`, `/optimized/dih3`) |
| Sortformer | [FluidInference/diar-streaming-sortformer-coreml](https://huggingface.co/FluidInference/diar-streaming-sortformer-coreml) |
| Kokoro ANE (7-stage) | [FluidInference/kokoro-82m-coreml/tree/main/ANE](https://huggingface.co/FluidInference/kokoro-82m-coreml/tree/main/ANE) (English: `/ANE`; Mandarin: `/ANE-zh`; shared G2P assets `G2PEncoder.mlmodelc`, `G2PDecoder.mlmodelc`, `g2p_vocab.json` at the repo root) |
| Kokoro ANE (v3 / legacy) | [FluidInference/kokoro-82m-coreml](https://huggingface.co/FluidInference/kokoro-82m-coreml): v3 English `ANE-v3`, Japanese `ANE-v3/ja`, Mandarin `ANE-v3/zh`; legacy `ANE`, `ANE-ja`, `ANE-zh`. Shared G2P assets remain at the repo root / language asset directories; SDK pins and cache paths are documented in [KokoroAne.md](TTS/KokoroAne.md). |
| PocketTTS | [FluidInference/pocket-tts-coreml](https://huggingface.co/FluidInference/pocket-tts-coreml) |
| StyleTTS2 (LibriTTS, iteration_3) | [FluidInference/StyleTTS-2-coreml/iteration_3/compiled](https://huggingface.co/FluidInference/StyleTTS-2-coreml/tree/main/iteration_3/compiled) (shared phonemizer assets pulled from [`FluidInference/kokoro-82m-coreml`](https://huggingface.co/FluidInference/kokoro-82m-coreml): `G2PEncoder.mlmodelc`, `G2PDecoder.mlmodelc`, `g2p_vocab.json`, `us_lexicon_cache.json`) |
| Supertonic-3 | [FluidInference/supertonic-3-coreml](https://huggingface.co/FluidInference/supertonic-3-coreml) |
Expand Down
4 changes: 3 additions & 1 deletion Documentation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,9 @@

## Text-to-Speech (TTS)

- [Kokoro ANE (7-stage)](TTS/KokoroAne.md)
- [Kokoro ANE v3 and legacy runtime](TTS/KokoroAne.md)
- [Kokoro v3 latency and compute-policy measurements](TTS/Benchmarks.md#kokoro-ane-v3-selected-m5-pro-measurements)
- [Kokoro ANE placement and telemetry](ANE_Profiler.md#kokoro-ane-v3)
- [PocketTTS](TTS/PocketTTS.md)
- [StyleTTS2](TTS/StyleTTS2.md)
- [SSML](TTS/SSML.md)
Expand Down
Loading
Loading