Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions Documentation/TTS/Benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,7 @@ Reference each language as `--corpus minimax-<lang>`:
| StyleTTS2 | `minimax-english` | `english` only (LibriTTS iteration_3, zero-shot from `--reference` audio) |
| Chatterbox (beta) | `minimax-english` | 18 of the 23 upstream languages via `--language` (zh/ja/he/ko/ru unported); built-in voice, macOS 15+ |
| Chatterbox Nano (beta) | `minimax-english` | `english` only; inline paralinguistic tags (`[laugh]`, `[chuckle]`, …); built-in voice, macOS 15+ |
| Paradee (beta) | `minimax-english` | `english` only (`af_heart`); `--variant int8` (default) or `fp32` |
| Supertonic-3 | `minimax-english` | 31 ISO codes minus `zh`: `english`, `korean`, `japanese`, `arabic`, `bulgarian`, `czech`, `danish`, `german`, `greek`, `spanish`, `estonian`, `finnish`, `french`, `hindi`, `croatian`, `hungarian`, `indonesian`, `italian`, `lithuanian`, `latvian`, `dutch`, `polish`, `portuguese`, `romanian`, `russian`, `slovak`, `slovenian`, `swedish`, `turkish`, `ukrainian`, `vietnamese`. Voice styling via `--voice-style <preset.json>` |

Lines beginning with `#` are comments. Custom corpora can still be
Expand Down
61 changes: 61 additions & 0 deletions Documentation/TTS/Paradee.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# Paradee (beta)

[Paradee-8M v1.0](https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0) (Sahil Mahendrakar, Apache-2.0) is
Kokoro-82M distilled into 8.07M parameters: Kokoro's own modules at smaller widths, one voice (`af_heart`),
American English, 24 kHz. Paper: [arXiv 2610.06817](https://arxiv.org/abs/2610.06817).

Core ML models: [FluidInference/paradee-8m-coreml](https://huggingface.co/FluidInference/paradee-8m-coreml)
(`int8/` 12 MB default, `fp32/` 34 MB).

> **Beta:** API, model artifacts and accuracy may change.

## Usage

```swift
let manager = ParadeeManager() // .int8, .cpuOnly
try await manager.initialize() // downloads models + English G2P assets
let samples = try await manager.synthesize(text: "Hello from FluidAudio.", speed: 1.0, noiseSeed: 0)
let phonemes = try await manager.phonemes(for: "Hello from FluidAudio.")
let fromPhonemes = try await manager.synthesize(phonemes: phonemes)
```

```bash
swift run fluidaudiocli tts "Hello from FluidAudio." --backend paradee --output out.wav
swift run fluidaudiocli tts "Hello." --backend paradee --variant fp32 --speed 1.2 --seed 7 --output out.wav
swift run fluidaudiocli tts-benchmark --backend paradee --corpus minimax-english
```

## Pipeline

```
text -> sentences -> NeMo TN -> Misaki lexicon / BART G2P -> ɾ→T, ʔ→t (misaki output form) -> ids
ParadeeText ids [1,T] -> duration [1,T], d [1,224,T], asr_tok [1,512,T]
host n = max(1, round(duration / speed)); repeat columns; N(0,1) source noise [1,1,600F]
ParadeeAcoustic en [1,224,F], asr [1,512,F], noise -> audio [1,600F]
```

- Text is synthesized one sentence at a time, as in the upstream package; sentences over 510 phonemes are split.
- The English frontend is KokoroAne's. Its lexicon stores misaki's raw flap `ɾ` and glottal stop `ʔ`; misaki
rewrites them to `T` / `t` before Kokoro v1.0 sees them, and Paradee was trained only on that output, so the
Paradee text path applies the same rewrite. Without it "kittens", "satellite" and "patterns" come out garbled.
- `noiseSeed` seeds the harmonic source noise; the same seed gives the same audio.
- Compute units: `.cpuOnly` (default) or `.cpuAndNeuralEngine`. `.all` / `.cpuAndGPU` are rejected because the
LSTMs abort in MPSGraph (`GPURNNOps … JIT not supported`).

## Benchmark

`tts-benchmark --backend paradee --corpus minimax-english`, 100 phrases, Parakeet TDT round trip,
M5 Pro, macOS 27, `.cpuOnly`:

| Variant | WER | CER | RTFx (audio / synth) | synth p50 / p95 | peak RSS |
|---|---:|---:|---:|---:|---:|
| int8 | 1.20 % | 0.14 % | 99.8× | 77 / 90 ms | 434 MB |
| fp32 | 1.20 % | 0.14 % | 100.3× | 77 / 90 ms | 433 MB |

Synth time covers the text frontend and both models. The models alone run ~150× real time; the
rest is the shared English G2P. The remaining WER is mostly text normalization ("will - power",
"five thousand" → "5,000", "spellbook").

Against upstream (Python, 10 sentences incl. a 28 s paragraph): 0 duration mismatches and identical lengths vs
PyTorch with matched noise; log-mel distance to the upstream ONNX equals the ONNX's own run-to-run noise; Whisper
transcripts identical to the ONNX's.
19 changes: 18 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ Want to convert your own model? Check [möbius](https://github.com/FluidInferenc

- **Automatic Speech Recognition (ASR)**: [Parakeet TDT v3](Documentation/Models.md#batch-transcription-near-real-time) (0.6b) and other TDT/CTC models for batch transcription supporting 25 European languages and Japanese, plus SenseVoice and Paraformer for Mandarin Chinese; [Parakeet EOU](Documentation/Models.md#streaming-transcription-true-real-time) (120m) for streaming ASR with end-of-utterance detection (English only). See all [ASR models](Documentation/Models.md#asr-models).
- **Inverse Text Normalization (ITN)**: Post-process ASR output to convert spoken-form to written-form ("two hundred" → "200"). See [text-processing-rs](https://github.com/FluidInference/text-processing-rs). Optional: ASR-only apps can drop the engine (~8 MB per slice) with `traits: []` (Swift 6.2+), see [PostProcessing.md](Documentation/ASR/PostProcessing.md#opting-out-of-the-engine)
- **Text-to-Speech (TTS)**: Kokoro (82m) for parallel synthesis with SSML and pronunciation control across 9 languages (EN, ES, FR, HI, IT, JA, PT, ZH); PocketTTS for streaming TTS with voice cloning support (EN, DE, ES, FR, IT, PT — 6L and 24L variants); Chatterbox Multilingual (520M, 18 languages) and Chatterbox Nano (110M, English with `[laugh]`/`[chuckle]` paralinguistic tags) in beta — see [Documentation/TTS/Chatterbox.md](Documentation/TTS/Chatterbox.md)
- **Text-to-Speech (TTS)**: Kokoro (82m) for parallel synthesis with SSML and pronunciation control across 9 languages (EN, ES, FR, HI, IT, JA, PT, ZH); PocketTTS for streaming TTS with voice cloning support (EN, DE, ES, FR, IT, PT — 6L and 24L variants); Chatterbox Multilingual (520M, 18 languages) and Chatterbox Nano (110M, English with `[laugh]`/`[chuckle]` paralinguistic tags) in beta — see [Documentation/TTS/Chatterbox.md](Documentation/TTS/Chatterbox.md); Paradee (8M Kokoro distill, English, 12 MB) in beta — see [Documentation/TTS/Paradee.md](Documentation/TTS/Paradee.md)
- **Speaker Diarization (Online + Offline)**: Speaker separation and identification across audio streams. Streaming pipeline for real-time processing and offline batch pipeline with advanced clustering.
- **Speaker Embedding Extraction**: Generate speaker embeddings for voice comparison and clustering, you can use this for speaker identification
- **Voice Activity Detection (VAD)**: Voice activity detection with Silero models
Expand Down Expand Up @@ -699,6 +699,21 @@ swift run fluidaudiocli tts "Hello from FluidAudio." --backend kokoroAne --outpu

Model assets are cached under `~/.cache/fluidaudio/Models/kokoro/`.

### Paradee (beta)

Kokoro-82M distilled into 8M parameters (one voice, `af_heart`, English), 12 MB of int8 weights,
~100× real time on CPU. Uses the Kokoro English frontend. See [Documentation/TTS/Paradee.md](Documentation/TTS/Paradee.md).

```swift
let manager = ParadeeManager()
try await manager.initialize()
let samples = try await manager.synthesize(text: "Hello from FluidAudio.") // 24 kHz Float32
```

```bash
swift run fluidaudiocli tts "Hello from FluidAudio." --backend paradee --output out.wav
```

## Continuous Integration

- `tests.yml`: Default build matrix covering SwiftPM tests and an iOS archive smoke test.
Expand Down Expand Up @@ -732,6 +747,8 @@ silero-vad: <https://github.com/snakers4/silero-vad>

Kokoro-82M: <https://huggingface.co/hexgrad/Kokoro-82M>

Paradee-8M: <https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0>

### Citation

If you use FluidAudio in your work, please cite:
Expand Down
32 changes: 32 additions & 0 deletions Sources/FluidAudio/ModelNames.swift
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,12 @@ public enum Repo: String, CaseIterable, Sendable {
/// Conversion lives in mobius (`models/tts/inflect-v2`).
case inflectMicro = "FluidInference/inflect-v2-coreml/micro"
case inflectNano = "FluidInference/inflect-v2-coreml/nano"
/// Paradee-8M (Kokoro-82M distilled to 8M, single voice). One repo with two
/// precision subdirectories (`int8/`, `fp32/`), each holding
/// `ParadeeText.mlmodelc` + `ParadeeAcoustic.mlmodelc` + `vocab.json`.
/// Conversion notes are on the model card.
case paradeeInt8 = "FluidInference/paradee-8m-coreml/int8"
case paradeeFp32 = "FluidInference/paradee-8m-coreml/fp32"
/// Chatterbox Multilingual (ResembleAI, 23 languages, **beta**) — T3 Llama-520M
/// AR speech-token generator (CFG batch 2, MLState KV decode) + S3Gen
/// flow-matching mel decoder + HiFT vocoder. Repo root holds the
Expand Down Expand Up @@ -232,6 +238,10 @@ public enum Repo: String, CaseIterable, Sendable {
return "inflect-v2-coreml/micro"
case .inflectNano:
return "inflect-v2-coreml/nano"
case .paradeeInt8:
return "paradee-8m-coreml/int8"
case .paradeeFp32:
return "paradee-8m-coreml/fp32"
}
}

Expand Down Expand Up @@ -262,6 +272,8 @@ public enum Repo: String, CaseIterable, Sendable {
return "FluidInference/StyleTTS-2-coreml"
case .inflectMicro, .inflectNano:
return "FluidInference/inflect-v2-coreml"
case .paradeeInt8, .paradeeFp32:
return "FluidInference/paradee-8m-coreml"
default:
return "FluidInference/\(name)"
}
Expand Down Expand Up @@ -319,6 +331,10 @@ public enum Repo: String, CaseIterable, Sendable {
return "micro"
case .inflectNano:
return "nano"
case .paradeeInt8:
return "int8"
case .paradeeFp32:
return "fp32"
default:
return nil
}
Expand Down Expand Up @@ -377,6 +393,10 @@ public enum Repo: String, CaseIterable, Sendable {
return "inflect-v2-coreml/micro"
case .inflectNano:
return "inflect-v2-coreml/nano"
case .paradeeInt8:
return "paradee-8m-coreml/int8"
case .paradeeFp32:
return "paradee-8m-coreml/fp32"
default:
return name.replacingOccurrences(of: "-coreml", with: "")
}
Expand Down Expand Up @@ -1437,6 +1457,16 @@ public enum ModelNames {
[encoderFile] + InflectConstants.frameBuckets.map { synthesizerFile(frames: $0) })
}

/// Paradee-8M model names. File names match
/// `FluidInference/paradee-8m-coreml/<int8|fp32>/`.
public enum Paradee {
public static let textFile = "ParadeeText.mlmodelc"
public static let acousticFile = "ParadeeAcoustic.mlmodelc"
public static let vocabFile = "vocab.json"

public static let requiredModels: Set<String> = [textFile, acousticFile, vocabFile]
}

/// LuxTTS (ZipVoice-Distill) model names. The HF repo publishes the same
/// text encoder + flow-matching decoder in two graph layouts:
/// - `gpu/` — original graph; fastest on Mac GPU (do NOT run on ANE:
Expand Down Expand Up @@ -1859,6 +1889,8 @@ public enum ModelNames {
return ModelNames.LuxTts.requiredFiles(variant: variant)
case .inflectMicro, .inflectNano:
return ModelNames.Inflect.requiredModels
case .paradeeInt8, .paradeeFp32:
return ModelNames.Paradee.requiredModels
}
}
}
121 changes: 121 additions & 0 deletions Sources/FluidAudio/TTS/Paradee/Assets/ParadeeModelStore.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,121 @@
@preconcurrency import CoreML
import Foundation

/// Actor store for one Paradee variant: `ParadeeText` + `ParadeeAcoustic`
/// CoreML bundles and the phoneme vocab, downloaded from
/// `FluidInference/paradee-8m-coreml/<variant>/` on first use.
///
/// - Note: Beta — this is a beta model conversion; API, model artifacts, and accuracy may change.
public actor ParadeeModelStore {

private let logger = AppLogger(category: "ParadeeModelStore")

private let variant: ParadeeVariant
private let directory: URL?
private let computeUnits: MLComputeUnits

private var textModel: MLModel?
private var acousticModel: MLModel?
private var vocab: KokoroAneVocab?

public init(
variant: ParadeeVariant = .int8,
directory: URL? = nil,
computeUnits: MLComputeUnits = .cpuOnly
) {
self.variant = variant
self.directory = directory
self.computeUnits = computeUnits
}

/// Download (if missing) and load both models + vocab.
public func loadIfNeeded() async throws {
if textModel != nil, acousticModel != nil, vocab != nil { return }
guard Self.isSupported(computeUnits) else {
throw ParadeeError.unsupportedComputeUnits(Self.describe(computeUnits))
}

let repoDir = try await ensureModels()
textModel = try loadModel(repoDir: repoDir, fileName: ModelNames.Paradee.textFile)
acousticModel = try loadModel(repoDir: repoDir, fileName: ModelNames.Paradee.acousticFile)
do {
vocab = try KokoroAneVocab.load(from: repoDir.appendingPathComponent(ModelNames.Paradee.vocabFile))
} catch {
throw ParadeeError.modelFileNotFound("\(ModelNames.Paradee.vocabFile): \(error.localizedDescription)")
}
logger.info("Paradee \(variant.rawValue) loaded from \(repoDir.path) (\(Self.describe(computeUnits)))")
}

public func models() throws -> (text: MLModel, acoustic: MLModel) {
guard let textModel, let acousticModel else { throw ParadeeError.notInitialized }
return (textModel, acousticModel)
}

public func vocabulary() throws -> KokoroAneVocab {
guard let vocab else { throw ParadeeError.notInitialized }
return vocab
}

public func unload() {
textModel = nil
acousticModel = nil
vocab = nil
}

/// `.all` and `.cpuAndGPU` abort in MPSGraph (`GPURNNOps … JIT not supported`).
static func isSupported(_ units: MLComputeUnits) -> Bool {
units == .cpuOnly || units == .cpuAndNeuralEngine
}

static func describe(_ units: MLComputeUnits) -> String {
switch units {
case .cpuOnly: return "cpuOnly"
case .cpuAndGPU: return "cpuAndGPU"
case .all: return "all"
case .cpuAndNeuralEngine: return "cpuAndNeuralEngine"
@unknown default: return "unknown"
}
}

// MARK: - Helpers

private func ensureModels() async throws -> URL {
let repo = variant.repo
let modelsRoot = try directory ?? Self.defaultCacheRoot()
let repoDir = modelsRoot.appendingPathComponent(repo.folderName)
let allPresent = ModelNames.Paradee.requiredModels.allSatisfy {
FileManager.default.fileExists(atPath: repoDir.appendingPathComponent($0).path)
}
if allPresent { return repoDir }

logger.info("Downloading Paradee \(variant.rawValue) CoreML models from HuggingFace…")
do {
try await ModelHub.download(repo, to: modelsRoot)
} catch {
throw ParadeeError.downloadFailed("\(error)")
}
return repoDir
}

private func loadModel(repoDir: URL, fileName: String) throws -> MLModel {
let url = repoDir.appendingPathComponent(fileName)
guard FileManager.default.fileExists(atPath: url.path) else {
throw ParadeeError.modelFileNotFound(fileName)
}
let config = MLModelConfiguration()
config.computeUnits = computeUnits
do {
return try MLModel(contentsOf: url, configuration: config)
} catch {
throw ParadeeError.corruptedModel(fileName, underlying: "\(error)")
}
}

private static func defaultCacheRoot() throws -> URL {
let root = try TtsCacheDirectory.ensure().appendingPathComponent("Models")
if !FileManager.default.fileExists(atPath: root.path) {
try FileManager.default.createDirectory(at: root, withIntermediateDirectories: true)
}
return root
}
}
30 changes: 30 additions & 0 deletions Sources/FluidAudio/TTS/Paradee/ParadeeConstants.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
import Foundation

/// Fixed pipeline parameters for the Paradee-8M CoreML backend. Values mirror
/// `FluidInference/paradee-8m-coreml/config.json` and the upstream
/// `paradee/tts.py`.
public enum ParadeeConstants {

/// Output sample rate (24 kHz mono).
public static let sampleRate = 24_000

/// Output samples per aligned frame: each frame carries two F0 steps of
/// 300 samples (`prod(upsample_rates) * istft_hop`).
public static let samplesPerFrame = 600

/// Upper bound of the acoustic model's frame axis (100 s at speed 1).
public static let maxFrames = 4_000

/// Phonemes per synthesis call; the text side has 512 positions, two of
/// which hold the pad token at each end.
public static let maxPhonemeLength = 510

/// Channels of the prosody features `d` (hidden 192 + style 32).
public static let prosodyChannels = 224

/// Channels of the text features `asr_tok` (the teacher decoder's width).
public static let textChannels = 512

/// Speech-rate multiplier. Durations scale by `1 / speed`.
public static let defaultSpeed: Float = 1.0
}
38 changes: 38 additions & 0 deletions Sources/FluidAudio/TTS/Paradee/ParadeeError.swift
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
import Foundation

/// Errors thrown by the Paradee TTS backend.
public enum ParadeeError: Error, LocalizedError {
case notInitialized
case downloadFailed(String)
case modelFileNotFound(String)
case corruptedModel(String, underlying: String)
/// `.all` / `.cpuAndGPU` abort in MPSGraph on the LSTMs.
case unsupportedComputeUnits(String)
case inputProcessingFailed(String)
/// A chunk's predicted duration exceeded the acoustic model's frame axis.
case durationOverflow(frames: Int, maxFrames: Int)
case predictionFailed(String)

public var errorDescription: String? {
switch self {
case .notInitialized:
return "Paradee backend is not initialized; call initialize() first."
case .downloadFailed(let detail):
return "Paradee model download failed: \(detail)"
case .modelFileNotFound(let name):
return "Paradee model file not found: \(name)"
case .corruptedModel(let name, let underlying):
return "Paradee model \(name) failed to load: \(underlying)"
case .unsupportedComputeUnits(let units):
return
"Paradee does not support compute units \(units): the LSTMs abort on the GPU. "
+ "Use .cpuOnly or .cpuAndNeuralEngine."
case .inputProcessingFailed(let detail):
return "Paradee input processing failed: \(detail)"
case .durationOverflow(let frames, let maxFrames):
return "Predicted duration \(frames) frames exceeds the model limit (\(maxFrames)); shorten the input."
case .predictionFailed(let detail):
return "Paradee CoreML prediction failed: \(detail)"
}
}
}
Loading
Loading