diff --git a/Documentation/TTS/Benchmarks.md b/Documentation/TTS/Benchmarks.md index 9bb72186e..a9f8bc170 100644 --- a/Documentation/TTS/Benchmarks.md +++ b/Documentation/TTS/Benchmarks.md @@ -65,6 +65,7 @@ Reference each language as `--corpus minimax-`: | StyleTTS2 | `minimax-english` | `english` only (LibriTTS iteration_3, zero-shot from `--reference` audio) | | Chatterbox (beta) | `minimax-english` | 18 of the 23 upstream languages via `--language` (zh/ja/he/ko/ru unported); built-in voice, macOS 15+ | | Chatterbox Nano (beta) | `minimax-english` | `english` only; inline paralinguistic tags (`[laugh]`, `[chuckle]`, …); built-in voice, macOS 15+ | +| Paradee (beta) | `minimax-english` | `english` only (`af_heart`); `--variant int8` (default) or `fp32` | | Supertonic-3 | `minimax-english` | 31 ISO codes minus `zh`: `english`, `korean`, `japanese`, `arabic`, `bulgarian`, `czech`, `danish`, `german`, `greek`, `spanish`, `estonian`, `finnish`, `french`, `hindi`, `croatian`, `hungarian`, `indonesian`, `italian`, `lithuanian`, `latvian`, `dutch`, `polish`, `portuguese`, `romanian`, `russian`, `slovak`, `slovenian`, `swedish`, `turkish`, `ukrainian`, `vietnamese`. Voice styling via `--voice-style ` | Lines beginning with `#` are comments. Custom corpora can still be diff --git a/Documentation/TTS/Paradee.md b/Documentation/TTS/Paradee.md new file mode 100644 index 000000000..96dec7846 --- /dev/null +++ b/Documentation/TTS/Paradee.md @@ -0,0 +1,61 @@ +# Paradee (beta) + +[Paradee-8M v1.0](https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0) (Sahil Mahendrakar, Apache-2.0) is +Kokoro-82M distilled into 8.07M parameters: Kokoro's own modules at smaller widths, one voice (`af_heart`), +American English, 24 kHz. Paper: [arXiv 2610.06817](https://arxiv.org/abs/2610.06817). + +Core ML models: [FluidInference/paradee-8m-coreml](https://huggingface.co/FluidInference/paradee-8m-coreml) +(`int8/` 12 MB default, `fp32/` 34 MB). + +> **Beta:** API, model artifacts and accuracy may change. + +## Usage + +```swift +let manager = ParadeeManager() // .int8, .cpuOnly +try await manager.initialize() // downloads models + English G2P assets +let samples = try await manager.synthesize(text: "Hello from FluidAudio.", speed: 1.0, noiseSeed: 0) +let phonemes = try await manager.phonemes(for: "Hello from FluidAudio.") +let fromPhonemes = try await manager.synthesize(phonemes: phonemes) +``` + +```bash +swift run fluidaudiocli tts "Hello from FluidAudio." --backend paradee --output out.wav +swift run fluidaudiocli tts "Hello." --backend paradee --variant fp32 --speed 1.2 --seed 7 --output out.wav +swift run fluidaudiocli tts-benchmark --backend paradee --corpus minimax-english +``` + +## Pipeline + +``` +text -> sentences -> NeMo TN -> Misaki lexicon / BART G2P -> ɾ→T, ʔ→t (misaki output form) -> ids +ParadeeText ids [1,T] -> duration [1,T], d [1,224,T], asr_tok [1,512,T] +host n = max(1, round(duration / speed)); repeat columns; N(0,1) source noise [1,1,600F] +ParadeeAcoustic en [1,224,F], asr [1,512,F], noise -> audio [1,600F] +``` + +- Text is synthesized one sentence at a time, as in the upstream package; sentences over 510 phonemes are split. +- The English frontend is KokoroAne's. Its lexicon stores misaki's raw flap `ɾ` and glottal stop `ʔ`; misaki + rewrites them to `T` / `t` before Kokoro v1.0 sees them, and Paradee was trained only on that output, so the + Paradee text path applies the same rewrite. Without it "kittens", "satellite" and "patterns" come out garbled. +- `noiseSeed` seeds the harmonic source noise; the same seed gives the same audio. +- Compute units: `.cpuOnly` (default) or `.cpuAndNeuralEngine`. `.all` / `.cpuAndGPU` are rejected because the + LSTMs abort in MPSGraph (`GPURNNOps … JIT not supported`). + +## Benchmark + +`tts-benchmark --backend paradee --corpus minimax-english`, 100 phrases, Parakeet TDT round trip, +M5 Pro, macOS 27, `.cpuOnly`: + +| Variant | WER | CER | RTFx (audio / synth) | synth p50 / p95 | peak RSS | +|---|---:|---:|---:|---:|---:| +| int8 | 1.20 % | 0.14 % | 99.8× | 77 / 90 ms | 434 MB | +| fp32 | 1.20 % | 0.14 % | 100.3× | 77 / 90 ms | 433 MB | + +Synth time covers the text frontend and both models. The models alone run ~150× real time; the +rest is the shared English G2P. The remaining WER is mostly text normalization ("will - power", +"five thousand" → "5,000", "spellbook"). + +Against upstream (Python, 10 sentences incl. a 28 s paragraph): 0 duration mismatches and identical lengths vs +PyTorch with matched noise; log-mel distance to the upstream ONNX equals the ONNX's own run-to-run noise; Whisper +transcripts identical to the ONNX's. diff --git a/README.md b/README.md index 4e4d80b49..e13cea35b 100644 --- a/README.md +++ b/README.md @@ -44,7 +44,7 @@ Want to convert your own model? Check [möbius](https://github.com/FluidInferenc - **Automatic Speech Recognition (ASR)**: [Parakeet TDT v3](Documentation/Models.md#batch-transcription-near-real-time) (0.6b) and other TDT/CTC models for batch transcription supporting 25 European languages and Japanese, plus SenseVoice and Paraformer for Mandarin Chinese; [Parakeet EOU](Documentation/Models.md#streaming-transcription-true-real-time) (120m) for streaming ASR with end-of-utterance detection (English only). See all [ASR models](Documentation/Models.md#asr-models). - **Inverse Text Normalization (ITN)**: Post-process ASR output to convert spoken-form to written-form ("two hundred" → "200"). See [text-processing-rs](https://github.com/FluidInference/text-processing-rs). Optional: ASR-only apps can drop the engine (~8 MB per slice) with `traits: []` (Swift 6.2+), see [PostProcessing.md](Documentation/ASR/PostProcessing.md#opting-out-of-the-engine) -- **Text-to-Speech (TTS)**: Kokoro (82m) for parallel synthesis with SSML and pronunciation control across 9 languages (EN, ES, FR, HI, IT, JA, PT, ZH); PocketTTS for streaming TTS with voice cloning support (EN, DE, ES, FR, IT, PT — 6L and 24L variants); Chatterbox Multilingual (520M, 18 languages) and Chatterbox Nano (110M, English with `[laugh]`/`[chuckle]` paralinguistic tags) in beta — see [Documentation/TTS/Chatterbox.md](Documentation/TTS/Chatterbox.md) +- **Text-to-Speech (TTS)**: Kokoro (82m) for parallel synthesis with SSML and pronunciation control across 9 languages (EN, ES, FR, HI, IT, JA, PT, ZH); PocketTTS for streaming TTS with voice cloning support (EN, DE, ES, FR, IT, PT — 6L and 24L variants); Chatterbox Multilingual (520M, 18 languages) and Chatterbox Nano (110M, English with `[laugh]`/`[chuckle]` paralinguistic tags) in beta — see [Documentation/TTS/Chatterbox.md](Documentation/TTS/Chatterbox.md); Paradee (8M Kokoro distill, English, 12 MB) in beta — see [Documentation/TTS/Paradee.md](Documentation/TTS/Paradee.md) - **Speaker Diarization (Online + Offline)**: Speaker separation and identification across audio streams. Streaming pipeline for real-time processing and offline batch pipeline with advanced clustering. - **Speaker Embedding Extraction**: Generate speaker embeddings for voice comparison and clustering, you can use this for speaker identification - **Voice Activity Detection (VAD)**: Voice activity detection with Silero models @@ -699,6 +699,21 @@ swift run fluidaudiocli tts "Hello from FluidAudio." --backend kokoroAne --outpu Model assets are cached under `~/.cache/fluidaudio/Models/kokoro/`. +### Paradee (beta) + +Kokoro-82M distilled into 8M parameters (one voice, `af_heart`, English), 12 MB of int8 weights, +~100× real time on CPU. Uses the Kokoro English frontend. See [Documentation/TTS/Paradee.md](Documentation/TTS/Paradee.md). + +```swift +let manager = ParadeeManager() +try await manager.initialize() +let samples = try await manager.synthesize(text: "Hello from FluidAudio.") // 24 kHz Float32 +``` + +```bash +swift run fluidaudiocli tts "Hello from FluidAudio." --backend paradee --output out.wav +``` + ## Continuous Integration - `tests.yml`: Default build matrix covering SwiftPM tests and an iOS archive smoke test. @@ -732,6 +747,8 @@ silero-vad: Kokoro-82M: +Paradee-8M: + ### Citation If you use FluidAudio in your work, please cite: diff --git a/Sources/FluidAudio/ModelNames.swift b/Sources/FluidAudio/ModelNames.swift index 84d1dd58d..b6a245c87 100644 --- a/Sources/FluidAudio/ModelNames.swift +++ b/Sources/FluidAudio/ModelNames.swift @@ -118,6 +118,12 @@ public enum Repo: String, CaseIterable, Sendable { /// Conversion lives in mobius (`models/tts/inflect-v2`). case inflectMicro = "FluidInference/inflect-v2-coreml/micro" case inflectNano = "FluidInference/inflect-v2-coreml/nano" + /// Paradee-8M (Kokoro-82M distilled to 8M, single voice). One repo with two + /// precision subdirectories (`int8/`, `fp32/`), each holding + /// `ParadeeText.mlmodelc` + `ParadeeAcoustic.mlmodelc` + `vocab.json`. + /// Conversion notes are on the model card. + case paradeeInt8 = "FluidInference/paradee-8m-coreml/int8" + case paradeeFp32 = "FluidInference/paradee-8m-coreml/fp32" /// Chatterbox Multilingual (ResembleAI, 23 languages, **beta**) — T3 Llama-520M /// AR speech-token generator (CFG batch 2, MLState KV decode) + S3Gen /// flow-matching mel decoder + HiFT vocoder. Repo root holds the @@ -232,6 +238,10 @@ public enum Repo: String, CaseIterable, Sendable { return "inflect-v2-coreml/micro" case .inflectNano: return "inflect-v2-coreml/nano" + case .paradeeInt8: + return "paradee-8m-coreml/int8" + case .paradeeFp32: + return "paradee-8m-coreml/fp32" } } @@ -262,6 +272,8 @@ public enum Repo: String, CaseIterable, Sendable { return "FluidInference/StyleTTS-2-coreml" case .inflectMicro, .inflectNano: return "FluidInference/inflect-v2-coreml" + case .paradeeInt8, .paradeeFp32: + return "FluidInference/paradee-8m-coreml" default: return "FluidInference/\(name)" } @@ -319,6 +331,10 @@ public enum Repo: String, CaseIterable, Sendable { return "micro" case .inflectNano: return "nano" + case .paradeeInt8: + return "int8" + case .paradeeFp32: + return "fp32" default: return nil } @@ -377,6 +393,10 @@ public enum Repo: String, CaseIterable, Sendable { return "inflect-v2-coreml/micro" case .inflectNano: return "inflect-v2-coreml/nano" + case .paradeeInt8: + return "paradee-8m-coreml/int8" + case .paradeeFp32: + return "paradee-8m-coreml/fp32" default: return name.replacingOccurrences(of: "-coreml", with: "") } @@ -1437,6 +1457,16 @@ public enum ModelNames { [encoderFile] + InflectConstants.frameBuckets.map { synthesizerFile(frames: $0) }) } + /// Paradee-8M model names. File names match + /// `FluidInference/paradee-8m-coreml//`. + public enum Paradee { + public static let textFile = "ParadeeText.mlmodelc" + public static let acousticFile = "ParadeeAcoustic.mlmodelc" + public static let vocabFile = "vocab.json" + + public static let requiredModels: Set = [textFile, acousticFile, vocabFile] + } + /// LuxTTS (ZipVoice-Distill) model names. The HF repo publishes the same /// text encoder + flow-matching decoder in two graph layouts: /// - `gpu/` — original graph; fastest on Mac GPU (do NOT run on ANE: @@ -1859,6 +1889,8 @@ public enum ModelNames { return ModelNames.LuxTts.requiredFiles(variant: variant) case .inflectMicro, .inflectNano: return ModelNames.Inflect.requiredModels + case .paradeeInt8, .paradeeFp32: + return ModelNames.Paradee.requiredModels } } } diff --git a/Sources/FluidAudio/TTS/Paradee/Assets/ParadeeModelStore.swift b/Sources/FluidAudio/TTS/Paradee/Assets/ParadeeModelStore.swift new file mode 100644 index 000000000..b9e9a9b5f --- /dev/null +++ b/Sources/FluidAudio/TTS/Paradee/Assets/ParadeeModelStore.swift @@ -0,0 +1,121 @@ +@preconcurrency import CoreML +import Foundation + +/// Actor store for one Paradee variant: `ParadeeText` + `ParadeeAcoustic` +/// CoreML bundles and the phoneme vocab, downloaded from +/// `FluidInference/paradee-8m-coreml//` on first use. +/// +/// - Note: Beta — this is a beta model conversion; API, model artifacts, and accuracy may change. +public actor ParadeeModelStore { + + private let logger = AppLogger(category: "ParadeeModelStore") + + private let variant: ParadeeVariant + private let directory: URL? + private let computeUnits: MLComputeUnits + + private var textModel: MLModel? + private var acousticModel: MLModel? + private var vocab: KokoroAneVocab? + + public init( + variant: ParadeeVariant = .int8, + directory: URL? = nil, + computeUnits: MLComputeUnits = .cpuOnly + ) { + self.variant = variant + self.directory = directory + self.computeUnits = computeUnits + } + + /// Download (if missing) and load both models + vocab. + public func loadIfNeeded() async throws { + if textModel != nil, acousticModel != nil, vocab != nil { return } + guard Self.isSupported(computeUnits) else { + throw ParadeeError.unsupportedComputeUnits(Self.describe(computeUnits)) + } + + let repoDir = try await ensureModels() + textModel = try loadModel(repoDir: repoDir, fileName: ModelNames.Paradee.textFile) + acousticModel = try loadModel(repoDir: repoDir, fileName: ModelNames.Paradee.acousticFile) + do { + vocab = try KokoroAneVocab.load(from: repoDir.appendingPathComponent(ModelNames.Paradee.vocabFile)) + } catch { + throw ParadeeError.modelFileNotFound("\(ModelNames.Paradee.vocabFile): \(error.localizedDescription)") + } + logger.info("Paradee \(variant.rawValue) loaded from \(repoDir.path) (\(Self.describe(computeUnits)))") + } + + public func models() throws -> (text: MLModel, acoustic: MLModel) { + guard let textModel, let acousticModel else { throw ParadeeError.notInitialized } + return (textModel, acousticModel) + } + + public func vocabulary() throws -> KokoroAneVocab { + guard let vocab else { throw ParadeeError.notInitialized } + return vocab + } + + public func unload() { + textModel = nil + acousticModel = nil + vocab = nil + } + + /// `.all` and `.cpuAndGPU` abort in MPSGraph (`GPURNNOps … JIT not supported`). + static func isSupported(_ units: MLComputeUnits) -> Bool { + units == .cpuOnly || units == .cpuAndNeuralEngine + } + + static func describe(_ units: MLComputeUnits) -> String { + switch units { + case .cpuOnly: return "cpuOnly" + case .cpuAndGPU: return "cpuAndGPU" + case .all: return "all" + case .cpuAndNeuralEngine: return "cpuAndNeuralEngine" + @unknown default: return "unknown" + } + } + + // MARK: - Helpers + + private func ensureModels() async throws -> URL { + let repo = variant.repo + let modelsRoot = try directory ?? Self.defaultCacheRoot() + let repoDir = modelsRoot.appendingPathComponent(repo.folderName) + let allPresent = ModelNames.Paradee.requiredModels.allSatisfy { + FileManager.default.fileExists(atPath: repoDir.appendingPathComponent($0).path) + } + if allPresent { return repoDir } + + logger.info("Downloading Paradee \(variant.rawValue) CoreML models from HuggingFace…") + do { + try await ModelHub.download(repo, to: modelsRoot) + } catch { + throw ParadeeError.downloadFailed("\(error)") + } + return repoDir + } + + private func loadModel(repoDir: URL, fileName: String) throws -> MLModel { + let url = repoDir.appendingPathComponent(fileName) + guard FileManager.default.fileExists(atPath: url.path) else { + throw ParadeeError.modelFileNotFound(fileName) + } + let config = MLModelConfiguration() + config.computeUnits = computeUnits + do { + return try MLModel(contentsOf: url, configuration: config) + } catch { + throw ParadeeError.corruptedModel(fileName, underlying: "\(error)") + } + } + + private static func defaultCacheRoot() throws -> URL { + let root = try TtsCacheDirectory.ensure().appendingPathComponent("Models") + if !FileManager.default.fileExists(atPath: root.path) { + try FileManager.default.createDirectory(at: root, withIntermediateDirectories: true) + } + return root + } +} diff --git a/Sources/FluidAudio/TTS/Paradee/ParadeeConstants.swift b/Sources/FluidAudio/TTS/Paradee/ParadeeConstants.swift new file mode 100644 index 000000000..dfcbed610 --- /dev/null +++ b/Sources/FluidAudio/TTS/Paradee/ParadeeConstants.swift @@ -0,0 +1,30 @@ +import Foundation + +/// Fixed pipeline parameters for the Paradee-8M CoreML backend. Values mirror +/// `FluidInference/paradee-8m-coreml/config.json` and the upstream +/// `paradee/tts.py`. +public enum ParadeeConstants { + + /// Output sample rate (24 kHz mono). + public static let sampleRate = 24_000 + + /// Output samples per aligned frame: each frame carries two F0 steps of + /// 300 samples (`prod(upsample_rates) * istft_hop`). + public static let samplesPerFrame = 600 + + /// Upper bound of the acoustic model's frame axis (100 s at speed 1). + public static let maxFrames = 4_000 + + /// Phonemes per synthesis call; the text side has 512 positions, two of + /// which hold the pad token at each end. + public static let maxPhonemeLength = 510 + + /// Channels of the prosody features `d` (hidden 192 + style 32). + public static let prosodyChannels = 224 + + /// Channels of the text features `asr_tok` (the teacher decoder's width). + public static let textChannels = 512 + + /// Speech-rate multiplier. Durations scale by `1 / speed`. + public static let defaultSpeed: Float = 1.0 +} diff --git a/Sources/FluidAudio/TTS/Paradee/ParadeeError.swift b/Sources/FluidAudio/TTS/Paradee/ParadeeError.swift new file mode 100644 index 000000000..f1afacff7 --- /dev/null +++ b/Sources/FluidAudio/TTS/Paradee/ParadeeError.swift @@ -0,0 +1,38 @@ +import Foundation + +/// Errors thrown by the Paradee TTS backend. +public enum ParadeeError: Error, LocalizedError { + case notInitialized + case downloadFailed(String) + case modelFileNotFound(String) + case corruptedModel(String, underlying: String) + /// `.all` / `.cpuAndGPU` abort in MPSGraph on the LSTMs. + case unsupportedComputeUnits(String) + case inputProcessingFailed(String) + /// A chunk's predicted duration exceeded the acoustic model's frame axis. + case durationOverflow(frames: Int, maxFrames: Int) + case predictionFailed(String) + + public var errorDescription: String? { + switch self { + case .notInitialized: + return "Paradee backend is not initialized; call initialize() first." + case .downloadFailed(let detail): + return "Paradee model download failed: \(detail)" + case .modelFileNotFound(let name): + return "Paradee model file not found: \(name)" + case .corruptedModel(let name, let underlying): + return "Paradee model \(name) failed to load: \(underlying)" + case .unsupportedComputeUnits(let units): + return + "Paradee does not support compute units \(units): the LSTMs abort on the GPU. " + + "Use .cpuOnly or .cpuAndNeuralEngine." + case .inputProcessingFailed(let detail): + return "Paradee input processing failed: \(detail)" + case .durationOverflow(let frames, let maxFrames): + return "Predicted duration \(frames) frames exceeds the model limit (\(maxFrames)); shorten the input." + case .predictionFailed(let detail): + return "Paradee CoreML prediction failed: \(detail)" + } + } +} diff --git a/Sources/FluidAudio/TTS/Paradee/ParadeeManager.swift b/Sources/FluidAudio/TTS/Paradee/ParadeeManager.swift new file mode 100644 index 000000000..fc80de8b3 --- /dev/null +++ b/Sources/FluidAudio/TTS/Paradee/ParadeeManager.swift @@ -0,0 +1,194 @@ +@preconcurrency import CoreML +import Foundation + +/// Public API for the Paradee-8M CoreML TTS backend. +/// +/// Paradee (Sahil Mahendrakar, Apache-2.0) is Kokoro-82M distilled into an +/// 8.07M-parameter single-voice (`af_heart`) English model, 24 kHz. Two CoreML +/// graphs (`ParadeeText`, `ParadeeAcoustic`) with host-side duration +/// expansion; see `FluidInference/paradee-8m-coreml`. +/// +/// Paradee uses Kokoro's phoneme vocabulary and was trained on misaki en-US +/// phonemes, so the text path reuses the KokoroAne English frontend (NeMo +/// text normalization, Misaki lexicon, per-word BART G2P fallback). Like the +/// upstream package, text is synthesized one sentence at a time. +/// +/// Run on `.cpuOnly` (default) or `.cpuAndNeuralEngine`; the LSTMs abort on +/// the GPU, so `.all` / `.cpuAndGPU` are rejected. +/// +/// - Note: Beta — this is a beta model conversion; API, model artifacts, and accuracy may change. +public actor ParadeeManager { + + private let logger = AppLogger(category: "ParadeeManager") + + private let store: ParadeeModelStore + private let englishLexiconCache = LexiconAssetCache() + private var englishPhonemizer: KokoroAneEnglishPhonemizer? + private var englishFrontendReady = false + + public nonisolated var sampleRate: Int { ParadeeConstants.sampleRate } + + public init( + variant: ParadeeVariant = .int8, + directory: URL? = nil, + computeUnits: MLComputeUnits = .cpuOnly + ) { + self.store = ParadeeModelStore(variant: variant, directory: directory, computeUnits: computeUnits) + } + + /// Download (if missing) and load both models, the vocab, and the + /// English frontend assets. + public func initialize() async throws { + try await store.loadIfNeeded() + guard !englishFrontendReady else { return } + // G2PModel.loadIfNeeded only reads from cache, so fetch the assets + // explicitly first. They live at the default kokoro cache path + // (G2PModel.shared hardcodes it), not the store's `directory`. + try await KokoroAneResourceDownloader.ensureG2PAssets(directory: nil) + try await G2PModel.shared.ensureModelsAvailable() + _ = await KokoroAneResourceDownloader.ensureEnglishLexicon(directory: nil) + englishFrontendReady = true + } + + public func isAvailable() async -> Bool { + (try? await store.models()) != nil + } + + // MARK: - Synthesis + + /// Text → 24 kHz mono Float32 PCM. Sentences are synthesized separately + /// and concatenated, as in the upstream `Paradee.__call__`. + /// + /// - Parameters: + /// - speed: speech-rate multiplier (> 1 is faster). + /// - noiseSeed: seed for the harmonic source noise; equal seeds give + /// identical audio. + public func synthesize( + text: String, + speed: Float = ParadeeConstants.defaultSpeed, + noiseSeed: UInt64 = 0 + ) async throws -> [Float] { + var phonemeChunks: [String] = [] + for sentence in Self.sentences(in: text) { + phonemeChunks.append(contentsOf: Self.chunks(try await phonemes(for: sentence))) + } + return try await synthesize(chunks: phonemeChunks, speed: speed, noiseSeed: noiseSeed) + } + + /// Misaki-style phonemes → audio, bypassing text normalization and G2P. + /// Input longer than 510 phonemes is split at punctuation/whitespace. + public func synthesize( + phonemes: String, + speed: Float = ParadeeConstants.defaultSpeed, + noiseSeed: UInt64 = 0 + ) async throws -> [Float] { + try await synthesize(chunks: Self.chunks(phonemes), speed: speed, noiseSeed: noiseSeed) + } + + /// The phoneme string ``synthesize(text:speed:noiseSeed:)`` feeds the + /// model for `text` (NeMo normalization, Misaki lexicon, BART G2P). + public func phonemes(for text: String) async throws -> String { + guard englishFrontendReady else { throw ParadeeError.notInitialized } + let phonemizer = await ensureEnglishPhonemizer() + let raw = try await phonemizer.phonemize(EnglishTextNormalizer.normalizeForFrontend(text)) { word in + try await G2PModel.shared.phonemize(word: word) + } + return Self.misakiOutputForm(raw) + } + + public func cleanup() async { + await store.unload() + englishPhonemizer = nil + } + + // MARK: - Helpers + + private func synthesize(chunks: [String], speed: Float, noiseSeed: UInt64) async throws -> [Float] { + let (text, acoustic) = try await store.models() + let vocab = try await store.vocabulary() + var samples: [Float] = [] + for (index, chunk) in chunks.enumerated() { + try Task.checkCancellation() + let ids = try vocab.encode(chunk) + guard ids.count > 2 else { continue } + samples += try ParadeeSynthesizer.synthesize( + inputIds: ids, speed: speed, noiseSeed: noiseSeed &+ UInt64(index), + text: text, acoustic: acoustic) + } + guard !samples.isEmpty else { + throw ParadeeError.inputProcessingFailed("no speakable phonemes in input") + } + return samples + } + + /// Split like upstream `re.split(r"(?<=[.!?…])\s+|\n+", text)`: after + /// sentence-final punctuation followed by whitespace, and at newlines. + static func sentences(in text: String) -> [String] { + var out: [String] = [] + var current = "" + var previous: Character? + var index = text.startIndex + while index < text.endIndex { + let ch = text[index] + if ch.isNewline || (ch.isWhitespace && previous.map { ".!?…".contains($0) } == true) { + out.append(current) + current = "" + // Consume the whole whitespace run. + while index < text.endIndex, text[index].isWhitespace { + index = text.index(after: index) + } + previous = nil + continue + } + current.append(ch) + previous = ch + index = text.index(after: index) + } + out.append(current) + return out.map { $0.trimmingCharacters(in: .whitespaces) }.filter { !$0.isEmpty } + } + + /// misaki's last step for Kokoro v1.0 (`G2P.__call__`): flap `ɾ` → `T`, + /// glottal stop `ʔ` → `t`. The lexicon stores the raw forms, and Paradee + /// never saw `ɾ`/`ʔ` in training ("kittens", "satellite" come out garbled). + static func misakiOutputForm(_ phonemes: String) -> String { + phonemes.replacingOccurrences(of: "ɾ", with: "T").replacingOccurrences(of: "ʔ", with: "t") + } + + static func chunks(_ phonemes: String) -> [String] { + PhonemeChunker.chunk( + phonemes, maxLength: ParadeeConstants.maxPhonemeLength, countsUnicodeScalars: true) + } + + /// Build (and cache) the English frontend from the model vocab and the + /// Misaki lexicon cache. Without the lexicon it returns a transient + /// G2P-only frontend so the lexicon is retried on the next call. + private func ensureEnglishPhonemizer() async -> KokoroAneEnglishPhonemizer { + if let cached = englishPhonemizer { return cached } + var lower: [String: [String]] = [:] + var caseSensitive: [String: [String]] = [:] + var punctuation: Set = [] + var lexiconLoaded = false + do { + let vocab = try await store.vocabulary() + punctuation = Set(vocab.map.keys.filter { !$0.isLetter && !$0.isNumber && !$0.isWhitespace }) + if let kokoroDir = await KokoroAneResourceDownloader.ensureEnglishLexicon(directory: nil) { + try await englishLexiconCache.ensureLoaded( + kokoroDirectory: kokoroDir, allowedTokens: Set(vocab.map.keys.map(String.init))) + let maps = await englishLexiconCache.lexicons() + lower = maps.word + caseSensitive = maps.caseSensitive + lexiconLoaded = true + } + } catch { + logger.warning("English lexicon unavailable (\(error.localizedDescription)); using BART G2P only") + } + let phonemizer = KokoroAneEnglishPhonemizer( + wordToPhonemes: lower, caseSensitiveWordToPhonemes: caseSensitive, + allowedPunctuation: punctuation) + if lexiconLoaded { + englishPhonemizer = phonemizer + } + return phonemizer + } +} diff --git a/Sources/FluidAudio/TTS/Paradee/ParadeeVariant.swift b/Sources/FluidAudio/TTS/Paradee/ParadeeVariant.swift new file mode 100644 index 000000000..5717c360d --- /dev/null +++ b/Sources/FluidAudio/TTS/Paradee/ParadeeVariant.swift @@ -0,0 +1,23 @@ +import Foundation + +/// Paradee weight precision. Both variants share the same graphs, vocab and +/// English frontend; only the weights and repo subdirectory differ. +/// +/// - Note: Beta — this is a beta model conversion; API, model artifacts, and accuracy may change. +public enum ParadeeVariant: String, Sendable, CaseIterable { + /// int8 per-channel weights (12 MB). Default, as upstream recommends its int8 build. + case int8 + /// fp32 weights (34 MB); matches PyTorch to rounding. + case fp32 + + /// Subdirectory under `FluidInference/paradee-8m-coreml/`. + public var subdirectory: String { rawValue } + + /// HuggingFace repo case for this variant's subdirectory. + public var repo: Repo { + switch self { + case .int8: return .paradeeInt8 + case .fp32: return .paradeeFp32 + } + } +} diff --git a/Sources/FluidAudio/TTS/Paradee/Pipeline/ParadeeSynthesizer.swift b/Sources/FluidAudio/TTS/Paradee/Pipeline/ParadeeSynthesizer.swift new file mode 100644 index 000000000..f004034a9 --- /dev/null +++ b/Sources/FluidAudio/TTS/Paradee/Pipeline/ParadeeSynthesizer.swift @@ -0,0 +1,143 @@ +@preconcurrency import CoreML +import Foundation + +/// Drives the two-model Paradee pipeline for one chunk: +/// `ParadeeText` → host duration rounding + column expansion + source noise → +/// `ParadeeAcoustic`. Both CoreML graphs are deterministic; the only +/// randomness is the harmonic source noise generated here. +/// +/// - Note: Beta — this is a beta model conversion; API, model artifacts, and accuracy may change. +enum ParadeeSynthesizer { + + /// Synthesize one chunk of `[0, ...ids, 0]` tokens. Returns 24 kHz mono PCM. + static func synthesize( + inputIds: [Int32], + speed: Float, + noiseSeed: UInt64, + text: MLModel, + acoustic: MLModel + ) throws -> [Float] { + let tokens = inputIds.count + guard tokens >= 3, tokens <= ParadeeConstants.maxPhonemeLength + 2 else { + throw ParadeeError.inputProcessingFailed( + "token count \(tokens) out of range (3...\(ParadeeConstants.maxPhonemeLength + 2))") + } + guard speed > 0 else { + throw ParadeeError.inputProcessingFailed("speed must be > 0, got \(speed)") + } + + let textOut = try predict(text, ["input_ids": try multiArray(inputIds, shape: [1, tokens])]) + let durations = try readChannelsByTime(textOut, "duration", channels: 1, time: tokens) + let d = try readChannelsByTime(textOut, "d", channels: ParadeeConstants.prosodyChannels, time: tokens) + let asrTok = try readChannelsByTime( + textOut, "asr_tok", channels: ParadeeConstants.textChannels, time: tokens) + + let counts = frameCounts(durations: durations, speed: speed) + let frames = counts.reduce(0, +) + guard frames <= ParadeeConstants.maxFrames else { + throw ParadeeError.durationOverflow(frames: frames, maxFrames: ParadeeConstants.maxFrames) + } + + let en = expand(d, channels: ParadeeConstants.prosodyChannels, counts: counts) + let asr = expand(asrTok, channels: ParadeeConstants.textChannels, counts: counts) + var noise = [Float](repeating: 0, count: frames * ParadeeConstants.samplesPerFrame) + var rng = InflectNoise(seed: noiseSeed) + rng.fill(&noise) + + let acousticOut = try predict( + acoustic, + [ + "en": try multiArray(en, shape: [1, ParadeeConstants.prosodyChannels, frames]), + "asr": try multiArray(asr, shape: [1, ParadeeConstants.textChannels, frames]), + "noise": try multiArray(noise, shape: [1, 1, noise.count]), + ]) + return try readChannelsByTime(acousticOut, "audio", channels: 1, time: noise.count) + } + + // MARK: - Host steps + + /// `max(1, round(duration / speed))` per token, rounding half to even like + /// `torch.round` in the upstream graph. + static func frameCounts(durations: [Float], speed: Float) -> [Int] { + durations.map { max(1, Int(($0 / speed).rounded(.toNearestOrEven))) } + } + + /// Repeat column `i` of a row-major `[channels, T]` matrix `counts[i]` + /// times → `[channels, sum(counts)]` (the upstream `x @ alignment`). + static func expand(_ source: [Float], channels: Int, counts: [Int]) -> [Float] { + let tokens = counts.count + let frames = counts.reduce(0, +) + var out = [Float](repeating: 0, count: channels * frames) + source.withUnsafeBufferPointer { src in + out.withUnsafeMutableBufferPointer { dst in + for c in 0.. MLFeatureProvider { + do { + let provider = try MLDictionaryFeatureProvider( + dictionary: inputs.mapValues { MLFeatureValue(multiArray: $0) }) + return try model.prediction(from: provider) + } catch { + throw ParadeeError.predictionFailed("\(error)") + } + } + + private static func multiArray(_ values: [Float], shape: [Int]) throws -> MLMultiArray { + let arr = try MLMultiArray(shape: shape.map(NSNumber.init), dataType: .float32) + let dst = arr.dataPointer.bindMemory(to: Float.self, capacity: values.count) + values.withUnsafeBufferPointer { dst.update(from: $0.baseAddress!, count: values.count) } + return arr + } + + private static func multiArray(_ values: [Int32], shape: [Int]) throws -> MLMultiArray { + let arr = try MLMultiArray(shape: shape.map(NSNumber.init), dataType: .int32) + let dst = arr.dataPointer.bindMemory(to: Int32.self, capacity: values.count) + values.withUnsafeBufferPointer { dst.update(from: $0.baseAddress!, count: values.count) } + return arr + } + + /// Read a float32 output shaped `[1, channels, time]` or `[1, time]` into a + /// row-major `[channels, time]` buffer, honouring the array's strides + /// (dynamic-shape outputs are not guaranteed to be densely packed). + private static func readChannelsByTime( + _ provider: MLFeatureProvider, _ name: String, channels: Int, time: Int + ) throws -> [Float] { + guard let arr = provider.featureValue(for: name)?.multiArrayValue else { + throw ParadeeError.predictionFailed("missing output '\(name)'") + } + guard arr.dataType == .float32 else { + throw ParadeeError.predictionFailed("output '\(name)' is not float32") + } + let shape = arr.shape.map(\.intValue) + let strides = arr.strides.map(\.intValue) + guard shape.last == time, shape.reduce(1, *) == channels * time else { + throw ParadeeError.predictionFailed("output '\(name)' has shape \(shape), expected \(channels)×\(time)") + } + let timeStride = strides[strides.count - 1] + let channelStride = shape.count >= 2 ? strides[strides.count - 2] : 0 + let src = arr.dataPointer.bindMemory(to: Float.self, capacity: 1) + var out = [Float](repeating: 0, count: channels * time) + for c in 0.. 0 ? audioSecs / synthS : 0 + + logger.info("Paradee synthesis complete") + logger.info(" Load: \(String(format: "%.3f", loadS))s") + logger.info(" Synthesis: \(String(format: "%.3f", synthS))s") + logger.info(" Audio: \(String(format: "%.3f", audioSecs))s") + logger.info(" RTFx: \(String(format: "%.2f", rtfx))x") + logger.info(" Total: \(String(format: "%.3f", totalS))s") + logger.info(" Output: \(outURL.path)") + + if let metricsPath { + let metricsDict: [String: Any] = [ + "backend": "paradee-\(variant.rawValue)", + "text": text, + "phonemes_mode": treatAsPhonemes, + "speed": Double(speed), + "seed": seed, + "output": outURL.path, + "model_load_time_s": loadS, + "inference_time_s": synthS, + "audio_duration_s": audioSecs, + "realtime_speed": rtfx, + "total_time_s": totalS, + ] + let artifactsRoot = try ensureArtifactsRoot() + let mURL = resolveOutputURL( + metricsPath, artifactsRoot: artifactsRoot, expectsDirectory: false) + try FileManager.default.createDirectory( + at: mURL.deletingLastPathComponent(), withIntermediateDirectories: true) + let json = try JSONSerialization.data( + withJSONObject: metricsDict, options: [.prettyPrinted]) + try json.write(to: mURL) + logger.info("Metrics saved: \(mURL.path)") + } + } catch { + logger.error("Paradee Error: \(error)") + print("Paradee failed: \(error)") + exit(1) } } @@ -1549,7 +1642,13 @@ public struct TTS { --backend TTS backend: kokoro-ane (default), pocket, styletts2, supertonic3, luxtts, neutts (beta), inflect (beta), chatterbox (beta, multilingual, macOS 15+), - chatterbox-nano (beta, English + [laugh]/[chuckle] tags, macOS 15+) + chatterbox-nano (beta, English + [laugh]/[chuckle] tags, macOS 15+), + paradee (beta, 8M English, voice af_heart) + Paradee: + --variant int8|fp32 weights (default int8) + --speed 1.0 speech rate (default 1.0) + --seed N source-noise seed (default 0) + --phonemes treat input as misaki phonemes Chatterbox (built-in voice, 18 languages): --lang de language code (default en) --seed N sampling seed (default 42) diff --git a/Sources/FluidAudioCLI/Commands/TtsBenchmarkCommand.swift b/Sources/FluidAudioCLI/Commands/TtsBenchmarkCommand.swift index 79dc10872..13bc79adf 100644 --- a/Sources/FluidAudioCLI/Commands/TtsBenchmarkCommand.swift +++ b/Sources/FluidAudioCLI/Commands/TtsBenchmarkCommand.swift @@ -341,6 +341,12 @@ public enum TtsBenchmarkCommand { phrases: phrases, corpusLabel: corpusLabel, preset: preset, outputJson: outputJson, audioDir: audioDir, asrChoice: asrChoice) + case .paradee: + try await runParadee( + phrases: phrases, corpusLabel: corpusLabel, + variant: ParadeeVariant(rawValue: variantArg?.lowercased() ?? "") ?? .int8, + preset: preset, outputJson: outputJson, audioDir: audioDir, + asrChoice: asrChoice) } } catch { logger.error("tts-benchmark failed: \(error)") @@ -837,6 +843,60 @@ public enum TtsBenchmarkCommand { } } + private static func runParadee( + phrases: [(category: String, text: String)], + corpusLabel: String, + variant: ParadeeVariant, + preset: TtsComputeUnitPreset, + outputJson: String?, + audioDir: String?, + asrChoice: AsrChoice + ) async throws { + // The LSTMs abort on the GPU, so only CPU and ANE routings exist. + let appliedPreset: TtsComputeUnitPreset = preset == .allAne ? .allAne : .cpuOnly + if appliedPreset != preset, preset != .default { + logger.warning("Paradee runs .cpuOnly or .cpuAndNeuralEngine; --compute-units \(preset.cliValue) ignored.") + } + let manager = ParadeeManager( + variant: variant, computeUnits: appliedPreset == .allAne ? .cpuAndNeuralEngine : .cpuOnly) + let coldStart = Date() + try await manager.initialize() + let coldStartS = Date().timeIntervalSince(coldStart) + logger.info(String(format: "Cold start (initialize): %.2fs", coldStartS)) + + let firstStart = Date() + _ = try await manager.synthesize(text: "Initialization warm-up.") + let firstSynthMs = Date().timeIntervalSince(firstStart) * 1000 + logger.info(String(format: "First synth: %.0f ms", firstSynthMs)) + + try await runPhraseLoop( + backendId: "paradee-\(variant.rawValue)", + voiceLabel: "af_heart", + corpusLabel: corpusLabel, + phrases: phrases, + preset: appliedPreset, + coldStartS: coldStartS, + firstSynthMs: firstSynthMs, + outputJson: outputJson, + audioDir: audioDir, + asrChoice: asrChoice, + normalizeWavs: false, + extraSummary: ["language": "en", "seed": 0] + ) { text in + let t0 = Date() + let samples = try await manager.synthesize(text: text) + let synthMs = Date().timeIntervalSince(t0) * 1000 + return BackendPhraseSample( + synthMs: synthMs, + ttftMs: synthMs, + samples: samples, + sampleRate: ParadeeConstants.sampleRate, + stageMs: [:], + extraFields: [:] + ) + } + } + /// Map `--language` or a `minimax-` corpus name onto a Chatterbox /// language code. Falls back to English. private static func resolveChatterboxLanguage(explicit: String?, corpus: String) -> String { @@ -1121,6 +1181,7 @@ public enum TtsBenchmarkCommand { case supertonic3 case chatterbox case chatterboxNano + case paradee var defaultCorpus: String { return "minimax-english" @@ -1141,6 +1202,8 @@ public enum TtsBenchmarkCommand { return .chatterbox case "chatterbox-nano", "chatterboxnano": return .chatterboxNano + case "paradee", "paradee-8m": + return .paradee default: logger.warning("Unknown backend '\(name)' — defaulting to kokoro-ane") return .kokoroAne diff --git a/Tests/FluidAudioTests/TTS/Paradee/ParadeeTests.swift b/Tests/FluidAudioTests/TTS/Paradee/ParadeeTests.swift new file mode 100644 index 000000000..4bdaf10b2 --- /dev/null +++ b/Tests/FluidAudioTests/TTS/Paradee/ParadeeTests.swift @@ -0,0 +1,86 @@ +import CoreML +import XCTest + +@testable import FluidAudio + +/// Paradee repo wiring and the host-side steps between the two CoreML graphs. +final class ParadeeTests: XCTestCase { + + // MARK: - Repo wiring + + func testRepoPathsAndSubdirectories() { + XCTAssertEqual(Repo.paradeeInt8.remotePath, "FluidInference/paradee-8m-coreml") + XCTAssertEqual(Repo.paradeeFp32.remotePath, "FluidInference/paradee-8m-coreml") + XCTAssertEqual(Repo.paradeeInt8.subPath, "int8") + XCTAssertEqual(Repo.paradeeFp32.subPath, "fp32") + XCTAssertEqual(Repo.paradeeInt8.folderName, "paradee-8m-coreml/int8") + XCTAssertEqual(Repo.paradeeFp32.folderName, "paradee-8m-coreml/fp32") + } + + func testVariantMapsToRepo() { + XCTAssertEqual(ParadeeVariant.int8.repo, .paradeeInt8) + XCTAssertEqual(ParadeeVariant.fp32.repo, .paradeeFp32) + } + + func testRequiredModels() { + let required: Set = ["ParadeeText.mlmodelc", "ParadeeAcoustic.mlmodelc", "vocab.json"] + XCTAssertEqual(ModelNames.Paradee.requiredModels, required) + XCTAssertEqual(ModelNames.getRequiredModelNames(for: .paradeeInt8, variant: nil), required) + XCTAssertEqual(ModelNames.getRequiredModelNames(for: .paradeeFp32, variant: nil), required) + } + + func testGpuComputeUnitsRejected() { + XCTAssertTrue(ParadeeModelStore.isSupported(.cpuOnly)) + XCTAssertTrue(ParadeeModelStore.isSupported(.cpuAndNeuralEngine)) + XCTAssertFalse(ParadeeModelStore.isSupported(.all)) + XCTAssertFalse(ParadeeModelStore.isSupported(.cpuAndGPU)) + } + + // MARK: - Durations + + func testFrameCountsRoundHalfToEvenLikeTorch() { + // torch.round: 0.5 → 0, 1.5 → 2, 2.5 → 2; then max(1, ·). + XCTAssertEqual( + ParadeeSynthesizer.frameCounts(durations: [0.5, 1.5, 2.5, 3.49, 0.1], speed: 1), + [1, 2, 2, 3, 1]) + } + + func testFrameCountsScaleWithSpeed() { + XCTAssertEqual(ParadeeSynthesizer.frameCounts(durations: [4, 6], speed: 2), [2, 3]) + XCTAssertEqual(ParadeeSynthesizer.frameCounts(durations: [4, 6], speed: 0.5), [8, 12]) + } + + // MARK: - Expansion + + func testExpandRepeatsColumns() { + // [2 channels, 3 tokens], counts [2, 1, 3] → [2, 6]. + let source: [Float] = [1, 2, 3, 10, 20, 30] + XCTAssertEqual( + ParadeeSynthesizer.expand(source, channels: 2, counts: [2, 1, 3]), + [1, 1, 2, 3, 3, 3, 10, 10, 20, 30, 30, 30]) + } + + // MARK: - Text frontend helpers + + func testSentenceSplitMatchesUpstreamRegex() { + XCTAssertEqual( + ParadeeManager.sentences(in: "Hello there. How are you? Fine!\nNew line… End"), + ["Hello there.", "How are you?", "Fine!", "New line…", "End"]) + // No split without whitespace after the punctuation ("3.5", "a.b"). + XCTAssertEqual(ParadeeManager.sentences(in: "Pi is 3.14 today."), ["Pi is 3.14 today."]) + } + + func testMisakiOutputFormRewritesFlapAndGlottalStop() { + XCTAssertEqual(ParadeeManager.misakiOutputForm("kˈɪʔnz sˈæɾᵊlˌIt"), "kˈɪtnz sˈæTᵊlˌIt") + } + + func testChunksRespectPhonemeLimit() { + let word = "həlˈO " + let long = String(repeating: word, count: 200) + let chunks = ParadeeManager.chunks(long) + XCTAssertGreaterThan(chunks.count, 1) + for chunk in chunks { + XCTAssertLessThanOrEqual(chunk.unicodeScalars.count, ParadeeConstants.maxPhonemeLength) + } + } +}