Skip to content

Repository files navigation

Blaizzy%2Fmlx-audio-swift | Trendshift

MLX Audio Swift

A modular Swift SDK for audio processing with MLX on Apple Silicon

Platform Swift License

Architecture

MLXAudio follows a modular design allowing you to import only what you need:

  • MLXAudioCore: Base types, protocols, and utilities
  • MLXAudioCodecs: Audio codec implementations (SNAC, Encodec, Vocos, Mimi, DACVAE, Descript DAC, Fish S1 DAC, S3TokenizerV2, MOSS Audio Tokenizer, Higgs Audio Tokenizer, Step-Audio-2 token-to-wav)
  • MLXAudioTTS: Text-to-Speech models (Breeze TTS 2, Qwen3-TTS, OmniVoice, Fish Audio S2 Pro, IndexTTS, Soprano, VyvoTTS, Orpheus, MOSS-TTS, Marvis TTS, Pocket TTS, Irodori TTS)
  • MLXAudioSTT: Speech-to-Text models (Qwen3-ASR, Qwen3-ForcedAligner, Voxtral Realtime, Cohere Transcribe, Parakeet, Nemotron ASR, GLMASR, FireRedASR2, SenseVoice, Granite Speech, Whisper, Canary, Moonshine, Wav2Vec2, MMS)
  • MLXAudioVAD: Voice Activity Detection & Speaker Diarization (Sortformer, SmartTurn, FSMN VAD, Silero VAD)
  • MLXAudioSTS: Speech-to-Speech models (LFM2.5-Audio, SAM-Audio, MossFormer2-SE, DeepFilterNet)
  • MLXAudioUI: SwiftUI components for audio interfaces

Installation

Add MLXAudio to your project using Swift Package Manager:

dependencies: [
    .package(url: "https://github.com/Blaizzy/mlx-audio-swift.git", branch: "main")
]

// Import only what you need
.product(name: "MLXAudioTTS", package: "mlx-audio-swift"),
.product(name: "MLXAudioCore", package: "mlx-audio-swift")

Quick Start

Text-to-Speech

import MLXAudioTTS
import MLXAudioCore

// Load a TTS model from HuggingFace
let model = try await SopranoModel.fromPretrained("mlx-community/Soprano-80M-bf16")

// Generate audio
let audio = try await model.generate(
    text: "Hello from MLX Audio Swift!",
    parameters: GenerateParameters(
        maxTokens: 200,
        temperature: 0.7,
        topP: 0.95
    )
)

// Save to file
try saveAudioArray(audio, sampleRate: Double(model.sampleRate), to: outputURL)

Speech-to-Text

import MLXAudioSTT
import MLXAudioCore

// Load audio file
let (sampleRate, audioData) = try loadAudioArray(from: audioURL)

// Load STT model
let model = try await GLMASRModel.fromPretrained("mlx-community/GLM-ASR-Nano-2512-4bit")

// Transcribe
let output = model.generate(audio: audioData)
print(output.text)

Speaker Diarization

import MLXAudioVAD
import MLXAudioCore

// Load audio file
let (sampleRate, audioData) = try loadAudioArray(from: audioURL)

// Load diarization model
let model = try await SortformerModel.fromPretrained(
    "mlx-community/diar_streaming_sortformer_4spk-v2.1-fp16"
)

// Detect who is speaking when
let output = try await model.generate(audio: audioData, threshold: 0.5)
for segment in output.segments {
    print("Speaker \(segment.speaker): \(segment.start)s - \(segment.end)s")
}

Streaming Generation

for try await event in model.generateStream(text: text, parameters: parameters) {
    switch event {
    case .token(let token):
        print("Generated token: \(token)")
    case .audio(let audio):
        print("Final audio shape: \(audio.shape)")
    case .info(let info):
        print(info.summary)
    }
}

Supported Models

TTS Models

Model Model README HuggingFace Repo
Breeze TTS 2 Breeze TTS 2 README mlx-community/Breeze-TTS-2-mlx-4bit
Qwen3-TTS Qwen3-TTS README mlx-community/Qwen3-TTS-12Hz-0.6B-Base-8bit
OmniVoice OmniVoice README mlx-community/OmniVoice
Fish Audio S2 Pro Fish Audio S2 Pro README mlx-community/fish-audio-s2-pro-8bit
Soprano Soprano README mlx-community/Soprano-80M-bf16
VyvoTTS VyvoTTS README mlx-community/VyvoTTS-EN-Beta-4bit
Orpheus Orpheus README mlx-community/orpheus-3b-0.1-ft-bf16
MOSS-TTS MOSS-TTS README OpenMOSS-Team/MOSS-TTS, OpenMOSS-Team/MOSS-TTSD-v1.0, OpenMOSS-Team/MOSS-TTS-Local-Transformer
IndexTTS — mlx-community/IndexTTS, mlx-community/IndexTTS-1.5
Marvis TTS Marvis TTS README Marvis-AI/marvis-tts-250m-v0.2-MLX-8bit
Pocket TTS Pocket TTS README mlx-community/pocket-tts
Irodori TTS Irodori TTS README mlx-community/Irodori-TTS-600M-v3-VoiceDesign-8bit
Spark-TTS Spark-TTS README mlx-community/Spark-TTS-0.5B-bf16

STT Models

Model Model README HuggingFace Repo
Qwen3-ASR Qwen3-ASR README mlx-community/Qwen3-ASR-1.7B-bf16
Qwen3-ForcedAligner Qwen3-ASR README mlx-community/Qwen3-ForcedAligner-0.6B-bf16
MOSS-Transcribe-Diarize MOSS-Transcribe-Diarize README OpenMOSS-Team/MOSS-Transcribe-Diarize
Voxtral Realtime Voxtral README mlx-community/Voxtral-Mini-4B-Realtime-2602-fp16
Cohere Transcribe Cohere Transcribe README beshkenadze/cohere-transcribe-03-2026-mlx-fp16
Parakeet Parakeet README mlx-community/parakeet-tdt-0.6b-v3
Nemotron ASR Nemotron ASR README Nemotron Speech Streaming, Nemotron 3.5 ASR
GLMASR GLMASR README mlx-community/GLM-ASR-Nano-2512-4bit
FireRedASR2 FireRedASR2 README Converted FireRedASR2-compatible MLX checkpoints
SenseVoice SenseVoice README Converted SenseVoice-compatible MLX checkpoints
Granite Speech Granite Speech README Converted Granite Speech-compatible MLX checkpoints
Whisper Whisper README openai/whisper-large-v3-turbo, mlx-community/whisper-large-v3-turbo, and every other openai/whisper-* / mlx-community/whisper-* size and .en variant
Canary — Mediform/canary-1b-v2-mlx-q8, Canary-compatible MLX/NeMo checkpoints
Moonshine — UsefulSensors/moonshine-tiny, Moonshine-compatible MLX checkpoints
Wav2Vec2 CTC — facebook/wav2vec2-base-960h, Wav2Vec2 CTC-compatible checkpoints
MMS — facebook/mms-1b-fl102, MMS adapter checkpoints

Audio Codecs

Codec Notes HuggingFace Repo
SNAC Neural audio codec with encode/decode support mlx-community/snac_24khz
Encodec Encodec-compatible audio codec runtime Converted Encodec-compatible MLX checkpoints
Vocos Vocoder/codec decode components Converted Vocos-compatible MLX checkpoints
Mimi Mimi encoder/decoder codec used by speech models Mimi-compatible MLX checkpoints
DACVAE DAC-style VAE audio codec Converted DACVAE-compatible MLX checkpoints
Descript DAC Descript DAC-compatible audio codec Descript DAC-compatible checkpoints
Fish S1 DAC Fish Speech S1 audio codec Fish S1 DAC-compatible checkpoints
S3TokenizerV2 S3 acoustic tokenizer exposed in MLXAudioCodecs mlx-community/S3TokenizerV2
MOSS Audio Tokenizer MOSS audio tokenizer runtime shared with MOSS TTS models mlx-community/MOSS-Audio-Tokenizer-Nano
Higgs Audio Tokenizer Higgs acoustic tokenizer decode and acoustic encode support bosonai/higgs-audio-v3-tts-4b bundled tokenizer weights
Step-Audio-2 Token2Wav Token-to-waveform stack for Step-Audio-2 style prompts mlx-community/Step-Audio-2-token2wav

STS Models

Model Model README HuggingFace Repo
LFM2.5-Audio LFM Audio README mlx-community/LFM2.5-Audio-1.5B-6bit
SAM-Audio SAM Audio README mlx-community/sam-audio-large-fp16
MossFormer2-SE — starkdmi/MossFormer2-SE-fp16
DeepFilterNet DeepFilterNet README mlx-community/DeepFilterNet-mlx

VAD / Speaker Diarization Models

Model Model README HuggingFace Repo
Sortformer Sortformer README mlx-community/diar_streaming_sortformer_4spk-v2.1-fp16
SmartTurn SmartTurn README mlx-community/smart-turn-v3
FSMN VAD — mlx-community/fsmn-vad
Silero VAD Silero VAD README Silero VAD-compatible MLX checkpoints

Features

  • Modular architecture for minimal app size - import only what you need
  • Automatic model downloading from HuggingFace Hub
  • Native async/await support for seamless Swift integration
  • Streaming audio generation for real-time TTS
  • Type-safe Swift API with comprehensive error handling
  • Optimized for Apple Silicon with MLX framework

Advanced Usage

Custom Generation Parameters

let parameters = GenerateParameters(
    maxTokens: 1200,
    temperature: 0.7,
    topP: 0.95,
    repetitionPenalty: 1.5,
    repetitionContextSize: 30
)

let audio = try await model.generate(text: "Your text here", parameters: parameters)

Audio Codec Usage

import MLXAudioCodecs

// Load SNAC codec
let snac = try await SNAC.fromPretrained("mlx-community/snac_24khz")

// Encode audio to tokens
let tokens = try snac.encode(audio)

// Decode tokens back to audio
let reconstructed = try snac.decode(tokens)

Voice Selection for Multi-Voice Models

// For models supporting multiple voices (like LlamaTTS/Orpheus)
let audio = try await model.generate(
    text: "Hello!",
    voice: "tara",  // Options: tara, leah, jess, leo, dan, mia, zac, zoe
    parameters: parameters
)

Requirements

  • macOS 14+ or iOS 17+
  • Apple Silicon (M1 or later) recommended for optimal performance
  • Xcode 15+
  • Swift 5.9+

Examples

Check out the Examples/VoicesApp directory for a complete SwiftUI application demonstrating:

  • Loading and running TTS models
  • Playing generated audio
  • UI components for model interaction

Additional usage examples can be found in the test files.

Contributing

See CONTRIBUTING.md for contribution guidelines.

Credits

License

MIT License - see LICENSE file for details.

About

A modular Swift SDK for audio processing with MLX on Apple Silicon

Topics

Resources

Code of conduct

Contributing

Stars

786 stars

Watchers

11 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages