Replies: 1 comment
|
| Engine | Languages | Notes |
|---|---|---|
| WhisperX (default) | ~100 languages | Dubbing, subtitles, word-level timing |
| Faster-Whisper | ~100 languages | General cross-platform transcription |
| MLX Whisper | ~100 languages | Apple Silicon only |
| PyTorch Whisper | ~100 languages | CUDA, MPS, CPU fallback |
| Parakeet TDT | 25+ European languages | Fast CPU/CUDA, EU-focused |
| FunASR (SenseVoice) | 50+ languages | VAD and inline diarization with cam++ |
| sherpa-onnx-asr | Model-dependent | Live streaming CPU dictation |
Backend API Schema (backend/api/routers/audiobook.py) - Current Behavior:
class AudiobookRequest(ExpressiveMixin):
text: str
language: str | None = None # None/"Auto" → profile language or autodetect
def _resolve_default_language(language: str | None, default_voice: str | None) -> str | None:
"""Language resolution logic (#505)"""
if language and language != "Auto":
return language
# Falls back to profile language if available, else None means autodetect
if default_voice:
prof_lang = get_profile_language(default_voice) # database call
if prof_lang and prof_lang != "Auto":
return prof_lang
# Return None for genuine autodetect (engine decides!)
return NoneThe engine receives None or None/Null/None when language=Auto, which tells the Whisper model to use its internal fallback language (which defaults to English, but can drift to any of ~100 languages based on training data distribution).
Database Schema (backend/core/db.py):
CREATE TABLE voice_profiles (
id TEXT PRIMARY KEY,
name TEXT NOT NULL,
ref_audio_path TEXT,
ref_text TEXT DEFAULT '',
instruct TEXT DEFAULT '',
language TEXT DEFAULT 'Auto', -- This is for TTS output language!
locked_audio_path TEXT DEFAULT '',
seed INTEGER DEFAULT NULL,
is_locked INTEGER DEFAULT 0,
created_at REAL
);
CREATE TABLE dub_history (
id TEXT PRIMARY KEY,
filename TEXT,
duration REAL,
segments_count INTEGER,
language TEXT, -- The autodetected/transcribed language
language_code TEXT, -- WhisperX-specific field
...Key Issue: The database stores the transcription result's language, not a constraint on which languages can be autodetected. The ASR model is given no restriction when language=None/Null/None (autodetect mode).
💡 Feature Requests for ASR Language Filtering
Option 1: ASR Whitelist Configuration (HIGH PRIORITY - Recommended)
Allow users to specify which languages the ASR engine should consider during autodetection. This prevents hallucinated transcripts in unrelated languages.
Configuration approaches:
A. Environment variables (simple, no UI change needed):
# Allow only these languages in order of priority
OMNIVOICE_ASR_LANGUAGE_WHITELIST=fr,es,pt,en,zh,ja # French, Spanish, Portuguese + common ones
# OR use blacklist to block certain languages:
# OMNIVOICE_ASR_LANGUAGE_BLACKLIST=ru,th,ar... (too specific, whitelist better)
# Optional per-engine settings if needed:
OMNIVOICE_WHISPERX_LANGUAGES=fr,es,pt,en,zh B. Database field (requires migration but persistent):
Add a new allowed_asr_languages column to relevant tables or store in user preferences:
-- Option A: User-specific preference table
CREATE TABLE IF NOT EXISTS user_preferences (
id TEXT PRIMARY KEY,
key TEXT UNIQUE, -- e.g., "asr_whitelist_languages"
value TEXT -- JSON array: ["fr", "es", "pt"] or comma-separated string
);
-- Option B: Add to voice_profiles for project-specific configs
ALTER TABLE voice_profiles ADD COLUMN allowed_asr_languages TEXT DEFAULT NULL;C. UI Settings (ideal user experience):
Add a new section in Settings → Models or a dedicated Settings → ASR Language Preferences:
- Whitelist mode: Only allow these languages during autodetection
- Dropdown/multi-select for common languages (or "all" to current behavior)
- Default fallback language when autodetect isn't confident
- Language history cache per project (remember last used language for each dub job)
D. Backend API request modification:
Extend the transcription/dubbing endpoints:
class TranscribeRequest(ExpressiveMixin):
model: str = "whisperx" # or faster-whisper, mlx-whisper, etc.
language: str | None = "Auto" # legacy for TTS compatibility
allowed_languages: list[str] = None # NEW: ["fr", "es", "pt"] - filter ASR engine's choices
fallback_language: str | None = None # NEW: Use if none detected or confidence low
# For dubbing specifically:
source_video_language: str | None = None # Extract from video metadata/first 10 seconds
def transcribe_audio_with_filters(
whisper_model,
audio_path: Path,
allowed_languages: list[str], # NEW parameter
source_hint: str, # Optional: hint from dubbing pipeline
) -> TranscriptResult:
"""Modify ASR initialization based on whitelist"""
lang_to_use = None
if source_hint: # From dubbing metadata (e.g., video language)
lang_to_use = source_hint
if allowed_languages: # Whitelist mode
# Initialize Whisper with specific language configuration
asr_config = {"language": allowed_languages} # pseudo-code, depends on engine wrapper
# Or filter the model's internal fallback behavior
lang_to_use = _pick_best_match_transcribe(audio_path, asr_config)
return transcribe_model(whisper_model, audio_path, language=lang_to_use or "Auto")Option 2: Smart Fallback with Confidence Threshold (MEDIUM PRIORITY)
When auto-detection is uncertain (< confidence threshold), use a configured fallback instead of random autodetection.
Example:
# Default to the most common language in the user's dub history or profile languages
DEFAULT_FALLBACK = "fr" # Or configurable per user
CONFIDENCE_THRESHOLD = 0.6 # Below this, use fallback instead of guessingThis ensures consistent behavior across different audio inputs without hallucinating unrelated languages when the model is unsure.
Option 3: Multilingual Content Handling (MEDIUM PRIORITY)
When content contains mixed languages (e.g., French + Spanish), offer options in UI:
- Keep the first detected language throughout
- Segment by language and ask user to confirm each segment
- Base on source video metadata or dubbing project settings
This addresses real multilingual use cases where users legitimately work with multiple related languages.
Option 4: Engine-Specific Language Constraints (LOW PRIORITY)
When switching ASR engines (e.g., from WhisperX to Moonshine), warn if the switch might affect language detection behavior and ask for confirmation.
Note: This is less impactful than Options 1-3 but provides transparency about engine changes affecting transcript quality.
Option 5: ASR Engine Per-Language Model Selection (ADVANCED)
Different ASR engines have different optimal language handling:
- WhisperX has best coverage (~100 languages, trained on multilingual data)
- Moonshine specializes in low-resource/rapid streaming, limited to English primarily
- FunASR with cam++ has 50+ languages including better European dialect support
Allow users to select which ASR models to use for specific language ranges, and automatically load/download only matching languages. This is more complex but could optimize performance:
# Pseudo-code: Match engine choice to allowed languages
def select_asr_engine_for_allowed_languages(allowed_list, user_hardware):
"""Choose best ASR engine based on allowed languages"""
# If only European languages (fr, de, es, pt, it, nl, ...)
if all(is_european_lang(l) for l in allowed_list):
return FunASR # Better coverage for EU languages
# Or if user primarily uses French/English/Spanish -> use WhisperX default
elif set(allowed_list).intersection({"fr", "en", "es", "pt"}):
return WhisperX
return FasterWhisper # General purpose fallback🎯 Impact Assessment
User Benefits:
| Use Case | Without Whitelist (Now) | With ASR Language Filtering (Requested) |
|---|---|---|
| Dictation widget | Multilingual meetings produce transcripts in random languages, many unreadable | Only transcribed in known/fallback language when confident OR user-set languages |
| Video dubbing | Auto-dub projects may hallucinate unrelated languages | Consistent translation/transcription workflow without cross-language corruption |
| Confidential/local first | Loading 646 TTS languages + ~100 ASR languages = unnecessary resource use | Only load matching language models, faster startup, better memory efficiency |
| Batch dubbing jobs | Wrong autodetection on low-quality audio or noisy recordings | Consistent transcription in expected languages |
| Performance benchmarks | GPU/CPU cycles wasted on irrelevant language detection runs | Optimized model selection and reduced computation |
Technical Benefits:
- ✅ Reduced memory footprint: Load only needed ASR models instead of downloading all ~100 Whisper variants
- ✅ Faster startup: Skip loading unused language models
- ✅ Better quality: Engine can focus on specific languages it's trained optimally for rather than guessing wrong ones
- ✅ Predictable workflow: Dubbing pipeline stays on track without unexpected language switches
Privacy Benefits:
- 🔒 When processing sensitive content, only transcribe in authorized languages
- 🔒 Less data leaves your machine when ASR models are constrained to whitelisted languages
🛠️ Implementation Considerations
Backend (FastAPI routes):
The main changes would be in these API endpoints:
/v1/audio/transcriptions- Main STT endpoint with newallowed_languagesparameter- Dubbing pipeline (
/dub/) - May need to auto-detect source language from video then pass to TTS
Note on ASR engine specifics:
- WhisperX: Has a
--modelflag for language-specific models in its original repo. VoiceStudio wraps this with thelanguageparameter sent during generation. Whenlanguage=None(Auto), uses English or default training language based on model weights.- Reference: https://github.com/m-bain/whisperx mentions loading multilingual models that can be biased toward certain languages when not configured properly
- Faster-Whisper: Uses similar multilingual Whisper but with int8 optimizations, same autodetection behavior
- FunASR: Has better multi-language support and is more forgiving with mixed-language content, which is why it's listed as "50+ languages" in the engine table
Frontend (Settings):
New UI section:
Settings → ASR Language Preferencesor integrate with existing language configuration- Mode selection:
Auto (all)/Whitelist only - Language picker/multi-select for whitelist
- Fallback language setting
- Per-project settings for dubbing workflows
Migration Path:
- Default behavior unchanged on first launch (backward compat)
- Users explicitly opt into whitelist mode via Settings or env vars
- No breaking changes to existing API endpoints, just additive parameters
📚 Related Issues / Documentation References
- Issue [Bug] Long-form (5+ min) OmniVoice TTS degrades: repeated/skipped/mispronounced words, Japanese voice clone #505: Language resolution logic in audiobook backend ("None/"Auto" → profile language, else autodetect")
- Current database defaults:
voice_profiles.language TEXT DEFAULT 'Auto'- This defaults the TTS output, not ASR input - API schema showing language parameter usage in
backend/api/routers/profiles.py,audiobook.py
🙏 Thank You!
P.S. If this is too technically complex to implement without architectural changes, please document the limitations explicitly (and any future roadmap timeline for ASR language filtering) in the FAQ or troubleshooting docs so users know what's possible.
Send this issue on https://github.com/debpalash/VoiceStudio/issues as a new feature request.
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
All reactions