AI-powered tool to turn long videos into short, viral-ready clips. Combines transcription, speaker diarization, scene detection & 9:16 resizing — perfect for creators & smart automation.
-
Updated
Apr 2, 2025 - Python
AI-powered tool to turn long videos into short, viral-ready clips. Combines transcription, speaker diarization, scene detection & 9:16 resizing — perfect for creators & smart automation.
speech to text gui for different (e.g. Whisper, Voxtral) models and backends, including whisper.cpp, crispasar, mlx-whisper, faster-whisper, ctranslate2; applies pyannote for diarization
ONNX implementation of Pyannote Speaker Diarization 3.1 pipeline.
A Tool to Transcribe Audio Files with Speaker Diarization
Audio Diarization and Classification to perform Classroom Activity Detection (detect teacher, student, multiple speakers) over an input video.
[Graduation Project] Flow: Enterprise Meeting Content Digitization & Intelligence Platform featuring Modular Monolith, Async AI Worker, Vector Semantic Search (pgvector), and Docker stack.
ONNX implementation of Pyannote Speaker Diarization Community-1 pipeline.
Speaker diarization that exports timeline-aligned tracks for video editing
A multimodal speaker diarization system using audio, video, and dialogue cues 🗣️💬
An intelligent Streamlit application to transcribe and analyze multi-speaker medical consultations. This tool automatically identifies who spoke when (diarization), transcribes their speech (ASR), and assigns their role (Clinician or Patient), even in conversations that mix English and other languages like Hindi or Tamil.
C++ port of the pyannote framework with exclusive diarization feature from community-1
This repository is to experiment the integration with @ggml-org/whisper.cpp for offline STT + pyannote/speaker-diarization-3.1
WebSocket based Python implementation that streams live audio to the Deepgram API for real-time transcription and speaker diarization.
AI-Powered Speech Recognition & Diarization: A robust Streamlit application leveraging WhisperX and Faster-Whisper for accurate transcription and speaker separation. Features dual-mode processing (Fast/Pro), automatic speaker identification, color-coded Word (.docx) export, and CPU-optimized Docker deployment on AWS EC2.
FastAPI app for podcast transcription with automatic speaker diarization using pyannote.audio and faster-whisper.
A simple protocol manager for your audios
Self-hosted meeting transcription with speaker attribution, searchable notes, and live streaming
Transcription audio locale via whisper.cpp — NestJS + Next.js + pyannote.audio
Event-driven pipeline for evaluating recruiting interviews (FastAPI, Redis Queue, PostgreSQL) with pluggable transcription (WhisperX / OpenAI), speaker diarization (Pyannote), structured LLM scoring via OpenRouter, and human-in-the-loop evidence validation.
Control a real Android phone from Claude Code via MCP + Shizuku—no root needed, no credentials on device.
To associate your repository with the pyannote-audio topic, visit your repo's landing page and select "manage topics."