The SaaS Cash Cow β A $0-cost AI voice agent that answers calls, books appointments, and updates a mock CRM. Built for your Upwork portfolio to win high-ticket clients.
Ava is a real-time, ultra-low-latency AI voice receptionist for Brightside Dental Clinic. When a "caller" connects via the web interface:
- Captures audio via WebRTC (browser microphone)
- Transcribes speech locally using OpenAI Whisper (
tinymodel) - Generates intelligent responses via Groq's LLaMA 3 (500+ tokens/sec)
- Speaks back using Microsoft Edge TTS (free, natural-sounding voice)
- Books appointments in a mock Google Calendar
- Updates CRM (mock HubSpot/Salesforce) with lead data
Total latency: ~300β800ms (well under the 500ms target for natural conversation).
βββββββββββββββββββ WebSocket βββββββββββββββββββββββββββββββββββββββ
β Browser β ββββββββββββββββββββΊ β FastAPI Backend (Python) β
β (WebRTC mic) β audio + text β β
βββββββββββββββββββ β βββββββββββββββ ββββββββββββββββ β
β β Whisper β β Groq LLM β β
β β (local) β β (API/Free) β β
β β STT β β llama3-8b β β
β βββββββββββββββ ββββββββββββββββ β
β β β β
β βββββββββββββββ ββββββββββββββββ β
β β Edge TTS β β Mock Calendarβ β
β β (free) β β + CRM β β
β βββββββββββββββ ββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββ
| Component | Technology | Cost |
|---|---|---|
| Phone/Audio Input | WebRTC (browser mic) | $0 |
| WebSocket Server | FastAPI + Uvicorn | $0 |
| Speech-to-Text | openai-whisper / faster-whisper (tiny) |
$0 (local CPU) |
| LLM Brain | Groq API (LLaMA 3 8B) | $0 (free tier: 1.5M tokens/day) |
| Text-to-Speech | edge-tts (Microsoft Edge voices) | $0 |
| Calendar | Mock Google Calendar (in-memory) | $0 |
| CRM | Mock HubSpot/Salesforce (in-memory) | $0 |
- Python 3.9+
- A free Groq API key β console.groq.com/keys
- Microphone access in your browser
cd voice-receptionistpython -m venv venv
# macOS/Linux
source venv/bin/activate
# Windows
venv\Scripts\activatepip install -r requirements.txtNote: The first run will download the Whisper
tinymodel (~39 MB). This is automatic.
cp .env.example .envEdit .env and add your Groq API key:
GROQ_API_KEY=gsk_your_actual_key_herepython app.pyYou should see:
π Loading Whisper model (tiny)...
β
Whisper model loaded.
INFO: Uvicorn running on http://0.0.0.0:8000
Navigate to: http://localhost:8000
-
Allow microphone access when prompted by your browser.
-
Press and hold the blue "Hold to Talk" button.
-
Speak naturally β e.g., "Hi, I'd like to book a teeth cleaning for tomorrow."
-
Release the button. Ava will:
- Transcribe your speech (Whisper)
- Think of a response (Groq LLM)
- Speak back to you (Edge TTS)
- If a booking is confirmed, update the dashboard
-
Watch the sidebar β appointments and CRM leads update in real time.
You: "Hello, I'd like to make an appointment."
Ava: "Of course! I'd be happy to help you book an appointment at Brightside Dental. May I have your name, please?"
You: "My name is David Chen."
Ava: "Nice to meet you, David. What service are you looking for? We offer general checkups, teeth cleaning, consultations, root canals, and whitening."
You: "I need a teeth cleaning."
Ava: "Great choice. When would you prefer to come in? We're open Monday through Friday, 9 AM to 5 PM."
You: "Tomorrow at 10 AM."
Ava: "Perfect! I've booked your teeth cleaning for tomorrow at 10 AM. Your appointment ID is APPT-0003. We look forward to seeing you at 123 Smile Street. Is there anything else I can help you with?"
voice-receptionist/
βββ app.py # FastAPI backend (WebSockets, AI pipeline, mock CRM/Calendar)
βββ static/
β βββ index.html # Professional dark-themed frontend (WebRTC, dashboard)
βββ requirements.txt # Python dependencies
βββ .env.example # Environment variable template
βββ .gitignore # Git ignore rules
βββ README.md # This file
Edit the SYSTEM_PROMPT in app.py (line ~90):
SYSTEM_PROMPT = """You are "Ava," a professional AI receptionist for [YOUR BUSINESS NAME]...
"""Update the services list inside SYSTEM_PROMPT:
Services offered:
- General Checkup ($50)
- Teeth Cleaning ($80)
- Consultation (Free)In app.py, change the voice variable in text_to_speech():
voice = "en-US-AriaNeural" # Female, US
voice = "en-US-GuyNeural" # Male, US
voice = "en-GB-SoniaNeural" # Female, UKOr use the settings dropdown in the UI.
For better accuracy (slightly slower):
# In app.py, line ~55
whisper_model = WhisperModel("base", device="cpu", compute_type="int8")
# Options: tiny (fastest), base, small, medium, large-v3 (most accurate)Replace the mock classes with real integrations:
Google Calendar:
from googleapiclient.discovery import build
# Use Google Calendar API to check real availabilityHubSpot CRM:
import hubspot
# Use HubSpot API to create real contacts and dealsTwilio (when ready to go live):
from twilio.rest import Client
# Route real phone numbers to your WebSocket endpointThis project is designed to be a client magnet on Upwork. Here's how to pitch it:
- "Build AI phone receptionist for my clinic"
- "Voice AI agent for appointment booking"
- "Twilio + OpenAI real-time voice assistant"
- "AI call center automation"
| Tier | Features | Monthly Price |
|---|---|---|
| Starter | 1 phone number, basic booking, email notifications | $300/mo |
| Pro | 3 numbers, CRM sync, calendar integration, analytics | $750/mo |
| Enterprise | Unlimited numbers, custom voice, priority support, SLA | $2,000+/mo |
- MVP Deployment: $2,000 β $5,000
- Custom Integrations (Salesforce, HubSpot, Calendly): +$1,500 each
- Custom Voice Cloning (ElevenLabs): +$1,000
- Twilio Setup + Number Porting: +$500
| Tech | Why It Was Chosen |
|---|---|
| FastAPI + WebSockets | Native async support, handles concurrent calls effortlessly |
| faster-whisper (tiny) | 39MB model, transcribes in <100ms on CPU, completely free |
| Groq (LLaMA 3 8B) | 500+ tokens/sec β fastest inference on the market, generous free tier |
| edge-tts | Uses Microsoft's Edge browser voices, no API key, sounds professional |
| WebRTC (browser) | Zero-cost alternative to Twilio for portfolio demos |
| Step | Time |
|---|---|
| Audio capture + send | ~50ms |
| Whisper transcription | ~100β200ms |
| Groq LLM generation | ~100β300ms |
| Edge TTS synthesis | ~200β400ms |
| Audio playback | ~50ms |
| Total Roundtrip | ~500β1000ms |
For production, switch to Groq's Realtime API or OpenAI's Realtime API with streaming to cut this to <300ms.
This happens because faster-whisper depends on PyAV (av), which needs FFmpeg system libraries.
Option A β Use openai-whisper (RECOMMENDED, easiest):
The project now auto-detects and falls back to openai-whisper if faster-whisper is not installed. Just run:
pip install -r requirements.txtThe requirements.txt uses openai-whisper by default β no system deps needed.
Option B β Install FFmpeg and use faster-whisper (fastest):
# Ubuntu / Debian
sudo apt-get update
sudo apt-get install -y libavformat-dev libavcodec-dev libavdevice-dev libavutil-dev libswscale-dev libswresample-dev libavfilter-dev pkg-config
# macOS
brew install ffmpeg pkg-config
# Then install faster-whisper
pip install faster-whisper==1.0.3The code will auto-detect faster-whisper and use it for ~2x faster transcription.
- Ensure you're on
localhostorhttps(browsers block mic on insecure origins) - Check browser permissions (click the lock icon in the address bar)
- Verify your
GROQ_API_KEYin.env - Check your Groq dashboard for rate limits
- Ensure you have ~500MB free disk space (for torch + tiny model)
- On first run, it auto-downloads from HuggingFace / PyTorch CDN
- The
tinyWhisper model is optimized for speed over accuracy - Upgrade to
baseorsmallfor cleaner transcription
MIT License β use it freely for client projects, portfolios, or SaaS products.
Built as a portfolio piece to demonstrate expertise in:
- Real-time AI systems
- WebRTC & WebSocket architecture
- LLM integration & prompt engineering
- Full-stack Python development
- SaaS product design
Ready to deploy? Swap WebRTC for Twilio, mock APIs for real ones, and you have a production-grade voice AI platform.
β Star this repo if it helps you land your next $5K client!