English | 简体中文
A personal agent that understands your context, keeps work moving through conversation, checks results, and speaks when it matters.
- 2026-10-09 · v0.3.1 — Smoother first run and upgrades: the quick start now powers typed chat out of the box, the Ubuntu package upgrades in place from 0.3.0, the macOS app bundle verifies cleanly, and the Feishu setup is fully available in English.
Earlier releases
- 2026-10-07 · v0.3.0 — A personal Workbench, source-backed memory and tasks with acceptance criteria. Realtime and cascaded voice support Qwen, OpenAI, Gemini, Volcengine and self-hosted services; availability depends on the selected pipeline.
- 2026-09-24 · v0.2.3 — Guided first run: one DashScope API Key is enough to start talking.
- 2026-09-21 · v0.2.2 — Ubuntu 22.04+ x64 desktop, plus the headless
nova-audio-agent-serverwith QR pairing. - 2026-09-21 · v0.2.0 — Configurable ASR / LLM / TTS pipelines, personal memory, custom MCP, wake words, bilingual desktop on macOS and Windows, and an iPhone client over Tailscale.
- 2026-08-31 · v0.1.0 — Always-on voice, background Codex tasks, live steering, workspace/session management, and selective progress updates.
Nova brings conversation, your authorized context and background execution into one personal agent. Talk through a goal, organize your todos, ideas and goals on the Workbench, and follow the work through to an evidence-backed result.
- Voice and text, together. Keep talking while work runs; clarify a request or change direction without starting over.
- Conversation beside your tasks. Todos, Ideas, Goals, Feeds, Tasks and Profile sit beside the conversation. Switch to the orb or let tasks run in background mode with the microphone off.
- Context from your own sources. Connect folders, mail, calendars and Feishu. Source access and permission for the configured model to process content are separate choices.
- Memory with sources. See what Nova has learned, trace it to its evidence, correct it, forget it or purge it. Imported document knowledge remains a separate searchable corpus.
- Tasks you can check. Clarify goals, confirm workspaces and sessions, then delegate to Codex. Refine the work as it runs; Nova checks evidence against acceptance criteria, and you can take over and return control.
- Thoughtfully proactive. Meaningful progress, camera events and optional daily briefs get your attention at an appropriate moment. Routine updates stay quiet; reminders respect your speaking turn.
Learn more about source permissions, task controls and when a voice agent should speak.
|
Ask Nova to watch for a condition and tell you when it occurs.
|
Connect your iPhone over Tailscale to Nova on macOS or an Ubuntu headless server, then talk and approve tasks from your phone.
|
Requirements: Node.js 22.13+, npm, Git, a logged-in codex executable (app-server is the only
Codex transport).
Install the release with npm.
npm install --global nova-audio-agent@latest
# open the shipped app; the first launch asks for one DashScope API key
novaaudio
# open the settings panel in the app
novaaudio config
# list the keys your voice pipeline needs; --online tests them
novaaudio doctorHeadless Ubuntu 22.04+: install nova-audio-agent-server with npm, configure it and initialize credentials; run novaaudio-server start and, in a second terminal, novaaudio-server pair wss://your-host.ts.net for a one-use QR. See the configuration guide.
Quit Nova, then run npm install --global nova-audio-agent@latest. To pin a version, use nova-audio-agent@0.3.1. Upgrades keep your local settings and data; moving back to an older version does not roll back data changes, so back up your Nova data first.
For headless Ubuntu 22.04+ x64, use npm install --global nova-audio-agent-server@latest.
For development from source:
git clone https://github.com/deepnovacore/NovaAudioAgent.git nova-audio-agent
cd nova-audio-agent
npm ci && cp .env.example .envGet an API key from DashScope and set DASHSCOPE_API_KEY. That one key runs voice, memory, the camera and web search. Search goes through Bailian until you add a Tavily TAVILY_API_KEY.
npm run start:client
# Open Workbench for this launch, overriding the saved startup preference
npm run start:workbenchThe client includes microphone, camera, sound, settings, and external MCP controls. Try hovering over the desktop orb to get surprised :) Also you may try build or run demo locally:
npm run build --workspace @nova-audio-agent/runtime
node runtime/dist/src/cli.js diagnose --json
node runtime/dist/src/cli.js demo allNative echo-cancelled capture (VoiceProcessingIO) is macOS-only. Wake detection uses that
capture when available; Windows, Linux source runs, and macOS fallback use Chromium
getUserMedia + AudioWorklet. While sleeping, microphone frames go only to the local
wake-word Worker; explicit mute stops wake detection. See
wake-word setup.
| Platform | Desktop app | Headless server | Connect iPhone |
|---|---|---|---|
| macOS arm64 | Yes, with native echo-cancelled capture | From source | From the desktop app |
| Windows x64 | Yes | — | — |
| Ubuntu 22.04+ x64 | Yes | npm package | Through the headless server |
| Layer | Providers | Default |
|---|---|---|
| Integrated voice | Qwen realtime, OpenAI realtime, Gemini Live, StepFun (preview) | Qwen qwen-audio-3.0-realtime-plus |
| Cascaded ASR / TTS | Volcengine Speech, Gemini, self-hosted (reference: Whisper / Breeze) | volc.seedasr.sauc.duration / seed-tts-2.0 |
| Cascaded LLM | DeepSeek, Qwen / DashScope, Volcengine Ark, OpenAI, Gemini, self-hosted | DeepSeek deepseek-flash |
| Vision | Qwen-VL family, Doubao Seed | Off in conversation; Loop Camera uses its own model |
| Coding executor | Codex | Codex |
| Sources | Local folders, Google via Composio, Apple Mail and Calendar (macOS), Feishu | None until you authorize them |
Models, credentials and per-provider limits: support matrix.
Conversation stays responsive, authorized context informs the work, and task results return as evidence. Speaking is a separate decision.
- Interaction: Workbench, voice orb and phone clients share conversation and task controls. Background mode turns off the desktop microphone while work continues.
- Context: authorized sources feed the personal memory ledger and source-grounded suggestions. Document knowledge is separate; neither store owns live task state or grants permissions.
- Coordination: the conversation clarifies goals; the host owns confirmations, permissions and task state. The Task loop executes, checks evidence, then completes, corrects or waits. You can take over and return control.
- Execution and expression: Codex performs background work; camera monitoring and selected tools have their own lifecycles. Runtime events feed Proactive's selection of optional updates, Floor coordinates the speaking opportunity, and the conversation model expresses the update.
The runtime blackboard holds causal conversation and execution state; personal memory retains source-backed facts; document knowledge searches imported files. They have different roles and authority.
See How Nova works for the product flow and the architecture reference for runtime internals.
| Read this | For |
|---|---|
| Architecture | Modules and boundaries |
| Glossary and invariants | Vocabulary and rules |
| Getting started | Setup and integrations |
| Workbench · Tasks · Sources and connectors | The main window, delegated work, and what Nova may read |
| When should a voice agent speak? | Voice interaction design |
| Node runtime migration archive | Migration-era plans in the history of tag v0.1.0 |
- v0.3.0: one main window for text and voice with Todos, Ideas, Goals, Feeds, Tasks and Profile; tasks with acceptance criteria, verified completion and takeover; memory-grounded suggestions; traceable, correctable and removable personal memory; user-authorized folders, email, calendars and Feishu conversations.
- v0.4.0 (in development): expand coding backends with Kimi Code and pi agent; add a GUI executor with AutoGLM as the first example, enabling collaboration across specialist agents.
Releases require feature and supported-platform acceptance. Ubuntu 22.04+ x64 desktop and headless npm packages are included in the candidate release checks.
npm ci && npm run check && npm run build && npm testLive integrations are credential- and hardware-dependent and never substitute for the deterministic tests. Security reports: SECURITY.md; contribution rules and invariants: CONTRIBUTING.md.
Copyright 2026 DeepNovaCore, Apache License 2.0.











