Capabilities
Everything you need to talk to AI, out loud
vtmate turns your terminal into a full voice interface for AI - natural conversation, agent debates, hands-free background mode, and realistic local voices, all running on your machine with the model of your choice.
Natural conversation
Talk the way you'd talk to a person. vtmate listens continuously and replies the moment you stop, or switches to push-to-talk when you want full control over what gets heard.
Interrupt with your voice
No buttons, no waiting for a pause - just start talking and the agent stops mid-sentence to listen, exactly like breaking into a real conversation.
Watch AI agents debate
Give two agents an opening subject and let them argue it out, complete with their own voices. Jump in and redirect the conversation whenever you want.
Customize your agents
Easily set per agent model, voice, voice speed and system prompt to adjust each agent personality and capabilities.
Switch agents mid-chat
Swap personalities and tune the voice speed live, without ever breaking the flow of the conversation.
Full control over the session
Interrupt a reply the instant you want to say something else, wipe the slate clean with a reset, or undo the last exchange if the conversation took a wrong turn.
Save every conversation
Export any conversation or debate as audio and text, or as a polished, playable HTML session you can revisit, replay or share.
Read anything aloud
Hand it any text file and listen back phrase by phrase with full navigation, or render straight to audio for your own scripts and pipelines.
Instant transcription
Turn a recording into accurate text in moments - no LLM, no voice synthesis, just fast speech-to-text when that's all you need.
Speech recognition, built in
Integrated whisper transcription works out of the box, offline, with nothing extra to install or configure.
Realistic voices, ready to go
Supertonic 3, Supertonic 2 and Kokoro come built in, with optional OpenTTS support for even more voices - no setup required.
Clone any voice
Train a brand new voice from a short recording, or refine an existing clone further, entirely offline.
Code stays on screen
Source code inside a reply is shown, never read aloud, so you're never stuck listening to a wall of syntax.
Bring your own model
Run fully local, connect to any OpenAI-compatible endpoint, plug in a hosted API from every major provider, or reuse a CLI subscription you already have.
Always on, everywhere
Background mode puts voice AI one shortcut away in any app - talk to an agent, ask about what's selected, hear text read aloud, or dictate straight into your cursor.
See it running in the showcase, or dive into the full walkthrough.