One button and one pipeline weren't enough. I wanted a fast local path and a careful one, a voice that reads back, my own logic anywhere in the pipe, no subscription, and a tool my agents can drive.
Voice is turning into one of the most powerful ways to control a computer, and more and more of what I do runs through agents. Yet the dictation tools I tried all looked the same: one button, one pipeline, one way of working, a monthly bill, and my voice sent to someone else’s server. Voice Pipes is the tool I wanted instead. Here’s why.
One hotkey isn’t enough
Hold-to-talk is perfect for a quick line. It’s miserable for a long passage: three minutes of thinking out loud with a finger pinned to a key. Sometimes I want to tap once, talk for as long as I need, and tap again.
So in Voice Pipes a pipeline (a track) can have as many hotkeys as you like, and each one is either hold (records while held) or toggle (press to start, press again to stop). Same pipeline, different trigger for the moment you’re in.
“Ship it after lunch.”
“So here's the plan for the quarter. First, we move the release checks into the pipeline, then…”
In the config file, that’s one setting on the track:
hotkeys = [
{ keys = "option+space", mode = "hold" },
{ keys = "option+shift+d", mode = "toggle" },
]
Two speeds, on purpose
Not all dictation is the same job. When I’m drafting an email, I don’t mind waiting a moment longer if what lands is clean, so I’m not running it through Grammarly or re-editing it afterwards. When I’m firing off quick notes or talking to a terminal, I’m trying to stay in flow, and any lag pulls me out of it.
So the pipelines are yours to shape. The fast one runs Parakeet on the Mac: no network, nothing to wait on. In the measurements behind the defaults it finished 43 to 140 ms after the recording ended, for 4 to 33 seconds of speech (on an M5 Max). The careful one sends the audio to a cloud transcription model and then to a language model for cleanup, through your own OpenRouter key. It takes longer, and for an email that’s a fair trade.
ON THIS MAC · works offline once the model is downloaded · 43–140 ms from release to text in our tests
CLOUD · two model calls through your OpenRouter key · slower, and polished
Local first, because I’m often offline
I travel a lot, and I’m often on bad hotel Wi-Fi or none at all. I need dictation that works anyway. I also don’t want all my voice data going up to the cloud when there’s no reason for it: for everyday dictation, a model on my own machine is faster.
So the fast path stays on the Mac. Parakeet for transcription and the Pocket TTS and Supertonic voices download once, the first time you use them, and then run with no connection. Fix words and History never leave the Mac either. The cloud is there for when I want something special, a bigger model or a web search, and only in the blocks I choose, with my own keys.
Speaking and listening belong together
Talking to my computer and having it talk back are the same idea, so why were they always two different tools? Voice Pipes reads to you, too: select text, press a key, and a voice on your Mac reads it. Press the same key to pause and again to carry on.
The same goes for agents. An agent can speak to me when a long job finishes, or ask me a question out loud and use my spoken answer, without me ever looking at the screen.
$ vp say "Tests passed."$ vp ask "Deploy now or after review?"question: Deploy now or after review?answer: after review
vp.A tool you own
I’m tired of subscriptions for everything. I like the old model: you get a tool, you own it, and it keeps working perfectly well until a new version is worth having. That’s the spirit Voice Pipes is built in. There’s no account to sign up for, the fast path needs nothing but your Mac, and when you do use cloud models you bring your own keys, so you pay those providers for what you use and nothing more.
My logic, anywhere in the pipe
The thing I missed most was a way to add my own step. I keep notes and dictations in my own system, and I want every take sent there too, without a second tool. In Voice Pipes that’s just one more block: an HTTP request to your endpoint, anywhere in the pipeline. Output blocks pass their text on, so a track can paste at the cursor and post to your notes.
[[track.step]]
type = "http"
method = "POST"
url = "https://notes.example.com/api/notes"
headers = { Authorization = "Bearer ${secret:notes}", "Content-Type" = "application/json" }
body = '{"text": {{input_json}}}'
Agent-first, all the way down
If voice is going to be a control surface for agentic work, the tool has to be one an agent can use and configure.
So everything (your tracks, their blocks, their hotkeys, your settings) lives in one TOML file,
~/.config/voice-pipes/config.toml, with a schema and a checker. And there’s vp, the command line. With it an
agent can run your tracks, speak to you, ask you something, transcribe a file without rolling its own speech
pipeline, edit your vocabulary, read your history, and open any window in the app to show you what it’s doing.
Most of all, I want the agent to build pipelines for me. Describe the track you want; it edits the config, checks it, and you watch each block appear in the editor as it’s added.
# the agent adds [[track]] voice-note$ vp config checkresult: ok$ vp open track voice-note# + transcribe, fix-words# + http (POST to your notes)$ vp run voice-note --text "…"
One line installs the app, vp and a skill that teaches Claude Code, Codex and other agents to use it:
curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash
macOS still asks you, not the script, for Microphone and Accessibility the first time. The install guide, the CLI reference and agents pages have the rest.
The word it keeps getting wrong
Every dictation tool has a word it can’t get right. For me it was my own surname. It’s a small thing, but it happens every single time, and it’s usually what pushes people to run every dictation through an LLM cleanup: slower, costlier and sometimes too clever, all to fix one word.
Voice Pipes fixes it the simple way. Fix words is a find-and-replace list, applied on your Mac in effectively no time: whenever transcription writes one of the “heard as” versions, you get the spelling you want. (I also tried boosting the speech model’s vocabulary instead. It was slower, and it started putting my terms where they didn’t belong.)
The hard part is knowing every way a word gets misheard, so you train it with takes. Say the word about five times; each take is replayed about 30 ways (faster, slower, quieter, noisier) through Parakeet, and you get every distinct way it came out. Then Jev, one of the new “System One” decision models from TypeSafe, judges each one: is this a garbled version of your word, safe to replace everywhere, or a real word you’d want left alone? When I tested it on my own name (with macOS voices standing in for me), five takes became 150 variations in about five seconds, and Parakeet got the name right only 4 times. To be clear about what this is: it doesn’t retrain the speech model, it learns how the model mishears you and fixes those spellings.
150 variations through Parakeet · right 6 times · Jev judges each result
Jev runs in the cloud with your TypeSafe key; without one, Voice Pipes ticks the likely mishearings with a simpler
rule, and everything else in training stays on your Mac. Agents can help here too: vp vocab add and
vp vocab test let one build up the list for you, and the takes are yours to give.
What it adds up to
A voice tool that bends to the job instead of the other way round: two speeds, hands-free when you need it, a voice that reads back, your own steps anywhere in the pipe, local first, no subscription, and all of it open to your agents. That’s Voice Pipes. Download it, or install it from a terminal with the line above.