voice | pipesDownload
← Blog

Why I built Voice Pipes

One button and one pipeline weren't enough. I wanted a fast local path and a careful one, a voice that reads back, my own logic anywhere in the pipe, no subscription, and a tool my agents can drive.

Voice is turning into one of the most powerful ways to control a computer, and more and more of what I do runs through agents. Yet the dictation tools I tried all looked the same: one button, one pipeline, one way of working, a monthly bill, and my voice sent to someone else’s server. Voice Pipes is the tool I wanted instead. Here’s why.

One hotkey isn’t enough

Hold-to-talk is perfect for a quick line. It’s miserable for a long passage: three minutes of thinking out loud with a finger pinned to a key. Sometimes I want to tap once, talk for as long as I need, and tap again.

So in Voice Pipes a pipeline (a track) can have as many hotkeys as you like, and each one is either hold (records while held) or toggle (press to start, press again to stop). Same pipeline, different trigger for the moment you’re in.

Fast dictation⌥Space hold⌥⇧D toggle
Hold ⌥Spacea quick line
REC 00:02›let go›OK 96ms

“Ship it after lunch.”

Tap ⌥⇧Da long passage, hands free
REC 02:41›tap again›OK 131ms

“So here's the plan for the quarter. First, we move the release checks into the pipeline, then…”

Illustrative recreation of the HUD, not a screenshot: one track with a hold hotkey and a second, toggle hotkey you could add. Text and timings are made up.

In the config file, that’s one setting on the track:

hotkeys = [
  { keys = "option+space", mode = "hold" },
  { keys = "option+shift+d", mode = "toggle" },
]

Two speeds, on purpose

Not all dictation is the same job. When I’m drafting an email, I don’t mind waiting a moment longer if what lands is clean, so I’m not running it through Grammarly or re-editing it afterwards. When I’m firing off quick notes or talking to a terminal, I’m trying to stay in flow, and any lag pulls me out of it.

So the pipelines are yours to shape. The fast one runs Parakeet on the Mac: no network, nothing to wait on. In the measurements behind the defaults it finished 43 to 140 ms after the recording ended, for 4 to 33 seconds of speech (on an M5 Max). The careful one sends the audio to a cloud transcription model and then to a language model for cleanup, through your own OpenRouter key. It takes longer, and for an email that’s a fair trade.

Fast dictation⌥Space hold
Mic›Parakeet›Fix words›Paste

ON THIS MAC · works offline once the model is downloaded · 43–140 ms from release to text in our tests

Clean dictation⌥⇧Space toggle
Mic›MAI-Transcribe-2›Fix words›Claude Haiku cleanup›Paste

CLOUD · two model calls through your OpenRouter key · slower, and polished

The two dictation tracks Voice Pipes starts with. Every block can be swapped for another.

Local first, because I’m often offline

I travel a lot, and I’m often on bad hotel Wi-Fi or none at all. I need dictation that works anyway. I also don’t want all my voice data going up to the cloud when there’s no reason for it: for everyday dictation, a model on my own machine is faster.

So the fast path stays on the Mac. Parakeet for transcription and the Pocket TTS and Supertonic voices download once, the first time you use them, and then run with no connection. Fix words and History never leave the Mac either. The cloud is there for when I want something special, a bigger model or a web search, and only in the blocks I choose, with my own keys.

Speaking and listening belong together

Talking to my computer and having it talk back are the same idea, so why were they always two different tools? Voice Pipes reads to you, too: select text, press a key, and a voice on your Mac reads it. Press the same key to pause and again to carry on.

The same goes for agents. An agent can speak to me when a long job finishes, or ask me a question out loud and use my spoken answer, without me ever looking at the screen.

Read aloud⌥R toggle
Selection›Speak · Pocket TTS
READ 42%⏸⏹
An agent, from the terminal
$ vp say "Tests passed."$ vp ask "Deploy now or after review?"question: Deploy now or after review?answer: after review
Illustrative recreation, not a screenshot or a real run: the read-aloud HUD, and an agent speaking and asking with vp.

A tool you own

I’m tired of subscriptions for everything. I like the old model: you get a tool, you own it, and it keeps working perfectly well until a new version is worth having. That’s the spirit Voice Pipes is built in. There’s no account to sign up for, the fast path needs nothing but your Mac, and when you do use cloud models you bring your own keys, so you pay those providers for what you use and nothing more.

My logic, anywhere in the pipe

The thing I missed most was a way to add my own step. I keep notes and dictations in my own system, and I want every take sent there too, without a second tool. In Voice Pipes that’s just one more block: an HTTP request to your endpoint, anywhere in the pipeline. Output blocks pass their text on, so a track can paste at the cursor and post to your notes.

Voice note⌥N toggle
Mic›Parakeet›Fix words›Paste›HTTP POST→ your notes service
A track you could build: dictation that also lands in your own notes. The token stays in the Keychain, not the file.
  [[track.step]]
  type = "http"
  method = "POST"
  url = "https://notes.example.com/api/notes"
  headers = { Authorization = "Bearer ${secret:notes}", "Content-Type" = "application/json" }
  body = '{"text": {{input_json}}}'

Agent-first, all the way down

If voice is going to be a control surface for agentic work, the tool has to be one an agent can use and configure. So everything (your tracks, their blocks, their hotkeys, your settings) lives in one TOML file, ~/.config/voice-pipes/config.toml, with a schema and a checker. And there’s vp, the command line. With it an agent can run your tracks, speak to you, ask you something, transcribe a file without rolling its own speech pipeline, edit your vocabulary, read your history, and open any window in the app to show you what it’s doing.

Most of all, I want the agent to build pipelines for me. Describe the track you want; it edits the config, checks it, and you watch each block appear in the editor as it’s added.

# the agent adds [[track]] voice-note$ vp config checkresult: ok$ vp open track voice-note# + transcribe, fix-words# + http (POST to your notes)$ vp run voice-note --text "…"
Voice note⌥N toggle
1Microphone
2Transcribe Parakeet · on this Mac
3Fix words
4HTTP request POST
OK 212ms
Illustrative recreation, not a screenshot or a real run: an agent building a track while the editor shows each new block. Timings are made up.

One line installs the app, vp and a skill that teaches Claude Code, Codex and other agents to use it:

curl -fsSL https://github.com/brancusi/voice-tools-releases/releases/latest/download/install.sh | bash

macOS still asks you, not the script, for Microphone and Accessibility the first time. The install guide, the CLI reference and agents pages have the rest.

The word it keeps getting wrong

Every dictation tool has a word it can’t get right. For me it was my own surname. It’s a small thing, but it happens every single time, and it’s usually what pushes people to run every dictation through an LLM cleanup: slower, costlier and sometimes too clever, all to fix one word.

Voice Pipes fixes it the simple way. Fix words is a find-and-replace list, applied on your Mac in effectively no time: whenever transcription writes one of the “heard as” versions, you get the spelling you want. (I also tried boosting the speech model’s vocabulary instead. It was slower, and it started putting my terms where they didn’t belong.)

The hard part is knowing every way a word gets misheard, so you train it with takes. Say the word about five times; each take is replayed about 30 ways (faster, slower, quieter, noisier) through Parakeet, and you get every distinct way it came out. Then Jev, one of the new “System One” decision models from TypeSafe, judges each one: is this a garbled version of your word, safe to replace everywhere, or a real word you’d want left alone? When I tested it on my own name (with macOS voices standing in for me), five takes became 150 variations in about five seconds, and Parakeet got the name right only 4 times. To be clear about what this is: it doesn’t retrain the speech model, it learns how the model mishears you and fixes those spellings.

Train “Kubernetes”take 1take 2take 3take 4take 5

150 variations through Parakeet · right 6 times · Jev judges each result

✓cube or netties41×91%
✓cuban eighties23×84%
✓cooper nettys9×72%
cuban4×8%
Added to Heard as:Kubernetes ← cube or netties, cuban eighties, cooper nettys
Illustrative recreation, not a screenshot: training a word from five takes. The results and Jev's scores are made up; a real word like “cuban” scores low and stays unticked.

Jev runs in the cloud with your TypeSafe key; without one, Voice Pipes ticks the likely mishearings with a simpler rule, and everything else in training stays on your Mac. Agents can help here too: vp vocab add and vp vocab test let one build up the list for you, and the takes are yours to give.

What it adds up to

A voice tool that bends to the job instead of the other way round: two speeds, hands-free when you need it, a voice that reads back, your own steps anywhere in the pipe, local first, no subscription, and all of it open to your agents. That’s Voice Pipes. Download it, or install it from a terminal with the line above.