Local English dictation with a talk-to-it refine loop
Install · Guide · Report a mishearing
Hold Ctrl+Win, say what you want to write, and let go. Flow pastes the words wherever you were typing: your editor, a terminal, Slack, a browser tab. If it heard something wrong, hold the keys again and say "change Tuesday to Thursday" or "scratch that", and it fixes the text it just pasted.
Flow sits in a small pill at the bottom of the screen, and the pill has two more modes. In
Refine, what you say is rewritten into a proper prompt for your project. In Ask, your question goes to the agent
CLI you already use (codex or claude) and the answer is read back to you.
I made Flow for people whose English has an accent that speech recognition keeps tripping over, often enough that they give up and type. Spanish, Indian, Russian and Japanese speakers are who it's designed and benchmarked for. Recognition runs on your own machine, there's no account or API key, and your audio never leaves your PC.
You need Windows 10 or 11 and a microphone. Download
flow-windows-x64.zip,
unzip it anywhere and run flow.exe. No Python required.
Note
The zip isn't code-signed, so the first launch shows a SmartScreen warning. Click More info, then Run anyway.
Or get the same build through Scoop, which makes updating one command:
scoop bucket add flow https://github.com/samartomar/scoop-flow
scoop install flow/flow # flow/ matters: Scoop's main bucket has a different "flow"If you already use uv, you can install from source instead. It fetches Python and Flow's three dependencies (faster-whisper, sounddevice and numpy):
uv tool install git+https://github.com/samartomar/flow
flowThe speech models aren't part of any of these. Flow downloads them on first run and shows you the progress.
Flow Lite, the version for macOS and Linux, is still in progress. Install it with uv, as above. It has the same recognition, corrections and agent features, but it can't paste into other apps or listen for global hotkeys. You hold the pill to talk, and Send copies the text for you to paste. The microphone is the only permission it needs. There's no Mac or Linux download yet. More about Lite.
Refine and Ask need codex or claude on your PATH, already signed in. Dictation and
spoken corrections work without either. The guide has the details.
- Hold Ctrl+Win (or hold the pill) and talk.
- Let go, and the text is pasted where your cursor was.
- Heard you wrong? Hold again and say "change Tuesday to Thursday" or "scratch that". Flow can fix its last paste until you type or click there.
- Pasted into the wrong window? Click the right one and press Alt+Shift+Z.
- Want an answer rather than text? Hold Ctrl+Alt+Win and ask.
Those last two are Wispr Flow's keys on Windows as well, so if you've used it, your hands already know them.
The first run walks you through the microphone, the model download, a 45-second tuning to your voice, and whether to keep a history. You can skip any step.
Tap the pill to cycle through its three modes. Its colour tells you which one it's in.
| Mode | Pill | When you let go |
|---|---|---|
| Type | white | The words are pasted. |
| Refine | gold | codex or claude turns them into a prompt for your project. Nothing is pasted until you press Send. |
| Ask | violet | The question goes to your agent CLI. The answer appears above the pill and is read aloud, and Continue in Flow opens the conversation in a window. |
Point Refine and Ask at a project folder (Flow Home ▸ Settings ▸ Workspaces, or --cwd)
and the answers are about your code. If you'd rather correct a draft by voice before
anything is sent, switch to the Classic pill in Settings.
Everything that isn't talking happens in one window, the one in the GIF above. Right-click
the pill and choose Open Flow, or run flow --home. Speech models download there with
a progress bar and swap in without a restart, and it's where you set the microphone,
shortcuts, workspaces, agent CLI and the voice that reads answers. Voice tunes Flow to how
you speak, History shows what you've dictated if you chose to keep it, and Conversations
is Ask in a full window.
By default nothing leaves your PC. Audio, recognition, corrections, your word list and your voice profile all stay local, and no API key is read, stored or passed anywhere in the code.
A few things do go out, and only when you use them:
- A draft you refine or a question you ask goes to that agent CLI's cloud, along with your workspace path if you set one.
- Pasting puts the text on the Windows clipboard for a moment, where clipboard managers and cloud sync can see it.
- If you choose one of Microsoft's natural voices to read answers, the text of each answer is sent to Microsoft. Every other voice runs locally.
History is kept only if you say so, in ~/.flow/history.jsonl. Flow Home is served on
127.0.0.1 to the window Flow opens and nothing else, with a new token every launch.
The exact boundary is in the architecture
doc.
- Flow does English only.
- Corrections have to be phrased as commands. "delete the bit about the standup" works; "I feel it should not contain the summary" gets typed as words.
- Your personal word list cuts both ways. It recovers 27–34% of the rare words the model
missed, but raises the error rate by 14–38% on speech that doesn't contain them
(measured on EdAcc with
small.en). Add the words you say often, not every word you know. - The live preview refreshes about once a second and is shown dimmed. The final text always replaces it.
- How accurate it is depends on your voice. Flow Home ▸ Voice ▸ How well Flow hears you measures yours with five sentences.
The rest are in the guide.
The guide is the manual: every flag, spoken command, mode and setting. product.md says who Flow is for and what it deliberately won't do, architecture.md how it fits together and the measurements behind its tuned numbers, roadmap.md what's left and the accent benchmark that tracks it, development.md how to work on it and cut a release, and decisions.md why things are the way they are.
The most useful thing you can send me is a mishearing. If you speak English with an accent, open an issue with what you said and what Flow typed. That's the one thing I can't measure on my own.
Pull requests are welcome for anything else. To run it from a clone:
git clone https://github.com/samartomar/flow && cd flow
uv sync && uv run flow # run it
uv run python -m unittest discover -s tests # ~3,200 tests, about 90 s, no mic neededThe GIF at the top is recorded by scripts/home_reel.py from a
scripted session, so it can be re-shot whenever the app changes.
MIT. Flow is built on faster-whisper and OpenAI's Whisper models, and reads answers aloud with Piper if you add it.
