Voice input
Last updated: 18 September 2026
What it does
Dictate into any terminal pane instead of typing. Click the π€ in the pane header, speak, and click it again. What you said is typed into the pane.
- Click, talk, click β a quick click starts listening and a second click stops it. A hint at the bottom of the pane says Listeningβ¦ while the mic is open.
- Words appear as you speak. StreamView re-transcribes what it has heard about once a second and types it into the pane, correcting earlier words as the sentence makes more sense. Watching a phrase rewrite itself is normal β Whisper revises with more context. Turn it off with Type while I speak in Settings if you would rather nothing appeared until you stop.
- Hold β hold the button (or
Ctrl+Shift+M) for push-to-talk; releasing stops it. Ctrl+Shift+Mworks the same way: tap to start and tap to stop, or hold. It only does anything once voice input is set up, so an accidental press still goes to the program.
It's useful mainly for the long, sentence-shaped prompts you give an AI agent, where talking is genuinely faster than typing. It works with any pane, because it just types: Claude Code, Cursor Agent, and plain shells all behave the same.
Typing as you speak
Live dictation types into the pane and takes characters back with backspaces when a later pass disagrees, exactly as if you had typed and corrected them yourself. Two limits keep that safe:
- It only takes back its own characters. StreamView tracks what this dictation typed and never backspaces past it.
- A big revision is appended, not rewritten. If a pass rewrites more than about 48 characters, StreamView keeps what is on screen and adds only what is new, rather than erasing half a sentence you may have started editing.
Either way nothing is submitted. The final pass is the most accurate one, since it sees the whole recording.
It types, it never sends
Dictated text is typed into the pane and left there. You press Enter.
That's deliberate and not configurable. Speech recognition mishears, and these panes are frequently running agents that execute what they're given β a misheard command that submitted itself is a genuinely bad outcome. Read it, fix it if needed, then send it.
Turning it on
Click the π€ in any pane header. The first time, StreamView offers to download the speech model (Small, about 466 MB, one time). Click OK; when the download finishes, voice input is on.
If a model is already downloaded, the first click just switches voice input on.
To pick a different model, or to switch voice input off, go to Settings > Data > Voice Input.
Nothing is downloaded and the microphone is never opened until you click the mic and agree. Voice input holds your microphone while recording and needs a few hundred megabytes of model on disk, and neither should happen to someone who never asked for dictation.
Models
Transcription runs entirely on your machine using Whisper. Models are downloaded on demand rather than shipped, since even the smallest is larger than StreamView itself.
| Model | Size | Good for |
|---|---|---|
| Tiny | ~75 MB | Short, plain sentences. Fastest, least accurate. |
| Base | ~141 MB | Faster and smaller. Fine for plain sentences; weaker on technical terms. |
| Small | ~466 MB | Recommended β the one-click setup downloads this. Most accurate, especially on code and technical terms; a second or two slower per prompt. |
Expect any of them to struggle with identifiers, file paths, and command flags. Dictation suits prose β "work out why the login test is flaky and propose a fix" β far better than it suits git rebase -i HEAD~3.
Models are stored in %LOCALAPPDATA%\StreamView\whisper-models, on this machine only. They are deliberately not kept in your data directory: they're large, replaceable binaries, and syncing them to cloud storage would recreate exactly the scan-and-sync churn that History Location exists to avoid. Delete one any time from Settings; you can re-download it later.
Privacy
Audio is captured, transcribed locally, and discarded. Nothing is uploaded β no speech service, no API key, no network call at transcription time. The only network access is the one-off model download.
The microphone is open only while you hold the button or hotkey, and there is a two-minute hard cap on any single recording so a missed key-up can't leave it listening.
Dictated text reaches the terminal as if typed, so your recording mode governs whether it lands in session history β same as anything else you type.
Troubleshooting
- No microphone button β this is a Tail File pane (no process to type into).
- Click the mic and nothing is typed β watch the hint at the bottom of the pane. Too short means the click was too quick to hold any speech: click, talk, then click again. Heard only silence means the Windows default input is muted or is the wrong device; StreamView records from whatever Windows Sound settings > Input has as the default.
- "No Whisper model is downloaded" β Settings > Data > Voice Input, then Download.
- "No microphone was found" β Windows reports no recording device. Check Windows Sound settings, then reopen Settings and click Recheck.
- Didn't catch any words β Whisper heard sound but no speech. Try again closer to the mic.
- Wrong words on technical terms β expected; try the Small model, or type the tricky parts.
Related
Still stuck? In the app, use Help β Contact support, or open a ticket here.