Premium dictation for macOS

Premium dictation.
State-of-the-art models.

Most dictation apps run one model. ZWhispr runs several of the best at the same time, reads the words on your screen, and learns your vocabulary. So the text that lands at your cursor is fast, clean, and correct.

$19/month. Every feature included. Nothing to configure.

Messages — Priya Nair
Hey, are we still good for the launch on Friday?
fn Hold to speak

Works wherever you type

SlackGmailNotionCursoriMessageLinearXcodeSafariNotesTerminalFigmaClaudeChatGPTMailObsidianVS Code SlackGmailNotionCursoriMessageLinearXcodeSafariNotesTerminalFigmaClaudeChatGPTMailObsidianVS Code

How it works

Why trust one model when you can ask several?

Every time you speak, ZWhispr sends your audio to multiple state-of-the-art transcription models at once. It intelligently merges their answers using the context of what you are working on, and never waits on the slowest one.

You say

“Schedule the Soniox sync for Thursday and loop in Priya.”

Model A0 ms

Model B0 ms

Model C0 ms

ZWhispr types

    The lineup right now

    Which models are behind ZWhispr today

    Updated September 2026. The lineup changes whenever something better ships, with no update on your side.

    • MicrosoftModel nameSpeech to text
    • MetaModel nameSpeech to text
    • SonioxModel nameSpeech to text
    1

    Click anywhere

    Put your cursor in any text field. Slack, Mail, Cursor, Notion, a terminal. Anywhere you would normally type.

    2

    Hold fn and talk

    Speak the way you think. Do not worry about punctuation, filler words, or changing your mind halfway through.

    3

    Let go

    Clean, correctly spelled text appears at your cursor. No copy, no paste, no fixing names afterwards.

    Zero compromise

    Premium where it counts: accuracy and speed.

    Most dictation apps pick one model and hope for the best. ZWhispr is built for people who cannot afford a wrong name in an email or a five-second stall in a meeting.

    Several models, intelligently merged. Fewer mistakes.

    Each speech model has its own blind spots. ZWhispr merges the output of several state-of-the-art models, using your dictionary and the context on your screen to decide which reading is right wherever they differ. One model’s slip gets corrected by the others. The result is a lower error rate than any single model can deliver on its own.

    Never wait on the slow one

    Because your audio is already with several models, a slow response from one provider simply gets left behind. That cuts the long tail that makes dictation feel laggy: your P95 and P99 latency drop, not just the average.

    Your screen is its cheat sheet

    With your permission, ZWhispr takes a quick look at the window you are working in and pulls out the distinctive words: names, tickets, product terms, code identifiers. Those hints go to the models with your audio, so “Kubernetes” and “Priya” come out spelled the way they appear on your screen.

    A dictionary that is yours

    Add the words that matter to you once: your team’s names, your company’s jargon, the acronyms only your industry uses. Correct a word and ZWhispr remembers. Your vocabulary follows you into every app.

    Sounds like you, reads like you meant it

    Filler words, false starts, and “wait, I mean” corrections are cleaned up automatically. You get natural sentences with proper punctuation, not a transcript of your thinking out loud.

    Always the best models. Automatically.

    You never pick a model or install an update to get a better one. ZWhispr uses the current state-of-the-art from leading third-party providers and swaps in stronger models as they ship. Your setup stays the same. Your accuracy keeps climbing.

    Privacy

    Your words are yours.

    ZWhispr processes what you say and what is on your screen to transcribe it, then lets it go. We are specific about what we keep and what we never keep.

    We never store

    • Your transcripts. The text you dictate is typed at your cursor and is not saved on our servers.
    • Your screen context. The words on your screen are read at the moment you speak, used for that transcription, and discarded.

    We do store

    • Your dictionary. The words you add or correct, so they follow you into every app.
    • Account and usage metadata. The basics needed to run your subscription and keep the service reliable.

    Pricing

    One plan.
    Everything in it.

    Running several premium models on every recording costs more than running one cheap one. We think your words are worth it. There is no light tier, no model add-ons, and no “pro” toggle. Every ZWhispr user gets the best we have.

    • Parallel transcription with multiple state-of-the-art models
    • Screen-aware vocabulary (optional, permission based)
    • Personal dictionary that learns from your corrections
    • Latency hedging so slow providers never hold you up
    • Automatic clean-up of fillers, repeats, and self-corrections
    • Works in every Mac app, pastes at your cursor
    ZWhispr for MacAll features

    $19per month

    Billed monthly. Cancel any time from your billing portal.

    Get ZWhispr for Mac

    Requires macOS. Secure checkout by Paddle, our merchant of record.

    FAQ

    Good questions.

    Doesn’t running several models make it slower?

    The opposite. The models run in parallel, not one after another, so the round trip is only as long as the fastest useful answer. When one provider has a slow moment, ZWhispr proceeds without it. That is what shrinks the worst-case P95 and P99 latency you actually feel.

    How does combining models reduce errors?

    Different models make different mistakes. ZWhispr merges their transcripts intelligently, using your dictionary and the context on your screen to pick the right reading wherever they differ. A slip from one model gets corrected instead of ending up in your email.

    What exactly does the screen context feature see?

    When you enable it, ZWhispr takes a screenshot of the window you are working in at the moment you start speaking and extracts the distinctive words on it. Those words are sent along with your audio as vocabulary hints. It is optional and you grant the permission in macOS.

    Do I have to choose or configure a model?

    No. ZWhispr always uses the current state-of-the-art models from leading third-party providers and upgrades them for you. There is nothing to select or tune.

    What do you store?

    We never store your transcripts or your screen context. Both are used to produce the transcription in front of you and then discarded. We do store your dictionary words and the account and usage metadata needed to run the service.

    How is billing handled?

    Subscriptions are $19 per month, billed through Paddle, our merchant of record. Paddle handles payment, invoices, and sales tax for your country. You can update your card or cancel any time from the billing portal link in your receipt. Questions go to hello@zwhispr.com.

    Which apps does it work in?

    Any Mac app with a text field. Hold the hotkey, speak, release, and the text is typed at your cursor in Slack, Mail, Notion, Cursor, Xcode, the terminal, a browser, or anything else.

    Stop typing. Start talking.

    Your hands are for the interesting parts of your work.

    Download ZWhispr for Mac