I want speech to text. I don’t need another writing assistant.
I started with Wispr Flow, then used Aqua Voice. At some point, paying for another subscription started to bother me. I’d also run into network outages. So I built my own speech-to-text app.
I use it every day. Mostly I talk to Claude and ChatGPT, but I use it for email and other things too.
The more I use it, the clearer my preference gets: I want to say something, see it on the screen, and decide what to do with it. After months of using it every day, I want speed and my own words. I don’t want another LLM rewriting them—or another bill for doing it.
I already have somewhere to edit
If I’m writing an email, I can read the draft and change it. If I want help polishing it, I can use the tools in my editor or ask an assistant.
If I’m talking to Claude, I’m already talking to something that can help me organize my thoughts. I don’t need to ask a second writing tool to prepare my sentences before the conversation starts.
Sometimes I want a polished paragraph. Sometimes I want to explain something in my own slightly rambling way. I’d like to choose.
The messy parts can contain the point
A lot of my prompts include corrections and qualifications. I’ll explain what I want, remember an exception, and add another sentence.
Those details aren’t always elegant. They can still be important.
I don’t want elegance to be the default goal of the tool taking down my words. I want to read what it produced and decide whether it needs editing.
Lots of people like automatic cleanup. I prefer to choose when it happens.
Here’s how I actually sound
When I was explaining this idea, I used Vokk to write:
“You don't actually need a speech to text to do interpretation or processing of your voice. What you just want is directly what you said as quickly as possible. The reason is, is that if, like for example, if you're, if you're writing an email in Gmail, Gmail has built-in grammar checking and auto correction and things like that. So you're much better off having exactly what you said and then being able to tweak it as you wish, as opposed to having an LLM decide which parts to keep. And again, this whole sentence I just, I just, everything I just said, I'm using Vokk right now. So everything I just said is straight from Vokk with no editing. Like it works amazingly well.”
That’s an excerpt from my own message, transcribed by Vokk without editing.
Would I trim the repetition before publishing an essay? Of course. Did it stop the LLM from understanding me? No. Not in the slightest.
Another one from a chat with ChatGPT:
Alright, this is maybe a strange thing to dive into, but I just want to look at all of the companies that IPO'd this year, and especially companies that are like, let's say, under $10 billion. And just kind of, let's start with basically a list and a sentence, like what the company is and what their story is. I don't know how hard that is to put together, but just as a starting point for conversations.
Again, that is straight from Vokk, in milliseconds, no editing applied.
For a conversation, I’m happy to sound like someone having a conversation.
Speech recognition still has to recognize speech
Vokk uses a speech-recognition model. It can mishear things. It produces punctuation. It isn’t a literal recording of every sound I make, and I don’t claim that every word will always be right.
The distinction is narrower: Vokk doesn’t take the resulting transcript and send it through another AI writing pass to polish the prose.
I check important details before sending. Then I move on.
That’s the app
Hold a key. Talk. Release. The text appears where you’re typing.
Vokk is $12 a year, or $1.50 paid monthly. It works on Apple Silicon Macs running macOS 13.3 or later. There’s a 14-day free trial, no card required.
I built it because I wanted it. If this is how you want to work too, give it a try.
Related: Using speech to text with coding assistants · Vokk vs. Wispr Flow · Vokk vs. Aqua Voice