Most speech to text mistakes start in the audio, not in the software. The biggest gains come from four things: the microphone close to the person speaking, little background noise, one person speaking at a time and the right language setting. After that, speak in full sentences at a steady pace, say names and numbers clearly, and read the text once before you use it. An AI rewrite can then clean up punctuation and filler words, but it cannot repair a word the transcription got wrong.
| Problem in the text | Most likely cause | What to change |
|---|---|---|
| Words missing or garbled | Microphone too far away, or noise | Move the phone closer, go somewhere quieter |
| Whole passages wrong | Wrong language setting, or two languages mixed | Set the spoken language; one language per recording |
| Names and brand words wrong | Rare words the system does not expect | Say them clearly, spell them once, fix them afterwards |
| Two voices blended | People talking over each other | Take turns; pause before answering |
| Nothing at all | Silent or extremely quiet recording | Check the input level; test with a short clip first |
The microphone and the room
Google's best practices for its Cloud Speech-to-Text service put it plainly: "Position the microphone as close as possible to the person that is speaking, particularly when background noise is present." In practice:
- Hold the phone close to your mouth when you dictate, about as close as during a phone call, and do not cover the microphone opening with your fingers or a case.
- Choose a quiet room. Turn off music, the TV and fans. A small room with soft furnishings echoes less than a kitchen or a stairwell.
- Avoid wind. Outdoors, turn your back to the wind or use a headset microphone.
- Do not shout. Google also advises to "avoid audio clipping". Speaking louder than normal close to the microphone distorts the signal.
- Record a meeting from the middle of the table, not from one end, so every voice reaches the phone at a similar level.
How you speak
- Full sentences at a steady pace. You do not need to speak slowly, but steady speech with short pauses between thoughts transcribes better than bursts.
- One language per recording. Switching languages mid-sentence confuses many systems. If you have to quote another language, keep it short.
- Say names and numbers clearly. For an unusual name, say it once slowly or spell it: "Siobhan, S I O B H A N". Say "fifteen" or "one five" so it cannot be heard as "fifty".
- Take turns. When two people talk at the same time, both usually come out wrong.
- Pause briefly after you tap record and before you stop, so the first and last words are not cut off.
Settings that matter
- Pick the spoken language where the tool asks for it. Gboard voice typing, WhatsApp transcripts and phone recorder apps each have their own language setting, and a wrong setting ruins the result. Our guide on WhatsApp transcription languages shows the WhatsApp setting.
- Add names and terms if your tool has a vocabulary or phrase list. Google's documentation recommends word and phrase hints "to add names and terms to the vocabulary".
- Keep the original file. Every re-recording or re-compression loses detail. Transcribe the file your recorder made, not a copy played through a speaker. For its own service, Google recommends lossless formats (FLAC or LINEAR16) and a sample rate of 16,000 Hz.
- Do not add noise reduction afterwards. Google advises against automatic gain control and noise reduction for its service and says its recognizer is designed to ignore background voices and noise without extra noise canceling.
Hard to hear audio
If a recording is already made and hard to understand:
- Transcribe the whole file once and see which parts fail.
- Listen to the failed parts at a slower speed with headphones and type them yourself.
- Mark what you cannot hear as [unclear] instead of guessing. A guessed word in a quote or a minute is worse than a gap.
- Do not expect miracles from very distant speech. If a voice is barely audible to you, software will not reliably recover it either.
How to test accuracy yourself
Speech to text accuracy is usually measured as word error rate (WER): WER = (S + D + I) / N, where S is the number of substituted words, D the deleted words, I the inserted words and N the number of words in a correct reference text. Lower is better.
A home version takes five minutes:
- Pick a paragraph of about 100 words with a few names and numbers.
- Read it aloud into your tool in your normal setting.
- Count the words that are wrong, missing or added, and divide by the number of words in the paragraph.
- Change one thing (distance, room, language setting) and repeat.
The result tells you what helps with your voice, your phone and your room. It does not compare tools in general: a published accuracy figure is only valid for the audio it was measured on.
What an AI rewrite fixes, and what it does not
Retexta, a cloud voice to text and AI writing app for Android, transcribes first and can then rewrite the text. The rewrite reliably handles things that make raw transcripts hard to read:
- punctuation, capitalization and paragraphs;
- filler words and false starts ("so, um, I mean");
- repeated phrases and broken sentences.
If you only want the transcript tidied up, the Clean up a text goal (under More templates) is described in the app as: keep your text, fix grammar, punctuation and flow. The writing goals, such as Write a message or Take notes, go further and restructure what you said.
What a rewrite cannot do is know that "Siobhan" was heard as "Shivon", or that "fifteen" should have been "fifty". It may even turn a misheard word into a fluent sentence that looks right. So check names, numbers, dates and amounts against the recording. Retexta labels results "AI-generated. Check important details."
Two app details that help with accuracy: Retexta detects the spoken language automatically, so a short clip with a full first sentence gives it more to work with than a single word. And if a recording is silent or extremely quiet, it says so instead of returning text ("Transcription finished, but no text was produced").
Retexta is free to try: new users get a few free Pro results to view in the app, and after them Transcript only also needs Retexta Pro or credits. Rewrite goals and Import audio are part of Retexta Pro, and copying, sharing or exporting results needs Retexta Pro or credits.
Keep reading
Recording a voice memo? See how to transcribe a voice memo on Android. Dictating instead of typing? Read how to dictate an email on Android. Two voices in one recording: how to transcribe an interview. File formats: MP3 to text. More about the app: Retexta, voice to text for Android.
Sources
Checked by FLAME Apps on September 25, 2026.
- Google Cloud, Best practices, Cloud Speech-to-Text (last updated September 18, 2026): microphone position, clipping, sample rate, lossless codecs, no gain control or noise reduction, word and phrase hints. These are recommendations for Google's own service.
- Wikipedia, Word error rate: definition WER = (S + D + I) / N.
- Retexta app code (Android, version 2026.1.23): automatic language detection (
RecordScreen.kt,TranscriptionLanguageCatalog.kt), Clean up a text goal description and error messages (app strings).
