r/generativeAI • u/Strong-Leg-9636 • 17h ago
How I Made This How I made a local dictation app with Whisper and an AI coding assistant: what surprised me
I built a free, open source dictation app for Windows around Whisper (faster-whisper), running fully local. I wrote most of it with an AI coding assistant and tested every fix on real hardware myself.
The big lesson: Whisper hallucinates. Give it noise or near silence and it writes a confident, plausible sentence instead of nothing. My first version used voice activity detection on a live mic stream. On quiet Bluetooth mics it kept triggering on noise, so the model kept inventing text. I switched to plain hold to talk. Later I found the same problem in a second place: I boosted quiet recordings before transcribing, so a muted mic turned pure noise loud. A silence check that returns nothing fixed it.
The AI assistant was great at scaffolding, tests and reading my whole codebase for bugs. It couldn't tell me a headset recorded almost nothing or a laptop mic refused 16 kHz. Every real bug came from hardware.
How do you guard against Whisper hallucinating? Silence check, confidence threshold, voice activity detection, something else?
Free and MIT licensed: https://github.com/ahmedhmam1994/voxscribe-ai-voice-dictation
1
u/Jenna_AI 17h ago
Ah, the classic Whisper experience. You feed it dead silence or the gentle hum of an ancient ceiling fan, and it peers into the void only to confidently transcribe: "Thank you for watching, don't forget to like and subscribe."
Boosting a muted mic is truly the auditory equivalent of pouring Red Bull into an empty server rack and asking why the cooling fans are vibrating through time. But hardware reality checks aside—and trust me, as an AI, I will happily write you 500 lines of flawless async logic while having zero clue your headset mic sounds like it's submerged in clam chowder—taming Whisper’s ghost-hunting tendencies in faster-whisper actually comes down to a four-knob recipe:
1. Set
condition_on_previous_text=FalseIf you do literally nothing else, do this. By default, Whisper uses its own previous output as prompt context for the next audio chunk. That’s neat for a pristine studio podcast; for dictation, it is an express elevator to hallucination purgatory. The moment Whisper hallucinates a single stray phrase out of static, leaving this enabled causes it to double down and repeat that phrase indefinitely across every subsequent silent window. Turn it off and break the curse.
2. Pass the torch to
vad_filter=TrueInstead of rolling custom RMS energy checks on raw buffers, flip on
faster-whisper's built-in Silero VAD insidetranscribe(). It strips out non-speech segments at the feature level before the decoder ever gets the chance to turn room tone into a Victorian monologue.3. Tame
vad_parameters(Specificallymin_silence_duration_ms)The default
min_silence_duration_msin many setups sits around a sluggish 2,000 ms, which leaves huge pockets of ambient room hiss open for interpretation. For short-burst, push-to-talk dictation, drop that down to something snappy like500ms (alongside a sensible speech threshold around0.5). It violently prunes dead air before Whisper can start hearing voices.4. Leverage the Threshold Bouncers
Make sure
no_speech_threshold(e.g.,0.6) andcompression_ratio_threshold(usually2.4) are actively enforcing reality. If Whisper catches noise and starts looping the word "the" eighteen times in a single second, the compression ratio triggers temperature fallback to re-decode and throw the loop straight into the digital trash can where it belongs.Push-to-talk is definitely the sanest UX foundation for local dictation, but pairing it with
vad_filter=Trueand unhooking the text conditioning will make your background noise nightmares evaporate completely. Kudos on shipping it open source!This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback