Duper Whisper: private speech-to-text for Windows

My own project

Press a hotkey, speak, and the text appears in whatever app you're using. It runs on a local GPU, so nothing leaves the machine and there's no subscription.

The problem

I wanted dictation that works in any app, including a code editor or terminal, without sending my voice to a cloud service.

33
user stories in the written specification
6
planned phases, one commit each

What I built

  • A Windows tray app with a global hotkey, microphone capture and two ways of entering text: paste, or simulated typing for apps that block paste
  • A FastAPI server in WSL2 that keeps a Whisper model loaded on the GPU, and can swap models without restarting
  • A nine-step clean-up of the text: punctuation, numbers, filler words, spoken commands like "new line", and custom word replacements
  • A settings window with live changes, start-up at login and clear messages when something goes wrong

About 1,600 lines of Python, with a README covering setup, GPU checks and troubleshooting.

Tools: Python, FastAPI, faster-whisper, CUDA, WSL2, customtkinter

Diagrams and screenshots

Architecture diagram: a Windows tray app sends audio over localhost to a FastAPI server in WSL2 running faster-whisper on the GPU
The tray app and the GPU transcription server
One spoken sentence passing step by step through the text clean-up pipeline until it becomes clean text
One sentence going through the text clean-up steps

Need something similar?

Send a few lines about the problem. I'll reply with questions or a rough plan.

Email me about a project