mirror of
https://github.com/yusufipk/dikte.git
synced 2026-09-11 10:56:10 +00:00
Checked against the release listings rather than guessed: whisper.cpp publishes Win32 and x64 for Windows and nothing else, while llama.cpp does publish bin-win-cpu-arm64.zip. So a Snapdragon machine gets a native cleanup model and an emulated transcriber, which is slow enough that the cloud is the better answer there, and neither the code nor the README said so. The test pins it, so that a whisper.cpp release which does start publishing an arm64 build turns the choice red rather than being quietly ignored.
3.1 KiB
3.1 KiB
Dikte on Windows
Press Ctrl+Space, talk, press again: what you said is transcribed, cleaned
up and pasted where your cursor is.
Requirements
- Windows 10/11
- Python 3.11+ with PyQt6 (
pip install PyQt6; install.ps1 installs it when it is missing) - ffmpeg for microphone capture:
winget install Gyan.FFmpeg
Installing
powershell -ExecutionPolicy Bypass -File install.ps1
This adds a Dikte entry to the Start Menu and a dikte command to the
terminal. Add -Autostart to also start it at sign-in; -Uninstall removes
all of it and leaves the repository and your settings alone.
To try it without installing anything:
python dikte.py
First run
- The tray icon appears and the Settings window opens.
- Under API and models, download a local whisper model (the whisper.cpp Windows build is fetched automatically) or enter an OpenAI, Groq or OpenRouter key.
- The shortcut defaults to
Ctrl+Spaceand is changed under Shortcuts. While Dikte runs, Windows' own hotkey service (RegisterHotKey) listens for it: nothing to install and no permission to grant.
What is different from Linux and macOS
- Meeting recording (microphone + speakers) is not supported yet. Windows does not offer what the speakers are playing as a capture device, so there is nothing to record the far side from. Everything else works, including transcribing audio and video files.
- The shortcut is swallowed: while Dikte holds
Ctrl+Space, the focused application does not see it. This is how macOS behaves too, and unlike the Linux listener, which shares the key. - No external tools for the clipboard or the key press: both go straight through the Windows API (the clipboard, SendInput).
- Settings live under
%APPDATA%\Dikte, models and recordings under%LOCALAPPDATA%\Dikte.
Performance
- The local install fetches whisper.cpp's OpenBLAS build, which transcribes about twice as fast as the stock one on a plain CPU. There is no GPU build to fetch for machines without an NVIDIA card, and none for Windows on ARM either: whisper.cpp publishes x64 only, so a Snapdragon machine runs it under emulation and the cloud is the faster option there.
- Setting Settings → API and models → Threads near your physical core count helps noticeably; the server's own default is 4.
- If speed matters more than accuracy,
ggml-smallandggml-baseare much faster;ggml-large-v3-turbo-q5_0transcribes best.
Troubleshooting
- Recording does not start: does
dikte doctorfind ffmpeg, and doesdikte deviceslist your microphone?devicesalso takes a fresh listing, which is what to run after plugging one in. - Nothing is pasted: a normal-privilege process cannot type into an elevated (administrator) window; run Dikte elevated too, or paste by hand. The text lands on the clipboard either way.
- The shortcut does nothing: another application already holds the combination. Dikte says so in a tray notification when it asks for the key; pick a different one under Settings → Shortcuts.