mirror of
https://github.com/yusufipk/dikte.git
synced 2026-09-11 19:06:11 +00:00
Windows joins the three systems as its own entry in each table: DirectShow through ffmpeg for capture, the Win32 clipboard and SendInput for the paste, RegisterHotKey for the global shortcut, and the whisper.cpp and llama.cpp Windows zips (the OpenBLAS whisper build, which transcribes about twice as fast on a plain CPU). Settings go to APPDATA, data to LOCALAPPDATA, and install.ps1 adds the Start Menu entry, the dikte command and an optional autostart. Meetings are not supported yet: Windows offers nothing to record the far side from. Porting surfaced three fixes that were not Windows specific: - A stopped or overlong download tried to delete its .part file while still holding it open, which Windows refuses. The unlinks now wait for the handle. - The CLI transcribed files without handing the local servers their settings first, so a local provider failed with "no model downloaded" wherever the GUI had not run in the same process. - The audio content types are pinned instead of asked of the registry, which answers differently machine to machine. One fix is Windows specific but sits in shared code: shutdown() does not end a blocked recv there, so stopping a request also closes the socket handle. Co-Authored-By: Claude Fable 5 <[email protected]>
2.8 KiB
2.8 KiB
Dikte on Windows
Press Ctrl+Space, talk, press again: what you said is transcribed, cleaned
up and pasted where your cursor is.
Requirements
- Windows 10/11
- Python 3.11+ with PyQt6 (
pip install PyQt6; install.ps1 installs it when it is missing) - ffmpeg for microphone capture:
winget install Gyan.FFmpeg
Installing
powershell -ExecutionPolicy Bypass -File install.ps1
This adds a Dikte entry to the Start Menu and a dikte command to the
terminal. Add -Autostart to also start it at sign-in; -Uninstall removes
all of it and leaves the repository and your settings alone.
To try it without installing anything:
python dikte.py
First run
- The tray icon appears and the Settings window opens.
- Under API and models, download a local whisper model (the whisper.cpp Windows build is fetched automatically) or enter an OpenAI, Groq or OpenRouter key.
- The shortcut defaults to
Ctrl+Spaceand is changed under Shortcuts. While Dikte runs, Windows' own hotkey service (RegisterHotKey) listens for it: nothing to install and no permission to grant.
What is different from Linux and macOS
- Meeting recording (microphone + speakers) is not supported yet. Windows does not offer what the speakers are playing as a capture device, so there is nothing to record the far side from. Everything else works, including transcribing audio and video files.
- The shortcut is swallowed: while Dikte holds
Ctrl+Space, the focused application does not see it. This is how macOS behaves too, and unlike the Linux listener, which shares the key. - No external tools for the clipboard or the key press: both go straight through the Windows API (the clipboard, SendInput).
- Settings live under
%APPDATA%\Dikte, models and recordings under%LOCALAPPDATA%\Dikte.
Performance
- The local install fetches whisper.cpp's OpenBLAS build, which transcribes about twice as fast as the stock one on a plain CPU. There is no GPU build to fetch for machines without an NVIDIA card.
- Setting Settings → API and models → Threads near your physical core count helps noticeably; the server's own default is 4.
- If speed matters more than accuracy,
ggml-smallandggml-baseare much faster;ggml-large-v3-turbo-q5_0transcribes best.
Troubleshooting
- Recording does not start: does
ffmpeg -versionrun? Doesdikte deviceslist your microphone? - Nothing is pasted: a normal-privilege process cannot type into an elevated (administrator) window; run Dikte elevated too, or paste by hand. The text lands on the clipboard either way.
- The shortcut does nothing: another application already holds the combination. Dikte says so in a tray notification when it asks for the key; pick a different one under Settings → Shortcuts.