Files
dikte/README.windows.md
T
yusufipek 23d56acfb5 Publish a Windows setup beside the AppImage and the disk image
The releases page had nothing for Windows, so the only way in was a
checkout, a Python and a pip install. What goes out now is one setup
program per release: PyInstaller's directory, the pinned ffmpeg the
disk image already uses, and Inno Setup around both. It installs for
the account alone, so no administrator is asked for.

Two executables over the one program there, because a windowed one on
Windows has no standard output at all: Dikte.exe for the Start Menu and
dikte.exe for the terminal, sharing everything they carry. The icon is
drawn by Dikte itself into an .ico, the way the Mac's .icns and Linux's
PNGs already are, so there is still no image file in the repository.

Starting at sign-in is a registry value rather than a Startup shortcut,
which is what lets the setup program, the uninstaller and
`dikte integrate` all mean the same thing: the wizard asks once, and
typing the command changes the answer later.

The three builds move into build.yml, which release.yml now calls
instead of holding its own copy, and which a pull request touching the
packaging runs on its own. A broken build is then a red pull request
rather than a failed release.
2026-08-18 16:08:47 +03:00

3.8 KiB

Dikte on Windows

Press Ctrl+Space, talk, press again: what you said is transcribed, cleaned up and pasted where your cursor is.

Requirements

Windows 10 or 11. The setup on the releases page carries everything else with it, and is x64, which an ARM machine runs emulated the way it runs whisper.cpp. A checkout wants:

  • Python 3.11+ with PyQt6 (pip install PyQt6; install.ps1 installs it when it is missing)
  • ffmpeg for microphone capture: winget install Gyan.FFmpeg

Installing

Dikte-<version>-x64-setup.exe from the releases page installs for your account alone, so no administrator is asked for, and puts down a Start Menu entry, a dikte command and, unless you untick it, a start at sign-in. It is signed with no certificate, so SmartScreen offers only Don't run until you press More info. Add/Remove Programs uninstalls it, and dikte integrate and dikte integrate --remove are the sign-in entry on its own, for changing your mind about that later.

From a checkout instead:

powershell -ExecutionPolicy Bypass -File install.ps1

This adds a Dikte entry to the Start Menu and a dikte command to the terminal. Add -Autostart to also start it at sign-in; -Uninstall removes all of it and leaves the repository and your settings alone.

To try it without installing anything:

python -m dikte

First run

  1. The tray icon appears and the Settings window opens.
  2. Under API and models, download a local whisper model (the whisper.cpp Windows build is fetched automatically) or enter an OpenAI, Groq or OpenRouter key.
  3. The shortcut defaults to Ctrl+Space and is changed under Shortcuts. While Dikte runs, Windows' own hotkey service (RegisterHotKey) listens for it: nothing to install and no permission to grant.

What is different from Linux and macOS

  • Meeting recording (microphone + speakers) is not supported yet. Windows does not offer what the speakers are playing as a capture device, so there is nothing to record the far side from. Everything else works, including transcribing audio and video files.
  • The shortcut is swallowed: while Dikte holds Ctrl+Space, the focused application does not see it. This is how macOS behaves too, and unlike the Linux listener, which shares the key.
  • No external tools for the clipboard or the key press: both go straight through the Windows API (the clipboard, SendInput).
  • Settings live under %APPDATA%\Dikte, models and recordings under %LOCALAPPDATA%\Dikte.

Performance

  • The local install fetches whisper.cpp's OpenBLAS build, which transcribes about twice as fast as the stock one on a plain CPU. There is no GPU build to fetch for machines without an NVIDIA card, and none for Windows on ARM either: whisper.cpp publishes x64 only, so a Snapdragon machine runs it under emulation and the cloud is the faster option there.
  • Setting Settings → API and models → Threads near your physical core count helps noticeably; the server's own default is 4.
  • If speed matters more than accuracy, ggml-small and ggml-base are much faster; ggml-large-v3-turbo-q5_0 transcribes best.

Troubleshooting

  • Recording does not start: does dikte doctor find ffmpeg, and does dikte devices list your microphone? devices also takes a fresh listing, which is what to run after plugging one in.
  • Nothing is pasted: a normal-privilege process cannot type into an elevated (administrator) window; run Dikte elevated too, or paste by hand. The text lands on the clipboard either way.
  • The shortcut does nothing: another application already holds the combination. Dikte says so in a tray notification when it asks for the key; pick a different one under Settings → Shortcuts.