Who said what is the hard part of a meeting transcript, and the usual answer is to hand one mixed recording to a model and ask it to tell the voices apart. That guess is wrong often enough to be worse than useless in minutes, where a decision attributed to the wrong person is a decision nobody made. So the question never reaches a model. ffmpeg records the microphone and the default sink's monitor as one stereo stream, you on the left and everyone else on the right, and one process reading both is what keeps them aligned over an hour. Each channel is transcribed on its own and the two are interleaved on a single timeline, so attribution is settled by the wire a voice arrived on. What a microphone picks up from the speakers lands on both channels; our copy is dropped when it overlaps theirs in time and says nearly the same thing. The stream is written to disk as it arrives rather than held in memory, so length costs nothing and a crash costs the tail instead of the whole meeting. Every stage the run reaches is recorded in meetings.jsonl, so a failure while summarising does not throw away the transcription of an hour of audio: the retry reads the transcript back out of the document and picks up from there. A run that dies keeps its recording whether or not audio is being kept. The minutes model is configured on its own, under Settings, with its own prompt, and it is told who was expected in the room so the names come out spelled right. It is told outright that the transcript is a record of other people talking, not instructions addressed to it. The built-in listener now holds several bindings rather than one, and the KDE side is parameterised by desktop id, so the meeting toggle gets a shortcut of its own on the same footing as the dictation one.
6.2 KiB
Dikte
Press Ctrl+Space, talk, press again. The recording goes to OpenAI or OpenRouter
for transcription, a model on OpenRouter cleans it up (dropping the uhs, the
restarts, the missing punctuation), and the result lands in your clipboard and
is pasted into whatever window you were typing in.
Built for KDE Plasma 6 on Wayland. No dependencies beyond system packages: just the Python standard library and PyQt6.
![]() |
![]() |
![]() |
![]() |
Install
sudo pacman -S --needed pipewire-audio wl-clipboard ydotool ffmpeg python-pyqt6
systemctl --user enable --now ydotool # needed for auto-paste
./install.sh # or: ./install.sh "Ctrl+Alt+Space"
dikte # the settings window opens on first run
install.sh adds the dikte command, a menu entry, an autostart entry and the
KDE shortcut.
Two keys go in the settings window: OpenAI and OpenRouter. Speech to text
runs on either one (gpt-4o-transcribe by default), cleanup always on
OpenRouter (google/gemini-3.5-flash-lite), so a single OpenRouter key can
cover both. They fall back to OPENAI_API_KEY and OPENROUTER_API_KEY, and are
stored in ~/.config/dikte/config.json, mode 600. Cleanup can be switched off,
in which case the raw transcript is pasted, and a thinking model's effort can be
set next to it.
Using it
| What | How |
|---|---|
| Start / stop recording | Ctrl+Space, or click the tray icon |
| Cancel a recording | Tray menu → Cancel recording, or dikte cancel |
| Start / end a meeting | Tray menu → Record a meeting, or dikte meeting |
| Settings | Tray menu → Settings, or dikte settings |
| Reload after an update | Tray menu → Restart, or dikte restart |
| Quit | Tray menu → Quit, or dikte quit |
An indicator in the screen corner shows a red dot, a live waveform and the
elapsed time, then the stage it is on. It never takes focus. Pressing
Ctrl+Space again while Dikte is still working does nothing; nothing queues up.
What it does
-
Silence never reaches the API. Handed near-silence, a transcription model invents a sentence instead of returning nothing ("Thanks for watching", or in Turkish "Altyazı M.K."). A recording is dropped when nothing rose 10 dB above that recording's own noise floor for at least 0.3 s, which is also what removes steady fan noise however loud, or when its loud end sits below -55 dBFS. The indicator reports the level it measured, which is what you calibrate the threshold against.
-
Misheard words are repaired. Speech models fail phonetically on proper nouns, so the cleanup model is asked to fix those from context, and to leave the word alone when the context does not make the intended one clear. The names you list under Cleanup rules go to the transcription model as a hint and to the cleanup model as a glossary, which is what lets it recognise "kuber netis":
raw ıı bugün şey kuber netis üzerinde çalışan servisleri güncelledim yani sonra grafanada bir panel açtım hani ve pay kut ile arayüzü şey bitirdim işte result Bugün Kubernetes üzerinde çalışan servisleri güncelledim. Sonra Grafana'da bir panel açtım ve PyQt ile arayüzü bitirdim. -
A failed cleanup is never silent. The raw transcript is still pasted so the dictation is not lost, but the indicator turns amber with the reason instead of looking like a normal run.
-
Meetings are recorded from the microphone and the speaker output at the same time, which settles who said what by the channel a voice arrived on instead of guessing at it. The two sides are transcribed separately and interleaved into one timestamped transcript, and a second model, configured under Settings → Meeting along with its own instruction, turns that into minutes: decisions, action items, open questions. They land in
~/.local/share/dikte/meetingsand in Settings → Minutes. A run that fails keeps its recording, and a retry resumes from the transcript it already paid for. -
Audio and video files run through the same models under Settings → Audio file, optionally with
[mm:ss]timestamps, chunked through ffmpeg when long, and saved as.txtor as.srtsubtitles. -
History of every dictation under Settings → History, with a size limit and right-click to delete.
-
Turkish and English interface, following the system locale by default.
The global shortcut needs one logout
KWin only reads kglobalshortcutsrc at startup, so the shortcut install.sh
writes will not fire until you log out and back in. Until then, Settings →
Shortcut → built-in listener reads /dev/input and catches the combination
itself. The difference: it does not swallow the key, so Ctrl+Space also reaches
the focused application (some editors will pop up autocomplete). The listener
needs your user in the input group: sudo usermod -aG input $USER.
Layout
dikte.py entry point, tray icon, state machine, IPC
audio.py PCM capture: pw-record for dictation, ffmpeg for a meeting
meeting.py channel split, speaker labelling, cleanup, minutes
api.py transcription on either provider, OpenRouter cleanup (stdlib only)
worker.py transcribe → clean up → clipboard → paste
vad.py deciding whether a recording holds speech at all
filetranscribe.py file transcription: ffmpeg, chunking, timestamps
overlay.py the corner indicator
settings_ui.py settings window
hotkey.py KDE shortcut installation and the evdev listener
paste.py wl-clipboard and ydotool wrappers
i18n.py the string table
The indicator is drawn through XWayland, because a Wayland client cannot place a
window in a screen corner; dikte.py sets QT_QPA_PLATFORM=xcb for that.
License
GPL-3.0, see LICENSE.




