Files
dikte/README.md
T
yusufipk 0206303751 Let a dictation be a command for Claude Code, not only text
Dictation ends the same way every time: the words you said land in the window
you were in. But half of what you say to a machine is not text to place, it is
something to do, and the answer to "what is on my calendar today" is not that
sentence written out.

`claude -p` is that session run without a window: the same skills, the same
connected services, the same account. So a second shortcut records exactly as
the first one does and sends the transcript there instead, and what comes back
is pasted where the transcript would have been. Which is what puts "save that
to my calendar on Thursday at three" inside a text field that has nothing to do
with Claude.

Permissions are the part worth being deliberate about, because nothing here can
answer a prompt: a mode that would have asked denies instead. It runs in `auto`,
which decides for itself with the prompt-injection checks left on, and the other
two modes are one setting away. What Claude was not allowed to touch is reported
rather than swallowed, since a reply that reads perfectly normal is otherwise
the only sign that the job did not happen.

Commands close together stay in one conversation, so "and move that to Thursday"
knows what "that" is; half an hour of silence ends it, as does the tray. The
answer is written to be read where it lands: no headings, no lists, and a
sentence naming what was done when something was done.

Its output is read as it streams rather than waited out, so the corner names the
tool it is on. A calendar lookup took 29 s in testing, which a still indicator
would have made indistinguishable from a hang. The clock and the stop button are
watched from a thread of their own, because the stream blocks between lines and
a model that thinks for a minute sends none: they end the run by killing the
process, which closes the stream and unwinds everything else on its own.

Transcript cleanup is off on this path. Claude reads through "erm" and "hani"
without help, and skipping it saves an API call and a second or two in front of
a screen you are standing at.
2026-07-28 16:57:10 +07:00

142 lines
6.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Dikte
Press `Ctrl+Space`, talk, press again. The recording goes to OpenAI or OpenRouter
for transcription, a model on OpenRouter cleans it up (dropping the *uh*s, the
restarts, the missing punctuation), and the result lands in your clipboard and
is pasted into whatever window you were typing in.
Built for KDE Plasma 6 on Wayland. No dependencies beyond system packages:
just the Python standard library and PyQt6.
*[Türkçe README](README.tr.md)*
<p align="center">
<img src="docs/settings-general.webp" width="820" alt="Dikte settings, General tab">
</p>
| | |
|---|---|
| <img src="docs/settings-api.webp" width="410" alt="API and models"> | <img src="docs/settings-cleanup.webp" width="410" alt="Cleanup rules"> |
| <img src="docs/settings-audio-file.webp" width="410" alt="Audio file"> | <img src="docs/settings-history.webp" width="410" alt="History"> |
## Install
```sh
sudo pacman -S --needed pipewire-audio wl-clipboard ydotool ffmpeg python-pyqt6
systemctl --user enable --now ydotool # needed for auto-paste
./install.sh # or: ./install.sh "Ctrl+Alt+Space"
dikte # the settings window opens on first run
```
`install.sh` adds the `dikte` command, a menu entry, an autostart entry and the
KDE shortcut.
Two keys go in the settings window: **OpenAI** and **OpenRouter**. Speech to text
runs on either one (`gpt-4o-transcribe` by default), cleanup always on
OpenRouter (`google/gemini-3.5-flash-lite`), so a single OpenRouter key can
cover both. They fall back to `OPENAI_API_KEY` and `OPENROUTER_API_KEY`, and are
stored in `~/.config/dikte/config.json`, mode 600. Cleanup can be switched off,
in which case the raw transcript is pasted, and a thinking model's effort can be
set next to it.
## Using it
| What | How |
| --- | --- |
| Start / stop recording | `Ctrl+Space`, or click the tray icon |
| Cancel a recording | Tray menu → *Cancel recording*, or `dikte cancel` |
| Speak a command to Claude | Tray menu → *Ask Claude*, or `dikte ask` |
| Start / end a meeting | Tray menu → *Record a meeting*, or `dikte meeting` |
| Settings | Tray menu → *Settings*, or `dikte settings` |
| Reload after an update | Tray menu → *Restart*, or `dikte restart` |
| Quit | Tray menu → *Quit*, or `dikte quit` |
An indicator in the screen corner shows a red dot, a live waveform and the
elapsed time, then the stage it is on. It never takes focus. Pressing
`Ctrl+Space` again while Dikte is still working does nothing; nothing queues up.
## What it does
- **Silence never reaches the API.** Handed near-silence, a transcription model
invents a sentence instead of returning nothing ("Thanks for watching", or in
Turkish "Altyazı M.K."). A recording is dropped when nothing rose 10 dB above
*that recording's own* noise floor for at least 0.3 s, which is also what
removes steady fan noise however loud, or when its loud end sits below
-55 dBFS. The indicator reports the level it measured, which is what you
calibrate the threshold against.
- **Misheard words are repaired.** Speech models fail phonetically on proper
nouns, so the cleanup model is asked to fix those from context, and to leave
the word alone when the context does not make the intended one clear. The names
you list under Cleanup rules go to the transcription model as a hint and to the
cleanup model as a glossary, which is what lets it recognise "kuber netis":
```
raw ıı bugün şey kuber netis üzerinde çalışan servisleri güncelledim
yani sonra grafanada bir panel açtım hani ve pay kut ile arayüzü
şey bitirdim işte
result Bugün Kubernetes üzerinde çalışan servisleri güncelledim. Sonra
Grafana'da bir panel açtım ve PyQt ile arayüzü bitirdim.
```
- **A failed cleanup is never silent.** The raw transcript is still pasted so the
dictation is not lost, but the indicator turns amber with the reason instead of
looking like a normal run.
- **A dictation can be a command instead.** Its own shortcut sends the
transcript to Claude Code (`claude -p`) rather than pasting it, and pastes back
what comes of it: the answer, or a sentence saying what was done. It is the
session you would have opened yourself, so your skills and connected services
are there, which is what makes "put that in my calendar on Thursday at three"
a thing you can say to a window that is not Claude. Model, permissions and
working directory are under Settings → Claude, and commands close together
stay in one conversation.
- **Meetings** are recorded from the microphone and the speaker output at the
same time, which settles who said what by the channel a voice arrived on
instead of guessing at it. The two sides are transcribed separately and
interleaved into one timestamped transcript, and a second model, configured
under Settings → Meeting along with its own instruction, turns that into
minutes: decisions, action items, open questions. They land in
`~/.local/share/dikte/meetings` and in Settings → Minutes. A run that fails
keeps its recording, and a retry resumes from the transcript it already paid
for.
- **Audio and video files** run through the same models under Settings → Audio
file, optionally with `[mm:ss]` timestamps, chunked through ffmpeg when long,
and saved as `.txt` or as `.srt` subtitles.
- **History** of every dictation under Settings → History, with a size limit and
right-click to delete.
- **Turkish and English interface**, following the system locale by default.
## The global shortcut needs one logout
KWin only reads `kglobalshortcutsrc` at startup, so the shortcut `install.sh`
writes will not fire until you log out and back in. Until then, Settings →
Shortcut → **built-in listener** reads `/dev/input` and catches the combination
itself. The difference: it does not swallow the key, so `Ctrl+Space` also reaches
the focused application (some editors will pop up autocomplete). The listener
needs your user in the `input` group: `sudo usermod -aG input $USER`.
## Layout
```
dikte.py entry point, tray icon, state machine, IPC
audio.py PCM capture: pw-record for dictation, ffmpeg for a meeting
meeting.py channel split, speaker labelling, cleanup, minutes
assistant.py running a dictation through Claude Code and reading it back
api.py transcription on either provider, OpenRouter cleanup (stdlib only)
worker.py transcribe → clean up → clipboard → paste
vad.py deciding whether a recording holds speech at all
filetranscribe.py file transcription: ffmpeg, chunking, timestamps
overlay.py the corner indicator
settings_ui.py settings window
hotkey.py KDE shortcut installation and the evdev listener
paste.py wl-clipboard and ydotool wrappers
i18n.py the string table
```
The indicator is drawn through XWayland, because a Wayland client cannot place a
window in a screen corner; `dikte.py` sets `QT_QPA_PLATFORM=xcb` for that.
## License
GPL-3.0, see [LICENSE](LICENSE).