mirror of
https://github.com/yusufipk/dikte.git
synced 2026-09-11 10:56:10 +00:00
Merge current local master and preserve automatic language detection
This commit is contained in:
@@ -116,11 +116,12 @@ Speech to text and cleanup each pick a provider in the settings window, and both
|
||||
run here by default, on models of your own. The cloud is the other option:
|
||||
speech to text on **OpenAI**, **Groq** or **OpenRouter** (`gpt-4o-transcribe`),
|
||||
cleanup on OpenRouter (`google/gemini-3.5-flash-lite`), on **Google AI Studio**
|
||||
(`gemini-3.5-flash-lite`) or, when one of them is installed, on Claude Code,
|
||||
Codex or Antigravity. The first two are a single HTTP request; the three CLIs
|
||||
each open a whole session to do it, which is where their few extra seconds go.
|
||||
The keys fall back to `OPENAI_API_KEY`, `GROQ_API_KEY`, `OPENROUTER_API_KEY`
|
||||
and `GEMINI_API_KEY`, and are stored in
|
||||
(`gemini-3.5-flash-lite`), on **OpenCode Go** (`deepseek-v4-flash`) or, when one
|
||||
of them is installed, on Claude Code, Codex or Antigravity. The first three are
|
||||
a single HTTP request; the three CLIs each open a whole session to do it, which
|
||||
is where their few extra seconds go. The keys fall back to `OPENAI_API_KEY`,
|
||||
`GROQ_API_KEY`, `OPENROUTER_API_KEY`, `GEMINI_API_KEY` and `OPENCODE_API_KEY`,
|
||||
and are stored in
|
||||
`~/.config/dikte/config.json`, mode 600, or in
|
||||
`~/Library/Application Support/Dikte` on a Mac. Cleanup can be switched off, in
|
||||
which case the raw transcript is pasted, and a thinking model's effort can be
|
||||
@@ -160,9 +161,14 @@ running.
|
||||
- **It all runs on this machine by default.** Speech to text on whisper.cpp and
|
||||
cleanup on llama.cpp, neither installed beforehand: the settings window fetches
|
||||
the program and the model, verifies the sha256 and refuses a download published
|
||||
without one, then keeps a server alive while you dictate. The graphics card is
|
||||
without one, then keeps a server alive while you dictate and hands the memory
|
||||
back once it has sat unused for ten minutes. The model list is
|
||||
grouped by model rather than by file size, and the row this machine's memory
|
||||
and graphics can take is marked. The graphics card is
|
||||
reached through CUDA, ROCm or Vulkan where the build allows. No key, no
|
||||
account, nothing leaving the machine.
|
||||
account, nothing leaving the machine. On x86_64 Linux the same button fetches
|
||||
a Vulkan build of whisper-server that Dikte publishes itself, because
|
||||
upstream's Linux archive is processor-only.
|
||||
- **Silence never reaches the API.** Handed near-silence, a transcription model
|
||||
invents a sentence instead of returning nothing ("Thanks for watching", or in
|
||||
Turkish "Altyazı M.K."). A recording is dropped when nothing rose 10 dB above
|
||||
@@ -195,10 +201,10 @@ running.
|
||||
a thing you can say to a window that is not Claude. Codex (`codex exec`) and
|
||||
Antigravity (`agy -p`) run the same way, though Antigravity takes neither a
|
||||
permission mode nor a sandbox from Dikte: what it may do without asking is
|
||||
whatever its own allow-rules say. OpenRouter is there as a plain
|
||||
question-and-answer fallback for a machine with no CLI on it. Provider, model,
|
||||
permissions and working directory are under Settings → Agent, and commands
|
||||
close together stay in one conversation.
|
||||
whatever its own allow-rules say. OpenRouter or OpenCode Go is there as a
|
||||
plain question-and-answer fallback for a machine with no CLI on it. Provider,
|
||||
model, permissions and working directory are under Settings → Agent, and
|
||||
commands close together stay in one conversation.
|
||||
- **Meetings** are recorded from the microphone and the speaker output at the
|
||||
same time, which settles who said what by the channel a voice arrived on
|
||||
instead of guessing at it. The two sides are transcribed separately and
|
||||
@@ -250,7 +256,7 @@ cli.py the command line: every verb, and what it answers with
|
||||
ipc.py one request and one reply over the local socket
|
||||
audio.py PCM capture: pw-record for dictation, ffmpeg for a meeting
|
||||
meeting.py channel split, speaker labelling, cleanup, minutes
|
||||
assistant.py handing a dictation to Claude Code, Codex, agy or OpenRouter
|
||||
assistant.py handing a dictation to Claude Code, Codex, agy or a chat model
|
||||
api.py transcription and cleanup requests (stdlib only)
|
||||
cleanup.py who rewrites the transcript: a hosted model, one here, a CLI
|
||||
ggml.py whisper.cpp and llama.cpp here: fetch, verify, keep serving
|
||||
|
||||
Reference in New Issue
Block a user