The button disappeared the moment anything landed, and nothing else on
the window asks for that download. So a machine that got the processor
build before its graphics driver was installed can never be given the
Vulkan one, and nobody can pick up a newer whisper.cpp either: both of
those are the same missing control.
It reads "Download again" once a copy is here, and stays hidden while a
system one is on the PATH, because that is the copy that would run.
The trigger listed dikte/ggml.py, both test files and the two READMEs,
which are among the files that change most often. A typo fix in the
README started a run that builds a container image, compiles every
Vulkan shader and starts two more containers, with a 45 minute timeout
on it because that is roughly what it costs.
Nothing is lost. What ties ggml.py to the release, the tag, version,
commit and reviewed digest, is asserted in tests/test_packaging.py, and
that file already runs on every pull request in milliseconds.
The install falls back to upstream's CPU archive whenever Dikte's own
release, the file in it, or its reviewed digest is not there, and until
the package is published by hand that is every download. It happened
without a word: the window said "Downloaded, version v1.9.3." either
way, and a graphics card sitting idle looks exactly like one being used.
The install record now carries which of the two builds landed, written
only where both were on offer, and the settings window says so on the
line that already reports the version.
Refusing the install on a Python without the extraction filters is a
change to how every archive on every platform is unpacked, and it has
nothing to do with shipping a Vulkan whisper-server. On those Pythons
the download stops working altogether, which is a worse answer than the
one that was there.
Worth doing on its own terms, in its own change, where the versions it
turns away can be argued about without a backend release riding on it.
A timestamped run on OpenRouter always asked openai/whisper-1 for the
segments, whatever model was picked for plain transcription. Not every
model there returns segment times, so the one to use is now its own
setting, openrouter_subtitle_model, shown in the speech-to-text box only
when OpenRouter is the provider. Empty keeps the old whisper-1 fallback.
Target carries the choice as subtitle_model and timestamp_model() reads
it; the other providers are unchanged.
Two things the first pass got wrong.
The card was named out of whichever listing came first, so a machine
with both a CUDA build and a Vulkan loader could have its card named
from the wrong one. Worse, whisper numbers every device it can see in
one sequence while the handle carries the backend's own index: a
processor-only build lent a Vulkan backend through GGML_BACKEND_PATH
reports "device 1: Vulkan0", and reading that listing by the handle's
digit names whatever sat in slot zero. The backend's own enumeration is
asked first now, because it is the only one indexed the way the handle
is, and whisper's listing is a fallback taken only when it names
exactly one card.
And "this build carries no graphics backend" left the reader with
nowhere to go. It is the one case with a fix worth naming: whisper.cpp
publishes no build that reaches a card on Linux, and program_path runs
a copy from the system ahead of the downloaded one, so installing one
is the whole remedy. Server.state() now says which copy is running,
and the line says so only when it is Dikte's own download; a system
build that cannot reach the card gets the shorter sentence, since
installing it again would change nothing.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019zqCqhmeaNT6m1GPp8ZWPp
"Use the graphics card" was a flag and nothing else: whisper got -ng
when it was off, llama got -ngl 99 or 0, and nobody looked at what
happened next. A build without a GPU backend runs on the processor
while the box stays ticked, which is what the release archives for
Linux do, every time. Nothing anywhere reported whether the local
server was even up.
Both servers already say where the model went, in the log Dikte
captures. It is read back once the server reports ready and turned
into a backend, a card and, for llama, the layers it offloaded. The
verdict comes from what whisper committed to -- "using X backend" and
the model buffer -- rather than from the devices it merely listed: a
card that is found and then fails to initialise sends it back to the
processor, and the listing alone would have called that a graphics
card. A log that says nothing stays "could not tell" instead of being
guessed at; a hand-built macOS whisper has Metal compiled in and
prints no backend line at all.
The state is then somewhere to be seen. Server.state() is a snapshot
of the process and what it settled on, and it reaches `dikte status`,
`dikte doctor` and a line under each local model box in the settings
window. doctor reads the log from disk when no instance is running, so
it still answers on a machine where Dikte is closed, and it says which
of the three it is: what the last run used, that the last run said
nothing, or that none ever ran here. Where the card was asked for and
not obtained, the line says which of the two it is -- none was found,
or this build carries none -- because only the second is worth
replacing a download over.
The tests grew a second isolation. They read ggml's own data
directory, which is the real one on the machine running them, and
program_path prefers a whisper-server on the PATH, so the suite
answered from whatever the developer happened to have installed. Both
are now the test's own.
Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_019zqCqhmeaNT6m1GPp8ZWPp
Review found the detected language was only switching the glossary rule, not
the base prompt: a detected Turkish recording with an English interface got the
English cleanup prompt with a Turkish glossary footnote. Both now follow the
spoken language, and the detected-language value is guarded against a server
that returns something other than a string. The record reply is covered by a
CLI test.
New installs start detecting instead of being locked to one language; a stored
value from before this default still wins. The dictation chain asks
transcribe_detected() in auto mode, records the detected code in history as
speech_language, hands it to the cleanup prompt (a detected Turkish recording
gets the Turkish prompt and glossary rule), and reports it on the socket reply.
The stale comment claiming whisper.cpp's -l auto does not detect is corrected.
The spoken language is only knowable when the transcription model reports it,
and only whisper.cpp does: the hosted endpoints accept auto but never say what
they heard. transcribe_detected() asks for detection exactly where it can be
answered, the local server in auto mode, by requesting verbose_json with
no_language_probabilities switched back on for that one request (the server
runs with -nlp, which keeps the sweep off everything else).
The agent and meeting tabs opened with an intro paragraph, and the
cleanup provider box showed a long tooltip comparing the providers.
None of them earned the space: the intro texts were unclear and the
tooltip restated what picking a provider already shows. The orphaned
Turkish translations go with them.
Master grew Google AI Studio and Antigravity as providers, a doctor that
names each provider's own key, and settings that fetch every hosted
model list as the window opens. OpenCode Go is folded into each: its
name joins SERVICES and the doctor's key table, its key row sits beside
Google's, its Fetch button follows the per-provider pattern Google's
uses, and _load_hosted_models fetches its catalog at open when a key is
on file, filling the cleanup and agent boxes alike.
Qt's xcb plugin links libxkbcommon and libxkbcommon-x11, and the x11 half
hands keymap objects to the core half to free, so the pair must come from
one build. The build machine has only the core half installed, so
PyInstaller bundled that one alone and the other kept loading from the
user's system; an Ubuntu 22.04 core freeing what a current x11 half
allocated is the segfault of issue #57, which a Turkish layout happened
to move to startup. Every desktop that can show a window carries both
libraries from one build, so the bundle now carries neither.
Fixes#57
Codex's list already arrived on its own, because reading a cache on disk
costs nothing. The other three lists were behind a Fetch button, which
meant the built-in ones aged in front of anyone who never pressed it. Now
OpenRouter's and Google's lists are fetched when the settings window
opens, and Antigravity's comes off `agy models`, which prints one
id-tab-name line per model and answers over the network in a couple of
seconds; all three run off the interface thread.
Nobody asked, so nothing is reported: a failure changes nothing on
screen, the built-in lists stay, and the Fetch buttons remain both the
retry and the place an error is worth explaining. A provider whose key
has not been given is not called at all, so opening Settings is not by
itself a request to two vendors.
Also drops an import the merge had left in twice.
Co-Authored-By: Claude Fable 5 <[email protected]>
Codex already refreshes its boxes from the source at open, so the
built-in list is never the whole truth for longer than a window takes to
build. OpenCode Go now gets the same courtesy: one request to /models on
a background thread, both of its boxes refilled with the picked model
kept, skipped entirely when there is no key to send.
Nothing was measured that put OpenCode Go beside it, so the tooltip
claims only what is known: OpenRouter is the quickest, and OpenCode Go
merely needs nothing installed either. The translation follows word for
word.
ox-alpha-free is not in the /models answer, and the comment claimed the
chat endpoint hides Grok, Luna, MiniMax and Qwen when its own catalog
lists them. The built-in list is a starting set; the Fetch button asks
the endpoint for the full catalog of the day.
The tooltip gained OpenCode Go but its translation was added without the
llama.cpp sentence the tooltip actually carries, so a Turkish window
showed the English text. The full entry now matches the tooltip word for
word, and the two shorter variants nothing looks up any more are gone.
The one button lived in the OpenRouter model row, which leaves the
screen whenever another provider is chosen, so the OpenCode fetch path
could never be clicked. The OpenCode row now carries its own button into
the same handler, the row as a whole is what hides, and the fetched list
lands in the box it belongs to without touching the meeting box.
The settings-storage paragraph goes back to the one sentence it was. The screen list now shows native resolutions rather than the scaled ones, the Turkish tab is named Ekran, and the repositioning comment says what a named screen changes.
Codex now asks itself for its model list, doctor grew ready flags and a
per-provider line, and the Groq key joined the masked ones; the OpenCode
additions are folded into each. The no-CLI cleanup test follows the
fake_run to fake_cli rename.
Both sides rewrote the doctor: master rebuilt it around ready flags so a
fully local setup stops reading as broken, while this branch taught it
that Google has a key of its own and that the local model has neither a
key nor a program. Kept master's structure and folded the branch in: the
two hosted providers answer for their keys, a CLI for its program, and
the JSON keeps both the provider-aware "key" and master's "ready".
SECRET_KEYS gained groq on master and gemini here; the resolution keeps
all four. _conclude was also changed by both: master stopped matching
stderr wording and asks _API_TROUBLE instead, this branch renamed its
last parameter to the provider's short name so a session is stored under
the name it is read back by; the rename now rides on master's body.
codex_models() and the settings loaders landed beside the new Gemini
ones, so both stay. The new cleanup tests called the CLI fake by its old
name, fake_run, which master had renamed to fake_cli.
The paste block was rewritten by both sides: master taught press() to
put the remembered application back in front (focus), this branch moved
the history write ahead of the paste and stopped restoring the old
clipboard over a transcript the key press refused. Kept this branch's
order and error handling, and handed press() the focus it now takes.
deleteLater destroyed a window whose model boxes a daemon thread was
still reporting into: a running download died with a RuntimeError, the
.part stayed on disk and the UI said nothing. Dropping the reference is
how the ordinary close path lets a window go, and the closures a thread
holds keep the object alive until it is done, so the rebuild now does
the same. While the transcriber or either model box is working the
rebuild is skipped altogether; _built_language keeps the old language,
so the next quiet save asks for it by itself.
The new window is also placed and turned to the old tab before it is
shown, so it no longer comes up at the default size and jumps. And
tests/test_ui.py now says that a save with a changed language emits
language_changed, and one without does not.