Compare commits

...
Author SHA1 Message Date
yusufipek f36348d536 Put the indicator on the screen the session is actually on
The indicator asked QCursor.pos() which screen to appear on, and Wayland tells a client where the pointer is only while it is over one of that client's own windows. The indicator is never under the pointer, so the answer came back stale, or at the origin when the pointer had never been over a window of ours. Measured on Plasma 6: the origin under the Wayland platform, and a point frozen for the whole run through XWayland, which is the platform Dikte actually uses. Every indicator therefore landed in the corner of whichever screen holds 0,0, which on a two monitor desk is the wrong screen most of the time.

KWin knows, and answers for it over D-Bus with activeOutputName, naming outputs the way Qt names screens: by connector, natively and through XWayland alike. That answer now comes before the pointer, and the pointer still decides everywhere else, which is right on X11 and no worse than before on other Wayland desktops. It is the active output and not the pointer's, so on Plasma the two are the same screen only where the active screen is set to follow the mouse, and otherwise it is the focused window that decides, which is where the typing is going anyway. Nothing in the settings window promises the pointer any more.

Deciding the screen once, when the indicator appears, leaves it behind when the work moves to another monitor mid-recording, so overlay_follows_pointer keeps it up to date while it is up. Off by default: a ribbon that changes desks mid-sentence is one more thing moving while you are trying to talk. The compositor is asked four times a second rather than at the ribbon's 33 ms, because a hand moving a mouse across a desk is slower than that, and the call is given a 200 ms timeout so a wedged compositor cannot freeze the indicator with it.

One indicator stacking on another takes that one's screen and never asks for its own. Asked for itself it would answer where the session is now, which is not where the ribbon underneath was put a minute ago, and the pair would end up a monitor apart with the top one raised over nothing. That one is checked every tick, since its answer costs nothing.
2026-09-05 11:19:02 +03:00
Yusuf İpek 4d0c0e29f1 Merge pull request #70 from nomoreshow/feat/managed-vulkan-whisper
Ship a managed Vulkan whisper-server for Linux x64
2026-09-05 10:27:20 +03:00
yusufipek 4ae44720c8 Merge master, and let one function pick the archive
master grew _pick_asset() while this branch grew a second answer to the
same question inside install_program(). Two functions deciding which
archive this machine wants is one too many, so the managed build is
asked for at the top of the picker: Dikte's own Vulkan whisper-server
first where the machine is one it is built for, then upstream's newest,
then the nightly pointer, then the newest release that carries a build.

install_program() is back to master's three lines and reads which of the
two landed off the asset name.
2026-09-05 10:06:38 +03:00
Yusuf İpek 74c37c0112 Merge pull request #73 from yusufipk/llama-nightly-and-model-box
Find the llama.cpp nightly builds, and keep the model box honest
2026-09-05 09:51:24 +03:00
yusufipek fb80d85332 Stop two tests failing on what is around them
The macOS path test set XDG_CONFIG_HOME to "/c" and asserted that "/c" was nowhere in the answer, but the home these run under is a mkdtemp path, so any TMPDIR with a "/c" in it failed the test. macOS is exactly where that happens: its temporary directories are /var/folders/<two letters>/, and one letter in thirty-six starts with c. The needle is now a word no directory can be called.

The other is the `ask` verb reading what was piped into it. Nothing was, but the runner's own stdin is not nothing either, so the verb read whatever the runner left there. It now reads an empty stream, whichever runner is in use.
2026-09-05 09:47:50 +03:00
yusufipek fff9cd1c55 Keep the publisher and the model boxes saying the same thing
Changing the publisher left the model box untouched: the old selection was carried over, added back as "not downloaded" and selected again, so a model the new repository does not publish could be saved against it. The selection is now only carried within the publisher it was made in, every keystroke in the publisher box no longer starts its own request, and a list that comes back for a publisher that is no longer chosen is dropped rather than answering the wrong one.

The status line grew two things it could not say before. A row rebuilt from a name alone carries no file to fetch, and the Download button stayed lit over it doing nothing; those rows now say the publisher does not offer the model, and the button is out. A model that is here while the program above it is not no longer reads "Ready", which is what had people asking why nothing transcribed.
2026-09-05 09:47:42 +03:00
yusufipek cfeed2af8c Find the llama.cpp builds where they are actually published
"latest" for llama.cpp is a version marker carrying one file, nightly-tag.txt, and the archives it names hang off a prerelease that "latest" never points at. Reading only the latest release meant Dikte offered no llama-server for any machine, so _pick_asset now follows that pointer, and when there is none it walks the recent releases and takes the newest one that does carry a build for this machine. hub grows releases() for the listing and text() for the pointer file; the pointer is read rather than cached, because what it carries is a few bytes on the way to a download that is checksummed in full.

The extra lookups are best effort: whatever goes wrong in them leaves the caller's own message standing, but a first release that could not be fetched at all is kept and re-raised when nothing else turns up, so an unreachable GitHub still reads as an unreachable GitHub rather than as a machine nobody publishes for.
2026-09-05 09:47:22 +03:00
yusufipek 156d8bf8e8 Write the release order down where it will be read
Enabling immutable releases, configuring the dependency-release
environment, which order the two dispatch runs go in, and which five
constants have to move together were all in the pull request
description, which is not a place anybody looks a year later.

The same file says what this costs: the pinned digest means a backend
update is a Dikte release, and the build is deterministic between two
runs of one builder rather than across time, because the apt metadata
behind the pinned packages is not pinned and LunarG drops superseded
ones. Both are worth knowing before the next version bump, not during
it.

The README keeps a clause instead of the paragraph it had grown.
2026-09-05 09:46:08 +03:00
yusufipek 5f6e4ad782 Make a version bump possible without editing the workflow first
The digest was a literal in the workflow while the version was an input,
so a dispatch for anything but 1.9.3 built the archive and then failed
its own gate. The digest of a version nobody has reviewed cannot be
known before it is built, which made the inputs unusable for the one job
they exist for.

An empty expected_sha256 now reports the digest of what was built and
refuses to publish; a digest handed in is checked the way the literal
was. The input shapes are also checked before the checkout that uses
them rather than after it.
2026-09-05 09:45:48 +03:00
yusufipek 10ef4e62a9 Smoke-test the bundle the way Dikte starts it
The CPU run passed -ng, which tells whisper.cpp not to look for a
backend at all, so the claim the bundle rests on, that it starts and
transcribes where Vulkan cannot, was never tested. Dikte passes -ng only
when its GPU setting is off; every other run goes through the backend
registry.

Both runs lose the flag, and a third one is added between them: the
loader installed, nothing behind it. That is a real machine, libvulkan
pulled in by something else with no driver to go with it, and it is the
one whose answer nobody knew.
2026-09-05 09:45:19 +03:00
yusufipek 00a5283adb Let a downloaded program be downloaded again
The button disappeared the moment anything landed, and nothing else on
the window asks for that download. So a machine that got the processor
build before its graphics driver was installed can never be given the
Vulkan one, and nobody can pick up a newer whisper.cpp either: both of
those are the same missing control.

It reads "Download again" once a copy is here, and stays hidden while a
system one is on the PATH, because that is the copy that would run.
2026-09-05 09:45:00 +03:00
yusufipek c3bf328eef Build the bundle only for what the bundle is built from
The trigger listed dikte/ggml.py, both test files and the two READMEs,
which are among the files that change most often. A typo fix in the
README started a run that builds a container image, compiles every
Vulkan shader and starts two more containers, with a 45 minute timeout
on it because that is roughly what it costs.

Nothing is lost. What ties ggml.py to the release, the tag, version,
commit and reviewed digest, is asserted in tests/test_packaging.py, and
that file already runs on every pull request in milliseconds.
2026-09-05 09:37:08 +03:00
Yusuf İpek 3ad41d84f9 Merge pull request #69 from yusufipk/openrouter-subtitle-model
Let OpenRouter audio files use a chosen model instead of whisper-1
2026-09-05 09:25:57 +03:00
yusufipek 59180b8eff Say when the processor build landed instead of the Vulkan one
The install falls back to upstream's CPU archive whenever Dikte's own
release, the file in it, or its reviewed digest is not there, and until
the package is published by hand that is every download. It happened
without a word: the window said "Downloaded, version v1.9.3." either
way, and a graphics card sitting idle looks exactly like one being used.

The install record now carries which of the two builds landed, written
only where both were on offer, and the settings window says so on the
line that already reports the version.
2026-09-05 09:16:59 +03:00
yusufipek 956c3eaf3c Leave the tar extraction fallback where it was
Refusing the install on a Python without the extraction filters is a
change to how every archive on every platform is unpacked, and it has
nothing to do with shipping a Vulkan whisper-server. On those Pythons
the download stops working altogether, which is a worse answer than the
one that was there.

Worth doing on its own terms, in its own change, where the versions it
turns away can be argued about without a backend release riding on it.
2026-09-05 09:16:49 +03:00
nomoreshow 1f57455a34 Fix cross-platform test assumptions 2026-09-04 01:04:36 +03:00
nomoreshow e0eae4d8fe Ship a managed Vulkan whisper-server for Linux x64 2026-09-04 00:34:55 +03:00
yusufipek eda1398a2b Call the OpenRouter subtitle model the audio file model
Timestamps are a file transcription option, so the model box is named
after the file rather than the format it ends up in.
2026-09-02 12:43:07 +03:00
yusufipek 6e307bd8d0 Let OpenRouter subtitles use a chosen model instead of whisper-1
A timestamped run on OpenRouter always asked openai/whisper-1 for the
segments, whatever model was picked for plain transcription. Not every
model there returns segment times, so the one to use is now its own
setting, openrouter_subtitle_model, shown in the speech-to-text box only
when OpenRouter is the provider. Empty keeps the old whisper-1 fallback.

Target carries the choice as subtitle_model and timestamp_model() reads
it; the other providers are unchanged.
2026-09-02 12:41:43 +03:00
yusufipek 310ef8d7cf Dikte 1.1.0 2026-08-27 16:38:59 +03:00
Yusuf İpek 4f304e3d94 Merge pull request #62 from yusufipk/claude/remove-tooltip-texts-356de5
Drop three explanatory texts from the settings window
2026-08-27 16:36:39 +03:00
yusufipek c1554c092e Drop three explanatory texts from the settings window
The agent and meeting tabs opened with an intro paragraph, and the
cleanup provider box showed a long tooltip comparing the providers.
None of them earned the space: the intro texts were unclear and the
tooltip restated what picking a provider already shows. The orphaned
Turkish translations go with them.
2026-08-27 16:35:01 +03:00
Yusuf İpek 632804e922 Merge pull request #56 from sudoeren/feat/opencode-go-provider
Add OpenCode Go as a cleanup and agent provider
2026-08-27 16:31:47 +03:00
Yusuf İpek ec57d154fa Merge pull request #61 from yusufipk/bundled-libxkbcommon
Leave both halves of libxkbcommon to the system
2026-08-27 16:31:03 +03:00
yusufipek 22a73bd939 Merge master into the OpenCode Go branch
Master grew Google AI Studio and Antigravity as providers, a doctor that
names each provider's own key, and settings that fetch every hosted
model list as the window opens. OpenCode Go is folded into each: its
name joins SERVICES and the doctor's key table, its key row sits beside
Google's, its Fetch button follows the per-provider pattern Google's
uses, and _load_hosted_models fetches its catalog at open when a key is
on file, filling the cleanup and agent boxes alike.
2026-08-27 16:27:57 +03:00
Yusuf İpek 5fd32d105f Merge pull request #59 from Oztturk/gemini-and-agy-cleanup
Clean up on Google AI Studio, or on Antigravity
2026-08-27 16:14:34 +03:00
yusufipekandClaude Fable 5 d331ccb877 Fetch every model list at open, without being asked
Codex's list already arrived on its own, because reading a cache on disk
costs nothing. The other three lists were behind a Fetch button, which
meant the built-in ones aged in front of anyone who never pressed it. Now
OpenRouter's and Google's lists are fetched when the settings window
opens, and Antigravity's comes off `agy models`, which prints one
id-tab-name line per model and answers over the network in a couple of
seconds; all three run off the interface thread.

Nobody asked, so nothing is reported: a failure changes nothing on
screen, the built-in lists stay, and the Fetch buttons remain both the
retry and the place an error is worth explaining. A provider whose key
has not been given is not called at all, so opening Settings is not by
itself a request to two vendors.

Also drops an import the merge had left in twice.

Co-Authored-By: Claude Fable 5 <[email protected]>
2026-08-27 16:12:18 +03:00
yusufipek fa3d72f5a7 Fetch OpenCode Go's catalog as the settings window opens
Codex already refreshes its boxes from the source at open, so the
built-in list is never the whole truth for longer than a window takes to
build. OpenCode Go now gets the same courtesy: one request to /models on
a background thread, both of its boxes refilled with the picked model
kept, skipped entirely when there is no key to send.
2026-08-27 16:04:30 +03:00
yusufipek 1ff18e46ab Call OpenRouter the quickest again
Nothing was measured that put OpenCode Go beside it, so the tooltip
claims only what is known: OpenRouter is the quickest, and OpenCode Go
merely needs nothing installed either. The translation follows word for
word.
2026-08-27 16:04:09 +03:00
yusufipek 664a6f53a4 Offer only models the OpenCode Go catalog actually lists
ox-alpha-free is not in the /models answer, and the comment claimed the
chat endpoint hides Grok, Luna, MiniMax and Qwen when its own catalog
lists them. The built-in list is a starting set; the Fetch button asks
the endpoint for the full catalog of the day.
2026-08-27 15:55:32 +03:00
yusufipek 0ecd4d251e Translate the cleanup provider tooltip again
The tooltip gained OpenCode Go but its translation was added without the
llama.cpp sentence the tooltip actually carries, so a Turkish window
showed the English text. The full entry now matches the tooltip word for
word, and the two shorter variants nothing looks up any more are gone.
2026-08-27 15:54:54 +03:00
Yusuf İpek 7b7df1df62 Merge pull request #60 from AydoganCan60/feature/display-selection
Add display selection for recording indicator
2026-08-27 15:54:30 +03:00
yusufipek 65055eeb56 Give OpenCode Go a reachable Fetch model list button
The one button lived in the OpenRouter model row, which leaves the
screen whenever another provider is chosen, so the OpenCode fetch path
could never be clicked. The OpenCode row now carries its own button into
the same handler, the row as a whole is what hides, and the fetched list
lands in the box it belongs to without touching the meeting box.
2026-08-27 15:53:44 +03:00
yusufipek 23243c5171 Trim the README and polish the display selection
The settings-storage paragraph goes back to the one sentence it was. The screen list now shows native resolutions rather than the scaled ones, the Turkish tab is named Ekran, and the repositioning comment says what a named screen changes.
2026-08-27 15:52:02 +03:00
yusufipek ccc8ba596e Merge master into the OpenCode Go branch
Codex now asks itself for its model list, doctor grew ready flags and a
per-provider line, and the Groq key joined the masked ones; the OpenCode
additions are folded into each. The no-CLI cleanup test follows the
fake_run to fake_cli rename.
2026-08-27 15:51:49 +03:00
yusufipek 1df2d03056 Merge master into the Gemini and Antigravity branch
Both sides rewrote the doctor: master rebuilt it around ready flags so a
fully local setup stops reading as broken, while this branch taught it
that Google has a key of its own and that the local model has neither a
key nor a program. Kept master's structure and folded the branch in: the
two hosted providers answer for their keys, a CLI for its program, and
the JSON keeps both the provider-aware "key" and master's "ready".

SECRET_KEYS gained groq on master and gemini here; the resolution keeps
all four. _conclude was also changed by both: master stopped matching
stderr wording and asks _API_TROUBLE instead, this branch renamed its
last parameter to the provider's short name so a session is stored under
the name it is read back by; the rename now rides on master's body.
codex_models() and the settings loaders landed beside the new Gemini
ones, so both stay. The new cleanup tests called the CLI fake by its old
name, fake_run, which master had renamed to fake_cli.
2026-08-27 15:50:48 +03:00
Aydogan e6208418f9 Add display selection for recording indicator 2026-08-27 07:46:34 +03:00
oztturkandClaude Opus 5 7ebc3bf825 Ask Google for the lowest rung it has rather than for none
Google's compatibility layer has no word for off. Sending
reasoning_effort "none" is refused outright, so choosing Thinking → Off
made every cleanup fail and paste the raw transcript instead:

  HTTP 400: Request contains an invalid argument. (INVALID_ARGUMENT)

Measured against gemini-3.5-flash-lite, asking it to reply "ok":

  nothing sent               50.59s
  reasoning_effort "none"    400
  reasoning_effort "minimal" 15.40s
  reasoning_effort "low"     53.35s
  reasoning_effort "high"    63.83s

So "none" lands on "minimal", which is both accepted and the quickest of
them, and quickest is what cleanup wants. The thinking_config route the
documentation offers is an SDK wrapper and is not a field this endpoint
knows: sending it is "Unknown name \"google\"".

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-26 16:40:27 +03:00
oztturkandClaude Opus 5 1ffc3cff9d Let an error body that is an array still be read
Google answers some failures with a JSON array holding the object every
other provider sends on its own. _extract_error called .get() on it and
raised AttributeError, which is not the ApiError every caller is holding,
so a 503 from Google took the whole dictation down instead of pasting the
raw transcript with the failure shown beside it.

Found by dictating against a Google AI Studio outage:

  HTTP Error 503: Service Unavailable
  AttributeError: 'list' object has no attribute 'get'

It runs while an exception is being raised, so it now ends in a string
whatever arrives.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-26 16:29:41 +03:00
oztturkandClaude Opus 5 1812461612 Clean up on Google AI Studio, or on Antigravity
OpenRouter's free tier rate-limits and carries no free Gemini model, and
cleaning up through Claude Code costs a fixed few seconds because it opens
a whole CLI session to drop three "uh"s. Google's own free tier suits a
short, frequent request, and its OpenAI-compatible endpoint answers
/chat/completions, so cleanup there is one request and the same code path
OpenRouter already takes.

The one thing that is not shared is the thinking level. Google reads
OpenAI's flat reasoning_effort rather than OpenRouter's object, and "none"
is how thinking is turned off, so it is sent rather than skipped: a Flash
model left to think spends exactly the second this provider was chosen to
save. Its top two rungs land on "high", which is as far as Google goes.

Speech to text stays where it was. That endpoint has no
/audio/transcriptions behind it, audio only goes in as base64 inside a
chat message, and what comes back has none of the segment times a subtitle
file or a meeting transcript is built out of.

Antigravity joins as well, on cleanup and as an agent. It is a CLI like the
other two and costs the same session, so it is here for people who already
pay for it rather than as an answer to the speed. It takes neither an empty
tool list nor a read-only sandbox, and cleanup.py now says so plainly
instead of implying parity; what it gets is a project of its own, the home
directory, and its slash commands off.

Three things were already wrong and are fixed on the way past, because the
new providers walk the same paths: doctor raised KeyError on the local
model, whose executable is ""; the history recorded Claude's model whoever
answered; and every agent row read "asked Claude".

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-08-26 16:23:04 +03:00
sudoeren 52880bd81a Document the OpenCode Go provider 2026-08-25 21:35:04 +03:00
sudoeren e6b3fcf48f Reach OpenCode Go from the command line
test-key accepts opencode and reports the model count its /models
endpoint offers, ask --provider accepts it and maps --model to its own
setting, doctor reports its key for cleanup on it, and config list
masks the key like the other two.
2026-08-25 21:05:23 +03:00
sudoeren 3c336e7178 Offer OpenCode Go in the settings window
A key row under Keys, a model box in the cleanup tab and a box in the
agent tab, each with its own model list of the models OpenCode Go
serves over /chat/completions. Fetch model list reads whichever cleanup
provider is on screen, and the Turkish strings cover the new rows.
2026-08-25 21:05:21 +03:00
sudoeren 0a6c6128a2 Add OpenCode Go as a cleanup and agent provider
OpenCode Go serves open coding models from an OpenAI-compatible endpoint. It cannot transcribe, so it joins
the two LLM jobs: transcript cleanup and the plain-chat agent. The key
falls back to OPENCODE_API_KEY, and api.chat() stops hardcoding the
OpenRouter name so the same call serves both.
2026-08-25 21:05:18 +03:00
37 changed files with 3492 additions and 244 deletions
+222
View File
@@ -0,0 +1,222 @@
name: whisper.cpp Vulkan bundle
# Only what the bundle is built from. Compiling the Vulkan shaders takes
# tens of minutes, and a README typo is not worth one: what ties ggml.py to
# this release is a handful of assertions in tests/test_packaging.py, and
# those run on every pull request in milliseconds.
on:
pull_request:
paths:
- packaging/whisper-vulkan/**
- .github/workflows/whisper-vulkan.yml
workflow_dispatch:
inputs:
whisper_version:
description: Upstream whisper.cpp version (without v)
required: true
default: "1.9.3"
type: string
whisper_commit:
description: Peeled commit SHA for that reviewed upstream tag
required: true
default: "371b5a7561823ab2bb32142d2751e35e7534727b"
type: string
expected_sha256:
description: >-
Reviewed archive SHA-256. Leave empty for a version this file has
not reviewed: the digest of what was built is reported instead of
being checked, and publishing is refused.
required: false
default: ""
type: string
publish:
description: Publish a Dikte dependency release
required: true
default: false
type: boolean
permissions:
contents: read
concurrency:
group: whisper-vulkan-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
env:
WHISPER_VERSION: ${{ inputs.whisper_version || '1.9.3' }}
WHISPER_COMMIT: ${{ inputs.whisper_commit || '371b5a7561823ab2bb32142d2751e35e7534727b' }}
REVIEWED_WHISPER_VERSION: "1.9.3"
REVIEWED_WHISPER_SHA256: c25ca76504144da488eb74441390a7b9aa7ce547e5f2f391cbd831253c9b54d8
jobs:
build:
runs-on: ubuntu-22.04
timeout-minutes: 45
steps:
- name: Check out Dikte
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
persist-credentials: false
- name: Validate source coordinates
shell: bash
run: |
set -euo pipefail
[[ "$WHISPER_VERSION" =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]]
[[ "$WHISPER_COMMIT" =~ ^[0-9a-f]{40}$ ]]
- name: Check out pinned whisper.cpp source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
repository: ggml-org/whisper.cpp
ref: ${{ env.WHISPER_COMMIT }}
path: vendor/whisper.cpp
fetch-depth: 0
persist-credentials: false
- name: Verify source version and commit
shell: bash
run: |
set -euo pipefail
test "$(git -C vendor/whisper.cpp rev-parse HEAD)" = "$WHISPER_COMMIT"
git -C vendor/whisper.cpp fetch --depth=1 origin \
"refs/tags/v$WHISPER_VERSION:refs/tags/v$WHISPER_VERSION"
test "$(git -C vendor/whisper.cpp rev-list -n1 "v$WHISPER_VERSION")" = "$WHISPER_COMMIT"
echo "SOURCE_DATE_EPOCH=$(git -C vendor/whisper.cpp show -s --format=%ct HEAD)" >> "$GITHUB_ENV"
- name: Build pinned build environment
run: docker build --pull=false -f packaging/whisper-vulkan/Dockerfile.build -t dikte-whisper-builder packaging/whisper-vulkan
- name: Build deterministic archive
run: |
docker run --rm \
-e WHISPER_VERSION -e WHISPER_COMMIT -e SOURCE_DATE_EPOCH \
-v "$PWD/vendor/whisper.cpp:/src:ro" \
-v "$PWD/packaging/whisper-vulkan:/packaging:ro" \
-v "$PWD/work:/work" \
dikte-whisper-builder \
bash /packaging/build-package.sh
mkdir -p dist
cp work/out/whisper-bin-ubuntu-vulkan-x64.* dist/
- name: Verify reviewed archive digest
shell: bash
env:
EXPECTED_SHA256: ${{ inputs.expected_sha256 }}
PUBLISH: ${{ inputs.publish }}
run: |
set -euo pipefail
read -r actual _ < dist/whisper-bin-ubuntu-vulkan-x64.tar.gz.sha256
echo "built archive sha256: $actual"
expected="$EXPECTED_SHA256"
if [ -z "$expected" ] \
&& [ "$WHISPER_VERSION" = "$REVIEWED_WHISPER_VERSION" ]; then
expected="$REVIEWED_WHISPER_SHA256"
fi
if [ -z "$expected" ]; then
# The digest of a version nobody has reviewed yet cannot be known
# before it is built. Reporting it is the whole point of the run;
# a release out of it is not.
if [ "${PUBLISH:-false}" = true ]; then
echo "refusing to publish an archive whose digest has not been reviewed" >&2
exit 1
fi
echo "::notice::no reviewed digest for $WHISPER_VERSION." \
"Review the one above, then dispatch again with expected_sha256."
exit 0
fi
test "$actual" = "$expected"
- name: Validate archive and ELF contract
run: OUT_DIR=dist packaging/whisper-vulkan/validate-package.sh
- name: Schema-validate CycloneDX 1.6 SBOM
run: |
docker run --rm \
-v "$PWD/dist/whisper-bin-ubuntu-vulkan-x64.cdx.json:/sbom.json:ro" \
cyclonedx/cyclonedx-cli@sha256:252c2e26f468c25fea1e63ecde1bc3198ad6e9dbb57f5ed3236bddcb2281b3a7 \
validate --input-file /sbom.json --input-format json \
--input-version v1_6 --fail-on-errors
- name: CPU fallback smoke test (no Vulkan loader)
run: OUT_DIR=dist packaging/whisper-vulkan/smoke-runtime.sh cpu
- name: Vulkan loader present, no device smoke test
run: OUT_DIR=dist packaging/whisper-vulkan/smoke-runtime.sh noicd
- name: Vulkan plugin-load smoke test (Mesa llvmpipe)
run: OUT_DIR=dist packaging/whisper-vulkan/smoke-runtime.sh vulkan
- name: Upload reviewed outputs
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: whisper-bin-ubuntu-vulkan-x64
path: dist/*
if-no-files-found: error
retention-days: 14
publish:
if: >-
github.event_name == 'workflow_dispatch' && inputs.publish &&
github.ref == 'refs/heads/master'
needs: build
runs-on: ubuntu-22.04
environment: dependency-release
permissions:
contents: write
id-token: write
attestations: write
artifact-metadata: write
steps:
- name: Download the exact tested outputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4
with:
name: whisper-bin-ubuntu-vulkan-x64
path: dist
- name: Verify digest sidecar
run: (cd dist && sha256sum --check whisper-bin-ubuntu-vulkan-x64.tar.gz.sha256)
- name: Attest build provenance
uses: actions/attest@1e69f48acb82d1966a394da916b4c1698aa569d6 # v4
with:
subject-path: dist/whisper-bin-ubuntu-vulkan-x64.tar.gz
- name: Attest SBOM to archive
uses: actions/attest@1e69f48acb82d1966a394da916b4c1698aa569d6 # v4
with:
subject-path: dist/whisper-bin-ubuntu-vulkan-x64.tar.gz
sbom-path: dist/whisper-bin-ubuntu-vulkan-x64.cdx.json
- name: Publish dependency release
env:
GH_TOKEN: ${{ github.token }}
RELEASE_TAG: whisper.cpp-v${{ inputs.whisper_version }}
RELEASE_TITLE: whisper.cpp v${{ inputs.whisper_version }} Vulkan bundle
RELEASE_NOTES: >-
Pinned source: ggml-org/whisper.cpp@${{ inputs.whisper_commit }}.
Verify with: gh attestation verify
whisper-bin-ubuntu-vulkan-x64.tar.gz
--repo ${{ github.repository }}
run: |
set -euo pipefail
if gh release view "$RELEASE_TAG" --repo "$GITHUB_REPOSITORY" >/dev/null 2>&1; then
echo "refusing to replace existing release $RELEASE_TAG" >&2
exit 1
fi
if gh api "repos/$GITHUB_REPOSITORY/git/ref/tags/$RELEASE_TAG" >/dev/null 2>&1; then
echo "refusing to replace existing tag $RELEASE_TAG" >&2
exit 1
fi
gh api --method POST "repos/$GITHUB_REPOSITORY/git/refs" \
-f ref="refs/tags/$RELEASE_TAG" \
-f sha="$GITHUB_SHA" >/dev/null
test "$(gh api "repos/$GITHUB_REPOSITORY/git/ref/tags/$RELEASE_TAG" \
--jq .object.sha)" = "$GITHUB_SHA"
gh release create "$RELEASE_TAG" dist/* \
--repo "$GITHUB_REPOSITORY" \
--verify-tag \
--prerelease \
--latest=false \
--title "$RELEASE_TITLE" \
--notes "$RELEASE_NOTES"
+19 -11
View File
@@ -115,9 +115,13 @@ or runs it on the spot.
Speech to text and cleanup each pick a provider in the settings window, and both
run here by default, on models of your own. The cloud is the other option:
speech to text on **OpenAI**, **Groq** or **OpenRouter** (`gpt-4o-transcribe`),
cleanup on OpenRouter (`google/gemini-3.5-flash-lite`) or, when either is
installed, on Claude Code or Codex. The keys fall back to `OPENAI_API_KEY`,
`GROQ_API_KEY` and `OPENROUTER_API_KEY`, and are stored in
cleanup on OpenRouter (`google/gemini-3.5-flash-lite`), on **Google AI Studio**
(`gemini-3.5-flash-lite`), on **OpenCode Go** (`deepseek-v4-flash`) or, when one
of them is installed, on Claude Code, Codex or Antigravity. The first three are
a single HTTP request; the three CLIs each open a whole session to do it, which
is where their few extra seconds go. The keys fall back to `OPENAI_API_KEY`,
`GROQ_API_KEY`, `OPENROUTER_API_KEY`, `GEMINI_API_KEY` and `OPENCODE_API_KEY`,
and are stored in
`~/.config/dikte/config.json`, mode 600, or in
`~/Library/Application Support/Dikte` on a Mac. Cleanup can be switched off, in
which case the raw transcript is pasted, and a thinking model's effort can be
@@ -159,7 +163,9 @@ running.
the program and the model, verifies the sha256 and refuses a download published
without one, then keeps a server alive while you dictate. The graphics card is
reached through CUDA, ROCm or Vulkan where the build allows. No key, no
account, nothing leaving the machine.
account, nothing leaving the machine. On x86_64 Linux the same button fetches
a Vulkan build of whisper-server that Dikte publishes itself, because
upstream's Linux archive is processor-only.
- **Silence never reaches the API.** Handed near-silence, a transcription model
invents a sentence instead of returning nothing ("Thanks for watching", or in
Turkish "Altyazı M.K."). A recording is dropped when nothing rose 10 dB above
@@ -189,11 +195,13 @@ running.
what comes of it: the answer, or a sentence saying what was done. It is the
session you would have opened yourself, so your skills and connected services
are there, which is what makes "put that in my calendar on Thursday at three"
a thing you can say to a window that is not Claude. Codex (`codex exec`) runs
the same way, and OpenRouter is there as a plain question-and-answer fallback
for a machine with neither CLI on it. Provider, model, permissions and working
directory are under Settings → Agent, and commands close together stay in one
conversation.
a thing you can say to a window that is not Claude. Codex (`codex exec`) and
Antigravity (`agy -p`) run the same way, though Antigravity takes neither a
permission mode nor a sandbox from Dikte: what it may do without asking is
whatever its own allow-rules say. OpenRouter or OpenCode Go is there as a
plain question-and-answer fallback for a machine with no CLI on it. Provider,
model, permissions and working directory are under Settings → Agent, and
commands close together stay in one conversation.
- **Meetings** are recorded from the microphone and the speaker output at the
same time, which settles who said what by the channel a voice arrived on
instead of guessing at it. The two sides are transcribed separately and
@@ -240,9 +248,9 @@ cli.py the command line: every verb, and what it answers with
ipc.py one request and one reply over the local socket
audio.py PCM capture: pw-record for dictation, ffmpeg for a meeting
meeting.py channel split, speaker labelling, cleanup, minutes
assistant.py running a dictation through Claude Code, Codex or OpenRouter
assistant.py handing a dictation to Claude Code, Codex, agy or a chat model
api.py transcription and cleanup requests (stdlib only)
cleanup.py who rewrites the transcript: OpenRouter, here, Claude or Codex
cleanup.py who rewrites the transcript: a hosted model, one here, a CLI
ggml.py whisper.cpp and llama.cpp here: fetch, verify, keep serving
hub.py what GitHub and Hugging Face have on offer today
update.py whether a newer release is out, and the page it is on
+17 -9
View File
@@ -112,10 +112,14 @@ Sesi yazıya çevirme ve temizleme, ayarlar penceresinde ayrı ayrı sağlayıc
seçer; ikisi de varsayılan olarak burada, kendi modellerinle çalışır. Bulutu
seçersen sesi yazıya çevirme **OpenAI**, **Groq** ya da **OpenRouter**'da
(varsayılan `gpt-4o-transcribe`), temizleme OpenRouter'da
(`google/gemini-3.5-flash-lite`) ya da kuruluysa Claude Code veya Codex'te
çalışır. Anahtarları boş bırakırsan `OPENAI_API_KEY`, `GROQ_API_KEY` ve
`OPENROUTER_API_KEY` kullanılır; anahtarlar `~/.config/dikte/config.json`
içinde, izinler 600, Mac'te ise `~/Library/Application Support/Dikte` altında.
(`google/gemini-3.5-flash-lite`), **Google AI Studio**'da
(`gemini-3.5-flash-lite`), **OpenCode Go**'da (`deepseek-v4-flash`) ya da
kuruluysa Claude Code, Codex veya Antigravity'de çalışır. İlk üçü tek bir HTTP
isteği; üç CLI ise bunun için birer oturum açar, fazladan giden birkaç saniye de
oradan gelir. Anahtarları boş bırakırsan `OPENAI_API_KEY`, `GROQ_API_KEY`,
`OPENROUTER_API_KEY`, `GEMINI_API_KEY` ve `OPENCODE_API_KEY` kullanılır;
anahtarlar `~/.config/dikte/config.json` içinde, izinler 600, Mac'te ise
`~/Library/Application Support/Dikte` altında.
Temizlemeyi tamamen kapatabilirsin, o zaman ham transkript yapıştırılır; modelin
yanındaki kutudan düşünme seviyesini de seçebilirsin.
@@ -156,6 +160,8 @@ olmasını ister.
checksum'suz yayınlanmış bir indirmeyi reddeder, sen dikte ettikçe sunucuyu
ayakta tutar. Derleme destekliyorsa ekran kartına CUDA, ROCm ya da Vulkan
üzerinden ulaşılır. Anahtar yok, hesap yok, makineden çıkan bir şey yok.
x86_64 Linux'ta aynı düğme, whisper-server'ın Dikte'nin kendi yayınladığı
Vulkan derlemesini indirir; upstream'in Linux arşivi yalnızca işlemci için.
- **Sessizlik API'ye gitmez.** Sessize yakın bir ses verildiğinde model boş dize
döndürmez, bir cümle uydurur ("Altyazı M.K.", "Thanks for watching"). *O
kaydın kendi* gürültü tabanının 10 dB üstüne en az 0,3 saniye çıkan bir şey
@@ -184,9 +190,11 @@ olmasını ister.
yapıştırır: cevabı ya da ne yapıldığını söyleyen bir cümle. Kendi açacağın
oturumun aynısıdır, yani skill'lerin ve bağlı servislerin oradadır; "bunu
perşembe üçe takvime koy" cümlesini Claude olmayan bir pencerede söyleyebilir
olmanı sağlayan da budur. Codex (`codex exec`) da aynı şekilde çalışır;
OpenRouter ise ikisi de kurulu olmayan bir makinede düz soru cevap için
duruyor. Sağlayıcı, model, izinler ve çalışma dizini Ayarlar → Ajan
olmanı sağlayan da budur. Codex (`codex exec`) ile Antigravity (`agy -p`) da
aynı şekilde çalışır; ama Antigravity'ye Dikte bir izin kipi ya da sandbox
veremiyor, sormadan ne yapabileceğini kendi allow-rule'ları belirliyor.
OpenRouter ya da OpenCode Go ise hiçbiri kurulu olmayan bir makinede düz soru
cevap için duruyor. Sağlayıcı, model, izinler ve çalışma dizini Ayarlar → Ajan
sekmesinde; arka arkaya verilen komutlar tek bir konuşmada kalır.
- **Toplantılar** mikrofonla hoparlör çıkışından aynı anda kaydedilir; kimin ne
dediği tahmin edilmez, sesin hangi kanaldan geldiğiyle belli olur. İki taraf
@@ -233,9 +241,9 @@ cli.py komut satırı: bütün fiiller ve verdikleri cevap
ipc.py yerel sokette bir istek, bir cevap
audio.py PCM kaydı: diktede pw-record, toplantıda ffmpeg
meeting.py kanal ayırma, konuşmacı etiketi, temizleme, tutanak
assistant.py dikteyi Claude Code, Codex ya da OpenRouter'dan geçirme
assistant.py dikteyi Claude Code, Codex, agy ya da sohbet modelinden geçirme
api.py transkript ve temizleme istekleri (yalnız stdlib)
cleanup.py transkripti kim temizler: OpenRouter, burası, Claude ya da Codex
cleanup.py transkripti kim temizler: bulutta bir model, burası, bir CLI
ggml.py whisper.cpp ve llama.cpp'yi indirip burada çalıştırma
hub.py GitHub ve Hugging Face'te bugün ne olduğu
update.py yeni sürüm çıkmış mı, çıkmışsa hangi sayfada
+1 -1
View File
@@ -10,4 +10,4 @@ business loading Qt to answer one question.
# both the .dmg's Info.plist and the AppImage's file name are built from it. A
# build off master rather than off a tag appends the commit to it, so that a
# bug report from someone running "latest" names a commit.
__version__ = "1.0.2"
__version__ = "1.1.0"
+107 -25
View File
@@ -1,10 +1,17 @@
"""OpenAI, Groq, OpenRouter and this machine, stdlib only.
"""OpenAI, Groq, OpenRouter, Google AI Studio and this machine, stdlib only.
Transcription runs on any of the four: Groq and OpenRouter both mirror OpenAI's
/audio/transcriptions endpoint field for field, and ggml.py starts whisper.cpp
on that same path, so one multipart request serves all of them and only the key,
the base URL and the model id change. llama.cpp answers /chat/completions the way
OpenRouter does, so cleanup here is the same request too.
Transcription runs on the first three and on this machine: Groq and OpenRouter
both mirror OpenAI's /audio/transcriptions endpoint field for field, and ggml.py
starts whisper.cpp on that same path, so one multipart request serves all of
them and only the key, the base URL and the model id change. llama.cpp answers
/chat/completions the way OpenRouter does, so cleanup here is the same request
too.
Google AI Studio is here for cleanup and nothing else. Its OpenAI-compatible
endpoint answers /chat/completions and /models, but there is no
/audio/transcriptions behind it: audio only goes in as base64 inside a chat
message, and what comes back has none of the segment times a subtitle file or a
meeting transcript is built out of.
What is on this machine has no key, and its base URL is not known until a server
is up, which is the one thing this module has to fill in for it.
@@ -31,6 +38,7 @@ USER_AGENT = f"dikte/1.0 (+{APP_URL})"
OPENAI_URL = "https://api.openai.com/v1"
GROQ_URL = "https://api.groq.com/openai/v1"
OPENROUTER_URL = "https://openrouter.ai/api/v1"
GEMINI_URL = "https://generativelanguage.googleapis.com/v1beta/openai"
# The floor for a local request. The timeouts elsewhere are sized for a hosted
# API, where a slow answer is a bill running; here the only thing being spent is
@@ -40,22 +48,33 @@ LOCAL_TIMEOUT = 3600
# Where a transcription request goes; built by config.Config.transcribe_target().
# `service` is the name the user sees in an error, `provider` the one the code
# branches on.
Target = collections.namedtuple("Target", "provider service api_key base_url model")
# branches on. `file_model` is what a timestamped run asks for instead of
# `model`, where the two differ; empty means the provider's own whisper.
Target = collections.namedtuple(
"Target", "provider service api_key base_url model file_model",
defaults=[""])
# What answers with segment times on OpenRouter when nothing else was chosen.
OPENROUTER_FILE_MODEL = "openai/whisper-1"
def timestamp_model(provider, selected=""):
def timestamp_model(provider, selected="", file_model=""):
"""Which model answers with segment times.
OpenAI keeps them to whisper-1 and OpenRouter namespaces that id. Everything
Groq transcribes with is a whisper, so the model already chosen does it and
the fallback is only for a provider left on its default. So is everything the
local server runs, whatever the file is called, and there asking for another
model would name one it has never heard of.
OpenAI keeps them to whisper-1. Everything Groq transcribes with is a
whisper, so the model already chosen does it and the fallback is only for a
provider left on its default. So is everything the local server runs,
whatever the file is called, and there asking for another model would name
one it has never heard of. OpenRouter fronts several models that do times
and several that do not, and a request to the wrong one gets a transcript
with no segments in it, so the one to use is a setting of its own
(`file_model`) and whisper-1 is only where that setting is left empty.
"""
if provider in ("groq", "local"):
return selected or "whisper-large-v3-turbo"
return "openai/whisper-1" if provider == "openrouter" else "whisper-1"
if provider == "openrouter":
return file_model or OPENROUTER_FILE_MODEL
return "whisper-1"
# What a gateway in front of the model answers of its own accord: the request
@@ -254,10 +273,23 @@ def _request(url, data, headers, timeout=120, aborter=None):
def _extract_error(body):
"""The line worth showing out of a failed request's body.
Whatever comes back, this has to end in a string: it is called while an
ApiError is being raised, and an exception thrown here would escape the
`except ApiError` every caller is holding and lose the dictation the raw
transcript would otherwise have been pasted from.
"""
try:
payload = json.loads(body)
except json.JSONDecodeError:
return body[:300]
if isinstance(payload, list):
# Google answers some failures with an array holding the object the
# other providers send on its own.
payload = next((item for item in payload if isinstance(item, dict)), None)
if not isinstance(payload, dict):
return body[:300]
err = payload.get("error")
if isinstance(err, dict):
return err.get("message") or json.dumps(err)[:300]
@@ -423,7 +455,8 @@ def transcribe_segments(target, audio_path, language="", prompt="", timeout=300,
aborter=None):
"""[(start_seconds, end_seconds, text)] using whisper-1's verbose response."""
data = _transcribe_request(
target._replace(model=timestamp_model(target.provider, target.model)),
target._replace(model=timestamp_model(target.provider, target.model,
target.file_model)),
audio_path, language, prompt, "verbose_json",
granularity="segment", timeout=timeout, aborter=aborter,
)
@@ -448,14 +481,24 @@ def transcribe_segments(target, audio_path, language="", prompt="", timeout=300,
return out
# The settings window offers OpenRouter's ladder, and Google has neither end of
# it: "none" is refused outright with a 400, and there is nothing above "high".
# Both ends land on the nearest rung that does exist, which costs the cleanup
# rather than the dictation when it is wrong. "minimal" is where "off" goes, and
# it is the quickest of them by a wide margin, which is what cleanup wants
# anyway.
GEMINI_EFFORT = {"none": "minimal", "xhigh": "high", "max": "high"}
def _thinking(payload, provider, reasoning):
"""Ask for as much thinking as this provider understands, or for none.
An empty level means "whatever the model does on its own", so nothing is
sent. The two mean opposite things by that, which is why the setting is kept
per provider: OpenRouter's cleanup models answer straight away, while a local
model that was trained to think will think, and cleanup is punctuation rather
than a job worth thinking about.
sent. The three mean opposite things by that, which is why the setting is
kept per provider: OpenRouter's cleanup models answer straight away, while a
local model that was trained to think will think, and a Gemini Flash left to
itself thinks too. Cleanup is punctuation rather than a job worth thinking
about.
"""
if not reasoning:
return
@@ -463,6 +506,13 @@ def _thinking(payload, provider, reasoning):
# What llama.cpp passes to the chat template. The models that think read
# it; the ones that do not ignore it.
payload["chat_template_kwargs"] = {"enable_thinking": reasoning != "none"}
elif provider == "gemini":
# Google's compatibility layer takes OpenAI's flat field rather than
# OpenRouter's object, and it has no word for off, so "none" is asked
# for as the lowest rung it has rather than skipped: a Flash model left
# to decide for itself thinks, and thinking about a comma is the second
# this provider was chosen to save.
payload["reasoning_effort"] = GEMINI_EFFORT.get(reasoning, reasoning)
elif reasoning != "none":
# The thinking itself is never shown, so ask for it to be left out.
payload["reasoning"] = {"effort": reasoning, "exclude": True}
@@ -524,15 +574,16 @@ def cleanup(text, api_key, model, system_prompt, reasoning="",
def chat(messages, api_key, model, system_prompt, reasoning="",
base_url=OPENROUTER_URL, timeout=180):
base_url=OPENROUTER_URL, timeout=180, provider="openrouter",
service="OpenRouter"):
"""A conversation, rather than one transcript rewritten.
The messages are the whole history and come back unchanged; the caller keeps
them, because there is no session on OpenRouter's side to resume.
them, because there is no session on the provider's side to resume.
"""
if not api_key:
raise ApiError(t("{service} API key is empty. Add it in Settings.",
service="OpenRouter"))
service=service))
payload = {
"model": model,
"messages": [{"role": "system", "content": system_prompt}] + list(messages),
@@ -543,11 +594,11 @@ def chat(messages, api_key, model, system_prompt, reasoning="",
data = _request(
f"{base_url.rstrip('/')}/chat/completions",
json.dumps(payload).encode("utf-8"),
_headers("openrouter", api_key, "application/json"),
_headers(provider, api_key, "application/json"),
timeout=timeout,
)
except ApiError as exc:
raise explain(exc, "OpenRouter") from None
raise explain(exc, service) from None
choices = data.get("choices") or []
if not choices:
raise ApiError(_extract_error(json.dumps(data)))
@@ -611,6 +662,37 @@ def openrouter_models(api_key="", transcription=False):
return sorted(m["id"] for m in models if m.get("id"))
# What a `gemini` id can be besides a model that answers a chat request: an
# embedding, a picture, or a voice. None of them is any use to cleanup.
NOT_CHAT = ("embedding", "-image", "-tts", "-audio")
def gemini_models(api_key, base_url=GEMINI_URL):
"""The Gemini models Google AI Studio will answer a chat request with.
Google serves its embedding, image and speech models out of the same list
and names them all `gemini` too, so the prefix alone is not the question;
none of those can clean up a sentence. The listing has also been known to
hand the ids back in their long form, `models/gemini-3.5-flash`, while a
request wants the short one; taking the prefix off costs nothing and is
right whichever form arrives.
"""
service = "Google AI Studio"
if not api_key:
raise ApiError(t("{service} API key is empty. Add it in Settings.",
service=service))
try:
data = _get_json(
f"{base_url.rstrip('/')}/models",
{"Authorization": f"Bearer {api_key}", "User-Agent": USER_AGENT},
)
except ApiError as exc:
raise explain(exc, service) from None
ids = [(m.get("id") or "").removeprefix("models/") for m in data.get("data", [])]
return sorted(i for i in ids
if i.startswith("gemini") and not any(w in i for w in NOT_CHAT))
def openai_models(api_key, base_url=OPENAI_URL, service="OpenAI"):
"""The audio models of anything that speaks OpenAI's /models, Groq included.
+10 -4
View File
@@ -163,11 +163,15 @@ class Dikte:
self.front_before = None
self._front_watch = None
self.overlay = Overlay(self.conf["overlay_corner"])
self.overlay = Overlay(self.conf["overlay_corner"],
screen_name=self.conf["overlay_screen"],
follow_pointer=self.conf["overlay_follows_pointer"])
# The agent's indicator sits on top of the dictation one when both are
# up, and drops into the corner when it is alone there.
self.ask_overlay = Overlay(self.conf["overlay_corner"], below=self.overlay,
dismissable=True)
dismissable=True,
screen_name=self.conf["overlay_screen"],
follow_pointer=self.conf["overlay_follows_pointer"])
self.recorder = audio.Recorder()
self.pipeline = Pipeline(self.conf)
self.ask_pipeline = Pipeline(self.conf)
@@ -1295,8 +1299,10 @@ class Dikte:
threading.Thread(target=warm, daemon=True).start()
def _apply_settings(self):
self.overlay.corner = self.conf["overlay_corner"]
self.ask_overlay.corner = self.conf["overlay_corner"]
for indicator in (self.overlay, self.ask_overlay):
indicator.corner = self.conf["overlay_corner"]
indicator.screen_name = self.conf["overlay_screen"]
indicator.follow_pointer = self.conf["overlay_follows_pointer"]
self._apply_local()
self._build_tray()
self._refresh_tray()
+187 -31
View File
@@ -1,22 +1,29 @@
"""Handing a dictation to an agent as a command, and pasting back its answer.
Three of them, because not everyone has the same one installed:
Five of them, because not everyone has the same one installed:
Claude Code `claude -p`, the session you would have opened yourself
Codex `codex exec`, the same idea from the other shop
Antigravity `agy -p`, Google's, with a browser of its own attached
OpenRouter a plain chat request, over the key that is already configured
OpenCode Go a plain chat request, over a subscription to open coding models
The first two are the whole machine: they run commands, read files, and reach
The first three are the whole machine: they run commands, read files, and reach
whatever skills and services you have connected, which is what makes "put that
in my calendar on Thursday" a thing you can say. OpenRouter cannot touch any of
that, and is there so that a question still gets an answer on a machine with
neither CLI installed.
in my calendar on Thursday" a thing you can say. The two chat requests cannot
touch any of that, and are there so that a question still gets an answer on a
machine with no CLI installed at all.
What each of the three is allowed to do without asking is settled where that
program keeps its own permissions, not here. Dikte hands Claude Code the mode
chosen in Settings because it has a flag for one; Codex gets a sandbox for the
same reason; Antigravity has neither, and reads its own allow-rules instead.
Whichever it is, the reply is pasted exactly where the transcript would have
been, and the conversation carries across dictations so that "and move that to
Friday" knows what "that" is.
The two CLIs are read as they stream rather than waited out. A command that
The three CLIs are read as they stream rather than waited out. A command that
reaches for the calendar or the web takes long enough that a still indicator is
indistinguishable from a hang, so every tool they pick up is named in the corner
while they work.
@@ -38,11 +45,16 @@ from . import paths
from .i18n import t
SESSION_FILE = cfg.DATA_DIR / "assistant.json"
PROVIDERS = ("claude", "codex", "openrouter")
PROVIDERS = ("claude", "codex", "agy", "openrouter", "opencode")
# How many messages of an OpenRouter conversation are carried forward. The two
# CLIs keep their own history and need no such number; here every turn is resent
# in full, so the window has to end somewhere.
# What each one is called where a person reads it: the tray, the corner of
# the screen, and the line an error is written in.
SERVICES = {"claude": "Claude", "codex": "Codex", "agy": "Antigravity",
"openrouter": "OpenRouter", "opencode": "OpenCode Go"}
# How many messages of a chat provider's conversation are carried forward. The
# two CLIs keep their own history and need no such number; here every turn is
# resent in full, so the window has to end somewhere.
MAX_HISTORY = 24
# What to say in the indicator for a tool, keyed by the name the CLI uses.
@@ -70,6 +82,29 @@ CODEX_ITEMS = {
"patch_apply": "Editing a file…",
"todo_list": "Planning…",
}
# Antigravity carries a browser around with it, so the handful of names below
# stand in for the couple of dozen browser_* tools it can pick up; being told
# which mouse button moved is not what the corner of the screen is for.
AGY_TOOLS = {
"run_command": "Running a command…",
"command_status": "Running a command…",
"send_command_input": "Running a command…",
"view_file": "Reading…",
"read_url_content": "Reading a web page…",
"list_dir": "Looking through files…",
"find_by_name": "Looking through files…",
"grep_search": "Searching the files…",
"search_web": "Searching the web…",
"replace_file_content": "Editing a file…",
"multi_replace_file_content": "Editing a file…",
"sed_file": "Editing a file…",
"notebook_edit": "Editing a file…",
"write_to_file": "Writing a file…",
"generate_image": "Drawing…",
"manage_task": "Planning…",
"invoke_subagent": "Handing it to a subagent…",
"browser_subagent": "Handing it to a subagent…",
}
# How hard to think, in each provider's own vocabulary. The setting is one
@@ -85,6 +120,12 @@ CLAUDE_EFFORT = {"none": "low", "minimal": "low", "low": "low",
CODEX_EFFORT = {"none": "low", "minimal": "low", "low": "low",
"medium": "medium", "high": "high", "xhigh": "high",
"max": "high"}
# agy has three rungs and no word for off, so the bottom of the ladder lands on
# "low" and the top two on "high". Shared with cleanup, which runs the same
# program for the smaller job.
AGY_EFFORT = {"none": "low", "minimal": "low", "low": "low",
"medium": "medium", "high": "high", "xhigh": "high",
"max": "high"}
class AssistantError(Exception):
@@ -102,12 +143,30 @@ def provider(conf):
def executable(name):
"""The CLI a provider runs, or "" when it needs none."""
return {"claude": "claude", "codex": "codex"}.get(name, "")
return {"claude": "claude", "codex": "codex", "agy": "agy"}.get(name, "")
def model(conf):
"""Which model answered, for the history to record.
Each provider keeps its own setting, and the one a CLI is left on has no id
to report, only a name — the same arrangement cleanup.model() makes.
"""
name = provider(conf)
if name == "codex":
return conf["assistant_codex_model"].strip() or "codex"
if name == "agy":
return conf["assistant_agy_model"].strip() or "agy"
if name == "openrouter":
return conf["assistant_openrouter_model"]
if name == "opencode":
return conf["assistant_opencode_model"]
return conf["assistant_model"]
def display_name(conf):
"""What to call the thing being asked, in the tray and in the corner."""
return {"claude": "Claude", "codex": "Codex"}.get(provider(conf), "OpenRouter")
return SERVICES.get(provider(conf), "OpenRouter")
# --- the conversation -----------------------------------------------------
@@ -197,8 +256,8 @@ def ask(prompt, conf, on_stage=None, should_stop=None):
one, and only the denial explains why it did not do what it was asked to.
"""
name = provider(conf)
if name == "openrouter":
return _ask_openrouter(prompt, conf, on_stage)
if name in ("openrouter", "opencode"):
return _ask_chat(name, SERVICES[name], prompt, conf, on_stage)
binary = executable(name)
if not shutil.which(binary):
@@ -207,7 +266,7 @@ def ask(prompt, conf, on_stage=None, should_stop=None):
"Settings → Agent.", binary=binary,
))
run = _ask_claude if name == "claude" else _ask_codex
run = {"claude": _ask_claude, "codex": _ask_codex, "agy": _ask_agy}[name]
session = read_session(name, conf["assistant_session_minutes"] * 60)
try:
return run(prompt, conf, session, on_stage, should_stop)
@@ -259,7 +318,7 @@ def _ask_claude(prompt, conf, session, on_stage, should_stop):
found["warning"] = _denial_warning(event)
code, stderr = _stream(cmd, conf, on_event, should_stop)
return _conclude(found, code, stderr, session, "Claude")
return _conclude(found, code, stderr, session, "claude")
def _claude_label(block):
@@ -327,7 +386,7 @@ def _ask_codex(prompt, conf, session, on_stage, should_stop):
else str(error)) or t("Codex ended with an error.")
code, stderr = _stream(cmd, conf, on_event, should_stop)
return _conclude(found, code, stderr, session, "Codex")
return _conclude(found, code, stderr, session, "codex")
def _codex_label(item):
@@ -366,29 +425,125 @@ def codex_models():
return [row["slug"] for row in rows]
# --- OpenRouter -----------------------------------------------------------
# --- Antigravity ----------------------------------------------------------
def _ask_openrouter(prompt, conf, on_stage):
"""No tools, no files, no calendar: a question and an answer.
def _ask_agy(prompt, conf, session, on_stage, should_stop):
# Antigravity takes no system prompt of its own either, so the instruction
# rides in front of the command, kept apart from it so the two are not read
# as one.
body = f"{conf.assistant_prompt()}\n\n---\n\n{prompt}"
cmd = [
"agy", "-p", body,
"--output-format", "stream-json",
# agy stops after five minutes unless it is told otherwise, which is
# shorter than the timeout this setting offers.
"--print-timeout", f"{conf['assistant_timeout']}s",
]
# One or the other, always: left with neither, agy picks up whichever
# project it was last in and works in that project's directory rather than
# the one _stream is about to start it in.
cmd += ["--conversation", session] if session else ["--new-project"]
if conf["assistant_agy_model"].strip():
cmd += ["--model", conf["assistant_agy_model"].strip()]
effort = AGY_EFFORT.get(conf["assistant_reasoning"], "")
if effort:
# Most of agy's own model ids carry the effort in their suffix already;
# this is for the ones that do not.
cmd += ["--effort", effort]
It is the fallback for a machine with neither CLI on it, so it says what it
knows and nothing else. The conversation is ours to keep here, since there
is no session on the other end to resume.
found = {"answer": "", "warning": "", "session": "", "failure": ""}
def on_event(event):
kind = event.get("event")
if kind == "init":
found["session"] = event.get("conversation_id") or found["session"]
elif kind == "step_update":
step = event.get("step_update") or {}
# A tool is reported twice, once when it starts and once when it is
# done; the corner wants the first of those.
if (on_stage and step.get("step_type") == "tool"
and step.get("state") == "ACTIVE"):
on_stage(_agy_label(step))
elif kind == "result":
result = event.get("result") or {}
found["session"] = result.get("conversation_id") or found["session"]
answer = (result.get("response") or "").strip()
if result.get("status") == "SUCCESS":
found["answer"] = answer
else:
found["failure"] = answer or t("{service} ended with an error.",
service="Antigravity")
code, stderr = _stream(cmd, conf, on_event, should_stop)
return _conclude(found, code, stderr, session, "agy")
def _agy_label(step):
name = step.get("tool_name", "")
if name in AGY_TOOLS:
return t(AGY_TOOLS[name])
if name.startswith("browser_") or name.startswith("capture_browser"):
return t("Working in the browser…")
if name == "call_mcp_tool":
server = (step.get("tool_info") or {}).get("parameters") or {}
return t("Using {name}", name=server.get("server") or "a tool")
return t("Using {name}", name=name or "a tool")
def agy_models():
"""The models Antigravity itself would offer right now, in its own order.
`agy models` prints one `id<TAB>display name` line per model, so the list
is as current as the account behind the CLI. Unlike Codex it asks Google
rather than a cache on disk, a couple of seconds the caller spends off the
interface thread. A machine without agy, or a call that fails, answers
with nothing and the caller keeps its built-in list.
"""
if not shutil.which("agy"):
return []
try:
proc = subprocess.run(["agy", "models"],
capture_output=True, text=True, timeout=30)
except (OSError, subprocess.SubprocessError):
return []
if proc.returncode != 0:
return []
ids = []
for line in (proc.stdout or "").splitlines():
model_id, tab, _ = line.partition("\t")
if tab and model_id.strip():
ids.append(model_id.strip())
return ids
# --- OpenRouter and OpenCode Go -------------------------------------------
def _ask_chat(name, service, prompt, conf, on_stage):
"""A plain question and answer, over a chat provider's key.
No tools, no files, no calendar. It is the fallback for a machine with
neither CLI on it, so it says what it knows and nothing else. The
conversation is ours to keep here, since there is no session on the other
end to resume.
"""
if on_stage:
on_stage(t("Thinking…"))
history = read_messages("openrouter", conf["assistant_session_minutes"] * 60)
history = read_messages(name, conf["assistant_session_minutes"] * 60)
messages = history + [{"role": "user", "content": prompt}]
model = (conf["assistant_openrouter_model"] if name == "openrouter"
else conf["assistant_opencode_model"])
base_url = (conf["openrouter_base_url"] if name == "openrouter"
else conf["opencode_base_url"])
key = conf.openrouter_key() if name == "openrouter" else conf.opencode_key()
try:
answer = api.chat(
messages, conf.openrouter_key(), conf["assistant_openrouter_model"],
conf.assistant_prompt(), reasoning=conf["assistant_reasoning"],
base_url=conf["openrouter_base_url"],
timeout=conf["assistant_timeout"],
messages, key, model, conf.assistant_prompt(),
reasoning=conf["assistant_reasoning"], base_url=base_url,
timeout=conf["assistant_timeout"], provider=name, service=service,
)
except api.ApiError as exc:
raise AssistantError(str(exc)) from exc
write_session("openrouter",
write_session(name,
messages=messages + [{"role": "assistant", "content": answer}])
return answer, ""
@@ -469,8 +624,9 @@ _API_TROUBLE = re.compile(
r"\b(401|403|429|5\d\d)\b")
def _conclude(found, code, stderr, session, service):
def _conclude(found, code, stderr, session, name):
"""Turn what the stream said into an answer, or into the reason there is none."""
service = SERVICES.get(name, name)
if code != 0 and not found["answer"]:
# A resumed run that died with nothing to show is treated as the
# session being gone, whatever the wording: this code used to look for
@@ -489,7 +645,7 @@ def _conclude(found, code, stderr, session, service):
if not found["answer"]:
raise AssistantError(t("{service} answered with nothing.", service=service))
if found["session"]:
write_session("claude" if service == "Claude" else "codex", found["session"])
write_session(name, found["session"])
return found["answer"], found["warning"]
+67 -7
View File
@@ -1,7 +1,8 @@
"""Who rewrites the transcript once it has been heard.
Normally a small model on OpenRouter: one request, a second, a few tenths of a
cent. A machine with Claude Code or Codex on it is already paying for a model
Normally a small model over one HTTP request: a second, and a few tenths of a
cent on OpenRouter or nothing at all on Google AI Studio's free tier. A machine
with Claude Code, Codex or Antigravity on it is already paying for a model
though, and the subscription that answers "put that in my calendar on Thursday"
can just as well take the "eee"s out of a sentence. No second key, no second
bill. It costs seconds rather than one, because a CLI opens a whole session to
@@ -10,7 +11,11 @@ do it, which is the trade.
Whoever does it, the job is the same one: no tools, no files, no memory of the
last dictation. There is nothing here to look up and nothing to carry over, and
a transcript is text from a microphone rather than an instruction, so the less
the agent can reach while it reads one, the better.
the agent can reach while it reads one, the better. Claude Code is handed an
empty tool list and Codex a read-only sandbox. Antigravity has neither switch,
and this is worth saying plainly rather than implying parity: there the
transcript is read by an agent that could go and do something. What can be done
is done a project of its own, the home directory, and its slash commands off.
"""
import os
@@ -24,7 +29,7 @@ from . import ggml
from . import paths
from .i18n import t
PROVIDERS = ("openrouter", "local", "claude", "codex")
PROVIDERS = ("openrouter", "gemini", "opencode", "local", "claude", "codex", "agy")
class CleanupError(api.ApiError):
@@ -43,7 +48,7 @@ def provider(conf):
def executable(name):
"""The CLI a provider runs, or "" when it needs none."""
return {"claude": "claude", "codex": "codex"}.get(name, "")
return {"claude": "claude", "codex": "codex", "agy": "agy"}.get(name, "")
def model(conf):
@@ -57,13 +62,20 @@ def model(conf):
# Codex is left on whatever it is set to unless a model is typed in, so
# here there is only the name of the thing that did it.
return conf["cleanup_codex_model"].strip() or "codex"
if name == "agy":
# The same arrangement as Codex, and the same reason for it.
return conf["cleanup_agy_model"].strip() or "agy"
if name == "gemini":
return conf["cleanup_gemini_model"]
if name == "opencode":
return conf["cleanup_opencode_model"]
return conf["cleanup_model"]
def run(text, conf, system_prompt, timeout=180, aborter=None):
"""Hand the transcript to whoever is set to clean it up.
`aborter` is only of use to the two that answer over HTTP; a CLI is stopped
`aborter` is only of use to the three that answer over HTTP; a CLI is stopped
between blocks instead, which is close enough when a block is seconds.
"""
name = provider(conf)
@@ -74,9 +86,23 @@ def run(text, conf, system_prompt, timeout=180, aborter=None):
base_url=conf["openrouter_base_url"], timeout=timeout,
aborter=aborter,
)
if name == "gemini":
return api.cleanup(
text, conf.gemini_key(), conf["cleanup_gemini_model"], system_prompt,
reasoning=conf["cleanup_reasoning"],
base_url=conf["gemini_base_url"], timeout=timeout,
provider="gemini", service="Google AI Studio", aborter=aborter,
)
if name == "opencode":
return api.cleanup(
text, conf.opencode_key(), conf["cleanup_opencode_model"], system_prompt,
reasoning=conf["cleanup_reasoning"],
base_url=conf["opencode_base_url"], timeout=timeout,
provider="opencode", service="OpenCode Go", aborter=aborter,
)
if name == "local":
return _local(text, conf, system_prompt, timeout, aborter)
runner = _claude if name == "claude" else _codex
runner = {"claude": _claude, "codex": _codex, "agy": _agy}[name]
return runner(text, conf, system_prompt, timeout)
@@ -172,6 +198,40 @@ def _codex(text, conf, system_prompt, timeout):
return answer
# --- Antigravity ----------------------------------------------------------
def _agy(text, conf, system_prompt, timeout):
# Antigravity takes no system prompt of its own either, so the rules ride in
# front of the transcript, kept apart from it so the two are not read as one.
body = f"{system_prompt}\n\n---\n\n{_wrap(text)}"
cmd = [
"agy", "-p", body,
"--output-format", "text", # the answer, and nothing around it
# Left to itself agy picks up whichever project it was last in and works
# in that project's directory rather than this one. A dictation belongs
# to no project, so each one starts on a project of its own.
"--new-project",
# A transcript that happens to begin with a slash is still a transcript.
"--disable-slash-commands",
# agy gives up after five minutes of its own accord, which would have it
# killed from outside rather than answering.
"--print-timeout", f"{timeout}s",
]
if conf["cleanup_agy_model"].strip():
cmd += ["--model", conf["cleanup_agy_model"].strip()]
effort = assistant.AGY_EFFORT.get(conf["cleanup_reasoning"], "")
if effort:
# agy's own model ids carry the effort in their suffix, so this only
# matters for the ones that do not, and for a model typed in by hand.
cmd += ["--effort", effort]
answer = _output(cmd, timeout, "Antigravity")
if not answer:
raise CleanupError(t("{service} answered with nothing.",
service="Antigravity"))
return answer
def _read(path):
try:
with open(path, encoding="utf-8", errors="replace") as fh:
+40 -11
View File
@@ -216,7 +216,9 @@ def cmd_ask(opts):
conf["assistant_provider"] = opts.provider
if opts.model:
key = {"claude": "assistant_model", "codex": "assistant_codex_model",
"openrouter": "assistant_openrouter_model"}[assistant.provider(conf)]
"openrouter": "assistant_openrouter_model",
"agy": "assistant_agy_model",
"opencode": "assistant_opencode_model"}[assistant.provider(conf)]
conf[key] = opts.model
if opts.dir:
conf["assistant_dir"] = opts.dir
@@ -255,7 +257,8 @@ def cmd_ask(opts):
"cleanup_error": warning,
"mode": "ask",
"question": text,
"assistant_model": conf["assistant_model"],
"assistant": assistant.provider(conf),
"assistant_model": assistant.model(conf),
"raw": text,
"text": answer,
})
@@ -508,7 +511,8 @@ def cmd_history_clear(opts):
# --- settings ---------------------------------------------------------------
SECRET_KEYS = ("openai_api_key", "groq_api_key", "openrouter_api_key")
SECRET_KEYS = ("openai_api_key", "groq_api_key", "openrouter_api_key",
"gemini_api_key", "opencode_api_key")
def _mask(key, value):
@@ -698,6 +702,22 @@ def cmd_test_key(opts):
results[name] = {"ok": True, "message": message}
except api.ApiError as exc:
results[name] = {"ok": False, "message": str(exc)}
if opts.which in ("gemini", "all"):
try:
count = len(api.gemini_models(conf.gemini_key(),
conf["gemini_base_url"]))
message = f"connection works, {count} models visible"
results["gemini"] = {"ok": True, "message": message}
except api.ApiError as exc:
results["gemini"] = {"ok": False, "message": str(exc)}
if opts.which in ("opencode", "all"):
try:
count = len(api.openai_models(conf.opencode_key(),
conf["opencode_base_url"], "OpenCode Go"))
results["opencode"] = {"ok": True,
"message": f"connection works, {count} models visible"}
except api.ApiError as exc:
results["opencode"] = {"ok": False, "message": str(exc)}
everything_ok = all(item["ok"] for item in results.values())
lines = [f"{'' if item['ok'] else ''} {name}: {item['message']}"
for name, item in results.items()]
@@ -862,8 +882,17 @@ def cmd_doctor(opts):
# and marking them by the key they do not use reported every fully local
# setup as broken.
transcribe_ready = conf.transcribe_ready()
if cleaner == "openrouter":
cleanup_ready = bool(conf.openrouter_key())
# Only the ones that answer over HTTP have a key worth looking at. A CLI
# has a program to find instead, and the model on this machine has neither,
# so "no key" there has to read as beside the point rather than as one that
# has gone missing.
cleanup_service, cleanup_key = {
"openrouter": ("OpenRouter", conf.openrouter_key()),
"gemini": ("Google AI Studio", conf.gemini_key()),
"opencode": ("OpenCode Go", conf.opencode_key()),
}.get(cleaner, ("", ""))
if cleanup_service:
cleanup_ready = bool(cleanup_key)
elif cleaner == "local":
cleanup_ready = conf.local_llm_ready()
else:
@@ -875,7 +904,7 @@ def cmd_doctor(opts):
"ready": transcribe_ready},
"cleanup": {"enabled": conf["cleanup_enabled"], "provider": cleaner,
"model": cleanup.model(conf),
"key": bool(conf.openrouter_key()),
"key": bool(cleanup_key) if cleanup_service else None,
"ready": cleanup_ready},
"agent": {"provider": assistant.provider(conf),
"directory": assistant.working_dir(conf)},
@@ -887,9 +916,9 @@ def cmd_doctor(opts):
else:
transcribe_line = (f"{'' if transcribe_ready else ''} {target.service} "
f"key, transcribing on {target.model}")
if cleaner == "openrouter":
cleanup_line = (f"{'' if cleanup_ready else ''} OpenRouter key, "
f"cleaning up on {conf['cleanup_model']}")
if cleanup_service:
cleanup_line = (f"{'' if cleanup_ready else ''} {cleanup_service} key, "
f"cleaning up on {cleanup.model(conf)}")
elif cleaner == "local":
cleanup_line = (f"{'' if cleanup_ready else ''} Local model, "
f"cleaning up on {conf['local_llm_model'] or 'no model yet'}")
@@ -990,7 +1019,7 @@ def build_parser():
ask = leaf(subs, "ask", "put a command to the agent")
ask.add_argument("text", nargs="*", help="the command; read from stdin, or "
"recorded when there is none")
ask.add_argument("--provider", choices=("claude", "codex", "openrouter"),
ask.add_argument("--provider", choices=assistant.PROVIDERS,
help="just for this run")
ask.add_argument("--model", help="just for this run")
ask.add_argument("--dir", help="working directory, just for this run")
@@ -1116,7 +1145,7 @@ def build_parser():
models.set_defaults(func=cmd_models)
test = leaf(subs, "test-key", "check the API keys")
test.add_argument("which", nargs="?", default="all",
choices=("all", *cfg.TRANSCRIBERS))
choices=("all", *cfg.TRANSCRIBERS, "gemini", "opencode"))
test.set_defaults(func=cmd_test_key)
leaf(subs, "doctor", "keys, programs, and what is missing").set_defaults(func=cmd_doctor)
+27 -2
View File
@@ -387,10 +387,19 @@ DEFAULTS = {
"groq_base_url": "https://api.groq.com/openai/v1",
"openrouter_api_key": "",
"openrouter_base_url": "https://openrouter.ai/api/v1",
"gemini_api_key": "",
# Google's OpenAI-compatible endpoint. Cleanup only: there is no
# /audio/transcriptions behind it, so it is not one of the TRANSCRIBERS.
"gemini_base_url": "https://generativelanguage.googleapis.com/v1beta/openai",
"opencode_api_key": "",
"opencode_base_url": "https://opencode.ai/zen/go/v1",
"transcribe_provider": "local", # "local", or a key of TRANSCRIBERS
"transcribe_model": "gpt-4o-transcribe", # used when provider is openai
"groq_transcribe_model": "whisper-large-v3-turbo",
"openrouter_transcribe_model": "openai/gpt-4o-transcribe",
# What a timestamped run (subtitles) asks OpenRouter for: not every model
# there returns segment times. Empty -> openai/whisper-1.
"openrouter_file_model": "",
"language": "tr",
"transcribe_prompt": "",
@@ -412,6 +421,9 @@ DEFAULTS = {
"cleanup_model": "google/gemini-3.5-flash-lite",
"cleanup_claude_model": "haiku", # Claude Code: an alias, or a full model id
"cleanup_codex_model": "", # empty -> whatever Codex is set to
"cleanup_gemini_model": "gemini-3.5-flash-lite",
"cleanup_agy_model": "", # empty -> whatever Antigravity is set to
"cleanup_opencode_model": "deepseek-v4-flash",
"cleanup_reasoning": "", # empty -> whatever the model does by default
# --- llama.cpp, on this machine -----------------------------------------
@@ -457,6 +469,10 @@ DEFAULTS = {
"pause_shortcut": "",
"evdev_hotkey": False,
"overlay_corner": "bottom-left",
"overlay_screen": "",
# Off, so that an indicator stays where it appeared unless it is asked to
# keep up with the pointer. Nothing to say when a screen is named above.
"overlay_follows_pointer": False,
"keep_audio": False,
"history_limit": 200,
# A look at the releases page once a day, and nothing more than a look:
@@ -484,12 +500,14 @@ DEFAULTS = {
# --- speaking a command to an agent -------------------------------------
"assistant_shortcut": "", # empty -> tray only
"assistant_provider": "claude", # claude | codex | openrouter
"assistant_provider": "claude", # claude | codex | agy | openrouter
"assistant_model": "sonnet", # Claude Code: an alias, or a full model id
"assistant_permission_mode": "auto",
"assistant_codex_model": "", # empty -> whatever Codex is set to
"assistant_codex_sandbox": "workspace-write",
"assistant_openrouter_model": "google/gemini-3.5-flash",
"assistant_agy_model": "", # empty -> whatever Antigravity is set to
"assistant_opencode_model": "deepseek-v4-flash",
"assistant_reasoning": "", # empty -> the model's own default
"assistant_dir": "", # empty -> the home directory
"assistant_prompt": "", # empty -> language-specific default
@@ -632,6 +650,12 @@ class Config:
def openrouter_key(self):
return self.api_key("openrouter_api_key")
def gemini_key(self):
return self.api_key("gemini_api_key")
def opencode_key(self):
return self.api_key("opencode_api_key")
def transcribe_target(self):
"""Key, endpoint and model for whichever provider does speech to text.
@@ -651,8 +675,9 @@ class Config:
# to land on rather than reading it from there.
name = "openai"
who = TRANSCRIBERS[name]
file_model = self["openrouter_file_model"] if name == "openrouter" else ""
return api.Target(name, who.service, self.api_key(who.key),
self[who.url], self[who.model])
self[who.url], self[who.model], file_model.strip())
def transcribe_ready(self):
"""Whether speech to text could run right now, without opening Settings."""
+135 -9
View File
@@ -76,6 +76,13 @@ Program = collections.namedtuple("Program", "name repo binary health")
WHISPER = Program("whisper", "ggml-org/whisper.cpp", "whisper-server", "")
LLAMA = Program("llama", "ggml-org/llama.cpp", "llama-server", "/health")
DIKTE_REPO = "yusufipk/dikte"
MANAGED_WHISPER_RELEASE = "whisper.cpp-v1.9.3"
MANAGED_WHISPER_VERSION = "v1.9.3"
MANAGED_WHISPER_VULKAN = "whisper-bin-ubuntu-vulkan-x64.tar.gz"
MANAGED_WHISPER_SHA256 = (
"c25ca76504144da488eb74441390a7b9aa7ce547e5f2f391cbd831253c9b54d8"
)
# Where the models are listed. Neither list is written into Dikte: a catalogue
# in the source means a release of Dikte for every model somebody else
@@ -83,6 +90,10 @@ LLAMA = Program("llama", "ggml-org/llama.cpp", "llama-server", "/health")
WHISPER_MODELS_REPO = "ggerganov/whisper.cpp"
LLM_AUTHOR = "ggml-org"
# The file llama.cpp attaches to its version releases in place of the binaries:
# a line naming the nightly tag those are published under.
NIGHTLY_TAG = "nightly-tag.txt"
# What the whisper repository holds besides models: Core ML encoders for Apple
# hardware and the odd loose file.
WHISPER_PREFIX = "ggml-"
@@ -277,6 +288,102 @@ def _wanted_assets(program):
return (f"bin-ubuntu-{arch}.tar.gz",)
def _managed_whisper(program, tag=""):
"""Whether this machine is one Dikte publishes its own whisper-server for.
Linux x86_64 with a Vulkan loader on it, and no version asked for by hand:
a pinned version is upstream's to answer.
"""
return (not tag and program is WHISPER and sys.platform == "linux"
and platform.machine().lower() in ("x86_64", "amd64")
and _has_vulkan())
def _managed_asset(refresh=False):
"""The Vulkan whisper-server Dikte builds itself, or None.
Taken only when the archive's digest is the reviewed one. Anything else,
a release that is not there yet, a GitHub that cannot be reached, a file
that is not the reviewed bytes, leaves upstream's processor build as the
answer, and the install record says which of the two landed.
"""
try:
_, assets = hub.release(DIKTE_REPO, MANAGED_WHISPER_RELEASE,
refresh=refresh)
except hub.HubError:
return None
return next((a for a in assets
if a.name.endswith(MANAGED_WHISPER_VULKAN)
and a.sha256 == MANAGED_WHISPER_SHA256), None)
def _matching_asset(program, assets):
"""The archive this machine wants out of one release's files, or None."""
for ending in _wanted_assets(program):
item = next((a for a in assets if a.name.endswith(ending)), None)
if item:
return item
return None
def _pick_asset(program, tag="", refresh=False):
"""(tag, Item) for the release archive to install. Item is None when there
is none for this machine.
Dikte's own Vulkan whisper-server comes before upstream's where this
machine is one it is built for, because whisper.cpp publishes no Vulkan
archive for Linux at all.
A named tag is taken as given. For the newest, what GitHub answers is not
always where the builds are: llama.cpp's latest release is a version marker
carrying a single nightly-tag.txt, which names the tag the archives are
actually attached to, and those are prereleases that "latest" never points
at. The pointer is followed when it is there, and when it is not, the newest
release that does carry a build for this machine is taken instead.
"""
if _managed_whisper(program, tag):
item = _managed_asset(refresh=refresh)
if item:
return MANAGED_WHISPER_VERSION, item
named = bool(tag) and tag != "latest"
missing = None
try:
tag, assets = hub.release(program.repo, tag or "latest", refresh=refresh)
except hub.HubError as exc:
# A release carrying no files at all is the case the search below exists
# for, not a reason to stop before it: the build for this machine may be
# attached to a prerelease that "latest" never points at. The failure is
# kept rather than dropped, because an unreachable GitHub arrives here
# the same way and that one is the message the caller wants.
if named:
raise
missing, assets = exc, []
item = _matching_asset(program, assets)
if item or named:
return tag, item
# Best effort from here on: a machine this project publishes nothing for is
# not a failed lookup, and the caller's message about that is the useful
# one. Whatever goes wrong while looking further leaves it standing.
try:
pointer = next((a for a in assets if a.name == NIGHTLY_TAG), None)
if pointer:
nightly = hub.text(pointer.url).strip()
if nightly:
found, assets = hub.release(program.repo, nightly, refresh=refresh)
item = _matching_asset(program, assets)
if item:
return found, item
for found, assets in hub.releases(program.repo, refresh=refresh):
item = _matching_asset(program, assets)
if item:
return found, item
except hub.HubError:
pass
if missing is not None:
raise missing
return tag, None
def _install_record(program):
return BIN_DIR / program.name / "installed.json"
@@ -300,6 +407,21 @@ def installed_version(program):
return _read_record(program).get("tag") or ""
def vulkan_missing(program):
"""Whether what Dikte installed is the processor build on a machine the
Vulkan one was fetched for.
The Vulkan whisper-server is a release of Dikte's own, published by hand
once the reviewed archive is built, and the install falls back to the
upstream processor build whenever that release, the file in it, or its
reviewed digest is not there. Nothing is wrong with the fallback except
that it is invisible: a graphics card sitting idle looks exactly like a
graphics card being used.
"""
return (bool(installed_program(program))
and _read_record(program).get("backend") == "processor")
def program_path(program, custom=""):
"""Which copy of the program to run, or "" when there is none.
@@ -372,18 +494,15 @@ def install_program(program, tag="", on_progress=None, should_stop=None,
`tag` is empty for whatever the project released last, which is the point:
a version pinned in Dikte's source would mean a release of Dikte every time
whisper.cpp has one.
whisper.cpp has one. The Linux Vulkan whisper-server is the exception, and
_pick_asset says why.
"""
managed = _managed_whisper(program, tag)
try:
tag, assets = hub.release(program.repo, tag or "latest", refresh=refresh)
tag, item = _pick_asset(program, tag, refresh=refresh)
except hub.HubError as exc:
raise LocalError(str(exc)) from exc
item = None
for ending in _wanted_assets(program):
item = next((a for a in assets if a.name.endswith(ending)), None)
if item:
break
if item is None:
# Nothing to download and nothing to install for you: whisper.cpp
# publishes no macOS binary, and Homebrew's whisper-cpp is configured
@@ -447,8 +566,15 @@ def install_program(program, tag="", on_progress=None, should_stop=None,
# Found under the sibling, run from the final directory.
binary = into / binary.relative_to(fresh)
# Written last, so the record never points at anything half-made.
_install_record(program).write_text(
json.dumps({"tag": tag, "binary": str(binary)}), encoding="utf-8")
record = {"tag": tag, "binary": str(binary)}
if managed:
# Which of the two builds this machine ended up with. Only written
# where both were on offer, so an install that never had the
# choice is not made to look like a fallback.
record["backend"] = (
"vulkan" if item.name.endswith(MANAGED_WHISPER_VULKAN)
else "processor")
_install_record(program).write_text(json.dumps(record), encoding="utf-8")
except OSError as exc:
raise LocalError(t("Could not install {name}: {error}",
name=program.name, error=exc)) from exc
+48 -4
View File
@@ -121,6 +121,13 @@ def _digest(value):
return value.split(":", 1)[1] if value.startswith("sha256:") else value
def _assets(data):
return [Item(a.get("name") or "", a.get("browser_download_url") or "",
int(a.get("size") or 0), _digest(a.get("digest")))
for a in (data.get("assets") or [])
if a.get("browser_download_url")]
def release(repo, tag="latest", refresh=False):
"""(tag, [Item]) for one GitHub release, newest when no tag is given."""
where = "latest" if tag in ("", "latest") else f"tags/{tag}"
@@ -128,10 +135,47 @@ def release(repo, tag="latest", refresh=False):
f"{GITHUB_API}/repos/{repo}/releases/{where}", refresh=refresh)
if not isinstance(data, dict) or not data.get("assets"):
raise HubError(t("{repo} has no downloadable release.", repo=repo))
assets = [Item(a.get("name") or "", a.get("browser_download_url") or "",
int(a.get("size") or 0), _digest(a.get("digest")))
for a in data["assets"] if a.get("browser_download_url")]
return data.get("tag_name") or tag, assets
return data.get("tag_name") or tag, _assets(data)
def releases(repo, limit=20, refresh=False):
"""[(tag, [Item])] for the recent releases, newest first, with their files.
"latest" is one release and this is the list behind it, prereleases
included: a project that attaches its builds to a prerelease is invisible
to release() above, and its newest usable build is in here.
"""
data = _fetch(f"gh-list-{repo}-{limit}",
f"{GITHUB_API}/repos/{repo}/releases?per_page={limit}",
refresh=refresh)
if not isinstance(data, list):
raise HubError(t("{repo} has no downloadable release.", repo=repo))
out = []
for entry in data:
tag, items = entry.get("tag_name") or "", _assets(entry)
if tag and items:
out.append((tag, items))
return out
def text(url, limit=4096, timeout=20):
"""A small text file from a release, as a string.
Not cached and not checksummed, because what it carries is a pointer: a few
bytes naming the release the actual archives are attached to, read once on
the way to a download that is checked in full.
"""
request = urllib.request.Request(url, headers={"User-Agent": USER_AGENT})
try:
with urllib.request.urlopen(request, timeout=timeout) as response:
return response.read(limit).decode("utf-8", "replace")
except urllib.error.HTTPError as exc:
exc.close()
raise HubError(t("{url} answered HTTP {code}.",
url=urllib.parse.urlsplit(url).netloc, code=exc.code)) from exc
except (urllib.error.URLError, OSError, ValueError) as exc:
raise HubError(t("Could not reach {url}: {error}",
url=urllib.parse.urlsplit(url).netloc, error=exc)) from exc
def newest_release(repo, refresh=False):
+57 -39
View File
@@ -40,8 +40,16 @@ def t(text, /, **kwargs):
# by the sentence, so it arrives already inflected. English takes the name as it
# is and puts the preposition in the sentence, where it belongs.
_TR_CASES = {
"dative": {"Claude": "Claude'a", "Codex": "Codex'e", "OpenRouter": "OpenRouter'a"},
"accusative": {"Claude": "Claude'u", "Codex": "Codex'i", "OpenRouter": "OpenRouter'ı"},
"dative": {
"Claude": "Claude'a", "Codex": "Codex'e", "OpenRouter": "OpenRouter'a",
"Google AI Studio": "Google AI Studio'ya", "Antigravity": "Antigravity'ye",
"OpenCode Go": "OpenCode Go'ya",
},
"accusative": {
"Claude": "Claude'u", "Codex": "Codex'i", "OpenRouter": "OpenRouter'ı",
"Google AI Studio": "Google AI Studio'yu", "Antigravity": "Antigravity'yi",
"OpenCode Go": "OpenCode Go'yu",
},
}
@@ -151,6 +159,7 @@ TR = {
# --- settings: tabs and general ------------------------------------
"Dikte Settings": "Dikte Ayarları",
"General": "Genel",
"Display": "Ekran",
"API and models": "API ve modeller",
"Cleanup rules": "Temizleme kuralları",
"Audio file": "Ses dosyası",
@@ -179,6 +188,11 @@ TR = {
"macOS bu ilk gönderildiğinde Erişilebilirlik izni ister.",
"Restore the previous clipboard after pasting":
"Yapıştırdıktan sonra eski pano içeriğini geri koy",
"Indicator screen": "Gösterge ekranı",
"Follow the active screen": "Etkin ekranı takip et",
"{name} (not connected)": "{name} (bağlı değil)",
"Move it when the active screen changes":
"Etkin ekran değiştiğinde göstergeyi de taşı",
"Indicator corner": "Gösterge köşesi",
"bottom-left": "sol-alt",
"bottom-right": "sağ-alt",
@@ -216,30 +230,34 @@ TR = {
"Transcript cleanup": "Transkripti temizleme",
"API key": "API anahtarı",
"Model": "Model",
"Audio file model": "Ses dosyası modeli",
"The model a timestamped audio file (subtitles) is sent to. Not every model on "
"OpenRouter returns segment times; empty means openai/whisper-1.":
"Zaman damgalı bir ses dosyasının (altyazı) gönderildiği model. OpenRouter'daki her "
"model segment zamanı döndürmez; boşsa openai/whisper-1 kullanılır.",
"Provider": "Sağlayıcı",
"sk-… (falls back to OPENAI_API_KEY)": "sk-… (boşsa OPENAI_API_KEY kullanılır)",
"gsk_… (falls back to GROQ_API_KEY)": "gsk_… (boşsa GROQ_API_KEY kullanılır)",
"sk-or-… (falls back to OPENROUTER_API_KEY)": "sk-or-… (boşsa OPENROUTER_API_KEY kullanılır)",
"(falls back to GEMINI_API_KEY)": "(boşsa GEMINI_API_KEY kullanılır)",
"(falls back to OPENCODE_API_KEY)": "(boşsa OPENCODE_API_KEY kullanılır)",
"Test": "Test et",
"Trying…": "Deneniyor…",
"Runs on OpenRouter.": "OpenRouter üzerinde çalışır.",
"Runs on Google AI Studio.": "Google AI Studio üzerinde çalışır.",
"Runs on OpenCode Go.": "OpenCode Go üzerinde çalışır.",
"Connection works. {count} audio models visible.":
"Bağlantı tamam. {count} ses modeli görünüyor.",
"Connection works. {count} models visible.":
"Bağlantı tamam. {count} model görünüyor.",
"Clean the transcript with a model": "Transkripti bir modelle temizle",
"OpenRouter is the quickest and the only one that needs nothing installed. "
"Claude Code and Codex clean up on the subscription you already have, "
"without a second key, and take a few seconds longer because each one opens "
"a session to do it.":
"En hızlısı OpenRouter'dır ve kurulu bir program istemeyen tek seçenektir. "
"Claude Code ile Codex, temizliği hâlihazırda ödediğin abonelik üzerinden "
"yapar, ikinci bir anahtar istemez; her biri bunun için bir oturum açtığından "
"birkaç saniye daha uzun sürer.",
"{binary} is not on your PATH, so cleanup would fail and the raw transcript "
"would be pasted. Install it, or pick another one above.":
"{binary} PATH'te değil; temizleme başarısız olur ve ham transkript "
"yapıştırılır. Kur ya da yukarıdan başka birini seç.",
"Thinking": "Düşünme",
"Model's own default": "Modelin kendi varsayılanı",
"Antigravity's own default": "Antigravity'nin kendi varsayılanı",
"Off": "Kapalı",
"Minimal": "En az",
"Low": "Düşük",
@@ -502,18 +520,6 @@ TR = {
# --- settings: the agent ------------------------------------------------
"Agent": "Ajan",
"This shortcut records the same way dictation does, but the transcript is "
"not what gets pasted. It goes to an agent as a command, and what comes "
"back is pasted instead: the answer to a question, or a sentence saying "
"what was done. Claude Code and Codex run as the session you would have "
"opened yourself, with your skills, your connected services and your "
"account.":
"Bu kısayol dikte ile aynı şekilde kaydeder, ama yapıştırılan şey "
"transkript değildir. Transkript bir ajana komut olarak gider ve yerine "
"oradan döneni yapıştırılır: bir sorunun cevabı ya da ne yapıldığını "
"söyleyen bir cümle. Claude Code ve Codex, kendi açacağın oturumun "
"aynısı olarak çalışır: skill'lerinle, bağlı servislerinle ve kendi "
"hesabınla.",
"How it runs": "Nasıl çalışıyor",
"Runs on": "Şunun üstünde çalışır",
"More thinking is slower, and you are standing in front of the screen while "
@@ -540,6 +546,24 @@ TR = {
"Yukarıdaki çalışma dizini ve izinler burada bir şey ifade etmez.",
"Needs no program installed, only the OpenRouter key.":
"Kurulu bir programa değil, yalnızca OpenRouter anahtarına ihtiyaç duyar.",
"A plain question and a plain answer, over the OpenCode Go key you already "
"have. It runs no commands, opens no files and reaches none of your "
"services, so it can tell you what the capital of Peru is but not what is "
"in your calendar. Working directory and permissions above mean nothing "
"here.":
"Elindeki OpenCode Go anahtarı üzerinden düz bir soru ve düz bir cevap. "
"Komut çalıştırmaz, dosya açmaz, servislerinin hiçbirine erişmez; yani "
"Peru'nun başkentini söyler ama takviminde ne olduğunu söyleyemez. "
"Yukarıdaki çalışma dizini ve izinler burada bir şey ifade etmez.",
"Needs no program installed, only an OpenCode Go key.":
"Kurulu bir programa değil, yalnızca bir OpenCode Go anahtarına ihtiyaç duyar.",
"Antigravity has neither a permission mode nor a sandbox to hand it, so "
"what it may do without asking is whatever its own allow-rules say. The "
"Permissions and Sandbox boxes above belong to the other two; the working "
"directory still applies.":
"Antigravity'ye verilebilecek bir izin kipi ya da sandbox yok; sormadan "
"ne yapabileceğini kendi allow-rule'ları belirler. Yukarıdaki İzinler ve "
"Sandbox kutuları diğer ikisine ait; çalışma dizini burada da geçerli.",
"{binary} is not on your PATH, so this cannot run yet. Install it, or pick "
"another one above.":
"{binary} PATH'te değil, dolayısıyla bu henüz çalışamaz. Kur ya da "
@@ -598,7 +622,7 @@ TR = {
"configuration already says.":
"Her komutla birlikte ajana söylenir, kendi yapılandırmanın zaten "
"söylediklerinin üstüne eklenir.",
" · asked Claude: {question}": " · Claude'a soruldu: {question}",
" · asked {who}: {question}": " · {who} soruldu: {question}",
# --- meetings: tray and pipeline ---------------------------------------
"Record a meeting": "Toplantı kaydet",
@@ -668,12 +692,6 @@ TR = {
# --- settings: meeting --------------------------------------------------
"Minutes": "Tutanaklar",
"A meeting is recorded from two devices at once: your microphone and "
"whatever comes out of your speakers. Nothing has to guess who was "
"speaking, because the two never share a channel.":
"Toplantı iki aygıttan aynı anda kaydedilir: mikrofonun ve hoparlöründen "
"çıkan ses. Kimin konuştuğunun tahmin edilmesi gerekmez, çünkü ikisi hiç "
"aynı kanala girmez.",
"Sound": "Ses",
"Same as dictation": "Diktedekiyle aynı",
"Current output": "Geçerli çıkış",
@@ -768,7 +786,11 @@ TR = {
"Local model": "Yerel model",
"Not installed.": "Kurulu değil.",
"Installed on the system: {path}": "Sistemde kurulu: {path}",
"Download again": "Yeniden indir",
"Downloaded, version {version}.": "İndirildi, sürüm {version}.",
"Downloaded, version {version}. There was no Vulkan build, "
"so this one runs on the processor.":
"İndirildi, sürüm {version}. Vulkan sürümü yoktu, bu sürüm işlemcide çalışıyor.",
"Fetching the model list…": "Model listesi çekiliyor…",
"Downloading…": "İndiriliyor…",
"Downloading: {done} of {total}{share}": "İndiriliyor: {done} / {total}{share}",
@@ -776,6 +798,12 @@ TR = {
"Ready: {name}.": "Hazır: {name}.",
"Nothing downloaded yet.": "Henüz bir şey indirilmedi.",
"{name} has not been downloaded yet.": "{name} henüz indirilmedi.",
"{name} is here, but the program above is not. Download it first.":
"{name} burada, ama yukarıdaki program değil. Önce onu indirin.",
"{name} is not on this machine and this publisher does not offer it. "
"Choose another model, or another publisher.":
"{name} bu makinede yok ve bu yayıncı da sunmuyor. Başka bir model, "
"ya da başka bir yayıncı seçin.",
"downloaded": "indirildi",
"not downloaded": "indirilmedi",
"Delete model": "Modeli sil",
@@ -801,16 +829,6 @@ TR = {
"Düşünmeye eğitilmiş bir model, aksi söylenmedikçe düşünür; bir virgül "
"için 300 token akıl yürütmek 300 token'lık bekleyiştir. Temizleme için "
"doğrusu Kapalı.",
"OpenRouter is the quickest and the only one that needs nothing "
"installed. llama.cpp runs here, on a model downloaded below. Claude Code "
"and Codex clean up on the subscription you already have, without a "
"second key, and take a few seconds longer because each one opens a "
"session to do it.":
"OpenRouter en hızlısıdır ve kurulum istemeyen tek seçenektir. "
"llama.cpp burada, aşağıda indirilen bir modelle çalışır. Claude Code "
"ve Codex, ikinci bir anahtar olmadan zaten sahip olduğun abonelikle "
"temizler; her biri bunun için bir oturum açtığından birkaç saniye "
"daha sürer.",
"whisper.cpp reaches the card through CUDA, ROCm or Vulkan when the build "
"it is running was made with one. A build without any of them runs on the "
"processor whatever this says.":
+108 -5
View File
@@ -1,6 +1,7 @@
"""The small recording indicator that appears in a screen corner without taking focus."""
import math
import os
import sys
from PyQt6.QtCore import Qt, QTimer, QRectF, QPointF
@@ -15,6 +16,7 @@ MIN_WIDTH = 210
MAX_WIDTH = 460
MARGIN = 28
GAP = 10 # between two indicators sharing a corner
FOLLOW_EVERY = 8 # ticks between two looks for the pointer: about four a second
BG = QColor(22, 24, 29, 238)
BORDER = QColor(255, 255, 255, 28)
@@ -37,14 +39,68 @@ STATE_COLORS = {"recording": REC, "asking": ASK, "meeting": REC, "busy": BUSY,
LIVE = ("recording", "asking", "meeting")
# KWin's interface, kept once one has been built. See _compositor_screen.
_kwin = None
def _compositor_screen():
"""The screen KWin says the session is on, or None where nothing says.
Wayland tells a client where the pointer is only while it is over one of
that client's own windows, and the indicator is never under the pointer, so
QCursor.pos() answers with a stale point or, when the pointer has never
been over a window of ours, with the origin. Either way the indicator lands
in the corner of whichever screen holds 0,0 instead of the one being worked
on, and on a two-monitor desk that is the wrong screen most of the time.
KWin does know, and it names outputs the way Qt names screens, by
connector, natively and through XWayland alike. No other Wayland desktop
answers this, so the rest are left with the pointer, which is right on X11
and wrong on Wayland exactly as before.
What it answers with is the active output, which is the one under the
pointer only where Plasma is set to let the active screen follow the mouse.
Under the default, click to focus, it is the focused window's screen, so
the indicator lands where the typing is going rather than where the mouse
was left. Which is why nothing here, and nothing in the settings window,
promises the pointer.
"""
global _kwin
if _kwin is None or not _kwin.isValid():
# Which also leaves macOS and Windows out, where nothing sets it and
# the pointer can be asked where it is like anywhere else.
desktop = os.environ.get("XDG_CURRENT_DESKTOP", "").lower()
if "kde" not in desktop and "plasma" not in desktop:
return None
try:
from PyQt6.QtDBus import QDBusConnection, QDBusInterface
_kwin = QDBusInterface("org.kde.KWin", "/KWin", "org.kde.KWin",
QDBusConnection.sessionBus())
except Exception:
return None
if not _kwin.isValid():
return None
# A compositor busy enough not to answer in a fifth of a second is one
# the indicator should stop waiting for, not one it should freeze with.
_kwin.setTimeout(200)
answer = _kwin.call("activeOutputName").arguments()
name = answer[0] if answer else ""
return next((item for item in QApplication.screens() if item.name() == name),
None)
class Overlay(QWidget):
"""One indicator. Give it `below` and it stacks on top of that one instead
of covering it, which is what lets a dictation and a command to the agent be
under way at the same time and still both be visible."""
def __init__(self, corner="bottom-left", below=None, dismissable=False):
def __init__(self, corner="bottom-left", below=None, dismissable=False,
screen_name="", follow_pointer=False):
super().__init__(None)
self.corner = corner
self.screen_name = screen_name
# Whether it goes on following the pointer once it is up, rather than
# settling on the screen it appeared on.
self.follow_pointer = follow_pointer
self.below = below
# A job that can run for ten minutes should not have to be watched for
# ten minutes. Clicking such an indicator puts the progress away; the
@@ -63,6 +119,8 @@ class Overlay(QWidget):
self.seconds = 0.0
self._phase = 0.0
self._concealed = True
self._shown_on = "" # the screen it was last put on, by name
self._looks = 0 # ticks since the pointer was last looked for
flags = (
Qt.WindowType.FramelessWindowHint
@@ -250,9 +308,52 @@ class Overlay(QWidget):
min(MAX_WIDTH, metrics.horizontalAdvance(self.message) + extra))
self.resize(width, HEIGHT)
def _screen(self):
"""The screen this indicator belongs on right now.
The one the settings name, or, when none is named or it is not plugged
in right now, where the user actually is. Names are connector names on
X11 and model names on macOS, where two identical monitors can share
one; the first then wins.
One stacking on another belongs on that one's screen and nowhere else.
Asked for itself it would answer where the user is now, which is not
where the ribbon it stacks on was put a minute ago, and the pair would
end up a monitor apart with this one raised over nothing.
"""
if self.below is not None and self.below.showing:
under = next((item for item in QApplication.screens()
if item.name() == self.below._shown_on), None)
if under is not None:
return under
named = next(
(item for item in QApplication.screens() if item.name() == self.screen_name),
None,
)
return (named or _compositor_screen()
or QApplication.screenAt(QCursor.pos())
or QApplication.primaryScreen())
def _wandered_off(self):
"""Whether the pointer has left the screen the indicator is on.
Only asked while it is following, and only every few ticks: the answer
costs a word with the compositor, and a hand moving a mouse across a
desk is slow next to a 33 ms ribbon. Every tick for one that stacks on
another, where the answer is free and waiting a third of a second for
it would leave the pair split over two monitors for that long.
"""
if not self.follow_pointer or self.screen_name:
return False
if self.below is None or not self.below.showing:
self._looks = (self._looks + 1) % FOLLOW_EVERY
if self._looks:
return False
return self._screen().name() != self._shown_on
def _reposition(self):
# On a multi-monitor setup, show up where the user actually is.
screen = QApplication.screenAt(QCursor.pos()) or QApplication.primaryScreen()
screen = self._screen()
self._shown_on = screen.name()
area = screen.availableGeometry()
left = "left" in self.corner
top = "top" in self.corner
@@ -268,8 +369,10 @@ class Overlay(QWidget):
def _tick(self):
self._phase += 0.12
# The one underneath can come and go while this one is up; drop back to
# the corner when it does rather than leaving a gap where it was.
if self.below is not None and self.below.showing != self._stacked:
# the corner when it does rather than leaving a gap where it was. And
# the screen under the pointer can change while it is up too.
moved = self.below is not None and self.below.showing != self._stacked
if moved or self._wandered_off():
self._reposition()
if self.state in LIVE and not self.paused:
# keep the ribbon moving even through a pause in speech
+469 -57
View File
@@ -6,7 +6,7 @@ import shutil
import sys
import threading
from PyQt6.QtCore import QEvent, QObject, QRect, Qt, QUrl, pyqtSignal
from PyQt6.QtCore import QEvent, QObject, QRect, Qt, QTimer, QUrl, pyqtSignal
from PyQt6.QtGui import QDesktopServices, QGuiApplication, QKeySequence, QShortcut
from PyQt6.QtWidgets import (
QAbstractItemView, QAbstractSpinBox, QCheckBox, QComboBox, QDialog,
@@ -60,12 +60,24 @@ CLEANUP_MODELS = [
"google/gemini-2.5-flash-lite", "anthropic/claude-haiku-4.5",
"openai/gpt-5-mini", "meta-llama/llama-3.3-70b-instruct",
]
# In the order they answer in. A request to OpenRouter is over in a second, a
# model here takes a little longer and costs nothing, and the two CLIs the agent
# can run on open a whole session to do the smaller job.
GEMINI_MODELS = [
"gemini-3.5-flash-lite", "gemini-3.1-flash-lite",
"gemini-2.5-flash-lite", "gemini-3.5-flash", "gemini-2.5-flash",
]
# agy's model ids carry the reasoning effort in their suffix, which is why one
# model appears here at more than one level. The same list seeds two boxes:
# cleanup, which wants the bottom rung, and the agent, which sometimes does not.
AGY_MODELS = [
"gemini-3.7-flash-low", "gemini-3.7-flash-medium", "gemini-3.7-flash-high",
"gemini-3.5-flash-low", "gemini-3.1-pro-low",
]
# In the order they answer in. The hosted requests are over in a second, a
# model here takes a little longer and costs nothing, and the three CLIs the
# agent can run on open a whole session to do the smaller job.
CLEANUP_PROVIDERS = [
("OpenRouter", "openrouter"), ("This machine (llama.cpp)", "local"),
("Claude Code", "claude"), ("Codex", "codex"),
("OpenRouter", "openrouter"), ("Google AI Studio", "gemini"),
("OpenCode Go", "opencode"), ("This machine (llama.cpp)", "local"),
("Claude Code", "claude"), ("Codex", "codex"), ("Antigravity", "agy"),
]
# Cleaning up a sentence is the lightest thing either of them will ever be
# asked, so the small model comes first.
@@ -77,7 +89,8 @@ MEETING_MODELS = [
"anthropic/claude-sonnet-5", "openai/gpt-5.4", "x-ai/grok-4.5",
]
ASSISTANT_PROVIDERS = [
("Claude Code", "claude"), ("Codex", "codex"), ("OpenRouter", "openrouter"),
("Claude Code", "claude"), ("Codex", "codex"), ("Antigravity", "agy"),
("OpenRouter", "openrouter"), ("OpenCode Go", "opencode"),
]
# Aliases resolve to the newest model of that name, so they age better than an
# id does; a full id can be typed in when a particular one is wanted.
@@ -92,6 +105,13 @@ ASSISTANT_OR_MODELS = [
"google/gemini-3.5-flash", "anthropic/claude-sonnet-5", "openai/gpt-5.4",
"x-ai/grok-4.5", "google/gemini-3.1-pro-preview",
]
# A starting set of the models OpenCode Go serves over /chat/completions; the
# Fetch button asks the endpoint itself for the full catalog of the day.
OPENCODE_MODELS = [
"deepseek-v4-flash", "deepseek-v4-pro", "glm-5.3", "glm-5.2", "glm-5.1",
"kimi-k3", "kimi-k2.7-code", "kimi-k2.6", "longcat-2.0",
"mimo-v2.5", "mimo-v2.5-pro", "hy3",
]
# What Claude Code may do without being able to ask. It cannot ask: there is no
# window to answer in, so a mode that would have prompted denies instead.
PERMISSION_MODES = [
@@ -227,6 +247,13 @@ class LocalModelBox(QGroupBox):
self._pending = False
self._stop = False
self._wanted = "" # the model to select once a list arrives
self._chosen_in = "" # the publisher the selected model is from
# Typing or arrowing through the publisher box changes its text a
# character at a time, and each of those would otherwise be a request.
self._later = QTimer(self)
self._later.setSingleShot(True)
self._later.setInterval(400)
self._later.timeout.connect(self._later_fetch)
form = QFormLayout(self)
@@ -308,6 +335,8 @@ class LocalModelBox(QGroupBox):
self._wanted = model
self._pending = True
self._show_program()
self._chosen_in = repo or (ggml.SUGGESTED_LLM[0] if self._repos is not None
else "")
if self._repos is not None:
self.repo.blockSignals(True)
self.repo.clear()
@@ -329,15 +358,30 @@ class LocalModelBox(QGroupBox):
path = ggml.program_path(self.program)
if not path:
self.program_label.setText(t("Not installed."))
self.install_button.setText(t("Download"))
self.install_button.setVisible(True)
return
self.install_button.setVisible(not ggml.installed_program(self.program)
and not ggml.system_program(self.program))
# A copy that is here is not a copy that is right. whisper.cpp releases
# every few weeks, and a graphics card installed after Dikte was
# changes which build this machine should be running; the button was
# hidden the moment anything landed, and nothing else on this window
# asks for the download again.
self.install_button.setText(t("Download again")
if ggml.installed_program(self.program)
else t("Download"))
self.install_button.setVisible(not ggml.system_program(self.program))
if ggml.system_program(self.program):
# Worth saying which one is running: a distribution package is built
# for this machine and may reach the graphics card, while the
# released binaries carry processor backends only.
self.program_label.setText(t("Installed on the system: {path}", path=path))
elif ggml.vulkan_missing(self.program):
# The download landed the processor build where the graphics card
# one belongs, and nothing else on this window would say so.
self.program_label.setText(
t("Downloaded, version {version}. There was no Vulkan build, "
"so this one runs on the processor.",
version=ggml.installed_version(self.program) or "?"))
else:
self.program_label.setText(
t("Downloaded, version {version}.",
@@ -347,11 +391,18 @@ class LocalModelBox(QGroupBox):
def _fill_repos(self, current):
def work():
self._listed.emit([("repos", ggml.llm_repos())], "")
self._listed.emit([("repos", ggml.llm_repos(), "")], "")
threading.Thread(target=work, daemon=True).start()
def _repo_changed(self):
if not self._downloading:
self._later.start()
def _later_fetch(self):
# A download that started inside the wait was not there to be seen when
# the timer went off, and rebuilding the rows underneath one is exactly
# what the guard above is for.
if not self._downloading:
self._fetch_models(self.repository())
@@ -361,18 +412,29 @@ class LocalModelBox(QGroupBox):
def work():
try:
found = self._models(repo) if self._repos is not None else self._models()
self._listed.emit([("models", found)], "")
self._listed.emit([("models", found, repo)], "")
except ggml.LocalError as exc:
self._listed.emit([], str(exc))
self._listed.emit([("models", [], repo)], str(exc))
threading.Thread(target=work, daemon=True).start()
def _on_listed(self, payload, error):
kind, found, repo = payload[0] if payload else ("repos", [], "")
# A publisher changed while its predecessor's list was still on the way
# would otherwise be answered with the wrong models, whichever request
# happened to come back last.
if kind == "models" and repo != self.repository():
return
if error:
# The list is the publisher's, so a failed one leaves the box no
# longer showing this publisher's models: emptying it is what keeps
# the two boxes saying the same thing. The message goes on after,
# because filling the box writes a status of its own.
if kind == "models":
self._fill_models([])
self._refresh_buttons()
self.status.setText(error)
self._refresh_buttons()
return
kind, found = payload[0]
if kind == "repos":
current = self.repo.currentText()
self.repo.blockSignals(True)
@@ -386,7 +448,12 @@ class LocalModelBox(QGroupBox):
def _fill_models(self, items):
"""One row per model, saying what it weighs and whether it is here."""
wanted = self._wanted or self.selected()
# The selection is only worth carrying over within the publisher it was
# made in. Carried across one, a model this repository does not publish
# would be added back as "not downloaded" and selected again, and
# changing the publisher would leave the model box looking untouched.
same = self._repos is None or self.repository() == self._chosen_in
wanted = self._wanted or (self.selected() if same else "")
here = [name for name in (self._model_path(i.name).name for i in items)]
self.model.blockSignals(True)
self.model.clear()
@@ -411,6 +478,7 @@ class LocalModelBox(QGroupBox):
self.model.blockSignals(False)
self._fit_popup(self.model)
self._wanted = ""
self._chosen_in = self.repository()
self._model_changed()
def _on_disk(self):
@@ -439,6 +507,9 @@ class LocalModelBox(QGroupBox):
self._show_program()
if error:
self.program_label.setText(error)
# The model line says whether the program is here, so installing one
# changes what it should read.
self._refresh_buttons()
self.changed.emit()
def _current_item(self):
@@ -525,15 +596,30 @@ class LocalModelBox(QGroupBox):
def _refresh_buttons(self):
name = self.selected()
here = bool(name) and ggml.have_model(self._model_path(name))
# A row carries what it takes to fetch it. The ones that do not are the
# models found on this disk and the one the settings name but the list
# does not offer: there is nothing to press Download for on those, and
# a button that can only do nothing is worse than one that is out.
item = self._current_item()
self.delete_button.setEnabled(here and not self._downloading)
self.download_button.setText(t("Stop") if self._downloading else t("Download"))
self.download_button.setEnabled(self._downloading or (bool(name) and not here))
self.download_button.setEnabled(self._downloading or (item is not None
and not here))
if self._downloading:
return
if not name:
self.status.setText(t("Nothing downloaded yet."))
elif here and not ggml.program_path(self.program):
# The model alone runs nothing, and "Ready" over a missing program
# reads as though it does.
self.status.setText(t("{name} is here, but the program above is "
"not. Download it first.", name=name))
elif here:
self.status.setText(t("Ready: {name}.", name=name))
elif item is None:
self.status.setText(t("{name} is not on this machine and this "
"publisher does not offer it. Choose another "
"model, or another publisher.", name=name))
else:
self.status.setText(t("{name} has not been downloaded yet.", name=name))
@@ -558,8 +644,13 @@ class SettingsWindow(QDialog):
return self._sources
_models_loaded = pyqtSignal(list, str)
_gemini_models_loaded = pyqtSignal(list, str)
_opencode_models_loaded = pyqtSignal(list, str)
_transcribe_models_loaded = pyqtSignal(list, str)
_codex_models_loaded = pyqtSignal(list)
_agy_models_loaded = pyqtSignal(list)
# Which hosted provider's list arrived on its own at open, and the list.
_hosted_models_loaded = pyqtSignal(str, list)
# Which key was tested, whether it worked, and what to write under it.
_test_done = pyqtSignal(str, bool, str)
# The release that was found, or None, and what went wrong instead.
@@ -594,6 +685,7 @@ class SettingsWindow(QDialog):
tabs = self.tabs = QTabWidget(self)
tabs.addTab(self._scrolled(self._general_tab()), t("General"))
tabs.addTab(self._scrolled(self._display_tab()), t("Display"))
self.api_tab_index = tabs.addTab(
self._scrolled(self._api_tab()), t("API and models"))
tabs.addTab(self._scrolled(self._prompt_tab()), t("Cleanup rules"))
@@ -617,8 +709,12 @@ class SettingsWindow(QDialog):
self._size_to_screen(680, 640)
self._models_loaded.connect(self._on_models_loaded)
self._gemini_models_loaded.connect(self._on_gemini_models_loaded)
self._transcribe_models_loaded.connect(self._on_transcribe_models_loaded)
self._codex_models_loaded.connect(self._on_codex_models_loaded)
self._opencode_models_loaded.connect(self._on_opencode_models_loaded)
self._agy_models_loaded.connect(self._on_agy_models_loaded)
self._hosted_models_loaded.connect(self._on_hosted_models_loaded)
self._test_done.connect(self._on_test_done)
self._update_checked.connect(self._on_update_checked)
self.transcriber.progress.connect(self._on_file_progress)
@@ -630,6 +726,8 @@ class SettingsWindow(QDialog):
self.meetings.failed.connect(self._on_minutes_failed)
self._load()
self._load_codex_models()
self._load_agy_models()
self._load_hosted_models()
# Connected after the load, so that filling the boxes in is not taken
# for the user ticking them.
self.file_timestamps.toggled.connect(self._remember_file_choices)
@@ -714,11 +812,6 @@ class SettingsWindow(QDialog):
self.restore_clipboard = QCheckBox(t("Restore the previous clipboard after pasting"))
form.addRow("", self.restore_clipboard)
self.corner = QComboBox()
for value in CORNERS:
self.corner.addItem(t(value), value)
form.addRow(t("Indicator corner"), self.corner)
self.max_seconds = QSpinBox()
self.max_seconds.setRange(10, 3600)
self.max_seconds.setSuffix(t(" s"))
@@ -769,6 +862,49 @@ class SettingsWindow(QDialog):
self.update_now))
return page
def _display_tab(self):
page = QWidget()
form = QFormLayout(page)
self.indicator_screen = QComboBox()
# The active screen rather than the pointer, for the reason in
# overlay._compositor_screen: it is what a compositor will answer for,
# and on Plasma the two are one screen only where the active screen is
# set to follow the mouse.
self.indicator_screen.addItem(t("Follow the active screen"), "")
for screen in QGuiApplication.screens():
# The native resolution, so that a scaled 4K screen reads
# 3840 × 2160 and not the 1920 × 1080 Qt sees through the scale.
area = screen.geometry()
ratio = screen.devicePixelRatio()
self.indicator_screen.addItem(
t("{name} ({width} × {height})", name=screen.name(),
width=round(area.width() * ratio),
height=round(area.height() * ratio)),
screen.name(),
)
form.addRow(t("Indicator screen"), self.indicator_screen)
# Only the screen it appeared on is decided when it appears; this is
# what makes it keep up with a session that moves to another one
# mid-recording. The active screen and not the pointer, because that is
# what a compositor will answer for: on Plasma the two are the same
# screen only where the active screen is set to follow the mouse, and
# otherwise it is the focused window that decides. Nothing to offer
# when a screen is named above, since that name is the whole answer.
self.follow_pointer = QCheckBox(t("Move it when the active screen changes"))
self.indicator_screen.currentIndexChanged.connect(self._sync_follow_pointer)
form.addRow("", self.follow_pointer)
self.corner = QComboBox()
for value in CORNERS:
self.corner.addItem(t(value), value)
form.addRow(t("Indicator corner"), self.corner)
return page
def _sync_follow_pointer(self):
self.follow_pointer.setEnabled(not self.indicator_screen.currentData())
def _api_tab(self):
page = QWidget()
outer = QVBoxLayout(page)
@@ -786,6 +922,12 @@ class SettingsWindow(QDialog):
self.openrouter_key = self._key_row(
keys_form, "openrouter", t("sk-or-… (falls back to OPENROUTER_API_KEY)"),
self._test_openrouter)
self.gemini_key = self._key_row(
keys_form, "gemini", t("(falls back to GEMINI_API_KEY)"),
self._test_gemini, service="Google AI Studio")
self.opencode_key = self._key_row(
keys_form, "opencode", t("(falls back to OPENCODE_API_KEY)"),
self._test_opencode, service="OpenCode Go")
outer.addWidget(keys)
stt = QGroupBox(t("Speech to text"))
@@ -806,6 +948,15 @@ class SettingsWindow(QDialog):
self.transcribe_model_row = self._row(self.transcribe_model,
self.refresh_transcribe_models)
stt_form.addRow(t("Model"), self.transcribe_model_row)
# OpenRouter only: which of its models a timestamped run asks for.
self.file_model = QComboBox()
self.file_model.setEditable(True)
self.file_model.lineEdit().setPlaceholderText(api.OPENROUTER_FILE_MODEL)
self.file_model.setToolTip(
t("The model a timestamped audio file (subtitles) is sent to. Not every model "
"on OpenRouter returns segment times; empty means openai/whisper-1."))
self.file_model_row = self._row(self.file_model)
stt_form.addRow(t("Audio file model"), self.file_model_row)
# A spanning row: in the narrow field column a wrapped label gets a
# height that fits one line, and the rest of the text is cut off.
self.transcribe_status = QLabel("")
@@ -857,13 +1008,6 @@ class SettingsWindow(QDialog):
self.cleanup_provider = QComboBox()
for label, value in CLEANUP_PROVIDERS:
self.cleanup_provider.addItem(t(label), value)
self.cleanup_provider.setToolTip(t(
"OpenRouter is the quickest and the only one that needs nothing "
"installed. llama.cpp runs here, on a model downloaded below. Claude "
"Code and Codex clean up on the subscription you already have, "
"without a second key, and take a few seconds longer because each "
"one opens a session to do it."
))
self.cleanup_provider.currentIndexChanged.connect(self._cleanup_provider_changed)
orr_form.addRow(t("Runs on"), self.cleanup_provider)
@@ -875,6 +1019,15 @@ class SettingsWindow(QDialog):
self.cleanup_model_row = self._row(self.cleanup_model, self.refresh_models)
orr_form.addRow(t("Model"), self.cleanup_model_row)
self.cleanup_gemini_model = QComboBox()
self.cleanup_gemini_model.setEditable(True)
self.cleanup_gemini_model.addItems(GEMINI_MODELS)
self.refresh_gemini_models = QPushButton(t("Fetch model list"))
self.refresh_gemini_models.clicked.connect(self._load_gemini_models)
self.cleanup_gemini_model_row = self._row(
self.cleanup_gemini_model, self.refresh_gemini_models)
orr_form.addRow(t("Model"), self.cleanup_gemini_model_row)
# One row per provider rather than one box that means a different thing
# in each: an OpenRouter id and a Claude alias do not belong in the same
# field, and only the row of whoever is chosen is on screen.
@@ -890,6 +1043,21 @@ class SettingsWindow(QDialog):
self.cleanup_codex_model.setToolTip(_typed_model_note("Codex"))
orr_form.addRow(t("Model"), self.cleanup_codex_model)
self.cleanup_opencode_model = QComboBox()
self.cleanup_opencode_model.setEditable(True)
self.cleanup_opencode_model.addItems(OPENCODE_MODELS)
self.cleanup_opencode_model.setToolTip(_typed_model_note("OpenCode Go"))
self.refresh_opencode_models = QPushButton(t("Fetch model list"))
self.refresh_opencode_models.clicked.connect(self._load_opencode_models)
self.cleanup_opencode_model_row = self._row(self.cleanup_opencode_model,
self.refresh_opencode_models)
orr_form.addRow(t("Model"), self.cleanup_opencode_model_row)
self.cleanup_agy_model = QComboBox()
self.cleanup_agy_model.setEditable(True)
self.cleanup_agy_model.addItems([t("Antigravity's own default")] + AGY_MODELS)
orr_form.addRow(t("Model"), self.cleanup_agy_model)
self.cleanup_reasoning = QComboBox()
for label, value in REASONING_LEVELS:
self.cleanup_reasoning.addItem(t(label), value)
@@ -969,17 +1137,6 @@ class SettingsWindow(QDialog):
def _assistant_tab(self):
page = QWidget()
layout = QVBoxLayout(page)
intro = QLabel(t(
"This shortcut records the same way dictation does, but the "
"transcript is not what gets pasted. It goes to an agent as a "
"command, and what comes back is pasted instead: the answer to a "
"question, or a sentence saying what was done. Claude Code and "
"Codex run as the session you would have opened yourself, with your "
"skills, your connected services and your account."
))
intro.setWordWrap(True)
layout.addWidget(intro)
self.assistant_found = QLabel("")
self.assistant_found.setWordWrap(True)
layout.addWidget(self.assistant_found)
@@ -1013,7 +1170,7 @@ class SettingsWindow(QDialog):
dir_note.setWordWrap(True)
how_form.addRow(dir_note)
# One scale for all three: how hard to think is one thing to want, and
# One scale for all four: how hard to think is one thing to want, and
# each provider is handed the nearest rung it actually has.
self.assistant_reasoning = QComboBox()
for label, value in REASONING_LEVELS:
@@ -1036,7 +1193,7 @@ class SettingsWindow(QDialog):
layout.addWidget(how)
# One box per provider, only the chosen one on screen: they have nothing
# in common past the model, and three sets of half-relevant fields would
# in common past the model, and four sets of half-relevant fields would
# be worse than none.
self.claude_box = QGroupBox(t("Claude Code"))
claude_form = QFormLayout(self.claude_box)
@@ -1087,6 +1244,41 @@ class SettingsWindow(QDialog):
or_form.addRow(or_note)
layout.addWidget(self.openrouter_box)
self.agy_box = QGroupBox(t("Antigravity"))
agy_form = QFormLayout(self.agy_box)
self.assistant_agy_model = QComboBox()
self.assistant_agy_model.setEditable(True)
self.assistant_agy_model.addItem(t("Antigravity's own default"), "")
for name in AGY_MODELS:
self.assistant_agy_model.addItem(name, name)
agy_form.addRow(t("Model"), self.assistant_agy_model)
agy_note = QLabel(t(
"Antigravity has neither a permission mode nor a sandbox to hand it, "
"so what it may do without asking is whatever its own allow-rules "
"say. The Permissions and Sandbox boxes above belong to the other "
"two; the working directory still applies."
))
agy_note.setWordWrap(True)
agy_form.addRow(agy_note)
layout.addWidget(self.agy_box)
self.opencode_box = QGroupBox("OpenCode Go")
og_form = QFormLayout(self.opencode_box)
self.assistant_opencode_model = QComboBox()
self.assistant_opencode_model.setEditable(True)
self.assistant_opencode_model.addItems(OPENCODE_MODELS)
og_form.addRow(t("Model"), self.assistant_opencode_model)
og_note = QLabel(t(
"A plain question and a plain answer, over the OpenCode Go key you "
"already have. It runs no commands, opens no files and reaches none "
"of your services, so it can tell you what the capital of Peru is "
"but not what is in your calendar. Working directory and permissions "
"above mean nothing here."
))
og_note.setWordWrap(True)
og_form.addRow(og_note)
layout.addWidget(self.opencode_box)
thread = QGroupBox(t("The conversation"))
thread_form = QFormLayout(thread)
self.assistant_session_minutes = QSpinBox()
@@ -1142,14 +1334,6 @@ class SettingsWindow(QDialog):
def _meeting_tab(self):
page = QWidget()
layout = QVBoxLayout(page)
intro = QLabel(t(
"A meeting is recorded from two devices at once: your microphone and "
"whatever comes out of your speakers. Nothing has to guess who was "
"speaking, because the two never share a channel."
))
intro.setWordWrap(True)
layout.addWidget(intro)
sources = QGroupBox(t("Sound"))
sources_form = QFormLayout(sources)
self.meeting_mic = QComboBox()
@@ -1558,7 +1742,7 @@ class SettingsWindow(QDialog):
box.lineEdit().setPlaceholderText(placeholder)
return box
def _key_row(self, form, provider, placeholder, tester):
def _key_row(self, form, provider, placeholder, tester, service=""):
"""A key field, its Test button and the line the answer lands on.
The field and the pair the answer needs are filed under the provider's
@@ -1572,7 +1756,8 @@ class SettingsWindow(QDialog):
button.clicked.connect(tester)
answer = QLabel("")
answer.setWordWrap(True)
form.addRow(cfg.TRANSCRIBERS[provider].service, self._row(field, button))
label = service or cfg.TRANSCRIBERS[provider].service
form.addRow(label, self._row(field, button))
form.addRow("", answer)
self._key_fields[provider] = field
self._testers[provider] = (button, answer)
@@ -1634,6 +1819,12 @@ class SettingsWindow(QDialog):
self.auto_paste.setChecked(conf["auto_paste"])
self.paste_shortcut.setCurrentText(conf["paste_shortcut"])
self.restore_clipboard.setChecked(conf["restore_clipboard"])
screen_name = conf["overlay_screen"]
if screen_name and self.indicator_screen.findData(screen_name) < 0:
self.indicator_screen.addItem(t("{name} (not connected)", name=screen_name), screen_name)
self._select_data(self.indicator_screen, screen_name)
self.follow_pointer.setChecked(conf["overlay_follows_pointer"])
self._sync_follow_pointer()
self._select_data(self.corner, conf["overlay_corner"])
self.max_seconds.setValue(conf["max_seconds"])
self.skip_silent.setChecked(conf["skip_silent"])
@@ -1646,9 +1837,12 @@ class SettingsWindow(QDialog):
for name, who in cfg.TRANSCRIBERS.items():
self._key_fields[name].setText(conf[who.key])
self._models[name] = conf[who.model]
self.gemini_key.setText(conf["gemini_api_key"])
self.opencode_key.setText(conf["opencode_api_key"])
self._shown_provider = ""
self._select_data(self.transcribe_provider, conf["transcribe_provider"])
self._provider_changed() # selecting index 0 fires no signal
self.file_model.setCurrentText(conf["openrouter_file_model"])
self.local_gpu.setChecked(conf["local_gpu"])
self.local_preload.setChecked(conf["local_preload"])
self.local_threads.setValue(int(conf["local_threads"]))
@@ -1656,10 +1850,17 @@ class SettingsWindow(QDialog):
self.cleanup_enabled.setChecked(conf["cleanup_enabled"])
self.cleanup_model.setCurrentText(conf["cleanup_model"])
self.cleanup_gemini_model.setCurrentText(
conf["cleanup_gemini_model"] or cfg.DEFAULTS["cleanup_gemini_model"]
)
self.cleanup_claude_model.setCurrentText(conf["cleanup_claude_model"])
self.cleanup_codex_model.setCurrentText(
conf["cleanup_codex_model"] or t("Codex's own default")
)
self.cleanup_agy_model.setCurrentText(
conf["cleanup_agy_model"] or t("Antigravity's own default")
)
self.cleanup_opencode_model.setCurrentText(conf["cleanup_opencode_model"])
self._select_data(self.cleanup_provider, conf["cleanup_provider"])
self._cleanup_provider_changed() # selecting index 0 fires no signal
self._select_data(self.cleanup_reasoning, conf["cleanup_reasoning"])
@@ -1690,6 +1891,8 @@ class SettingsWindow(QDialog):
self.assistant_codex_model.setCurrentText(conf["assistant_codex_model"])
self._select_data(self.assistant_codex_sandbox, conf["assistant_codex_sandbox"])
self.assistant_openrouter_model.setCurrentText(conf["assistant_openrouter_model"])
self.assistant_agy_model.setCurrentText(conf["assistant_agy_model"])
self.assistant_opencode_model.setCurrentText(conf["assistant_opencode_model"])
self._assistant_provider_changed() # selecting index 0 fires no signal
self._select_data(self.assistant_reasoning, conf["assistant_reasoning"])
self.assistant_dir.setText(conf["assistant_dir"])
@@ -1741,6 +1944,10 @@ class SettingsWindow(QDialog):
conf["auto_paste"] = self.auto_paste.isChecked()
conf["paste_shortcut"] = self.paste_shortcut.currentText().strip()
conf["restore_clipboard"] = self.restore_clipboard.isChecked()
conf["overlay_screen"] = self.indicator_screen.currentData() or ""
# Read even while it is greyed out, so that naming a screen and taking
# the name back again does not clear a preference nobody touched.
conf["overlay_follows_pointer"] = self.follow_pointer.isChecked()
conf["overlay_corner"] = self.corner.currentData() or "bottom-left"
conf["max_seconds"] = self.max_seconds.value()
conf["skip_silent"] = self.skip_silent.isChecked()
@@ -1756,6 +1963,9 @@ class SettingsWindow(QDialog):
for name, who in cfg.TRANSCRIBERS.items():
conf[who.key] = self._key_fields[name].text().strip()
conf[who.model] = self._models[name].strip() or cfg.DEFAULTS[who.model]
conf["openrouter_file_model"] = self.file_model.currentText().strip()
conf["gemini_api_key"] = self.gemini_key.text().strip()
conf["opencode_api_key"] = self.opencode_key.text().strip()
conf["local_model"] = self.local_whisper.selected()
conf["local_gpu"] = self.local_gpu.isChecked()
conf["local_preload"] = self.local_preload.isChecked()
@@ -1764,12 +1974,25 @@ class SettingsWindow(QDialog):
conf["cleanup_enabled"] = self.cleanup_enabled.isChecked()
conf["cleanup_provider"] = self.cleanup_provider.currentData() or "openrouter"
conf["cleanup_model"] = self.cleanup_model.currentText().strip()
conf["cleanup_gemini_model"] = (
self.cleanup_gemini_model.currentText().strip()
or cfg.DEFAULTS["cleanup_gemini_model"]
)
conf["cleanup_claude_model"] = (self.cleanup_claude_model.currentText().strip()
or cfg.DEFAULTS["cleanup_claude_model"])
codex_cleanup_model = self.cleanup_codex_model.currentText().strip()
conf["cleanup_codex_model"] = (
"" if codex_cleanup_model == t("Codex's own default") else codex_cleanup_model
)
agy_cleanup_model = self.cleanup_agy_model.currentText().strip()
conf["cleanup_agy_model"] = (
"" if agy_cleanup_model == t("Antigravity's own default")
else agy_cleanup_model
)
conf["cleanup_opencode_model"] = (
self.cleanup_opencode_model.currentText().strip()
or cfg.DEFAULTS["cleanup_opencode_model"]
)
conf["cleanup_reasoning"] = self.cleanup_reasoning.currentData() or ""
conf["local_llm_model"] = self.local_llm.selected()
conf["local_llm_repo"] = self.local_llm.repository()
@@ -1808,6 +2031,14 @@ class SettingsWindow(QDialog):
self.assistant_openrouter_model.currentText().strip()
or cfg.DEFAULTS["assistant_openrouter_model"]
)
agy_model = self.assistant_agy_model.currentText().strip()
conf["assistant_agy_model"] = (
"" if agy_model == t("Antigravity's own default") else agy_model
)
conf["assistant_opencode_model"] = (
self.assistant_opencode_model.currentText().strip()
or cfg.DEFAULTS["assistant_opencode_model"]
)
conf["assistant_reasoning"] = self.assistant_reasoning.currentData() or ""
conf["assistant_dir"] = self.assistant_dir.text().strip()
conf["assistant_timeout"] = self.assistant_timeout.value()
@@ -1906,6 +2137,7 @@ class SettingsWindow(QDialog):
self._shown_provider = provider
local = provider == "local"
self.stt_form.setRowVisible(self.transcribe_model_row, not local)
self.stt_form.setRowVisible(self.file_model_row, provider == "openrouter")
self.stt_form.setRowVisible(self.transcribe_status, not local)
self.stt_form.setRowVisible(self.local_whisper, local)
self.stt_form.setRowVisible(self.local_options, local)
@@ -1914,8 +2146,16 @@ class SettingsWindow(QDialog):
self.transcribe_model.clear()
self.transcribe_model.addItems(TRANSCRIBE_MODELS[provider])
self.transcribe_model.setCurrentText(self._models[provider])
if provider == "openrouter":
self._fill_file_models(TRANSCRIBE_MODELS[provider])
self.transcribe_status.setText("")
def _fill_file_models(self, models):
current = self.file_model.currentText()
self.file_model.clear()
self.file_model.addItems(models)
self.file_model.setCurrentText(current)
def _load_transcribe_models(self):
"""The model list of whichever provider is selected."""
provider = self.transcribe_provider.currentData() or "openai"
@@ -1944,6 +2184,8 @@ class SettingsWindow(QDialog):
self.transcribe_model.clear()
self.transcribe_model.addItems(models)
self.transcribe_model.setCurrentText(current)
if self._shown_provider == "openrouter":
self._fill_file_models(models)
self.transcribe_status.setText(t("{count} models loaded.", count=len(models)))
def _load_models(self):
@@ -1971,6 +2213,30 @@ class SettingsWindow(QDialog):
combo.setCurrentText(current)
self.models_label.setText(t("{count} models loaded.", count=len(models)))
def _load_gemini_models(self):
self.refresh_gemini_models.setEnabled(False)
self.models_label.setText(t("Fetching model list…"))
key, base = self._typed_key("gemini")
def work():
try:
self._gemini_models_loaded.emit(api.gemini_models(key, base), "")
except api.ApiError as exc:
self._gemini_models_loaded.emit([], str(exc))
threading.Thread(target=work, daemon=True).start()
def _on_gemini_models_loaded(self, models, error):
self.refresh_gemini_models.setEnabled(True)
if error:
self.models_label.setText(t("Could not fetch the list: {error}", error=error))
return
current = self.cleanup_gemini_model.currentText()
self.cleanup_gemini_model.clear()
self.cleanup_gemini_model.addItems(models)
self.cleanup_gemini_model.setCurrentText(current)
self.models_label.setText(t("{count} models loaded.", count=len(models)))
def _load_codex_models(self):
"""Ask Codex which models it offers, off the interface thread.
@@ -1998,6 +2264,112 @@ class SettingsWindow(QDialog):
combo.addItem(name, name)
combo.setCurrentText(current)
def _load_opencode_models(self):
self.refresh_opencode_models.setEnabled(False)
self.models_label.setText(t("Fetching model list…"))
key, base = self._typed_key("opencode")
def work():
try:
self._opencode_models_loaded.emit(
api.openai_models(key, base, "OpenCode Go"), "")
except api.ApiError as exc:
self._opencode_models_loaded.emit([], str(exc))
threading.Thread(target=work, daemon=True).start()
def _on_opencode_models_loaded(self, models, error):
self.refresh_opencode_models.setEnabled(True)
if error:
self.models_label.setText(t("Could not fetch the list: {error}", error=error))
return
self._fill_opencode_boxes(models)
self.models_label.setText(t("{count} models loaded.", count=len(models)))
def _fill_opencode_boxes(self, models):
# The agent runs on the same key and catalog, so its box is refilled
# from the same list.
for combo in (self.cleanup_opencode_model, self.assistant_opencode_model):
current = combo.currentText()
combo.clear()
combo.addItems(models)
combo.setCurrentText(current)
def _load_agy_models(self):
"""Ask Antigravity which models it offers, off the interface thread.
The same arrangement as Codex, except agy answers over the network
rather than from a cache, so the couple of seconds it takes are spent
where nobody is waiting. Skipped when agy is not installed, which is
also when the built-in list stays on screen and nobody is running
Antigravity anyway.
"""
if not shutil.which("agy"):
return
def work():
found = assistant.agy_models()
if found:
self._agy_models_loaded.emit(found)
threading.Thread(target=work, daemon=True).start()
def _on_agy_models_loaded(self, models):
for combo in (self.cleanup_agy_model, self.assistant_agy_model):
current = combo.currentText()
combo.clear()
combo.addItem(t("Antigravity's own default"), "")
for name in models:
combo.addItem(name, name)
combo.setCurrentText(current)
def _load_hosted_models(self):
"""Fetch the hosted model lists at open, without being asked.
The Fetch buttons stay: they are the retry, and the place a failure is
worth explaining. Here nobody asked, so an error changes nothing on
screen and the built-in lists remain, and a provider whose key has not
been given yet is not called at all.
"""
jobs = []
openrouter_key = self.conf.openrouter_key()
if openrouter_key:
jobs.append(("openrouter",
lambda: api.openrouter_models(openrouter_key)))
gemini_key = self.conf.gemini_key()
gemini_base = self.conf["gemini_base_url"]
if gemini_key:
jobs.append(("gemini",
lambda: api.gemini_models(gemini_key, gemini_base)))
opencode_key = self.conf.opencode_key()
opencode_base = self.conf["opencode_base_url"]
if opencode_key:
jobs.append(("opencode",
lambda: api.openai_models(opencode_key, opencode_base,
"OpenCode Go")))
for provider, fetch in jobs:
def work(provider=provider, fetch=fetch):
try:
found = fetch()
except api.ApiError:
return
if found:
self._hosted_models_loaded.emit(provider, found)
threading.Thread(target=work, daemon=True).start()
def _on_hosted_models_loaded(self, provider, models):
if provider == "opencode":
self._fill_opencode_boxes(models)
return
combos = ((self.cleanup_model, self.meeting_model)
if provider == "openrouter" else (self.cleanup_gemini_model,))
for combo in combos:
current = combo.currentText()
combo.clear()
combo.addItems(models)
combo.setCurrentText(current)
def _test_openai(self):
key, base = self._typed_key("openai")
self._test_key("openai", lambda: t(
@@ -2016,11 +2388,30 @@ class SettingsWindow(QDialog):
key, _ = self._typed_key("openrouter")
self._test_key("openrouter", lambda: api.openrouter_key_status(key))
def _test_gemini(self):
key, base = self._typed_key("gemini")
self._test_key("gemini", lambda: t(
"Connection works. {count} models visible.",
count=len(api.gemini_models(key, base)),
))
def _test_opencode(self):
key, base = self._typed_key("opencode")
self._test_key("opencode", lambda: t(
"Connection works. {count} models visible.",
count=len(api.openai_models(key, base, "OpenCode Go")),
))
def _typed_key(self, provider):
"""(key, base URL) for a provider, preferring what is in the field now."""
who = cfg.TRANSCRIBERS[provider]
if provider in cfg.TRANSCRIBERS:
who = cfg.TRANSCRIBERS[provider]
key_setting, url_setting = who.key, who.url
else:
key_setting = f"{provider}_api_key"
url_setting = f"{provider}_base_url"
typed = self._key_fields[provider].text().strip()
return typed or self.conf.api_key(who.key), self.conf[who.url]
return typed or self.conf.api_key(key_setting), self.conf[url_setting]
def _test_key(self, provider, ask):
"""Run `ask` off the interface thread and write its answer under the key.
@@ -2244,10 +2635,16 @@ class SettingsWindow(QDialog):
provider = self.cleanup_provider.currentData() or "openrouter"
self.cleanup_form.setRowVisible(self.cleanup_model_row,
provider == "openrouter")
self.cleanup_form.setRowVisible(self.cleanup_gemini_model_row,
provider == "gemini")
self.cleanup_form.setRowVisible(self.cleanup_claude_model,
provider == "claude")
self.cleanup_form.setRowVisible(self.cleanup_codex_model,
provider == "codex")
self.cleanup_form.setRowVisible(self.cleanup_opencode_model_row,
provider == "opencode")
self.cleanup_form.setRowVisible(self.cleanup_agy_model,
provider == "agy")
self.cleanup_form.setRowVisible(self.cleanup_reasoning,
provider != "local")
self.cleanup_form.setRowVisible(self.local_llm, provider == "local")
@@ -2256,6 +2653,10 @@ class SettingsWindow(QDialog):
found = shutil.which(binary) if binary else ""
if provider == "local":
self.models_label.setText(t("Runs on this machine, on llama.cpp."))
elif provider == "gemini":
self.models_label.setText(t("Runs on Google AI Studio."))
elif provider == "opencode":
self.models_label.setText(t("Runs on OpenCode Go."))
elif not binary:
self.models_label.setText(t("Runs on OpenRouter."))
elif found:
@@ -2272,6 +2673,8 @@ class SettingsWindow(QDialog):
self.claude_box.setVisible(provider == "claude")
self.codex_box.setVisible(provider == "codex")
self.openrouter_box.setVisible(provider == "openrouter")
self.agy_box.setVisible(provider == "agy")
self.opencode_box.setVisible(provider == "opencode")
self._refresh_assistant_status()
def _refresh_assistant_status(self):
@@ -2279,9 +2682,14 @@ class SettingsWindow(QDialog):
binary = assistant.executable(provider)
found = shutil.which(binary) if binary else ""
if not binary:
self.assistant_found.setText(
t("Needs no program installed, only the OpenRouter key.")
)
if provider == "opencode":
self.assistant_found.setText(
t("Needs no program installed, only an OpenCode Go key.")
)
else:
self.assistant_found.setText(
t("Needs no program installed, only the OpenRouter key.")
)
elif found:
self.assistant_found.setText(t("Found: {path}", path=found))
else:
@@ -2400,7 +2808,11 @@ class SettingsWindow(QDialog):
# The text of an answer says nothing about what was asked, and
# out of that context half of them read like non sequiturs.
asked = (row.get("question") or row.get("raw") or "").replace("\n", " ")
header += t(" · asked Claude: {question}",
# Rows written before the provider was recorded are all Claude's,
# because it was the only one the history could name.
who = assistant.SERVICES.get(row.get("assistant"), "Claude")
header += t(" · asked {who}: {question}",
who=i18n.name(who, "dative"),
question=asked[:60] + ("" if len(asked) > 60 else ""))
item = QListWidgetItem(f"{header}\n{preview}")
item.setData(Qt.ItemDataRole.UserRole, row)
+2 -1
View File
@@ -181,7 +181,8 @@ class Pipeline(QObject):
"cleanup_error": warning,
"mode": "ask" if ask else "",
"question": question,
"assistant_model": conf["assistant_model"] if ask else "",
"assistant": assistant.provider(conf) if ask else "",
"assistant_model": assistant.model(conf) if ask else "",
"raw": raw,
"text": text,
}
+44
View File
@@ -0,0 +1,44 @@
FROM ubuntu@sha256:2edbbc5dc405e9612ba3584ce95480277e3eb374407b5505fe26f17df77c7dbc
ARG DEBIAN_FRONTEND=noninteractive
ARG CMAKE_VERSION=3.31.6
ARG CMAKE_SHA256=5a1133ff103c71eb5120e2cc3de922733e7d8a26a98ae716397e8676adb367bf
COPY lunarg-signing-key-pub.asc /tmp/lunarg.asc
RUN set -eux; \
test "$(sha256sum /tmp/lunarg.asc | cut -d' ' -f1)" = aa1c3c29673140e77f0d6a9aaeed5d9b5621e305ead51c59fae4458bbb4df92b; \
apt-get update; \
apt-get install --no-install-recommends -y \
build-essential=12.9ubuntu3 \
ca-certificates \
curl \
file \
git \
gnupg \
ninja-build=1.10.1-1 \
patchelf=0.14.3-1 \
python3 \
xz-utils; \
install -d -m 0755 /usr/share/keyrings; \
gpg --dearmor -o /usr/share/keyrings/lunarg.gpg /tmp/lunarg.asc; \
printf '%s\n' 'deb [signed-by=/usr/share/keyrings/lunarg.gpg] https://packages.lunarg.com/vulkan jammy main' \
> /etc/apt/sources.list.d/lunarg-vulkan.list; \
apt-get update; \
apt-get install --no-install-recommends -y \
libvulkan-dev=1.4.313.0~rc1-1lunarg22.04-1 \
vulkan-headers=1.4.313.0~rc1-1lunarg22.04-1 \
shaderc=2025.2~rc1-1lunarg22.04-1 \
spirv-headers=1.6.1+1.4.313.0~rc1-1lunarg22.04-1; \
curl --fail --location --retry 3 \
"https://github.com/Kitware/CMake/releases/download/v${CMAKE_VERSION}/cmake-${CMAKE_VERSION}-linux-x86_64.tar.gz" \
-o /tmp/cmake.tar.gz; \
test "$(sha256sum /tmp/cmake.tar.gz | cut -d' ' -f1)" = "$CMAKE_SHA256"; \
tar -xzf /tmp/cmake.tar.gz --strip-components=1 -C /usr/local; \
rm -rf /var/lib/apt/lists/* /tmp/cmake.tar.gz /tmp/lunarg.asc; \
cmake --version; \
glslc --version; \
test -f /usr/include/vulkan/vulkan.h; \
test -f /usr/share/cmake/SPIRV-Headers/SPIRV-HeadersConfig.cmake
WORKDIR /work
@@ -0,0 +1,6 @@
FROM ubuntu@sha256:2edbbc5dc405e9612ba3584ce95480277e3eb374407b5505fe26f17df77c7dbc
ARG DEBIAN_FRONTEND=noninteractive
RUN apt-get update \
&& apt-get install --no-install-recommends -y ca-certificates curl libstdc++6 \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /bundle
@@ -0,0 +1,9 @@
FROM ubuntu@sha256:2edbbc5dc405e9612ba3584ce95480277e3eb374407b5505fe26f17df77c7dbc
ARG DEBIAN_FRONTEND=noninteractive
# The loader and nothing behind it: the machine that has libvulkan because
# something else pulled it in, and no driver to go with it.
RUN apt-get update \
&& apt-get install --no-install-recommends -y \
ca-certificates curl libstdc++6 libvulkan1 \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /bundle
@@ -0,0 +1,7 @@
FROM ubuntu@sha256:2edbbc5dc405e9612ba3584ce95480277e3eb374407b5505fe26f17df77c7dbc
ARG DEBIAN_FRONTEND=noninteractive
RUN apt-get update \
&& apt-get install --no-install-recommends -y \
ca-certificates curl libstdc++6 libvulkan1 mesa-vulkan-drivers vulkan-tools \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /bundle
+49
View File
@@ -0,0 +1,49 @@
# The Vulkan whisper-server bundle
whisper.cpp publishes a CPU-only archive for Linux, so the graphics card on a
Linux machine is out of reach through the Download button. This directory
builds the archive upstream does not: `whisper-server` with a dynamic Vulkan
backend next to the CPU ones, for x86_64, against the Ubuntu 22.04 runtime
contract.
It is published as a release of Dikte's own, `whisper.cpp-v<version>`, marked
as a prerelease and kept off Latest so that neither the update check nor the
download page picks it up. `dikte/ggml.py` fetches it by tag and installs it
only when the archive's digest is the reviewed one; anything else falls back
to upstream's CPU archive, and the settings window says when it did.
## Publishing a new bundle
1. Enable GitHub's immutable releases setting for the repository, and give the
`dependency-release` environment a required reviewer. Both are repository
settings, not something this workflow can do for itself.
2. Run **whisper.cpp Vulkan bundle** on `master` with the new version and its
peeled commit, `expected_sha256` empty and `publish: false`. The run builds
the archive and reports its digest; without a reviewed digest it refuses to
publish, which is what the first run is for.
3. Review that digest against a build of your own, then run the workflow again
with the same version and commit, `expected_sha256` set to it, and
`publish: true`. Approve the environment when it asks.
4. Write the same version, tag and digest into `MANAGED_WHISPER_RELEASE`,
`MANAGED_WHISPER_VERSION` and `MANAGED_WHISPER_SHA256` in `dikte/ggml.py`,
and into `REVIEWED_WHISPER_VERSION` and `REVIEWED_WHISPER_SHA256` in the
workflow. `tests/test_packaging.py` holds the two sides together.
5. Ship a Dikte release. Until one goes out, nobody's Dikte knows the new
bundle exists.
## What this costs, and what it does not promise
The digest lives in Dikte's source, so a backend update is a Dikte release.
Linux x86_64 machines with a Vulkan loader stay on the pinned whisper.cpp
version until step 5 happens, while every other platform follows upstream's
newest release on its own. That is the deliberate trade: an executable Dikte
downloads is not allowed to change without a reviewed digest behind it.
The build is deterministic between two runs of the same builder, not across
time. The base image, the CMake tarball, the LunarG packages and the direct
apt packages are pinned by digest or version, but the Ubuntu and LunarG
repository metadata behind them is not, and LunarG drops superseded packages.
A rebuild months later can fail to resolve, or resolve to something that
produces a different digest. Treat the published archive as the artifact, not
as something reproducible on demand: a version bump means building,
validating, reviewing the new digest and updating the pinned tuple together.
+128
View File
@@ -0,0 +1,128 @@
#!/usr/bin/env bash
set -euo pipefail
shopt -s nullglob
: "${SOURCE_DIR:=/src}"
: "${OUT_DIR:=/work/out}"
: "${WHISPER_VERSION:=1.9.3}"
: "${WHISPER_COMMIT:=371b5a7561823ab2bb32142d2751e35e7534727b}"
: "${SOURCE_DATE_EPOCH:=1787219223}"
export SOURCE_DATE_EPOCH TZ=UTC LC_ALL=C LANG=C
asset=whisper-bin-ubuntu-vulkan-x64
build=/work/build
source_copy=/work/source
root="$OUT_DIR/root/$asset"
rm -rf "$build" "$source_copy" "$OUT_DIR"
mkdir -p "$build" "$root/LICENSES"
# Upstream configures bindings/javascript/package.json in the source directory.
# Build a private copy so the checked-out, verified source remains untouched.
cp -a "$SOURCE_DIR" "$source_copy"
chmod -R u+w "$source_copy"
git config --global --add safe.directory "$source_copy"
cmake -S "$source_copy" -B "$build" -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_BUILD_RPATH='$ORIGIN' \
-DCMAKE_INSTALL_RPATH='$ORIGIN' \
-DCMAKE_BUILD_WITH_INSTALL_RPATH=ON \
-DCMAKE_C_FLAGS="-ffile-prefix-map=$source_copy=. -fdebug-prefix-map=$source_copy=. -fmacro-prefix-map=$source_copy=." \
-DCMAKE_CXX_FLAGS="-ffile-prefix-map=$source_copy=. -fdebug-prefix-map=$source_copy=. -fmacro-prefix-map=$source_copy=." \
-DBUILD_SHARED_LIBS=ON \
-DGGML_BACKEND_DL=ON \
-DGGML_CPU_ALL_VARIANTS=ON \
-DGGML_NATIVE=OFF \
-DGGML_CCACHE=OFF \
-DGGML_OPENMP=OFF \
-DGGML_VULKAN=ON \
-DWHISPER_BUILD_EXAMPLES=ON \
-DWHISPER_BUILD_SERVER=ON \
-DWHISPER_BUILD_TESTS=OFF \
-DWHISPER_BUILD_IS_DEV=OFF \
-DWHISPER_CURL=OFF \
-DWHISPER_SDL2=OFF \
-DWHISPER_COMMON_FFMPEG=OFF \
-DWHISPER_BUILD_COMMIT="$WHISPER_COMMIT" \
-DWHISPER_BUILD_NUMBER=0
cmake --build "$build" --target whisper-server --parallel "$(nproc)"
# Package an allowlist, not everything examples/ happens to build in the future.
cp -a "$build/bin/whisper-server" "$root/"
cp -a "$build/bin"/libwhisper.so* "$root/"
cp -a "$build/bin"/libggml.so* "$root/"
cp -a "$build/bin"/libggml-base.so* "$root/"
cp -a "$build/bin"/libggml-cpu*.so* "$root/"
cp -a "$build/bin"/libggml-vulkan.so* "$root/"
# Strip real ELF files only; preserve the SONAME symlink chains.
while IFS= read -r -d '' file; do
if file "$file" | grep -q ELF; then
strip --strip-unneeded "$file"
patchelf --set-rpath '$ORIGIN' "$file"
fi
done < <(find "$root" -type f -print0)
cp "$SOURCE_DIR/LICENSE" "$root/LICENSES/whisper.cpp-MIT.txt"
cp /packaging/licenses/cpp-httplib-MIT.txt "$root/LICENSES/"
cp /packaging/licenses/nlohmann-json-MIT.txt "$root/LICENSES/"
cat > "$root/BUILD-INFO.json" <<EOF
{
"asset": "$asset.tar.gz",
"source": "https://github.com/ggml-org/whisper.cpp",
"source_version": "v$WHISPER_VERSION",
"source_commit": "$WHISPER_COMMIT",
"source_date_epoch": $SOURCE_DATE_EPOCH,
"build_platform": "ubuntu-22.04-x86_64",
"base_image": "ubuntu@sha256:2edbbc5dc405e9612ba3584ce95480277e3eb374407b5505fe26f17df77c7dbc",
"cmake": "3.31.6",
"cmake_flags": [
"BUILD_SHARED_LIBS=ON",
"C/CXX_FILE_PREFIX_MAP=/work/source=.",
"GGML_BACKEND_DL=ON",
"GGML_CPU_ALL_VARIANTS=ON",
"GGML_NATIVE=OFF",
"GGML_CCACHE=OFF",
"GGML_OPENMP=OFF",
"GGML_VULKAN=ON",
"WHISPER_BUILD_EXAMPLES=ON",
"WHISPER_BUILD_SERVER=ON",
"WHISPER_BUILD_TESTS=OFF",
"WHISPER_BUILD_IS_DEV=OFF",
"WHISPER_CURL=OFF",
"WHISPER_SDL2=OFF",
"WHISPER_COMMON_FFMPEG=OFF"
],
"runtime_contract": {
"minimum_glibc": "2.34",
"minimum_glibcxx": "3.4.30",
"required": ["x86_64 Linux", "glibc", "libstdc++.so.6", "libgcc_s.so.1"],
"optional_gpu": ["libvulkan.so.1", "a working Vulkan ICD"],
"cpu_fallback": "dynamic CPU backends are included; -ng forces CPU"
}
}
EOF
# A deterministic CycloneDX sidecar generated from the files actually shipped.
ROOT="$root" VERSION="$WHISPER_VERSION" COMMIT="$WHISPER_COMMIT" EPOCH="$SOURCE_DATE_EPOCH" \
python3 /packaging/make-sbom.py > "$root/$asset.cdx.json"
(
cd "$root"
find . -type f ! -name SHA256SUMS -print0 \
| sort -z \
| xargs -0 sha256sum
) > "$root/SHA256SUMS"
mkdir -p "$OUT_DIR"
tar --sort=name --owner=0 --group=0 --numeric-owner \
--mtime="@$SOURCE_DATE_EPOCH" \
--pax-option=delete=atime,delete=ctime \
-C "$OUT_DIR/root" -cf - "$asset" \
| gzip -n -9 > "$OUT_DIR/$asset.tar.gz"
(
cd "$OUT_DIR"
sha256sum "$asset.tar.gz" > "$asset.tar.gz.sha256"
)
cp "$root/$asset.cdx.json" "$OUT_DIR/$asset.cdx.json"
@@ -0,0 +1,21 @@
The MIT License (MIT)
Copyright (c) 2017 yhirose
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2013-2022 Niels Lohmann
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
@@ -0,0 +1,31 @@
-----BEGIN PGP PUBLIC KEY BLOCK-----
mQENBFuOrjYBCADT5MjShtbeSsWHADqVP7PZIp+m/wWkSUA7/FX/qrixhQE9DFyt
XtKSbBdwh+Jg5nsttCUiePtdrrRD1tcyowG256Tus3vOysZzvpfjWA4gcVmTjJXn
gwezKsPZLQi0wvjwQD8ByxnM1i2eiJC4xcMjT21uZkwDfgLTzVO4InWlVyZDB/da
PLJl4r1MqnsI603RKalMQmZzs43YUDssdeOiGOpXvb1Rj0XcsOOqnAEvIwyUWGku
1Hr+b6C9Nj6wksD7TCB10IdOeuwBqFgrVDzicG4fijwnpzA+UUfncIhKYdI/oIvj
mcAPobWzcBkM3uc+Yf/CxlBahzu6jv7AFdT1ABEBAAG0VEx1bmFyRyBTaWduaW5n
IEtleSAoS2V5IHVzZWQgYnkgTHVuYXJHIHRvIHNpZ24gcGFja2FnZXMpIDxsaW51
eC1wYWNrYWdlc0BsdW5hcmcuY29tPokBTgQTAQoAOBYhBAP11iGjcQ+pWpPYm6qE
UggOOD9+BQJbjq42AhsDBQsJCAcDBRUKCQgLBRYCAwEAAh4BAheAAAoJEKqEUggO
OD9+ECgH/Ro6LVB08FifApBS235v0Af3dsJlZGE0miKu2hR12qAvWackE6//E5GN
5xKSNpgLzV6kyylBntQDhcFzW3hLt/AsMLOXvuxYNFcLes2y10DrqVekNeJiR95V
KiTPI2jP8m4eFpcSnY0riHk2MmstN1icehQhYrWFyUtt3VxSsRWiRDeNUfCHC6YP
MjOXonmTWfH7T+UA2IqLFrt9dAsYGiCtMKVgzaZaZwm727c0aqy0e43nsWqjWxmE
EsEA1RvzjKKyzyixwpnzIyQ8dqL8sH0G3E2OYTlS7A8//yfgykRQVHwg2TsTBKfG
LlTmKj7RCT6GqISo+rbYYo/hZ6l2hH25AQ0EW46uNgEIANZfPWerTPzmvswWqp0P
iQvW+0qTBxZH3gQlwq5s6ahpY1pIebfrL/SAYJUGyjJVcjkG+HBXRGyRxtWFDE+D
+WEuziBfKd3aBUXb5DnvWdCiXeyQnFfwUVYNXhU5PlpAB5M409a30p9gGOrYy3Ah
g4VHhpM9wzGUAOzTwQ4WaC2WkR84sZYyqdKoo6C3m4IR4KHMYXF9nRlPSNEckL9U
MZe6I2uvor9FOPIfIOAI8lN+gbj/anf3lfy0ZYPyUtl3EWveGpWAPvdw3LMKg5QN
B8bR9TkPk0YZyQQcWkmN7gLUg0Vba+PYHH9DRlG8w1rH4TKxXJV3wmHo2aZRF1kc
30kAEQEAAYkBNgQYAQoAIBYhBAP11iGjcQ+pWpPYm6qEUggOOD9+BQJbjq42AhsM
AAoJEKqEUggOOD9+MEUH/2pm2QOttjd7DmEaS4LGvaTlEif0xtymRAh3axGuqQhl
KCZbw0jwsQlo/DwMRZwZHYCj1A/5H8mEg9qNGjF35GEpQTFSQI6Mt7F2DK69J86w
61v8tjxs4eO201ndhy+DRwDwG8vryFldx3f0nEdlE7IusgiUdvkcJPc8rX7p0MJJ
istTREAq8bRnvWYJzd4k3tgwHglEDxyjBRwLtqZyQ19XZb3V/aVKygqvZbwdJyXO
RHAZxK81p9Gp/8VkogJHLx6+3V8UlDepJg9/8MUCBQ9wWkdF0Pfqzgu7xtIHSxvW
62EF4nxqVuC946OIeITgXpd4F+iTFVII8w0P+nyCzac=
=nXAe
-----END PGP PUBLIC KEY BLOCK-----
+116
View File
@@ -0,0 +1,116 @@
#!/usr/bin/env python3
import datetime
import hashlib
import json
import os
import uuid
from pathlib import Path
root = Path(os.environ["ROOT"])
version = os.environ["VERSION"]
commit = os.environ["COMMIT"]
epoch = int(os.environ["EPOCH"])
asset = "whisper-bin-ubuntu-vulkan-x64"
sbom_path = root / f"{asset}.cdx.json"
def digest(path):
h = hashlib.sha256()
with path.open("rb") as stream:
for block in iter(lambda: stream.read(1024 * 1024), b""):
h.update(block)
return h.hexdigest()
files = []
for path in sorted(root.rglob("*")):
if path != sbom_path and path.is_file() and not path.is_symlink():
rel = path.relative_to(root).as_posix()
files.append({
"type": "file",
"bom-ref": f"file:{rel}",
"name": rel,
"hashes": [{"alg": "SHA-256", "content": digest(path)}],
})
ts = datetime.datetime.fromtimestamp(
epoch, datetime.timezone.utc,
).isoformat().replace("+00:00", "Z")
root_ref = f"pkg:github/ggml-org/whisper.cpp@{version}?commit={commit}"
ggml_ref = "pkg:github/ggml-org/[email protected]"
httplib_ref = "pkg:github/yhirose/[email protected]"
json_ref = "pkg:github/nlohmann/[email protected]"
sbom = {
"bomFormat": "CycloneDX",
"specVersion": "1.6",
"serialNumber": f"urn:uuid:{uuid.uuid5(uuid.NAMESPACE_URL, root_ref)}",
"version": 1,
"metadata": {
"timestamp": ts,
"tools": {"components": [
{"type": "application", "name": "make-sbom.py", "version": "1"},
{"type": "application", "name": "CMake", "version": "3.31.6"},
{"type": "application", "name": "glslc", "version": "2025.2"},
]},
"component": {
"type": "application",
"bom-ref": root_ref,
"group": "ggml-org",
"name": "whisper-server",
"version": version,
"purl": root_ref,
"licenses": [{"expression": "MIT"}],
"externalReferences": [{
"type": "vcs",
"url": f"https://github.com/ggml-org/whisper.cpp/tree/{commit}",
}],
"properties": [
{"name": "dikte:asset-name", "value": f"{asset}.tar.gz"},
{"name": "dikte:source-commit", "value": commit},
{"name": "dikte:runtime:glibc-minimum", "value": "2.34"},
{"name": "dikte:runtime:glibcxx-minimum", "value": "3.4.30"},
{"name": "dikte:runtime:vulkan-loader", "value": "optional; libvulkan.so.1"},
],
},
},
"components": [
{
"type": "library",
"bom-ref": ggml_ref,
"group": "ggml-org",
"name": "ggml",
"version": "0.20.2",
"purl": ggml_ref,
"licenses": [{"expression": "MIT"}],
"properties": [{
"name": "dikte:source",
"value": "vendored by the pinned whisper.cpp commit",
}],
},
{
"type": "library",
"bom-ref": httplib_ref,
"group": "yhirose",
"name": "cpp-httplib",
"version": "0.20.0",
"purl": httplib_ref,
"licenses": [{"expression": "MIT"}],
},
{
"type": "library",
"bom-ref": json_ref,
"group": "nlohmann",
"name": "json",
"version": "3.11.2",
"purl": json_ref,
"licenses": [{"expression": "MIT"}],
},
*files,
],
"dependencies": [{
"ref": root_ref,
"dependsOn": [ggml_ref, httplib_ref, json_ref]
+ [item["bom-ref"] for item in files],
}],
}
json.dump(sbom, fp=os.sys.stdout, indent=2, sort_keys=True)
print()
+76
View File
@@ -0,0 +1,76 @@
#!/usr/bin/env bash
set -euo pipefail
mode=${1:?usage: smoke-runtime.sh cpu|noicd|vulkan}
: "${OUT_DIR:=work/out}"
: "${FIXTURE_SOURCE:=vendor/whisper.cpp}"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
OUT_DIR="$(realpath "$OUT_DIR")"
FIXTURE_SOURCE="$(realpath "$FIXTURE_SOURCE")"
asset=whisper-bin-ubuntu-vulkan-x64
case "$mode" in
cpu) dockerfile=Dockerfile.runtime-cpu; image=dikte-whisper-runtime-cpu:spike ;;
noicd) dockerfile=Dockerfile.runtime-noicd; image=dikte-whisper-runtime-noicd:spike ;;
vulkan) dockerfile=Dockerfile.runtime-vulkan; image=dikte-whisper-runtime-vulkan:spike ;;
*) echo "unknown mode: $mode" >&2; exit 2 ;;
esac
tmp=$(mktemp -d)
trap 'rm -rf "$tmp"' EXIT
tar -xzf "$OUT_DIR/$asset.tar.gz" -C "$tmp"
docker build --pull=false -f "$SCRIPT_DIR/$dockerfile" -t "$image" "$SCRIPT_DIR"
args=("/bundle/$asset/whisper-server" -m /fixtures/model.bin
--host 127.0.0.1 --port 8080
--inference-path /v1/audio/transcriptions -l auto -sns -nlp)
env_args=()
# No -ng anywhere: Dikte passes it only when its GPU setting is off, so the
# run that has to survive a missing loader or a missing device is this one,
# where the backend registry actually goes looking for them.
if [[ "$mode" == vulkan ]]; then
env_args=(-e LIBGL_ALWAYS_SOFTWARE=1
-e VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/lvp_icd.x86_64.json)
fi
docker run --rm --name "dikte-whisper-$mode-smoke" \
-e SMOKE_MODE="$mode" \
"${env_args[@]}" \
-v "$tmp/$asset:/bundle/$asset:ro" \
-v "$FIXTURE_SOURCE/models/for-tests-ggml-base.en.bin:/fixtures/model.bin:ro" \
-v "$FIXTURE_SOURCE/samples/jfk.wav:/fixtures/jfk.wav:ro" \
"$image" bash -ec '
if [ "$SMOKE_MODE" = cpu ] && ldconfig -p | grep -q libvulkan.so.1; then
echo "CPU smoke image unexpectedly has a Vulkan loader" >&2
exit 1
fi
if [ "$SMOKE_MODE" = noicd ]; then
if ! ldconfig -p | grep -q libvulkan.so.1; then
echo "no-ICD smoke image has no Vulkan loader to load" >&2
exit 1
fi
if compgen -G "/usr/share/vulkan/icd.d/*.json" >/dev/null; then
echo "no-ICD smoke image has a driver after all" >&2
exit 1
fi
fi
"$@" >/tmp/server.log 2>&1 &
pid=$!
trap "kill $pid 2>/dev/null || true" EXIT
for _ in $(seq 1 120); do
kill -0 "$pid" 2>/dev/null || { cat /tmp/server.log; exit 1; }
if curl --silent --show-error --fail --max-time 180 \
-F file=@/fixtures/jfk.wav -F response_format=json \
http://127.0.0.1:8080/v1/audio/transcriptions >/tmp/response.json; then
grep -q "\"text\"" /tmp/response.json
if [ "$SMOKE_MODE" = vulkan ]; then
grep -q "loaded Vulkan backend" /tmp/server.log
fi
cat /tmp/response.json
cat /tmp/server.log
exit 0
fi
sleep 1
done
cat /tmp/server.log
exit 1
' bash "${args[@]}"
+145
View File
@@ -0,0 +1,145 @@
#!/usr/bin/env bash
set -euo pipefail
: "${OUT_DIR:=work/out}"
: "${SOURCE_DIR:=whisper.cpp}"
asset=whisper-bin-ubuntu-vulkan-x64
archive="$OUT_DIR/$asset.tar.gz"
tmp=$(mktemp -d)
trap 'rm -rf "$tmp"' EXIT
test -s "$archive"
(cd "$OUT_DIR" && sha256sum --check "$asset.tar.gz.sha256")
ARCHIVE="$archive" ASSET="$asset" python3 - <<'PY'
import os
import posixpath
import tarfile
archive = os.environ["ARCHIVE"]
asset = os.environ["ASSET"]
def under_root(name):
normalized = posixpath.normpath(name)
return (not posixpath.isabs(normalized)
and normalized != ".."
and not normalized.startswith("../")
and normalized.split("/", 1)[0] == asset)
with tarfile.open(archive, "r:gz") as bundle:
for member in bundle:
if not under_root(member.name):
raise SystemExit(f"unsafe archive member: {member.name}")
if member.isdev() or member.isfifo():
raise SystemExit(f"special archive member: {member.name}")
if not (member.isdir() or member.isfile()
or member.issym() or member.islnk()):
raise SystemExit(f"unsupported archive member: {member.name}")
if member.issym():
target = posixpath.join(posixpath.dirname(member.name),
member.linkname)
if not under_root(target):
raise SystemExit(f"unsafe symlink: {member.name}")
if member.islnk() and not under_root(member.linkname):
raise SystemExit(f"unsafe hardlink: {member.name}")
PY
tar -xzf "$archive" -C "$tmp"
root="$tmp/$asset"
test -x "$root/whisper-server"
test -f "$root/libwhisper.so"
test -f "$root/libggml.so"
test -f "$root/libggml-base.so"
test -f "$root/libggml-vulkan.so"
compgen -G "$root/libggml-cpu-*.so" >/dev/null
test -f "$root/LICENSES/whisper.cpp-MIT.txt"
test -f "$root/LICENSES/cpp-httplib-MIT.txt"
test -f "$root/LICENSES/nlohmann-json-MIT.txt"
(cd "$root" && sha256sum --check SHA256SUMS)
# All shipped ELF objects must be relocatable and must not remember /work.
while IFS= read -r -d '' file; do
file "$file" | grep -q ELF || continue
dynamic=$(readelf -d "$file")
if ! grep -Fq 'Library runpath: [$ORIGIN]' <<<"$dynamic"; then
echo "runpath is not \$ORIGIN in $file" >&2
exit 1
fi
if grep -Eq '/(home|tmp|work)/' <<<"$dynamic"; then
echo "build path remains in $file" >&2
exit 1
fi
done < <(find "$root" -type f -print0)
# Vulkan remains a plugin dependency. The executable must start without a loader.
if readelf -d "$root/whisper-server" | grep -q 'libvulkan.so'; then
echo "whisper-server links Vulkan instead of loading it as a plugin" >&2
exit 1
fi
readelf -d "$root/libggml-vulkan.so" | grep -q 'libvulkan.so.1'
# Ubuntu 22.04 establishes the glibc ceiling promised by this artifact.
ROOT="$root" python3 - <<'PY'
import os, pathlib, re, subprocess
root = pathlib.Path(os.environ['ROOT'])
seen = {'GLIBC': set(), 'GLIBCXX': set(), 'CXXABI': set()}
external = {
'libc.so.6', 'libgcc_s.so.1', 'libm.so.6', 'libstdc++.so.6',
'libvulkan.so.1', 'ld-linux-x86-64.so.2',
}
for path in root.iterdir():
if not path.is_file() or path.is_symlink():
continue
header = subprocess.run(['readelf', '-h', path], text=True,
stdout=subprocess.PIPE,
stderr=subprocess.DEVNULL).stdout
if not header:
continue
if 'Machine: Advanced Micro Devices X86-64' not in header:
raise SystemExit(f'wrong ELF architecture: {path.name}')
dynamic = subprocess.run(['readelf', '-d', path], text=True,
stdout=subprocess.PIPE,
stderr=subprocess.DEVNULL).stdout
needed = re.findall(r'\(NEEDED\).*\[(.*?)\]', dynamic)
unexpected = [name for name in needed
if name not in external
and not re.fullmatch(
r'lib(?:whisper|ggml(?:-base)?)\.so\.\d+', name)]
if unexpected:
raise SystemExit(
f'unexpected DT_NEEDED in {path.name}: {unexpected}')
if path.name != 'libggml-vulkan.so' and 'libvulkan.so.1' in needed:
raise SystemExit(f'Vulkan is not plugin-only in {path.name}')
contents = path.read_bytes()
for marker in (b'/home/', b'/tmp/', b'/work/'):
if marker in contents:
raise SystemExit(
f'build path {marker!r} remains in {path.name}')
text = subprocess.run(['objdump', '-T', path], text=True,
stdout=subprocess.PIPE, stderr=subprocess.DEVNULL).stdout
for family in seen:
pattern = rf'{family}_([0-9]+(?:\.[0-9]+)+)'
seen[family].update(tuple(map(int, version.split('.')))
for version in re.findall(pattern, text))
assert seen['GLIBC'] and max(seen['GLIBC']) <= (2, 34), max(seen['GLIBC'])
assert seen['GLIBCXX'] and max(seen['GLIBCXX']) <= (3, 4, 30), max(seen['GLIBCXX'])
assert seen['CXXABI'] and max(seen['CXXABI']) <= (1, 3, 13), max(seen['CXXABI'])
for family, versions in seen.items():
print(f'maximum {family} symbol:', '.'.join(map(str, max(versions))))
PY
python3 - "$root/$asset.cdx.json" <<'PY'
import json, sys
with open(sys.argv[1], encoding='utf-8') as stream:
doc = json.load(stream)
assert doc['bomFormat'] == 'CycloneDX'
assert doc['specVersion'] == '1.6'
assert doc['metadata']['component']['name'] == 'whisper-server'
assert len(doc['components']) >= 3
print('SBOM components:', len(doc['components']))
PY
LD_LIBRARY_PATH='' "$root/whisper-server" --help >/dev/null 2>&1
echo "structure: PASS"
+95
View File
@@ -53,6 +53,16 @@ class TimestampModel(unittest.TestCase):
self.assertEqual(api.timestamp_model("openai", "gpt-4o-transcribe"),
"whisper-1")
def test_openrouter_takes_the_file_model_that_was_set(self):
self.assertEqual(
api.timestamp_model("openrouter", "openai/gpt-4o-transcribe",
"openai/whisper-large-v3"),
"openai/whisper-large-v3")
def test_openrouter_with_no_file_model_falls_back_to_whisper(self):
self.assertEqual(api.timestamp_model("openrouter", "openai/gpt-4o-transcribe", ""),
"openai/whisper-1")
class Explain(DikteTest):
def error(self, status):
@@ -119,9 +129,22 @@ class ExtractError(unittest.TestCase):
body = json.dumps({"error": {"code": 42}})
self.assertIn("42", api._extract_error(body))
def test_an_error_wrapped_in_an_array(self):
"""Google's 503 arrives this way, and .get() on a list raises."""
body = json.dumps([{"error": {"code": 503,
"message": "The model is overloaded."}}])
self.assertEqual(api._extract_error(body), "The model is overloaded.")
def test_a_body_that_is_not_json(self):
self.assertEqual(api._extract_error("<html>502</html>"), "<html>502</html>")
def test_no_shape_at_all_still_comes_back_as_a_string(self):
"""It runs while an ApiError is being raised: throwing here would
escape the `except ApiError` holding the raw transcript."""
for body in ("[]", "[1, 2]", '"a string"', "null", "17"):
with self.subTest(body=body):
self.assertIsInstance(api._extract_error(body), str)
def test_a_wall_of_html_is_cut_short(self):
self.assertEqual(len(api._extract_error("x" * 5000)), 300)
@@ -305,6 +328,13 @@ class TranscribeSegments(DikteTest):
api.transcribe_segments(OPENROUTER, self.wav)
self.assertEqual(multipart_fields(calls[0])["model"], "openai/whisper-1")
def test_openrouter_asks_for_the_file_model_when_one_is_set(self):
target = OPENROUTER._replace(file_model="mistralai/voxtral-mini-transcribe")
with fake_urlopen(self.reply([{"start": 0, "end": 1, "text": "hi"}])) as calls:
api.transcribe_segments(target, self.wav)
self.assertEqual(multipart_fields(calls[0])["model"],
"mistralai/voxtral-mini-transcribe")
def test_groq_stays_on_the_model_it_was_given(self):
target = GROQ._replace(model="whisper-large-v3")
with fake_urlopen(self.reply([{"start": 0, "end": 1, "text": "hi"}])) as calls:
@@ -384,6 +414,37 @@ class Cleanup(DikteTest):
self.assertEqual(sent_json(calls[0])["reasoning"],
{"effort": "high", "exclude": True})
def test_gemini_takes_openai_s_flat_field_rather_than_the_object(self):
_, calls = self.call(chat_reply("Hello."), reasoning="low",
provider="gemini", service="Google AI Studio")
payload = sent_json(calls[0])
self.assertEqual(payload["reasoning_effort"], "low")
self.assertNotIn("reasoning", payload)
def test_off_is_asked_for_as_the_lowest_rung_google_actually_has(self):
"""Sending "none" is a 400, and Flash left alone thinks."""
_, calls = self.call(chat_reply("Hello."), reasoning="none",
provider="gemini", service="Google AI Studio")
self.assertEqual(sent_json(calls[0])["reasoning_effort"], "minimal")
def test_a_rung_google_does_not_have_lands_on_the_nearest_one(self):
for asked in ("xhigh", "max"):
with self.subTest(asked=asked):
_, calls = self.call(chat_reply("Hello."), reasoning=asked,
provider="gemini", service="Google AI Studio")
self.assertEqual(sent_json(calls[0])["reasoning_effort"], "high")
def test_gemini_left_on_the_model_s_own_default_is_told_nothing(self):
_, calls = self.call(chat_reply("Hello."), provider="gemini",
service="Google AI Studio")
self.assertNotIn("reasoning_effort", sent_json(calls[0]))
def test_a_missing_gemini_key_says_google_ai_studio(self):
with self.assertRaises(api.ApiError) as caught:
api.cleanup("hello", "", "gemini-3.5-flash-lite", "prompt",
provider="gemini", service="Google AI Studio")
self.assertIn("Google AI Studio", str(caught.exception))
def test_a_local_base_url(self):
_, calls = self.call(chat_reply("Hello."), base_url="http://localhost:1234/v1")
self.assertEqual(calls[0].full_url, "http://localhost:1234/v1/chat/completions")
@@ -513,6 +574,40 @@ class ModelLists(DikteTest):
api.openai_models("", api.GROQ_URL, "Groq")
self.assertIn("Groq", str(caught.exception))
def test_gemini_keeps_only_the_models_that_answer_a_chat_request(self):
with fake_urlopen({"data": [{"id": "gemini-3.5-flash"},
{"id": "text-embedding-004"},
{"id": "imagen-4.0"},
{"id": "gemini-2.5-flash-lite"}]}) as calls:
models = api.gemini_models("AIza-test")
self.assertEqual(calls[0].full_url,
"https://generativelanguage.googleapis.com/v1beta/openai/models")
self.assertEqual(models, ["gemini-2.5-flash-lite", "gemini-3.5-flash"])
def test_the_long_form_of_an_id_is_shortened_to_what_a_request_wants(self):
with fake_urlopen({"data": [{"id": "models/gemini-3.5-flash-lite"}]}):
self.assertEqual(api.gemini_models("AIza-test"),
["gemini-3.5-flash-lite"])
def test_a_gemini_id_that_is_not_a_chat_model_is_left_out(self):
"""Google names its pictures and its voices `gemini` too."""
with fake_urlopen({"data": [{"id": "gemini-3.5-flash"},
{"id": "gemini-embedding-001"},
{"id": "gemini-2.5-flash-image"},
{"id": "gemini-2.5-flash-preview-tts"},
{"id": "gemini-2.5-native-audio"}]}):
self.assertEqual(api.gemini_models("AIza-test"), ["gemini-3.5-flash"])
def test_gemini_sends_the_key_as_a_bearer_token(self):
with fake_urlopen({"data": []}) as calls:
api.gemini_models("AIza-test")
self.assertEqual(calls[0].get_header("Authorization"), "Bearer AIza-test")
def test_a_missing_gemini_key_says_google_ai_studio(self):
with self.assertRaises(api.ApiError) as caught:
api.gemini_models("")
self.assertIn("Google AI Studio", str(caught.exception))
if __name__ == "__main__":
unittest.main()
+237 -24
View File
@@ -97,15 +97,38 @@ class Provider(DikteTest):
def test_what_each_one_runs(self):
self.assertEqual(assistant.executable("claude"), "claude")
self.assertEqual(assistant.executable("codex"), "codex")
self.assertEqual(assistant.executable("agy"), "agy")
self.assertEqual(assistant.executable("openrouter"), "")
self.assertEqual(assistant.executable("opencode"), "")
def test_the_model_recorded_is_the_one_that_answered(self):
"""The history used to write Claude's setting whoever had answered."""
self.assertEqual(assistant.model(self.config()), "sonnet")
self.assertEqual(
assistant.model(self.config(assistant_provider="codex")), "codex")
self.assertEqual(
assistant.model(self.config(assistant_provider="agy")), "agy")
self.assertEqual(
assistant.model(self.config(assistant_provider="agy",
assistant_agy_model="gemini-3.1-pro-low")),
"gemini-3.1-pro-low")
self.assertEqual(
assistant.model(self.config(assistant_provider="openrouter")),
"google/gemini-3.5-flash")
def test_what_each_one_is_called(self):
self.assertEqual(assistant.display_name(self.config()), "Claude")
self.assertEqual(
assistant.display_name(self.config(assistant_provider="codex")), "Codex")
self.assertEqual(
assistant.display_name(self.config(assistant_provider="openrouter")),
"OpenRouter")
for name, called in (("codex", "Codex"), ("agy", "Antigravity"),
("openrouter", "OpenRouter"),
("opencode", "OpenCode Go")):
with self.subTest(name=name):
self.assertEqual(
assistant.display_name(self.config(assistant_provider=name)),
called)
def test_every_provider_has_a_name_to_be_called_by(self):
"""_conclude writes its errors in it, so a gap here is a bare id."""
self.assertEqual(set(assistant.SERVICES), set(assistant.PROVIDERS))
class Effort(unittest.TestCase):
@@ -113,21 +136,26 @@ class Effort(unittest.TestCase):
def test_the_scales_cover_the_same_settings(self):
self.assertEqual(set(assistant.CLAUDE_EFFORT), set(assistant.CODEX_EFFORT))
self.assertEqual(set(assistant.CLAUDE_EFFORT), set(assistant.AGY_EFFORT))
def test_codex_has_no_rung_above_high(self):
self.assertEqual(assistant.CODEX_EFFORT["xhigh"], "high")
self.assertEqual(assistant.CODEX_EFFORT["max"], "high")
def test_neither_codex_nor_agy_has_a_rung_above_high(self):
for scale in (assistant.CODEX_EFFORT, assistant.AGY_EFFORT):
self.assertEqual(scale["xhigh"], "high")
self.assertEqual(scale["max"], "high")
def test_neither_one_asks_for_a_rung_below_low(self):
def test_none_of_them_asks_for_a_rung_below_low(self):
# Claude has none; Codex has one, but calls it "minimal" on the older
# models and "none" on the newer ones, and refuses the wrong word.
for scale in (assistant.CLAUDE_EFFORT, assistant.CODEX_EFFORT):
# models and "none" on the newer ones, and refuses the wrong word; agy
# has three rungs and no word for off at all.
for scale in (assistant.CLAUDE_EFFORT, assistant.CODEX_EFFORT,
assistant.AGY_EFFORT):
self.assertEqual(scale["none"], "low")
self.assertEqual(scale["minimal"], "low")
def test_an_empty_setting_asks_for_nothing(self):
self.assertEqual(assistant.CLAUDE_EFFORT.get("", ""), "")
self.assertEqual(assistant.CODEX_EFFORT.get("", ""), "")
for scale in (assistant.CLAUDE_EFFORT, assistant.CODEX_EFFORT,
assistant.AGY_EFFORT):
self.assertEqual(scale.get("", ""), "")
class Session(DikteTest):
@@ -276,25 +304,32 @@ class Conclude(DikteTest):
def test_an_answer_and_its_session(self):
answer, warning = assistant._conclude(
self.found(answer="done", session="abc"), 0, "", "", "Claude")
self.found(answer="done", session="abc"), 0, "", "", "claude")
self.assertEqual(answer, "done")
self.assertEqual(warning, "")
self.assertEqual(assistant.read_session("claude", 1800), "abc")
def test_codex_stores_under_its_own_name(self):
assistant._conclude(self.found(answer="done", session="t-1"), 0, "",
"", "Codex")
self.assertEqual(assistant.read_session("codex", 1800), "t-1")
def test_each_one_stores_under_its_own_name(self):
for name, session in (("codex", "t-1"), ("agy", "c-9")):
with self.subTest(name=name):
assistant._conclude(self.found(answer="done", session=session),
0, "", "", name)
self.assertEqual(assistant.read_session(name, 1800), session)
def test_the_error_is_written_in_the_provider_s_own_name(self):
with self.assertRaises(assistant.AssistantError) as caught:
assistant._conclude(self.found(), 1, "", "", "agy")
self.assertIn("Antigravity", str(caught.exception))
def test_a_non_zero_exit_with_nothing_to_show_for_it(self):
with self.assertRaises(assistant.AssistantError) as caught:
assistant._conclude(self.found(), 1, "it all went wrong\n", "", "Claude")
assistant._conclude(self.found(), 1, "it all went wrong\n", "", "claude")
self.assertIn("it all went wrong", str(caught.exception))
def test_a_session_that_is_gone_is_raised_apart(self):
with self.assertRaises(assistant._SessionGone):
assistant._conclude(self.found(), 1, "session abc not found",
"abc", "Claude")
"abc", "claude")
def test_the_recovery_no_longer_hangs_on_the_words_the_cli_chose(self):
# The complaint used to be matched by substring, which a CLI update or
@@ -321,11 +356,11 @@ class Conclude(DikteTest):
def test_a_session_that_is_gone_only_matters_when_one_was_resumed(self):
with self.assertRaises(assistant.AssistantError):
assistant._conclude(self.found(), 1, "session abc not found",
"", "Claude")
"", "claude")
def test_an_answer_survives_a_non_zero_exit(self):
answer, _ = assistant._conclude(self.found(answer="done"), 1, "noise",
"", "Claude")
"", "claude")
self.assertEqual(answer, "done")
def test_an_answer_on_a_resumed_session_is_kept_rather_than_retried(self):
@@ -336,12 +371,12 @@ class Conclude(DikteTest):
def test_a_reported_failure_with_no_answer(self):
with self.assertRaises(assistant.AssistantError) as caught:
assistant._conclude(self.found(failure="the model refused"), 0, "",
"", "Claude")
"", "claude")
self.assertIn("refused", str(caught.exception))
def test_a_run_that_said_nothing_at_all(self):
with self.assertRaises(assistant.AssistantError) as caught:
assistant._conclude(self.found(), 0, "", "", "Codex")
assistant._conclude(self.found(), 0, "", "", "codex")
self.assertIn("Codex", str(caught.exception))
@@ -565,6 +600,109 @@ class AskCodex(DikteTest):
self.assertIn("quota", str(caught.exception))
class AskAgy(DikteTest):
"""agy's stream is shaped nothing like the other two: the key is `event`,
the answer arrives whole in `result.response`, and the conversation to
resume is named in the first line rather than the last."""
def run_ask(self, conf=None, events=None, session=""):
conf = conf or self.config(assistant_provider="agy")
proc = FakeCli(events or [
{"event": "init", "conversation_id": "c-9", "init": {"cwd": "/home"}},
{"event": "result",
"result": {"conversation_id": "c-9", "status": "SUCCESS",
"response": " done "}},
])
stages = []
with only_these_tools("agy"), \
mock.patch.object(subprocess, "Popen", return_value=proc) as popen:
result = assistant._ask_agy("book it", conf, session,
stages.append, None)
return result, popen.call_args.args[0], stages
def test_the_answer_comes_back_stripped(self):
(answer, warning), _, _ = self.run_ask()
self.assertEqual(answer, "done")
self.assertEqual(warning, "")
def test_the_instruction_is_kept_apart_from_the_command(self):
"""agy takes no system prompt, so the two must not read as one."""
conf = self.config(assistant_provider="agy")
_, cmd, _ = self.run_ask(conf)
body = cmd[cmd.index("-p") + 1]
self.assertTrue(body.startswith(conf.assistant_prompt()))
self.assertIn("\n\n---\n\n", body)
self.assertTrue(body.endswith("book it"))
def test_a_first_command_starts_a_project_of_its_own(self):
"""Without it agy works in whichever project it was last in."""
_, cmd, _ = self.run_ask()
self.assertIn("--new-project", cmd)
self.assertNotIn("--conversation", cmd)
def test_a_second_command_carries_the_conversation_rather_than_starting_one(self):
_, cmd, _ = self.run_ask(session="c-9")
self.assertEqual(cmd[cmd.index("--conversation") + 1], "c-9")
self.assertNotIn("--new-project", cmd)
def test_the_conversation_is_kept_under_agy_s_own_name(self):
self.run_ask()
self.assertEqual(assistant.read_session("agy", 1800), "c-9")
def test_it_is_not_left_to_give_up_before_the_caller_does(self):
conf = self.config(assistant_provider="agy", assistant_timeout=90)
_, cmd, _ = self.run_ask(conf)
self.assertEqual(cmd[cmd.index("--print-timeout") + 1], "90s")
def test_no_model_named_means_whatever_agy_is_set_to(self):
_, cmd, _ = self.run_ask()
self.assertNotIn("--model", cmd)
def test_a_model_of_your_own(self):
_, cmd, _ = self.run_ask(
self.config(assistant_provider="agy",
assistant_agy_model="gemini-3.1-pro-low"))
self.assertEqual(cmd[cmd.index("--model") + 1], "gemini-3.1-pro-low")
def test_a_tool_is_named_in_the_corner_as_it_starts(self):
_, _, stages = self.run_ask(events=[
{"event": "step_update",
"step_update": {"step_type": "tool", "state": "ACTIVE",
"tool_name": "run_command"}},
{"event": "step_update",
"step_update": {"step_type": "tool", "state": "DONE",
"tool_name": "run_command"}},
{"event": "result",
"result": {"status": "SUCCESS", "response": "done"}},
])
self.assertEqual(stages, ["Running a command…"])
def test_the_two_dozen_browser_tools_are_one_line_between_them(self):
_, _, stages = self.run_ask(events=[
{"event": "step_update",
"step_update": {"step_type": "tool", "state": "ACTIVE",
"tool_name": "browser_click_element"}},
{"event": "result",
"result": {"status": "SUCCESS", "response": "done"}},
])
self.assertEqual(stages, ["Working in the browser…"])
def test_a_turn_that_did_not_succeed_is_a_failure_rather_than_an_answer(self):
with self.assertRaises(assistant.AssistantError) as caught:
self.run_ask(events=[
{"event": "result",
"result": {"status": "ERROR", "response": "the model refused"}},
])
self.assertIn("refused", str(caught.exception))
def test_a_failure_with_nothing_to_say_is_still_named(self):
with self.assertRaises(assistant.AssistantError) as caught:
self.run_ask(events=[
{"event": "result", "result": {"status": "ERROR"}},
])
self.assertIn("Antigravity", str(caught.exception))
class AskOpenRouter(DikteTest):
def test_a_question_and_an_answer(self):
conf = self.config(assistant_provider="openrouter",
@@ -600,6 +738,41 @@ class AskOpenRouter(DikteTest):
assistant.ask("when is it", conf)
class AskOpenCode(DikteTest):
def test_a_question_and_an_answer(self):
conf = self.config(assistant_provider="opencode",
opencode_api_key="opencode-test-key")
with fake_urlopen({"choices": [{"message": {"content": "on Thursday"}}]}):
answer, warning = assistant.ask("when is it", conf)
self.assertEqual(answer, "on Thursday")
self.assertEqual(warning, "")
def test_the_conversation_is_ours_to_keep(self):
conf = self.config(assistant_provider="opencode",
opencode_api_key="opencode-test-key")
with fake_urlopen({"choices": [{"message": {"content": "on Thursday"}}]}):
assistant.ask("when is it", conf)
stored = assistant.read_messages("opencode", 1800)
self.assertEqual([row["content"] for row in stored],
["when is it", "on Thursday"])
def test_the_model_and_endpoint_are_opencode_s_own(self):
conf = self.config(assistant_provider="opencode",
opencode_api_key="opencode-test-key",
assistant_opencode_model="glm-5.3")
with fake_urlopen({"choices": [{"message": {"content": "on Thursday"}}]}) as calls:
assistant.ask("when is it", conf)
sent = json.loads(calls[0].data.decode("utf-8"))
self.assertEqual(sent["model"], "glm-5.3")
self.assertIn("https://opencode.ai/zen/go/v1/chat/completions",
calls[0].full_url)
def test_an_api_failure_reads_as_an_assistant_failure(self):
conf = self.config(assistant_provider="opencode")
with self.assertRaises(assistant.AssistantError):
assistant.ask("when is it", conf)
class Ask(DikteTest):
def test_a_cli_that_is_not_installed_says_where_to_change_it(self):
with only_these_tools(), \
@@ -710,5 +883,45 @@ class CodexModels(DikteTest):
self.assertEqual(assistant.codex_models(), [])
class AgyModels(DikteTest):
"""The model list read off `agy models`: one id, a tab, a display name."""
LISTING = ("gemini-4-flash-high\tGemini 4 Flash (High)\n"
"gemini-4-flash-low\tGemini 4 Flash (Low)\n"
"a line with no tab is not a model\n"
"\ta tab with no id in front of it is not one either\n")
def models(self, reply, code=0):
with only_these_tools("agy"), \
mock.patch.object(subprocess, "run",
return_value=FakeCompleted(
returncode=code, stdout=reply)) as run:
found = assistant.agy_models()
self.run_call = run
return found
def test_the_listing_arrives_in_agy_s_own_order(self):
found = self.models(self.LISTING)
self.assertEqual(found, ["gemini-4-flash-high", "gemini-4-flash-low"])
self.assertEqual(self.run_call.call_args.args[0], ["agy", "models"])
def test_an_agy_that_is_not_installed_is_not_run(self):
with only_these_tools(), \
mock.patch.object(subprocess, "run") as run:
self.assertEqual(assistant.agy_models(), [])
run.assert_not_called()
def test_a_call_that_failed_answers_with_nothing(self):
self.assertEqual(self.models("error: not logged in", code=1), [])
self.assertEqual(self.models(""), [])
def test_an_agy_that_hangs_is_given_up_on(self):
with only_these_tools("agy"), \
mock.patch.object(subprocess, "run",
side_effect=subprocess.TimeoutExpired(
["agy"], 30)):
self.assertEqual(assistant.agy_models(), [])
if __name__ == "__main__":
unittest.main()
+139
View File
@@ -57,7 +57,10 @@ class Provider(DikteTest):
def test_what_each_one_runs(self):
self.assertEqual(cleanup.executable("claude"), "claude")
self.assertEqual(cleanup.executable("codex"), "codex")
self.assertEqual(cleanup.executable("agy"), "agy")
self.assertEqual(cleanup.executable("openrouter"), "")
self.assertEqual(cleanup.executable("gemini"), "")
self.assertEqual(cleanup.executable("opencode"), "")
def test_the_model_named_in_the_history_is_the_one_that_did_it(self):
self.assertEqual(cleanup.model(self.config(cleanup_model="some/model")),
@@ -73,6 +76,19 @@ class Provider(DikteTest):
self.assertEqual(
cleanup.model(self.config(cleanup_provider="codex",
cleanup_codex_model="gpt-5.4")), "gpt-5.4")
self.assertEqual(
cleanup.model(self.config(cleanup_provider="gemini")),
"gemini-3.5-flash-lite")
# Antigravity is left on its own default the way Codex is.
self.assertEqual(
cleanup.model(self.config(cleanup_provider="agy")), "agy")
self.assertEqual(
cleanup.model(self.config(cleanup_provider="agy",
cleanup_agy_model="gemini-3.7-flash-low")),
"gemini-3.7-flash-low")
self.assertEqual(
cleanup.model(self.config(cleanup_provider="opencode",
cleanup_opencode_model="glm-5.3")), "glm-5.3")
class OpenRouter(DikteTest):
@@ -94,6 +110,72 @@ class OpenRouter(DikteTest):
self.assertEqual(calls, [])
class OpenCode(DikteTest):
def test_it_is_one_request_with_the_settings_as_they_were(self):
conf = self.config(cleanup_provider="opencode",
opencode_api_key="opencode-test-key",
cleanup_opencode_model="some/model",
cleanup_reasoning="low")
with mock.patch.object(api, "cleanup", return_value="Done.") as call:
self.assertEqual(cleanup.run("uh, done", conf, "the rules"), "Done.")
text, key, model, prompt = call.call_args.args
self.assertEqual((text, key, model, prompt),
("uh, done", "opencode-test-key", "some/model", "the rules"))
self.assertEqual(call.call_args.kwargs["reasoning"], "low")
self.assertEqual(call.call_args.kwargs["provider"], "opencode")
self.assertEqual(call.call_args.kwargs["service"], "OpenCode Go")
self.assertEqual(call.call_args.kwargs["base_url"],
"https://opencode.ai/zen/go/v1")
def test_no_cli_is_started_for_it(self):
conf = self.config(cleanup_provider="opencode",
opencode_api_key="opencode-test-key")
patcher, calls = fake_cli(stdout="never")
with patcher, mock.patch.object(api, "cleanup", return_value="Done."):
cleanup.run("uh, done", conf, "the rules")
self.assertEqual(calls, [])
class GoogleAiStudio(DikteTest):
"""Cleanup over Google's OpenAI-compatible endpoint: one request, no CLI."""
def setUp(self):
super().setUp()
self.conf = self.config(cleanup_provider="gemini",
gemini_api_key="AIza-test")
def test_it_goes_to_google_with_the_settings_as_they_were(self):
self.conf["cleanup_reasoning"] = "none"
with fake_urlopen(chat_reply("Done.")) as calls:
self.assertEqual(cleanup.run("uh, done", self.conf, "the rules"),
"Done.")
self.assertEqual(
calls[0].full_url,
"https://generativelanguage.googleapis.com/v1beta/openai/chat/completions")
payload = sent_json(calls[0])
self.assertEqual(payload["model"], "gemini-3.5-flash-lite")
self.assertEqual(payload["reasoning_effort"], "minimal")
self.assertIn("uh, done", payload["messages"][1]["content"])
def test_the_key_travels_as_a_bearer_token(self):
with fake_urlopen(chat_reply("Done.")) as calls:
cleanup.run("uh, done", self.conf, "the rules")
self.assertEqual(calls[0].get_header("Authorization"), "Bearer AIza-test")
def test_a_missing_key_names_google_rather_than_openrouter(self):
self.conf["gemini_api_key"] = ""
with mock.patch.dict(os.environ, {}, clear=True), \
self.assertRaises(api.ApiError) as caught:
cleanup.run("uh, done", self.conf, "the rules")
self.assertIn("Google AI Studio", str(caught.exception))
def test_no_cli_is_started_for_it(self):
patcher, calls = fake_cli(stdout="never")
with patcher, fake_urlopen(chat_reply("Done.")):
cleanup.run("uh, done", self.conf, "the rules")
self.assertEqual(calls, [])
class ClaudeCode(DikteTest):
def setUp(self):
super().setUp()
@@ -221,6 +303,63 @@ class Codex(DikteTest):
self.run_cleanup(stdout="tokens used 400", last_message="")
class Antigravity(DikteTest):
def setUp(self):
super().setUp()
self.conf = self.config(cleanup_provider="agy")
self.patch_attr(cleanup.shutil, "which", lambda name: f"/usr/bin/{name}")
def run_cleanup(self, text="uh, book it", **kwargs):
patcher, calls = fake_cli(**kwargs)
with patcher:
answer = cleanup.run(text, self.conf, "the rules")
return answer, calls[0]
def test_the_rules_ride_in_front_of_the_transcript(self):
answer, cmd = self.run_cleanup(stdout="Book it.\n")
self.assertEqual(answer, "Book it.")
self.assertEqual(cmd[0], "agy")
self.assertEqual(cmd[cmd.index("-p") + 1],
"the rules\n\n---\n\n<transcript>\nuh, book it\n</transcript>")
def test_it_starts_somewhere_of_its_own_and_takes_no_slash_commands(self):
"""Without --new-project agy works in whichever project it was last in."""
_, cmd = self.run_cleanup(stdout="Book it.")
self.assertIn("--new-project", cmd)
self.assertIn("--disable-slash-commands", cmd)
self.assertEqual(cmd[cmd.index("--output-format") + 1], "text")
def test_it_is_not_left_to_give_up_before_the_caller_does(self):
_, cmd = self.run_cleanup(stdout="Book it.")
self.assertEqual(cmd[cmd.index("--print-timeout") + 1], "180s")
def test_the_model_is_left_alone_until_one_is_typed_in(self):
_, cmd = self.run_cleanup(stdout="Book it.")
self.assertNotIn("--model", cmd)
self.conf["cleanup_agy_model"] = "gemini-3.7-flash-low"
_, cmd = self.run_cleanup(stdout="Book it.")
self.assertEqual(cmd[cmd.index("--model") + 1], "gemini-3.7-flash-low")
def test_the_thinking_setting_lands_on_the_nearest_rung_agy_has(self):
self.conf["cleanup_reasoning"] = "max"
_, cmd = self.run_cleanup(stdout="Book it.")
self.assertEqual(cmd[cmd.index("--effort") + 1], "high")
def test_no_thinking_setting_means_no_flag(self):
_, cmd = self.run_cleanup(stdout="Book it.")
self.assertNotIn("--effort", cmd)
def test_an_answer_of_nothing_is_a_failure_rather_than_an_empty_paste(self):
with self.assertRaises(cleanup.CleanupError):
self.run_cleanup(stdout=" ")
def test_a_program_that_is_not_installed_says_so_before_running_anything(self):
self.patch_attr(cleanup.shutil, "which", lambda name: "")
with self.assertRaises(cleanup.CleanupError) as caught:
self.run_cleanup(stdout="Book it.")
self.assertIn("agy", str(caught.exception))
if __name__ == "__main__":
unittest.main()
+45
View File
@@ -9,12 +9,14 @@ socket is faked, and everything that runs locally runs for real.
import contextlib
import io
import json
import sys
import unittest
import webbrowser
from typing import ClassVar
from unittest import mock
from dikte import audio
from dikte import cleanup
from dikte import cli
from dikte import config as cfg
from dikte import ggml
@@ -423,6 +425,16 @@ class Providers(DikteTest):
self.assertIn("groq", out)
self.assertIn("Groq", out)
def test_opencode_is_a_choice_and_reports_under_its_own_name(self):
parser = cli.build_parser()
self.assertEqual(
parser.parse_args(["test-key", "opencode"]).which, "opencode")
self.write_config({"opencode_api_key": "opencode-test"})
with fake_urlopen({"data": [{"id": "deepseek-v4-flash"}]}):
code, out, _ = self.run_cmd(cli.cmd_test_key, which="opencode")
self.assertEqual(code, 0)
self.assertIn("opencode: connection works, 1 models visible", out)
class Updates(DikteTest):
"""`dikte update` looks, says what it found, and installs nothing."""
@@ -503,6 +515,35 @@ class Doctor(DikteTest):
self.assertIn("OpenRouter key, cleaning up on some/model",
self.run_doctor(as_json=False, cleanup_model="some/model"))
def test_cleanup_on_opencode_is_a_question_about_its_own_key(self):
reply = self.run_doctor(cleanup_provider="opencode",
cleanup_opencode_model="glm-5.3")
self.assertEqual(reply["cleanup"]["provider"], "opencode")
self.assertEqual(reply["cleanup"]["model"], "glm-5.3")
self.assertIn("OpenCode Go key, cleaning up on glm-5.3",
self.run_doctor(as_json=False, cleanup_provider="opencode",
cleanup_opencode_model="glm-5.3"))
def test_it_survives_every_provider_cleanup_can_be_set_to(self):
"""It used to raise KeyError on the local model, whose executable is ""."""
for name in cleanup.PROVIDERS:
with self.subTest(provider=name):
reply = self.run_doctor(cleanup_provider=name)
self.assertEqual(reply["cleanup"]["provider"], name)
self.run_doctor(as_json=False, cleanup_provider=name)
def test_a_provider_with_no_key_to_check_says_so_rather_than_no(self):
"""A CLI needs none, so `false` there would read as one gone missing."""
self.assertIsNone(self.run_doctor(cleanup_provider="claude")["cleanup"]["key"])
self.assertIsNone(self.run_doctor(cleanup_provider="local")["cleanup"]["key"])
self.assertIs(self.run_doctor(cleanup_provider="gemini")["cleanup"]["key"],
False)
def test_cleanup_on_google_is_a_question_about_its_own_key(self):
line = self.run_doctor(as_json=False, cleanup_provider="gemini",
cleanup_gemini_model="gemini-2.5-flash")
self.assertIn("Google AI Studio key, cleaning up on gemini-2.5-flash", line)
def test_it_asks_after_the_programs_this_desktop_actually_uses(self):
"""A missing ydotool on a Mac is a red mark with nothing behind it."""
with mock.patch.object(cli.paste, "desktop", return_value=paste.MACOS):
@@ -628,7 +669,11 @@ class WithoutAnInstance(DikteTest):
def run_verb(self, argv):
# launch_gui replaces this process with the application, so it never
# comes back in real use and must not be allowed to here.
# `ask` with no text reads what was piped in, and the runner's own
# stdin is not that: under pytest it is an object that refuses to be
# read at all.
with mock.patch.object(ipc, "send", return_value=None), \
mock.patch.object(sys, "stdin", io.StringIO()), \
mock.patch.object(cli, "launch_gui") as launch, \
captured() as (out, err):
code = cli.run(argv)
+30
View File
@@ -187,6 +187,10 @@ class Keys(DikteTest):
def test_every_provider_falls_back_to_the_variable_of_its_own_name(self):
with mock.patch.dict(os.environ, {"GROQ_API_KEY": "gsk-env"}):
self.assertEqual(cfg.Config().groq_key(), "gsk-env")
with mock.patch.dict(os.environ, {"GEMINI_API_KEY": "AIza-env"}):
self.assertEqual(cfg.Config().gemini_key(), "AIza-env")
with mock.patch.dict(os.environ, {"OPENCODE_API_KEY": "opencode-env"}):
self.assertEqual(cfg.Config().opencode_key(), "opencode-env")
class TranscribeTarget(DikteTest):
@@ -216,6 +220,19 @@ class TranscribeTarget(DikteTest):
self.assertEqual(target.service, "OpenRouter")
self.assertEqual(target.api_key, "sk-or-test")
self.assertEqual(target.model, "openai/whisper-1")
self.assertEqual(target.file_model, "")
def test_openrouter_carries_its_file_model(self):
conf = self.config(transcribe_provider="openrouter",
openrouter_api_key="sk-or-test",
openrouter_file_model=" openai/whisper-large-v3 ")
self.assertEqual(conf.transcribe_target().file_model,
"openai/whisper-large-v3")
def test_only_openrouter_has_a_file_model(self):
conf = self.config(transcribe_provider="openai", openai_api_key="sk-test",
openrouter_file_model="openai/whisper-large-v3")
self.assertEqual(conf.transcribe_target().file_model, "")
def test_groq_when_it_is_picked(self):
conf = self.config(transcribe_provider="groq", groq_api_key="gsk-test",
@@ -574,6 +591,19 @@ class Defaults(unittest.TestCase):
def test_the_keys_ship_empty(self):
self.assertEqual(cfg.DEFAULTS["openai_api_key"], "")
self.assertEqual(cfg.DEFAULTS["openrouter_api_key"], "")
self.assertEqual(cfg.DEFAULTS["gemini_api_key"], "")
self.assertEqual(cfg.DEFAULTS["opencode_api_key"], "")
def test_google_ai_studio_is_a_cleanup_provider_and_not_a_transcriber(self):
"""Its compatible endpoint has no /audio/transcriptions behind it."""
self.assertNotIn("gemini", cfg.TRANSCRIBERS)
self.assertIn("gemini", cleanup.PROVIDERS)
def test_opencode_ships_on_its_own_endpoint(self):
self.assertEqual(cfg.DEFAULTS["opencode_base_url"],
"https://opencode.ai/zen/go/v1")
self.assertEqual(cfg.DEFAULTS["cleanup_opencode_model"], "deepseek-v4-flash")
self.assertEqual(cfg.DEFAULTS["assistant_opencode_model"], "deepseek-v4-flash")
def test_every_language_specific_prompt_has_both_languages(self):
for name in ("CLEANUP_PROMPT", "FILE_CLEANUP_PROMPT", "MEETING_PROMPT",
+158
View File
@@ -205,6 +205,7 @@ class InstallProgram(Local):
# These fixtures are Ubuntu release archives. Keep checking that path
# on every host, including the Mac that checks the macOS backend.
self.patch_attr(sys, "platform", "linux")
self.patch_attr(ggml.platform, "machine", lambda: "x86_64")
# Built once, because the release listing has to publish its checksum
# and a tarball is not the same bytes twice.
self.archive = tarball({
@@ -221,6 +222,7 @@ class InstallProgram(Local):
def install(self, *names, archive=None):
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: False)
blob = self.archive if archive is None else archive
with serving(self.release(*names, archive=blob), blob) as calls:
path = ggml.install_program(ggml.WHISPER)
@@ -238,6 +240,162 @@ class InstallProgram(Local):
"whisper-bin-ubuntu-x64.tar.gz")
self.assertTrue(urls[1].endswith("whisper-bin-ubuntu-x64.tar.gz"))
def test_the_nightly_pointer_is_followed_to_where_the_builds_are(self):
"""llama.cpp's latest release carries a tag name, not the binaries."""
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: False)
marker = self.release(ggml.NIGHTLY_TAG)
nightly = dict(self.release("llama-b10809-bin-ubuntu-x64.tar.gz"),
tag_name="b10809")
def opener(request, timeout=None):
url = request.full_url
if url.endswith("/releases/latest"):
return json_body(marker)
if url.endswith("/releases/tags/b10809"):
return json_body(nightly)
if url.endswith(ggml.NIGHTLY_TAG):
return body(b"b10809\n")
return body(self.archive)
with mock.patch("urllib.request.urlopen", side_effect=opener):
tag, found = ggml._pick_asset(ggml.LLAMA)
self.assertEqual(tag, "b10809")
self.assertEqual(found.name, "llama-b10809-bin-ubuntu-x64.tar.gz")
def test_without_a_pointer_the_newest_release_that_has_a_build_is_taken(self):
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: False)
marker = self.release("source.zip")
listing = [dict(self.release("llama-b2-bin-win-cpu-x64.zip"), tag_name="b2"),
dict(self.release("llama-b1-bin-ubuntu-x64.tar.gz"), tag_name="b1")]
def opener(request, timeout=None):
url = request.full_url
return json_body(listing if "per_page" in url else marker)
with mock.patch("urllib.request.urlopen", side_effect=opener):
tag, found = ggml._pick_asset(ggml.LLAMA)
self.assertEqual(tag, "b1")
self.assertEqual(found.name, "llama-b1-bin-ubuntu-x64.tar.gz")
def test_linux_x64_with_vulkan_takes_diktes_accelerated_build(self):
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: True)
listing = self.release("whisper-bin-ubuntu-vulkan-x64.tar.gz")
listing["tag_name"] = "whisper.cpp-v1.9.3"
managed_sha = hashlib.sha256(self.archive).hexdigest()
with mock.patch.object(ggml, "MANAGED_WHISPER_SHA256", managed_sha,
create=True):
with fake_urlopen(listing, body(self.archive)) as calls:
path = ggml.install_program(ggml.WHISPER)
urls = [call.full_url for call in calls]
self.assertIn(
"/repos/yusufipk/dikte/releases/tags/whisper.cpp-v1.9.3",
urls[0],
)
self.assertTrue(urls[1].endswith(
"whisper-bin-ubuntu-vulkan-x64.tar.gz"))
self.assertTrue(os.path.isfile(path))
self.assertEqual("v1.9.3", ggml.installed_version(ggml.WHISPER))
self.assertFalse(ggml.vulkan_missing(ggml.WHISPER))
def test_an_explicit_whisper_version_still_comes_from_upstream(self):
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: True)
listing = self.release("whisper-bin-ubuntu-x64.tar.gz")
with fake_urlopen(listing, body(self.archive)) as calls:
ggml.install_program(ggml.WHISPER, tag="v1.9.1")
self.assertIn(
"/repos/ggml-org/whisper.cpp/releases/tags/v1.9.1",
calls[0].full_url,
)
def test_linux_arm64_keeps_using_the_upstream_cpu_build(self):
self.patch_attr(ggml, "_arch", lambda: "arm64")
self.patch_attr(ggml.platform, "machine", lambda: "aarch64")
self.patch_attr(ggml, "_has_vulkan", lambda: True)
listing = self.release("whisper-bin-ubuntu-arm64.tar.gz")
with fake_urlopen(listing, body(self.archive)) as calls:
ggml.install_program(ggml.WHISPER)
self.assertIn(
"/repos/ggml-org/whisper.cpp/releases/latest",
calls[0].full_url,
)
def test_linux_non_x86_does_not_try_the_managed_x64_build(self):
self.patch_attr(ggml, "_has_vulkan", lambda: True)
listing = self.release("whisper-bin-ubuntu-arm64.tar.gz")
with mock.patch("platform.machine", return_value="ppc64le"):
with fake_urlopen(listing, listing) as calls:
with self.assertRaises(ggml.LocalError):
ggml.install_program(ggml.WHISPER)
self.assertIn(
"/repos/ggml-org/whisper.cpp/releases/latest",
calls[0].full_url,
)
def test_a_missing_managed_build_falls_back_to_upstream_cpu(self):
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: True)
managed = self.release("Dikte-1.1.0-x86_64.AppImage")
managed["tag_name"] = "whisper.cpp-v1.9.3"
upstream = self.release("whisper-bin-ubuntu-x64.tar.gz")
with fake_urlopen(managed, upstream, body(self.archive)) as calls:
path = ggml.install_program(ggml.WHISPER)
urls = [call.full_url for call in calls]
self.assertIn(
"/repos/yusufipk/dikte/releases/tags/whisper.cpp-v1.9.3",
urls[0],
)
self.assertIn("/repos/ggml-org/whisper.cpp/releases/latest", urls[1])
self.assertTrue(urls[2].endswith("whisper-bin-ubuntu-x64.tar.gz"))
self.assertTrue(os.path.isfile(path))
def test_a_managed_build_with_an_unreviewed_digest_falls_back(self):
self.patch_attr(ggml, "_has_vulkan", lambda: True)
managed = self.release("whisper-bin-ubuntu-vulkan-x64.tar.gz")
managed["assets"][0]["digest"] = "sha256:" + "0" * 64
upstream = self.release("whisper-bin-ubuntu-x64.tar.gz")
with fake_urlopen(managed, upstream, body(self.archive)) as calls:
try:
path = ggml.install_program(ggml.WHISPER)
except ggml.LocalError as exc:
self.fail(f"unreviewed digest did not fall back: {exc}")
urls = [call.full_url for call in calls]
self.assertEqual(3, len(urls))
self.assertTrue(urls[2].endswith("whisper-bin-ubuntu-x64.tar.gz"))
self.assertTrue(os.path.isfile(path))
def test_an_unavailable_managed_release_falls_back_to_upstream_cpu(self):
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: True)
upstream = self.release("whisper-bin-ubuntu-x64.tar.gz")
with fake_urlopen(http_error(404), upstream,
body(self.archive)) as calls:
path = ggml.install_program(ggml.WHISPER)
self.assertEqual(3, len(calls))
self.assertTrue(calls[2].full_url.endswith(
"whisper-bin-ubuntu-x64.tar.gz"))
self.assertTrue(os.path.isfile(path))
def test_a_fallback_to_the_processor_build_is_there_to_be_shown(self):
"""Until the Vulkan package is published every download lands the
processor build, and a graphics card sitting idle looks exactly like
one being used. The window asks this and says so."""
self.patch_attr(ggml, "_arch", lambda: "x64")
self.patch_attr(ggml, "_has_vulkan", lambda: True)
managed = self.release("Dikte-1.1.0-x86_64.AppImage")
managed["tag_name"] = "whisper.cpp-v1.9.3"
upstream = self.release("whisper-bin-ubuntu-x64.tar.gz")
with fake_urlopen(managed, upstream, body(self.archive)):
ggml.install_program(ggml.WHISPER)
self.assertTrue(ggml.vulkan_missing(ggml.WHISPER))
def test_a_machine_with_no_vulkan_is_not_told_it_is_missing_one(self):
# Nothing was on offer to fall back from, so there is nothing to say.
self.install("whisper-bin-ubuntu-x64.tar.gz")
self.assertFalse(ggml.vulkan_missing(ggml.WHISPER))
def test_a_release_with_nothing_for_this_machine_says_so(self):
self.patch_attr(ggml, "_arch", lambda: "x64")
with fake_urlopen(self.release("whisper-bin-Win32.zip")):
+217
View File
@@ -0,0 +1,217 @@
"""The release build that makes Linux Vulkan a one-click install."""
import hashlib
import io
import json
import os
import pathlib
import shutil
import subprocess
import sys
import tarfile
import tempfile
import unittest
from dikte import ggml
ROOT = pathlib.Path(__file__).parents[1]
PACKAGING = ROOT / "packaging" / "whisper-vulkan"
WORKFLOW = ROOT / ".github" / "workflows" / "whisper-vulkan.yml"
class WhisperVulkanPackaging(unittest.TestCase):
@unittest.skipUnless(sys.platform != "win32" and shutil.which("bash"),
"bash syntax check is unavailable")
def test_the_release_scripts_parse_as_shell(self):
for name in ("build-package.sh", "validate-package.sh",
"smoke-runtime.sh"):
script = PACKAGING / name
checked = subprocess.run(
["bash", "-n", script], capture_output=True, text=True,
)
self.assertEqual("", checked.stderr)
self.assertEqual(0, checked.returncode)
def test_the_workflow_builds_validates_smokes_and_publishes(self):
workflow = WORKFLOW.read_text(encoding="utf-8")
for step in ("Build deterministic archive",
"Verify reviewed archive digest",
"Validate archive and ELF contract",
"CPU fallback smoke test (no Vulkan loader)",
"Vulkan loader present, no device smoke test",
"Vulkan plugin-load smoke test (Mesa llvmpipe)",
"Publish dependency release"):
self.assertIn(step, workflow)
self.assertNotRegex(workflow, r"uses: [^\n]+@v\d+(?:\s|$)")
def test_publish_is_safe_for_dikte_and_limited_to_reviewed_master(self):
workflow = WORKFLOW.read_text(encoding="utf-8")
self.assertGreaterEqual(workflow.count("persist-credentials: false"), 2)
self.assertIn("github.ref == 'refs/heads/master'", workflow)
self.assertIn("--prerelease", workflow)
self.assertIn("--latest=false", workflow)
self.assertIn("--verify-tag", workflow)
self.assertIn("refusing to replace existing tag", workflow)
self.assertIn("^[0-9]+\\.[0-9]+\\.[0-9]+$", workflow)
self.assertIn("^[0-9a-f]{40}$", workflow)
publish_script = workflow.split(" - name: Publish dependency release", 1)[1]
publish_script = publish_script.split(" run: |", 1)[1]
self.assertNotIn("${{ inputs.", publish_script)
def test_bundle_ci_runs_only_for_what_the_bundle_is_built_from(self):
"""A 45 minute build on a README typo is a tax on every other change.
What ties ggml.py to the release is checked in this file instead, and
this file runs on every pull request in milliseconds."""
workflow = WORKFLOW.read_text(encoding="utf-8")
trigger = workflow.split("workflow_dispatch:", 1)[0]
self.assertIn("- packaging/whisper-vulkan/**", trigger)
self.assertIn("- .github/workflows/whisper-vulkan.yml", trigger)
for path in ("dikte/ggml.py", "tests/test_ggml.py",
"tests/test_packaging.py", "README.md", "README.tr.md"):
self.assertNotIn(f"- {path}", trigger)
def test_the_smoke_tests_run_what_dikte_runs(self):
"""-ng is what Dikte passes when its GPU setting is off, and a run
with it never asks for a backend at all. The three runs that have to
hold are the ones without it: no loader, a loader with nothing behind
it, and a working device."""
script = (PACKAGING / "smoke-runtime.sh").read_text(encoding="utf-8")
code = "\n".join(line for line in script.splitlines()
if not line.lstrip().startswith("#"))
self.assertNotIn("-ng", code)
for mode in ("cpu)", "noicd)", "vulkan)"):
self.assertIn(mode, script)
self.assertTrue((PACKAGING / "Dockerfile.runtime-noicd").is_file())
def test_an_unreviewed_version_is_reported_and_never_published(self):
"""The digest of a version nobody has reviewed cannot be known before
it is built, so the gate cannot be the only way through."""
workflow = WORKFLOW.read_text(encoding="utf-8")
self.assertIn("expected_sha256", workflow)
self.assertIn(
"refusing to publish an archive whose digest has not been reviewed",
workflow)
def test_the_shape_of_the_inputs_is_checked_before_they_are_used(self):
workflow = WORKFLOW.read_text(encoding="utf-8")
self.assertLess(workflow.index("- name: Validate source coordinates"),
workflow.index("- name: Check out pinned whisper.cpp"))
def test_the_validator_checks_tar_links_before_extraction(self):
validator = (PACKAGING / "validate-package.sh").read_text(
encoding="utf-8")
for check in ("member.issym()", "member.islnk()", "member.isdev()"):
self.assertIn(check, validator)
@unittest.skipUnless(sys.platform == "linux" and shutil.which("bash"),
"Linux packaging test is unavailable")
def test_the_validator_rejects_an_escaping_symlink(self):
asset = "whisper-bin-ubuntu-vulkan-x64"
with tempfile.TemporaryDirectory() as temporary:
output = pathlib.Path(temporary)
archive = output / f"{asset}.tar.gz"
with tarfile.open(archive, "w:gz") as bundle:
link = tarfile.TarInfo(f"{asset}/whisper-server")
link.type = tarfile.SYMTYPE
link.linkname = "/etc/passwd"
bundle.addfile(link, io.BytesIO())
digest = hashlib.sha256(archive.read_bytes()).hexdigest()
(output / f"{asset}.tar.gz.sha256").write_text(
f"{digest} {asset}.tar.gz\n", encoding="utf-8",
)
checked = subprocess.run(
["bash", PACKAGING / "validate-package.sh"],
env=os.environ | {"OUT_DIR": str(output)},
capture_output=True, text=True,
)
self.assertNotEqual(0, checked.returncode)
self.assertIn("unsafe symlink", checked.stderr)
def test_the_validator_checks_elf_architecture_dependencies_and_paths(self):
validator = (PACKAGING / "validate-package.sh").read_text(
encoding="utf-8")
for check in ("Advanced Micro Devices X86-64", "unexpected DT_NEEDED",
"path.read_bytes()"):
self.assertIn(check, validator)
def test_the_builder_and_its_downloads_are_pinned(self):
dockerfile = (PACKAGING / "Dockerfile.build").read_text(
encoding="utf-8")
self.assertRegex(dockerfile, r"FROM ubuntu@sha256:[0-9a-f]{64}")
self.assertIn("CMAKE_SHA256=", dockerfile)
self.assertIn("libvulkan-dev=", dockerfile)
self.assertIn("shaderc=", dockerfile)
key = (PACKAGING / "lunarg-signing-key-pub.asc").read_bytes()
key = key.replace(b"\r\n", b"\n")
self.assertEqual(
"aa1c3c29673140e77f0d6a9aaeed5d9b5621e305ead51c59fae4458bbb4df92b",
hashlib.sha256(key).hexdigest(),
)
def test_the_bundle_has_portable_dynamic_backends(self):
script = (PACKAGING / "build-package.sh").read_text(
encoding="utf-8")
for flag in ("GGML_BACKEND_DL=ON", "GGML_CPU_ALL_VARIANTS=ON",
"GGML_NATIVE=OFF", "GGML_OPENMP=OFF",
"GGML_VULKAN=ON"):
self.assertIn(flag, script)
self.assertIn("libggml-cpu*.so", script)
self.assertIn("libggml-vulkan.so", script)
def test_the_dependency_release_matches_the_installer(self):
workflow = WORKFLOW.read_text(encoding="utf-8")
script = (PACKAGING / "build-package.sh").read_text(
encoding="utf-8")
self.assertEqual("whisper.cpp-v1.9.3",
ggml.MANAGED_WHISPER_RELEASE)
self.assertEqual("v1.9.3", ggml.MANAGED_WHISPER_VERSION)
self.assertIn("RELEASE_TAG: whisper.cpp-v${{ inputs.whisper_version }}",
workflow)
self.assertIn("WHISPER_VERSION:=1.9.3", script)
commit = "371b5a7561823ab2bb32142d2751e35e7534727b"
self.assertIn(f"WHISPER_COMMIT:={commit}", script)
self.assertIn(commit, workflow)
self.assertIn(ggml.MANAGED_WHISPER_VULKAN, workflow)
self.assertIn(ggml.MANAGED_WHISPER_SHA256, workflow)
def test_the_bundle_carries_metadata_and_all_required_licenses(self):
script = (PACKAGING / "build-package.sh").read_text(
encoding="utf-8")
for name in ("BUILD-INFO.json", "SHA256SUMS", ".cdx.json"):
self.assertIn(name, script)
for name in ("cpp-httplib-MIT.txt", "nlohmann-json-MIT.txt"):
self.assertTrue((PACKAGING / "licenses" / name).is_file())
def _make_test_sbom(self):
with tempfile.TemporaryDirectory() as temporary:
root = pathlib.Path(temporary)
(root / "whisper-server").write_bytes(b"elf")
sbom = root / "whisper-bin-ubuntu-vulkan-x64.cdx.json"
environment = os.environ | {
"ROOT": str(root),
"VERSION": "1.9.3",
"COMMIT": "371b5a7561823ab2bb32142d2751e35e7534727b",
"EPOCH": "1787219223",
}
with sbom.open("w", encoding="utf-8") as output:
subprocess.run(
[sys.executable, PACKAGING / "make-sbom.py"],
env=environment, stdout=output, check=True,
)
return json.loads(sbom.read_text(encoding="utf-8")), sbom.name
def test_the_sbom_does_not_record_the_file_being_written(self):
document, sbom_name = self._make_test_sbom()
names = {component["name"] for component in document["components"]}
self.assertNotIn(sbom_name, names)
def test_the_sbom_lists_ggml(self):
document, _ = self._make_test_sbom()
names = {component["name"] for component in document["components"]}
self.assertIn("ggml", names)
if __name__ == "__main__":
unittest.main()
+5 -2
View File
@@ -42,9 +42,12 @@ class Directories(unittest.TestCase):
def test_a_mac_does_not_read_the_xdg_variables(self):
"""A Mac with them set from some other tool still stores in one place."""
with mock.patch.dict(os.environ, {"XDG_CONFIG_HOME": "/c"}):
# Something no temporary directory can be called: the home this runs
# under is a mkdtemp path, and a two-letter needle matched the "/c" in
# somebody's TMPDIR rather than the variable being read.
with mock.patch.dict(os.environ, {"XDG_CONFIG_HOME": "/xdg-elsewhere"}):
config_dir, _ = paths.directories("darwin")
self.assertNotIn("/c", config_dir.as_posix())
self.assertNotIn("xdg-elsewhere", config_dir.as_posix())
def test_windows_keeps_the_models_out_of_the_roaming_profile(self):
"""Settings roam with the account; several gigabytes must not."""
+397 -2
View File
@@ -6,8 +6,10 @@ save, so a setting added to one half and not the other is silently reset the
next time anybody presses Save. That is the failure this catches.
"""
import json
import os
import sys
import time
import unittest
from typing import ClassVar
from unittest import mock
@@ -21,13 +23,20 @@ from dikte import cleanup
from dikte import config as cfg
from dikte import ggml
from dikte import hotkey
from dikte import hub
from dikte import ipc
from dikte import overlay as overlay_module
from dikte import paste
from dikte import settings_ui
from dikte import update
from dikte.i18n import t
from tests.support import DikteTest, only_these_tools
# The harness below replaces this method on the class so that opening a window
# in a test never calls anybody; taken here, before any test runs, so the two
# tests about what it does when called still have the real one.
REAL_LOAD_HOSTED_MODELS = settings_ui.SettingsWindow._load_hosted_models
# One application for the whole run; Qt allows no second one.
_app = QApplication.instance() or QApplication([])
@@ -42,6 +51,8 @@ CHANGED = {
"paste_shortcut": "ctrl+shift+v",
"restore_clipboard": True,
"overlay_corner": "top-right",
"overlay_screen": "DP-1",
"overlay_follows_pointer": True,
"max_seconds": 120,
"skip_silent": False,
"silence_db": -42.0,
@@ -50,6 +61,7 @@ CHANGED = {
"openai_api_key": "sk-test-key",
"groq_api_key": "gsk-test-key",
"openrouter_api_key": "sk-or-test-key",
"gemini_api_key": "AIza-test-key",
"transcribe_provider": "openrouter",
"transcribe_model": "whisper-1",
"groq_transcribe_model": "whisper-large-v3",
@@ -59,6 +71,8 @@ CHANGED = {
"cleanup_model": "some/other-model",
"cleanup_claude_model": "opus",
"cleanup_codex_model": "gpt-5",
"cleanup_gemini_model": "gemini-2.5-flash",
"cleanup_agy_model": "gemini-3.1-pro-low",
"cleanup_reasoning": "high",
"local_model": "ggml-small.bin",
"local_gpu": False,
@@ -78,6 +92,7 @@ CHANGED = {
"assistant_codex_model": "gpt-5",
"assistant_codex_sandbox": "read-only",
"assistant_openrouter_model": "some/agent-model",
"assistant_agy_model": "gemini-3.1-pro-low",
"assistant_reasoning": "high",
"assistant_dir": "/tmp",
"assistant_timeout": 600,
@@ -135,6 +150,10 @@ class Settings(DikteTest):
"_load_transcribe_models"))
self.enterContext(mock.patch.object(settings_ui.SettingsWindow,
"_load_codex_models"))
self.enterContext(mock.patch.object(settings_ui.SettingsWindow,
"_load_agy_models"))
self.enterContext(mock.patch.object(settings_ui.SettingsWindow,
"_load_hosted_models"))
# The local model boxes fetch their own list the moment they are shown,
# from a thread, which is nobody's test failing but a real request.
self.enterContext(mock.patch.object(settings_ui.LocalModelBox,
@@ -161,7 +180,7 @@ class Settings(DikteTest):
def test_the_window_opens_with_every_tab_on_it(self):
window = self.window(cfg.Config())
tabs = window.findChildren(settings_ui.QTabWidget)[0]
self.assertEqual(tabs.count(), 9)
self.assertEqual(tabs.count(), 10)
self.assertEqual(window.windowTitle(), "Dikte Settings")
def test_no_tab_can_stretch_the_window_past_a_small_screen(self):
@@ -261,6 +280,7 @@ class Settings(DikteTest):
"""An OpenRouter id and a Claude alias are not the same field."""
window = self.window(cfg.Config())
boxes = {"openrouter": window.cleanup_model_row,
"opencode": window.cleanup_opencode_model_row,
"claude": window.cleanup_claude_model,
"codex": window.cleanup_codex_model}
for provider, box in boxes.items():
@@ -284,6 +304,90 @@ class Settings(DikteTest):
self.assertEqual(window.cleanup_codex_model.currentText(),
"my-own-model")
def test_opencode_answering_refills_both_of_its_boxes(self):
"""The fetched catalog replaces the built-in list in the cleanup and
agent boxes alike, and neither loses what was picked."""
conf = self.config(cleanup_opencode_model="my-own-model")
window = self.window(conf)
window._on_opencode_models_loaded(["glm-9", "kimi-k9"], "")
for combo in (window.cleanup_opencode_model,
window.assistant_opencode_model):
with self.subTest(combo=combo.objectName() or "combo"):
offered = [combo.itemText(i) for i in range(combo.count())]
self.assertEqual(offered, ["glm-9", "kimi-k9"])
self.assertEqual(window.cleanup_opencode_model.currentText(),
"my-own-model")
def test_opencode_s_list_arriving_at_open_leaves_the_other_boxes_alone(self):
conf = self.config(cleanup_opencode_model="my-own-model",
meeting_model="some/meeting-model")
window = self.window(conf)
before = [window.meeting_model.itemText(i)
for i in range(window.meeting_model.count())]
window._on_hosted_models_loaded("opencode", ["glm-5.3", "kimi-k3"])
combo = window.cleanup_opencode_model
offered = [combo.itemText(i) for i in range(combo.count())]
self.assertEqual(offered, ["glm-5.3", "kimi-k3"])
self.assertEqual(combo.currentText(), "my-own-model")
self.assertEqual([window.meeting_model.itemText(i)
for i in range(window.meeting_model.count())], before)
def test_opencode_cleanup_offers_a_fetch_button_of_its_own(self):
"""The OpenRouter button leaves the screen with its box, so OpenCode Go
carries its own."""
window = self.window(cfg.Config())
window._select_data(window.cleanup_provider, "opencode")
self.assertFalse(window.cleanup_opencode_model_row.isHidden())
self.assertTrue(window.cleanup_model_row.isHidden())
def test_agy_answering_refills_both_of_its_boxes(self):
"""The same arrangement as Codex: both boxes, nothing chosen is lost."""
conf = self.config(cleanup_agy_model="my-own-model")
window = self.window(conf)
window._on_agy_models_loaded(["gemini-4-flash-low", "gemini-4-pro-low"])
for combo in (window.cleanup_agy_model, window.assistant_agy_model):
with self.subTest(combo=combo.objectName() or "combo"):
offered = [combo.itemText(i) for i in range(combo.count())]
self.assertEqual(offered[1:],
["gemini-4-flash-low", "gemini-4-pro-low"])
self.assertEqual(window.cleanup_agy_model.currentText(), "my-own-model")
def test_openrouter_s_list_arriving_at_open_refills_cleanup_and_meetings(self):
conf = self.config(cleanup_model="my/own-model")
window = self.window(conf)
window._on_hosted_models_loaded("openrouter", ["a/one", "b/two"])
for combo in (window.cleanup_model, window.meeting_model):
with self.subTest(combo=combo.objectName() or "combo"):
offered = [combo.itemText(i) for i in range(combo.count())]
self.assertEqual(offered, ["a/one", "b/two"])
self.assertEqual(window.cleanup_model.currentText(), "my/own-model")
def test_google_s_list_arriving_at_open_refills_its_own_box_only(self):
window = self.window(self.config(cleanup_gemini_model="gemini-x"))
before = window.cleanup_model.count()
window._on_hosted_models_loaded("gemini", ["gemini-4-flash"])
offered = [window.cleanup_gemini_model.itemText(i)
for i in range(window.cleanup_gemini_model.count())]
self.assertEqual(offered, ["gemini-4-flash"])
self.assertEqual(window.cleanup_gemini_model.currentText(), "gemini-x")
self.assertEqual(window.cleanup_model.count(), before)
def test_no_key_no_call_home_at_open(self):
"""Opening Settings is not consent to be talked about to two vendors."""
window = self.window(self.config())
with mock.patch.dict(os.environ, {}, clear=True), \
mock.patch.object(settings_ui.threading, "Thread") as thread:
REAL_LOAD_HOSTED_MODELS(window)
thread.assert_not_called()
def test_a_key_on_file_is_fetched_with_at_open(self):
window = self.window(self.config(openrouter_api_key="sk-or-x",
gemini_api_key="AIza-x",
opencode_api_key="opencode-x"))
with mock.patch.object(settings_ui.threading, "Thread") as thread:
REAL_LOAD_HOSTED_MODELS(window)
self.assertEqual(thread.call_count, 3)
def test_the_update_line_names_the_version_that_is_running(self):
window = self.window(cfg.Config())
self.assertIn(settings_ui.__version__, window.update_status.text())
@@ -405,6 +509,20 @@ class Settings(DikteTest):
self.assertEqual(conf["transcribe_model"], "gpt-4o-transcribe")
self.assertEqual(conf["groq_transcribe_model"], "whisper-large-v3")
def test_the_file_model_is_saved_and_only_shown_for_openrouter(self):
self.write_config({"transcribe_provider": "openrouter",
"openrouter_file_model": "openai/whisper-large-v3"})
conf = cfg.Config()
window = self.window(conf)
self.assertEqual(window.file_model.currentText(), "openai/whisper-large-v3")
self.assertTrue(window.stt_form.isRowVisible(window.file_model_row))
window.file_model.setCurrentText(" deepgram/nova-3 ")
window._save()
self.assertEqual(conf["openrouter_file_model"], "deepgram/nova-3")
window.transcribe_provider.setCurrentIndex(
window.transcribe_provider.findData("openai"))
self.assertFalse(window.stt_form.isRowVisible(window.file_model_row))
def test_the_provider_box_offers_every_provider_config_knows(self):
window = self.window(cfg.Config())
offered = [window.transcribe_provider.itemData(i)
@@ -941,6 +1059,132 @@ class Overlay(DikteTest):
widget.show_recording()
widget._reposition()
def test_a_named_screen_is_used_instead_of_the_pointer_screen(self):
screen = mock.Mock()
screen.name.return_value = "DP-1"
screen.availableGeometry.return_value = settings_ui.QRect(1920, 0, 1920, 1080)
widget = self.overlay(screen_name="DP-1")
with mock.patch.object(QApplication, "screens", return_value=[screen]), \
mock.patch.object(QApplication, "screenAt") as screen_at:
widget._reposition()
screen_at.assert_not_called()
self.assertEqual(widget.pos(), QPoint(1948, 995))
def _screen(self, name, area):
screen = mock.Mock()
screen.name.return_value = name
screen.availableGeometry.return_value = area
return screen
def _kwin(self, *answer):
kwin = mock.Mock()
kwin.isValid.return_value = True
kwin.call.return_value.arguments.return_value = list(answer)
return kwin
def test_the_compositor_says_which_screen_the_pointer_is_on(self):
"""Wayland tells a client where the pointer is only while it is over one
of that client's own windows, so QCursor.pos() comes back at the origin
and every indicator lands on whichever screen holds it. KWin knows."""
screens = [self._screen("DP-1", settings_ui.QRect(0, 0, 1920, 1080)),
self._screen("DP-2", settings_ui.QRect(1920, 0, 1920, 1080))]
widget = self.overlay()
with mock.patch.object(overlay_module, "_kwin", self._kwin("DP-2")), \
mock.patch.object(QApplication, "screens", return_value=screens), \
mock.patch.object(QApplication, "screenAt") as screen_at:
widget._reposition()
screen_at.assert_not_called()
self.assertEqual(widget.pos(), QPoint(1948, 995))
def test_the_pointer_decides_when_the_compositor_will_not_say(self):
"""Every desktop but Plasma, and Plasma while KWin is being replaced."""
screens = [self._screen("DP-1", settings_ui.QRect(0, 0, 1920, 1080))]
widget = self.overlay()
with mock.patch.object(overlay_module, "_kwin", self._kwin()), \
mock.patch.object(QApplication, "screens", return_value=screens), \
mock.patch.object(QApplication, "screenAt",
return_value=screens[0]) as screen_at:
widget._reposition()
screen_at.assert_called()
self.assertEqual(widget.pos(), QPoint(28, 995))
def _two_screens(self):
return [self._screen("DP-1", settings_ui.QRect(0, 0, 1920, 1080)),
self._screen("DP-2", settings_ui.QRect(1920, 0, 1920, 1080))]
def _ticks_on(self, widget, screens, kwin):
"""Run the ribbon long enough for one look at where the pointer is."""
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens), \
mock.patch.object(QApplication, "screenAt", return_value=screens[0]):
for _ in range(overlay_module.FOLLOW_EVERY):
widget._tick()
def test_it_can_be_told_to_keep_up_with_the_pointer(self):
"""The screen it started on is not always the screen you end up on."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay(follow_pointer=True)
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
self.assertEqual(widget.pos(), QPoint(1948, 995))
kwin.call.return_value.arguments.return_value = ["DP-1"]
self._ticks_on(widget, screens, kwin)
self.assertEqual(widget.pos(), QPoint(28, 995))
def test_it_stays_where_it_appeared_unless_it_was_told_otherwise(self):
"""Left off, because an indicator that jumps desks mid-sentence is one
more thing moving while you are trying to talk."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay()
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
kwin.call.return_value.arguments.return_value = ["DP-1"]
self._ticks_on(widget, screens, kwin)
self.assertEqual(widget.pos(), QPoint(1948, 995))
def test_a_named_screen_is_never_left_for_the_pointer(self):
"""Naming one is the whole answer; following it would undo the naming."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay(screen_name="DP-1", follow_pointer=True)
with mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
self._ticks_on(widget, screens, kwin)
kwin.call.assert_not_called()
self.assertEqual(widget.pos(), QPoint(28, 995))
def test_the_one_on_top_goes_where_the_one_underneath_is(self):
"""Asking for itself would put the pair on two monitors, with this one
raised over a ribbon that is not underneath it."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
first = self.overlay()
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
first.show_recording()
kwin.call.return_value.arguments.return_value = ["DP-1"]
second = self.overlay(below=first)
second.show_busy("Asking Claude…")
self.assertEqual(first.pos(), QPoint(1948, 995))
self.assertEqual(second.pos(), QPoint(1948, 929))
def test_the_compositor_is_asked_only_now_and_then(self):
"""Every tick would be thirty conversations a second about a hand
moving a mouse."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay(follow_pointer=True)
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
kwin.call.reset_mock()
self._ticks_on(widget, screens, kwin)
self.assertEqual(kwin.call.call_count, 1)
def test_a_warning_and_an_error_both_show(self):
widget = self.overlay()
widget.show_warning("cleanup failed")
@@ -1001,7 +1245,11 @@ class MeetingSources(DikteTest):
mock.patch.object(settings_ui.SettingsWindow,
"_load_transcribe_models"), \
mock.patch.object(settings_ui.SettingsWindow,
"_load_codex_models"):
"_load_codex_models"), \
mock.patch.object(settings_ui.SettingsWindow,
"_load_agy_models"), \
mock.patch.object(settings_ui.SettingsWindow,
"_load_hosted_models"):
window = settings_ui.SettingsWindow(cfg.Config())
self.addCleanup(window.deleteLater)
self.addCleanup(window.close)
@@ -1033,6 +1281,10 @@ class LocalModels(DikteTest):
# And one with Codex on it would ask it for its model list.
self.enterContext(mock.patch.object(settings_ui.SettingsWindow,
"_load_codex_models"))
self.enterContext(mock.patch.object(settings_ui.SettingsWindow,
"_load_agy_models"))
self.enterContext(mock.patch.object(settings_ui.SettingsWindow,
"_load_hosted_models"))
def window(self, conf):
window = settings_ui.SettingsWindow(conf)
@@ -1096,6 +1348,123 @@ class LocalModels(DikteTest):
for row in range(box.repo.count()))
self.assertGreaterEqual(view.minimumWidth(), widest)
@staticmethod
def _item(name, size=1 << 20):
return hub.Item(name, f"https://example.invalid/{name}", size, "")
def test_a_row_with_nothing_to_fetch_does_not_offer_a_download(self):
# The model the settings name is not in the list any more, so its row
# was rebuilt from the name alone and carries no file to fetch. The
# button stayed lit and the press did nothing at all.
box = self.window(self.config(local_llm_model="gone.gguf")).local_llm
box.load("gone.gguf", "ggml-org/SmolLM3-3B-GGUF")
self.assertEqual(box.selected(), "gone.gguf")
self.assertFalse(box.download_button.isEnabled())
self.assertIn("gone.gguf", box.status.text())
self.assertIn("publisher", box.status.text())
def test_a_model_without_its_program_does_not_say_it_is_ready(self):
# The model runs on the program above it, and "Ready" over a missing
# one is what had people asking why nothing transcribed.
box = self.window(cfg.Config()).local_whisper
path = ggml.whisper_model_path("ggml-small.bin")
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(b"not really a model")
box.load("ggml-small.bin")
self.assertFalse(ggml.program_path(ggml.WHISPER))
self.assertNotIn("Ready", box.status.text())
self.assertIn("program", box.status.text())
def test_changing_the_publisher_changes_the_model(self):
# The model chosen under the old publisher is not published by the new
# one. Carried over, it was added back as "not downloaded" and selected
# again, and the box looked as though the change had not taken.
box = self.window(self.config(local_llm_model="gemma-3-4b-it-Q4_K_M.gguf",
local_llm_repo="ggml-org/gemma-3-4b-it-GGUF")).local_llm
box.load("gemma-3-4b-it-Q4_K_M.gguf", "ggml-org/gemma-3-4b-it-GGUF")
box.repo.blockSignals(True)
box.repo.setCurrentText("ggml-org/SmolLM3-3B-GGUF")
box.repo.blockSignals(False)
box._on_listed([("models", [self._item("SmolLM3-Q4_K_M.gguf")],
"ggml-org/SmolLM3-3B-GGUF")], "")
self.assertEqual(box.selected(), "SmolLM3-Q4_K_M.gguf")
self.assertEqual(box.model.count(), 1)
def test_a_list_for_a_publisher_that_is_no_longer_chosen_is_dropped(self):
# Every change starts its own request, and they do not come back in the
# order they went out.
box = self.window(cfg.Config()).local_llm
box.load("", "ggml-org/SmolLM3-3B-GGUF")
box.repo.blockSignals(True)
box.repo.setCurrentText("ggml-org/SmolLM3-3B-GGUF")
box.repo.blockSignals(False)
box._on_listed([("models", [self._item("SmolLM3-Q4_K_M.gguf")],
"ggml-org/SmolLM3-3B-GGUF")], "")
box._on_listed([("models", [self._item("gemma-3-4b-it-Q4_K_M.gguf")],
"ggml-org/gemma-3-4b-it-GGUF")], "")
self.assertEqual(box.selected(), "SmolLM3-Q4_K_M.gguf")
def test_the_publisher_box_is_not_asked_on_every_keystroke(self):
box = self.window(cfg.Config()).local_llm
with mock.patch.object(box, "_fetch_models") as fetch:
for text in ("g", "gg", "ggm", "ggml-org/SmolLM3-3B-GGUF"):
box.repo.setCurrentText(text)
fetch.assert_not_called()
box._later.setInterval(0)
box._later.start()
_app.processEvents()
time.sleep(0.05)
_app.processEvents()
self.assertEqual(fetch.call_count, 1)
def test_a_processor_build_where_the_vulkan_one_belongs_says_so(self):
# The Vulkan whisper-server is published by hand, and until it is
# there the download lands upstream's processor build. Said nowhere,
# an idle graphics card looks exactly like one that is being used.
binary = self.path("bin/whisper/v1.9.3/whisper-server")
binary.parent.mkdir(parents=True)
binary.write_text("")
binary.chmod(0o755)
self.path("bin/whisper/installed.json").write_text(json.dumps(
{"tag": "v1.9.3", "binary": str(binary), "backend": "processor"}))
# A whisper-server on this machine's PATH would win over the download.
self.patch_attr(ggml.shutil, "which", lambda name: None)
label = self.window(cfg.Config()).local_whisper.program_label.text()
self.assertIn("v1.9.3", label)
self.assertIn("Vulkan", label)
def test_an_ordinary_install_is_reported_without_a_word_about_vulkan(self):
binary = self.path("bin/whisper/v1.9.3/whisper-server")
binary.parent.mkdir(parents=True)
binary.write_text("")
binary.chmod(0o755)
self.path("bin/whisper/installed.json").write_text(json.dumps(
{"tag": "v1.9.3", "binary": str(binary)}))
self.patch_attr(ggml.shutil, "which", lambda name: None)
label = self.window(cfg.Config()).local_whisper.program_label.text()
self.assertIn("v1.9.3", label)
self.assertNotIn("Vulkan", label)
def test_a_downloaded_program_can_still_be_asked_for_again(self):
# The button used to disappear the moment anything landed, which left
# no way to pick up a newer whisper.cpp, or the Vulkan build on a
# machine whose driver was installed after Dikte was.
binary = self.path("bin/whisper/v1.9.3/whisper-server")
binary.parent.mkdir(parents=True)
binary.write_text("")
binary.chmod(0o755)
self.path("bin/whisper/installed.json").write_text(json.dumps(
{"tag": "v1.9.3", "binary": str(binary)}))
self.patch_attr(ggml.shutil, "which", lambda name: None)
box = self.window(cfg.Config()).local_whisper
self.assertTrue(box.install_button.isVisibleTo(box))
self.assertEqual(box.install_button.text(), t("Download again"))
def test_a_system_copy_is_not_offered_for_download(self):
# Nothing Dikte downloads would be run while one is on the PATH.
self.patch_attr(ggml.shutil, "which", lambda name: "/usr/bin/" + name)
box = self.window(cfg.Config()).local_whisper
self.assertFalse(box.install_button.isVisibleTo(box))
def test_only_the_chosen_transcriber_is_on_screen(self):
window = self.window(self.config(transcribe_provider="openai"))
self.assertTrue(window.stt_form.isRowVisible(window.transcribe_model_row))
@@ -1113,3 +1482,29 @@ class LocalModels(DikteTest):
self.assertFalse(window.cleanup_form.isRowVisible(window.cleanup_model_row))
# Its own thinking box, because the two default to opposite things.
self.assertFalse(window.cleanup_form.isRowVisible(window.cleanup_reasoning))
def test_each_cleaner_brings_its_own_model_row_and_no_other(self):
window = self.window(cfg.Config())
rows = {"openrouter": window.cleanup_model_row,
"gemini": window.cleanup_gemini_model_row,
"claude": window.cleanup_claude_model,
"codex": window.cleanup_codex_model,
"agy": window.cleanup_agy_model}
for chosen, row in rows.items():
with self.subTest(provider=chosen):
window._select_data(window.cleanup_provider, chosen)
for name, other in rows.items():
self.assertEqual(window.cleanup_form.isRowVisible(other),
name == chosen)
def test_each_agent_brings_its_own_box_and_no_other(self):
window = self.window(cfg.Config())
boxes = {"claude": window.claude_box, "codex": window.codex_box,
"agy": window.agy_box, "openrouter": window.openrouter_box}
for chosen, box in boxes.items():
with self.subTest(provider=chosen):
window._select_data(window.assistant_provider, chosen)
for name, other in boxes.items():
# isHidden rather than isVisible: the window itself is never
# shown in a test, so nothing in it is ever visible.
self.assertEqual(other.isHidden(), name != chosen)