Compare commits

...
Author SHA1 Message Date
yusufipek 70bc4c16fa Give a local model room to think without spending the answer on it
llama.cpp counts the thinking towards max_tokens along with the answer it
precedes, and the local ceiling was sized for the answer alone. Turning
Thinking up therefore came out of the reply rather than being added to
it, and on a short dictation the 512 floor is the whole budget, so the
model spent it in the think block and came back with nothing to paste.

Each rung of the ladder now carries its own budget, doubling from 256 at
"minimal" to 8192 at "maximum", added on top of the answer's share
rather than taken out of it. The rungs are small because cleanup is
punctuation and locally every one of these tokens is also a second of
somebody standing in front of the screen. "Off" keeps the old tight
ceiling untouched, and an empty setting is given a middling amount,
since a template that can think thinks by default and there is no way to
ask which kind of model this is.

The ceiling is also held under what the server was started with. Above
the context it is not a ceiling at all: the runaway it exists to stop
would run to the end of the context instead, which on CPU is minutes of
waiting. The prompt keeps its share at two characters to the token,
which is under any tokeniser's rate for natural language and so reserves
too much rather than promising room that is not there.

Separately, a reply cut off at somebody's ceiling was returned as if it
were whole. Half a sentence looks like a cleaned-up transcript and is
not one, so finish_reason is now read in both cleanup and chat. The
callers already keep the transcript they started with, which is the
better of the two. This one is not local-only: a hosted provider
stopping at its own output limit was silently pasted the same way.
2026-09-05 12:02:28 +03:00
Yusuf İpek b13b08fc38 Merge pull request #77 from yusufipk/claude/download-start-indicator-4e6835
Say the download started before the first byte arrives
2026-09-05 11:39:48 +03:00
yusufipek e282e6b0cf Say the download started before the first byte arrives
Opening the connection takes ten or twenty seconds, and the byte counts
under the model box only start after it. Until then the line read "X has
not been downloaded yet" beside a button that had just turned into Stop,
so a download that was running looked like a click that had not landed.

The stop had the same gap the other way round: should_stop is read
between blocks, and the wait for the server to answer is not between
blocks, so pressing Stop during it changed nothing on screen either.
2026-09-05 11:37:19 +03:00
Yusuf İpek 06e578d901 Merge pull request #76 from yusufipk/claude/model-selection-ui-organization-62dc90
Group the model lists and say which row this machine should take
2026-09-05 11:33:03 +03:00
Yusuf İpek 84c79b2d68 Merge pull request #75 from yusufipk/claude/mouse-cursor-tracking-issue-490b6d
Put the indicator on the screen the session is actually on
2026-09-05 11:24:32 +03:00
yusufipek f36348d536 Put the indicator on the screen the session is actually on
The indicator asked QCursor.pos() which screen to appear on, and Wayland tells a client where the pointer is only while it is over one of that client's own windows. The indicator is never under the pointer, so the answer came back stale, or at the origin when the pointer had never been over a window of ours. Measured on Plasma 6: the origin under the Wayland platform, and a point frozen for the whole run through XWayland, which is the platform Dikte actually uses. Every indicator therefore landed in the corner of whichever screen holds 0,0, which on a two monitor desk is the wrong screen most of the time.

KWin knows, and answers for it over D-Bus with activeOutputName, naming outputs the way Qt names screens: by connector, natively and through XWayland alike. That answer now comes before the pointer, and the pointer still decides everywhere else, which is right on X11 and no worse than before on other Wayland desktops. It is the active output and not the pointer's, so on Plasma the two are the same screen only where the active screen is set to follow the mouse, and otherwise it is the focused window that decides, which is where the typing is going anyway. Nothing in the settings window promises the pointer any more.

Deciding the screen once, when the indicator appears, leaves it behind when the work moves to another monitor mid-recording, so overlay_follows_pointer keeps it up to date while it is up. Off by default: a ribbon that changes desks mid-sentence is one more thing moving while you are trying to talk. The compositor is asked four times a second rather than at the ribbon's 33 ms, because a hand moving a mouse across a desk is slower than that, and the call is given a 200 ms timeout so a wedged compositor cannot freeze the indicator with it.

One indicator stacking on another takes that one's screen and never asks for its own. Asked for itself it would answer where the session is now, which is not where the ribbon underneath was put a minute ago, and the pair would end up a monitor apart with the top one raised over nothing. That one is checked every tick, since its answer costs nothing.
2026-09-05 11:19:02 +03:00
10 changed files with 442 additions and 25 deletions
+62 -6
View File
@@ -518,22 +518,65 @@ def _thinking(payload, provider, reasoning):
payload["reasoning"] = {"effort": reasoning, "exclude": True}
def local_ceiling(text):
# Room for the thinking on this machine, one budget per rung of the settings
# ladder. llama.cpp counts the thinking towards max_tokens along with the answer
# it precedes, so a ceiling sized for the answer alone leaves a model that
# thinks nothing to answer with. The rungs double, starting where a small model
# lands when it barely thinks at all: cleanup is punctuation, and locally every
# one of these tokens is also a second of somebody standing in front of the
# screen, so the low rungs are the ones meant to be used.
THINKING_ROOM = {
"minimal": 256, "low": 512, "medium": 1024,
"high": 2048, "xhigh": 4096, "max": 8192,
}
# An empty setting leaves it to the model, and the templates that can think
# think by default. Room for a middling amount of it, since there is no way to
# ask which kind of model this is.
DEFAULT_THINKING_ROOM = THINKING_ROOM["medium"]
def local_ceiling(text, reasoning="", context=0, prompt=""):
"""How much of a reply is worth waiting for from a model on this machine.
Cleanup gives back what it was given, near enough, so a reply several times
the length of the transcript is a model that has lost the thread rather than
one doing the job. A small one will happily repeat the transcript until the
context is full, and every one of those tokens is a second of somebody
waiting. A hosted model is left alone: there the same runaway is rare, and a
ceiling would cut the minutes short instead.
waiting, with only the hour-long local timeout underneath. A hosted model is
left alone: there the same runaway is rare, and a ceiling would cut the
minutes short instead.
The answer's share is the transcript's length in characters spent as a
budget in tokens, so what it really allows is two to four times the
transcript depending on how well the language tokenises. Turkish sits at the
tight end of that and still has room to spare for a reply that is meant to
come back the same length it went in.
Thinking is added on top of that share rather than taken out of it. Sharing
one budget is what makes turning thinking up quietly cost the answer, and on
a short dictation the 512 floor is the whole budget, so the answer is what
goes missing first.
`context` is what the server was started with, and the whole of it is the
real limit whatever is asked for here: a ceiling above it is not a ceiling,
because the runaway it exists to stop would run to the end of the context
instead. So the ceiling is held below what the prompt leaves. Two characters
to the token is under any tokeniser's rate for natural language, Turkish
included, which makes the reserve an over-estimate rather than a promise of
room that is not there.
"""
return max(512, len(text))
answer = max(512, len(text))
if reasoning != "none":
answer += THINKING_ROOM.get(reasoning, DEFAULT_THINKING_ROOM)
context = int(context or 0)
if not context:
return answer
return max(256, min(answer, context - (len(prompt) + len(text)) // 2))
def cleanup(text, api_key, model, system_prompt, reasoning="",
base_url=OPENROUTER_URL, timeout=180, provider="openrouter",
service="OpenRouter", aborter=None):
service="OpenRouter", aborter=None, context=0):
if not api_key and provider != "local-llm":
raise ApiError(t("{service} API key is empty. Add it in Settings.",
service=service))
@@ -546,7 +589,8 @@ def cleanup(text, api_key, model, system_prompt, reasoning="",
],
}
if provider == "local-llm":
payload["max_tokens"] = local_ceiling(text)
payload["max_tokens"] = local_ceiling(text, reasoning, context,
system_prompt)
_thinking(payload, provider, reasoning)
try:
data = _request(
@@ -570,6 +614,13 @@ def cleanup(text, api_key, model, system_prompt, reasoning="",
raise ApiError(t("The cleanup model spent its whole reply on "
"thinking. Set Thinking to \u201cOff\u201d."))
raise ApiError(t("The cleanup model returned an empty reply."))
if choices[0].get("finish_reason") == "length":
# Cut off at somebody's ceiling: ours locally, the provider's otherwise.
# What came back is a sentence that stops mid-word, and cleanup is meant
# to hand back the whole dictation, so the half is refused rather than
# returned. The callers keep the transcript they started with, which is
# the better of the two.
raise ApiError(t("The cleanup model was cut off before it finished."))
return content
@@ -605,6 +656,11 @@ def chat(messages, api_key, model, system_prompt, reasoning="",
content = ((choices[0].get("message") or {}).get("content") or "").strip()
if not content:
raise ApiError(t("The model returned an empty reply."))
if choices[0].get("finish_reason") == "length":
# An answer that stops mid-sentence reads like a whole one once it has
# been pasted, so it is refused here for the same reason cleanup refuses
# a half transcript.
raise ApiError(t("The model was cut off before it finished."))
return content
+8 -6
View File
@@ -164,12 +164,14 @@ class Dikte:
self._front_watch = None
self.overlay = Overlay(self.conf["overlay_corner"],
screen_name=self.conf["overlay_screen"])
screen_name=self.conf["overlay_screen"],
follow_pointer=self.conf["overlay_follows_pointer"])
# The agent's indicator sits on top of the dictation one when both are
# up, and drops into the corner when it is alone there.
self.ask_overlay = Overlay(self.conf["overlay_corner"], below=self.overlay,
dismissable=True,
screen_name=self.conf["overlay_screen"])
screen_name=self.conf["overlay_screen"],
follow_pointer=self.conf["overlay_follows_pointer"])
self.recorder = audio.Recorder()
self.pipeline = Pipeline(self.conf)
self.ask_pipeline = Pipeline(self.conf)
@@ -1297,10 +1299,10 @@ class Dikte:
threading.Thread(target=warm, daemon=True).start()
def _apply_settings(self):
self.overlay.corner = self.conf["overlay_corner"]
self.overlay.screen_name = self.conf["overlay_screen"]
self.ask_overlay.corner = self.conf["overlay_corner"]
self.ask_overlay.screen_name = self.conf["overlay_screen"]
for indicator in (self.overlay, self.ask_overlay):
indicator.corner = self.conf["overlay_corner"]
indicator.screen_name = self.conf["overlay_screen"]
indicator.follow_pointer = self.conf["overlay_follows_pointer"]
self._apply_local()
self._build_tray()
self._refresh_tray()
+3
View File
@@ -121,6 +121,9 @@ def _local(text, conf, system_prompt, timeout, aborter=None):
base_url=api.serving(ggml.llm),
timeout=max(timeout, api.LOCAL_TIMEOUT),
provider="local-llm", service=service, aborter=aborter,
# The ceiling is only a ceiling while it sits under what the server
# was started with; above that the context is what stops the reply.
context=ggml.llm.settings()["context"],
)
except api.ApiError as exc:
# A server that died mid-request would otherwise report only that the
+3
View File
@@ -470,6 +470,9 @@ DEFAULTS = {
"evdev_hotkey": False,
"overlay_corner": "bottom-left",
"overlay_screen": "",
# Off, so that an indicator stays where it appeared unless it is asked to
# keep up with the pointer. Nothing to say when a screen is named above.
"overlay_follows_pointer": False,
"keep_audio": False,
"history_limit": 200,
# A look at the releases page once a day, and nothing more than a look:
+8 -1
View File
@@ -189,8 +189,10 @@ TR = {
"Restore the previous clipboard after pasting":
"Yapıştırdıktan sonra eski pano içeriğini geri koy",
"Indicator screen": "Gösterge ekranı",
"Follow the mouse pointer": "Fare imlecini takip et",
"Follow the active screen": "Etkin ekranı takip et",
"{name} (not connected)": "{name} (bağlı değil)",
"Move it when the active screen changes":
"Etkin ekran değiştiğinde göstergeyi de taşı",
"Indicator corner": "Gösterge köşesi",
"bottom-left": "sol-alt",
"bottom-right": "sağ-alt",
@@ -791,6 +793,7 @@ TR = {
"İndirildi, sürüm {version}. Vulkan sürümü yoktu, bu sürüm işlemcide çalışıyor.",
"Fetching the model list…": "Model listesi çekiliyor…",
"Downloading…": "İndiriliyor…",
"Starting the download…": "İndirme başlatılıyor…",
"Downloading: {done} of {total}{share}": "İndiriliyor: {done} / {total}{share}",
"Download stopped.": "İndirme durduruldu.",
"Ready: {name}.": "Hazır: {name}.",
@@ -956,6 +959,10 @@ TR = {
"“Off”.":
"Temizleme modeli bütün yanıtını düşünmeye harcadı. Düşünme'yi "
"“Kapalı” yap.",
"The cleanup model was cut off before it finished.":
"Temizleme modeli bitiremeden kesildi.",
"The model was cut off before it finished.":
"Model bitiremeden kesildi.",
# --- this pass's new messages ---------------------------------------
"Audio recorder stopped before receiving sound":
+104 -10
View File
@@ -1,6 +1,7 @@
"""The small recording indicator that appears in a screen corner without taking focus."""
import math
import os
import sys
from PyQt6.QtCore import Qt, QTimer, QRectF, QPointF
@@ -15,6 +16,7 @@ MIN_WIDTH = 210
MAX_WIDTH = 460
MARGIN = 28
GAP = 10 # between two indicators sharing a corner
FOLLOW_EVERY = 8 # ticks between two looks for the pointer: about four a second
BG = QColor(22, 24, 29, 238)
BORDER = QColor(255, 255, 255, 28)
@@ -37,16 +39,68 @@ STATE_COLORS = {"recording": REC, "asking": ASK, "meeting": REC, "busy": BUSY,
LIVE = ("recording", "asking", "meeting")
# KWin's interface, kept once one has been built. See _compositor_screen.
_kwin = None
def _compositor_screen():
"""The screen KWin says the session is on, or None where nothing says.
Wayland tells a client where the pointer is only while it is over one of
that client's own windows, and the indicator is never under the pointer, so
QCursor.pos() answers with a stale point or, when the pointer has never
been over a window of ours, with the origin. Either way the indicator lands
in the corner of whichever screen holds 0,0 instead of the one being worked
on, and on a two-monitor desk that is the wrong screen most of the time.
KWin does know, and it names outputs the way Qt names screens, by
connector, natively and through XWayland alike. No other Wayland desktop
answers this, so the rest are left with the pointer, which is right on X11
and wrong on Wayland exactly as before.
What it answers with is the active output, which is the one under the
pointer only where Plasma is set to let the active screen follow the mouse.
Under the default, click to focus, it is the focused window's screen, so
the indicator lands where the typing is going rather than where the mouse
was left. Which is why nothing here, and nothing in the settings window,
promises the pointer.
"""
global _kwin
if _kwin is None or not _kwin.isValid():
# Which also leaves macOS and Windows out, where nothing sets it and
# the pointer can be asked where it is like anywhere else.
desktop = os.environ.get("XDG_CURRENT_DESKTOP", "").lower()
if "kde" not in desktop and "plasma" not in desktop:
return None
try:
from PyQt6.QtDBus import QDBusConnection, QDBusInterface
_kwin = QDBusInterface("org.kde.KWin", "/KWin", "org.kde.KWin",
QDBusConnection.sessionBus())
except Exception:
return None
if not _kwin.isValid():
return None
# A compositor busy enough not to answer in a fifth of a second is one
# the indicator should stop waiting for, not one it should freeze with.
_kwin.setTimeout(200)
answer = _kwin.call("activeOutputName").arguments()
name = answer[0] if answer else ""
return next((item for item in QApplication.screens() if item.name() == name),
None)
class Overlay(QWidget):
"""One indicator. Give it `below` and it stacks on top of that one instead
of covering it, which is what lets a dictation and a command to the agent be
under way at the same time and still both be visible."""
def __init__(self, corner="bottom-left", below=None, dismissable=False,
screen_name=""):
screen_name="", follow_pointer=False):
super().__init__(None)
self.corner = corner
self.screen_name = screen_name
# Whether it goes on following the pointer once it is up, rather than
# settling on the screen it appeared on.
self.follow_pointer = follow_pointer
self.below = below
# A job that can run for ten minutes should not have to be watched for
# ten minutes. Clicking such an indicator puts the progress away; the
@@ -65,6 +119,8 @@ class Overlay(QWidget):
self.seconds = 0.0
self._phase = 0.0
self._concealed = True
self._shown_on = "" # the screen it was last put on, by name
self._looks = 0 # ticks since the pointer was last looked for
flags = (
Qt.WindowType.FramelessWindowHint
@@ -252,16 +308,52 @@ class Overlay(QWidget):
min(MAX_WIDTH, metrics.horizontalAdvance(self.message) + extra))
self.resize(width, HEIGHT)
def _reposition(self):
# The screen the settings name, or, when none is named or it is not
# plugged in right now, where the user actually is. Names are connector
# names on X11 and model names on macOS, where two identical monitors
# can share one; the first then wins.
screen = next(
def _screen(self):
"""The screen this indicator belongs on right now.
The one the settings name, or, when none is named or it is not plugged
in right now, where the user actually is. Names are connector names on
X11 and model names on macOS, where two identical monitors can share
one; the first then wins.
One stacking on another belongs on that one's screen and nowhere else.
Asked for itself it would answer where the user is now, which is not
where the ribbon it stacks on was put a minute ago, and the pair would
end up a monitor apart with this one raised over nothing.
"""
if self.below is not None and self.below.showing:
under = next((item for item in QApplication.screens()
if item.name() == self.below._shown_on), None)
if under is not None:
return under
named = next(
(item for item in QApplication.screens() if item.name() == self.screen_name),
None,
)
screen = screen or QApplication.screenAt(QCursor.pos()) or QApplication.primaryScreen()
return (named or _compositor_screen()
or QApplication.screenAt(QCursor.pos())
or QApplication.primaryScreen())
def _wandered_off(self):
"""Whether the pointer has left the screen the indicator is on.
Only asked while it is following, and only every few ticks: the answer
costs a word with the compositor, and a hand moving a mouse across a
desk is slow next to a 33 ms ribbon. Every tick for one that stacks on
another, where the answer is free and waiting a third of a second for
it would leave the pair split over two monitors for that long.
"""
if not self.follow_pointer or self.screen_name:
return False
if self.below is None or not self.below.showing:
self._looks = (self._looks + 1) % FOLLOW_EVERY
if self._looks:
return False
return self._screen().name() != self._shown_on
def _reposition(self):
screen = self._screen()
self._shown_on = screen.name()
area = screen.availableGeometry()
left = "left" in self.corner
top = "top" in self.corner
@@ -277,8 +369,10 @@ class Overlay(QWidget):
def _tick(self):
self._phase += 0.12
# The one underneath can come and go while this one is up; drop back to
# the corner when it does rather than leaving a gap where it was.
if self.below is not None and self.below.showing != self._stacked:
# the corner when it does rather than leaving a gap where it was. And
# the screen under the pointer can change while it is up too.
moved = self.below is not None and self.below.showing != self._stacked
if moved or self._wandered_off():
self._reposition()
if self.state in LIVE and not self.paused:
# keep the ribbon moving even through a pause in speech
+33 -1
View File
@@ -701,12 +701,21 @@ class LocalModelBox(QGroupBox):
def _download(self):
if self._downloading:
self._stop = True
# The flag is only read between blocks, and the wait for the server
# to answer is not between blocks: a click during it changes
# nothing on screen for as long as the connection takes.
self.status.setText(t("Stopping…"))
return
item = self._current_item()
if item is None:
return
self._downloading, self._stop = True, False
self._refresh_buttons()
# Opening the connection can take ten or twenty seconds, and the first
# byte counts are what the line below would otherwise wait for. Left
# saying "not downloaded yet" beside a button that now reads Stop, a
# download that started looks like a click that did not register.
self.status.setText(t("Starting the download…"))
def work():
try:
@@ -1068,7 +1077,11 @@ class SettingsWindow(QDialog):
form = QFormLayout(page)
self.indicator_screen = QComboBox()
self.indicator_screen.addItem(t("Follow the mouse pointer"), "")
# The active screen rather than the pointer, for the reason in
# overlay._compositor_screen: it is what a compositor will answer for,
# and on Plasma the two are one screen only where the active screen is
# set to follow the mouse.
self.indicator_screen.addItem(t("Follow the active screen"), "")
for screen in QGuiApplication.screens():
# The native resolution, so that a scaled 4K screen reads
# 3840 × 2160 and not the 1920 × 1080 Qt sees through the scale.
@@ -1082,12 +1095,26 @@ class SettingsWindow(QDialog):
)
form.addRow(t("Indicator screen"), self.indicator_screen)
# Only the screen it appeared on is decided when it appears; this is
# what makes it keep up with a session that moves to another one
# mid-recording. The active screen and not the pointer, because that is
# what a compositor will answer for: on Plasma the two are the same
# screen only where the active screen is set to follow the mouse, and
# otherwise it is the focused window that decides. Nothing to offer
# when a screen is named above, since that name is the whole answer.
self.follow_pointer = QCheckBox(t("Move it when the active screen changes"))
self.indicator_screen.currentIndexChanged.connect(self._sync_follow_pointer)
form.addRow("", self.follow_pointer)
self.corner = QComboBox()
for value in CORNERS:
self.corner.addItem(t(value), value)
form.addRow(t("Indicator corner"), self.corner)
return page
def _sync_follow_pointer(self):
self.follow_pointer.setEnabled(not self.indicator_screen.currentData())
def _api_tab(self):
page = QWidget()
outer = QVBoxLayout(page)
@@ -2006,6 +2033,8 @@ class SettingsWindow(QDialog):
if screen_name and self.indicator_screen.findData(screen_name) < 0:
self.indicator_screen.addItem(t("{name} (not connected)", name=screen_name), screen_name)
self._select_data(self.indicator_screen, screen_name)
self.follow_pointer.setChecked(conf["overlay_follows_pointer"])
self._sync_follow_pointer()
self._select_data(self.corner, conf["overlay_corner"])
self.max_seconds.setValue(conf["max_seconds"])
self.skip_silent.setChecked(conf["skip_silent"])
@@ -2126,6 +2155,9 @@ class SettingsWindow(QDialog):
conf["paste_shortcut"] = self.paste_shortcut.currentText().strip()
conf["restore_clipboard"] = self.restore_clipboard.isChecked()
conf["overlay_screen"] = self.indicator_screen.currentData() or ""
# Read even while it is greyed out, so that naming a screen and taking
# the name back again does not clear a preference nobody touched.
conf["overlay_follows_pointer"] = self.follow_pointer.isChecked()
conf["overlay_corner"] = self.corner.currentData() or "bottom-left"
conf["max_seconds"] = self.max_seconds.value()
conf["skip_silent"] = self.skip_silent.isChecked()
+37 -1
View File
@@ -463,6 +463,29 @@ class Cleanup(DikteTest):
with fake_urlopen(chat_reply(" ")), self.assertRaises(api.ApiError):
api.cleanup("hello", "k", "m", "p")
def test_a_reply_cut_off_at_a_ceiling_is_refused_rather_than_pasted(self):
# Half a sentence looks like a cleaned-up transcript and is not one. The
# caller keeps what it was given, which is the whole dictation.
reply = {"choices": [{"message": {"content": "Hello, and then the"},
"finish_reason": "length"}]}
with fake_urlopen(reply), self.assertRaises(api.ApiError) as caught:
api.cleanup("hello", "k", "m", "p")
self.assertIn("cut off", str(caught.exception))
def test_a_reply_that_stopped_on_its_own_is_kept(self):
reply = {"choices": [{"message": {"content": "Hello."},
"finish_reason": "stop"}]}
with fake_urlopen(reply):
self.assertEqual(api.cleanup("hello", "k", "m", "p"), "Hello.")
def test_all_thinking_is_named_before_the_ceiling_it_was_cut_at(self):
"""Both are true at once, and only one of them says what to change."""
reply = {"choices": [{"message": {"content": "", "reasoning": "hmm"},
"finish_reason": "length"}]}
with fake_urlopen(reply), self.assertRaises(api.ApiError) as caught:
api.cleanup("hello", "k", "m", "p")
self.assertIn("Thinking", str(caught.exception))
def test_a_rate_limit_is_explained(self):
with fake_urlopen(http_error(429)), \
self.assertRaises(api.ApiError) as caught:
@@ -471,6 +494,14 @@ class Cleanup(DikteTest):
class Chat(DikteTest):
def test_an_answer_cut_off_at_a_ceiling_is_refused_rather_than_pasted(self):
# Half an answer reads like a whole one once it is on the screen.
reply = {"choices": [{"message": {"content": "Booked it for the"},
"finish_reason": "length"}]}
with fake_urlopen(reply), self.assertRaises(api.ApiError) as caught:
api.chat([{"role": "user", "content": "book it"}], "k", "m", "p")
self.assertIn("cut off", str(caught.exception))
def test_the_history_is_sent_after_the_system_prompt(self):
history = [{"role": "user", "content": "book it"},
{"role": "assistant", "content": "done"}]
@@ -616,11 +647,13 @@ if __name__ == "__main__":
class FakeServer:
"""A ggml.Server as far as api.py is concerned."""
def __init__(self, url="http://127.0.0.1:9999/v1", fails="", log=""):
def __init__(self, url="http://127.0.0.1:9999/v1", fails="", log="",
context=8192):
self.url = url
self.fails = fails
self.log = log
self.starts = 0
self.context = context
def serve(self):
self.starts += 1
@@ -631,6 +664,9 @@ class FakeServer:
def error(self):
return self.log
def settings(self):
return {"context": self.context}
LOCAL = api.Target("local", "Local whisper", "", "", "ggml-base.bin")
+46
View File
@@ -411,6 +411,52 @@ class Here(DikteTest):
cleanup.run("uh, done", self.conf, "the rules")
self.assertEqual(sent_json(calls[0])["max_tokens"], 512)
def test_thinking_is_given_room_of_its_own_rather_than_the_answer_s(self):
# llama.cpp counts the thinking towards the same ceiling, so a rung that
# took its budget out of the answer would leave a short dictation with
# nothing to reply with. On a context roomy enough that the clamp the
# top rung would otherwise meet is not what is being measured.
self.patch_attr(ggml, "llm", FakeServer(context=32768))
for rung, room in api.THINKING_ROOM.items():
with self.subTest(rung=rung):
self.conf["local_llm_reasoning"] = rung
with fake_urlopen(chat_reply("Done.")) as calls:
cleanup.run("uh, done", self.conf, "the rules")
self.assertEqual(sent_json(calls[0])["max_tokens"], 512 + room)
def test_each_rung_of_the_ladder_thinks_longer_than_the_one_below(self):
rungs = [api.THINKING_ROOM[name] for name in
("minimal", "low", "medium", "high", "xhigh", "max")]
self.assertEqual(rungs, sorted(rungs))
self.assertEqual(len(set(rungs)), len(rungs))
def test_the_models_own_default_is_given_room_to_think_in_too(self):
# Nothing is sent, so a template that thinks will think, and the ceiling
# has to survive that as well.
self.conf["local_llm_reasoning"] = ""
with fake_urlopen(chat_reply("Done.")) as calls:
cleanup.run("uh, done", self.conf, "the rules")
self.assertEqual(sent_json(calls[0])["max_tokens"],
512 + api.DEFAULT_THINKING_ROOM)
def test_the_ceiling_stays_under_the_context_the_server_was_started_with(self):
# Above the context there is no ceiling at all: the runaway would run to
# the end of the context instead of stopping where this says.
self.patch_attr(ggml, "llm", FakeServer(context=2048))
self.conf["local_llm_reasoning"] = "max"
with fake_urlopen(chat_reply("Done.")) as calls:
cleanup.run("uh, done", self.conf, "the rules")
self.assertLess(sent_json(calls[0])["max_tokens"], 2048)
def test_the_prompt_keeps_its_share_of_a_small_context(self):
self.patch_attr(ggml, "llm", FakeServer(context=2048))
self.conf["local_llm_reasoning"] = "max"
with fake_urlopen(chat_reply("Done.")) as calls:
cleanup.run("x" * 2000, self.conf, "the rules")
# 2048 less half the characters of prompt and transcript together.
self.assertEqual(sent_json(calls[0])["max_tokens"],
2048 - (len("the rules") + 2000) // 2)
def test_a_reply_that_was_all_thinking_names_the_setting_that_fixes_it(self):
reply = {"choices": [{"message": {"content": "", "reasoning": "hmm"}}]}
with fake_urlopen(reply), self.assertRaises(api.ApiError) as caught:
+138
View File
@@ -52,6 +52,7 @@ CHANGED = {
"restore_clipboard": True,
"overlay_corner": "top-right",
"overlay_screen": "DP-1",
"overlay_follows_pointer": True,
"max_seconds": 120,
"skip_silent": False,
"silence_db": -42.0,
@@ -1069,6 +1070,121 @@ class Overlay(DikteTest):
screen_at.assert_not_called()
self.assertEqual(widget.pos(), QPoint(1948, 995))
def _screen(self, name, area):
screen = mock.Mock()
screen.name.return_value = name
screen.availableGeometry.return_value = area
return screen
def _kwin(self, *answer):
kwin = mock.Mock()
kwin.isValid.return_value = True
kwin.call.return_value.arguments.return_value = list(answer)
return kwin
def test_the_compositor_says_which_screen_the_pointer_is_on(self):
"""Wayland tells a client where the pointer is only while it is over one
of that client's own windows, so QCursor.pos() comes back at the origin
and every indicator lands on whichever screen holds it. KWin knows."""
screens = [self._screen("DP-1", settings_ui.QRect(0, 0, 1920, 1080)),
self._screen("DP-2", settings_ui.QRect(1920, 0, 1920, 1080))]
widget = self.overlay()
with mock.patch.object(overlay_module, "_kwin", self._kwin("DP-2")), \
mock.patch.object(QApplication, "screens", return_value=screens), \
mock.patch.object(QApplication, "screenAt") as screen_at:
widget._reposition()
screen_at.assert_not_called()
self.assertEqual(widget.pos(), QPoint(1948, 995))
def test_the_pointer_decides_when_the_compositor_will_not_say(self):
"""Every desktop but Plasma, and Plasma while KWin is being replaced."""
screens = [self._screen("DP-1", settings_ui.QRect(0, 0, 1920, 1080))]
widget = self.overlay()
with mock.patch.object(overlay_module, "_kwin", self._kwin()), \
mock.patch.object(QApplication, "screens", return_value=screens), \
mock.patch.object(QApplication, "screenAt",
return_value=screens[0]) as screen_at:
widget._reposition()
screen_at.assert_called()
self.assertEqual(widget.pos(), QPoint(28, 995))
def _two_screens(self):
return [self._screen("DP-1", settings_ui.QRect(0, 0, 1920, 1080)),
self._screen("DP-2", settings_ui.QRect(1920, 0, 1920, 1080))]
def _ticks_on(self, widget, screens, kwin):
"""Run the ribbon long enough for one look at where the pointer is."""
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens), \
mock.patch.object(QApplication, "screenAt", return_value=screens[0]):
for _ in range(overlay_module.FOLLOW_EVERY):
widget._tick()
def test_it_can_be_told_to_keep_up_with_the_pointer(self):
"""The screen it started on is not always the screen you end up on."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay(follow_pointer=True)
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
self.assertEqual(widget.pos(), QPoint(1948, 995))
kwin.call.return_value.arguments.return_value = ["DP-1"]
self._ticks_on(widget, screens, kwin)
self.assertEqual(widget.pos(), QPoint(28, 995))
def test_it_stays_where_it_appeared_unless_it_was_told_otherwise(self):
"""Left off, because an indicator that jumps desks mid-sentence is one
more thing moving while you are trying to talk."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay()
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
kwin.call.return_value.arguments.return_value = ["DP-1"]
self._ticks_on(widget, screens, kwin)
self.assertEqual(widget.pos(), QPoint(1948, 995))
def test_a_named_screen_is_never_left_for_the_pointer(self):
"""Naming one is the whole answer; following it would undo the naming."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay(screen_name="DP-1", follow_pointer=True)
with mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
self._ticks_on(widget, screens, kwin)
kwin.call.assert_not_called()
self.assertEqual(widget.pos(), QPoint(28, 995))
def test_the_one_on_top_goes_where_the_one_underneath_is(self):
"""Asking for itself would put the pair on two monitors, with this one
raised over a ribbon that is not underneath it."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
first = self.overlay()
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
first.show_recording()
kwin.call.return_value.arguments.return_value = ["DP-1"]
second = self.overlay(below=first)
second.show_busy("Asking Claude…")
self.assertEqual(first.pos(), QPoint(1948, 995))
self.assertEqual(second.pos(), QPoint(1948, 929))
def test_the_compositor_is_asked_only_now_and_then(self):
"""Every tick would be thirty conversations a second about a hand
moving a mouse."""
screens = self._two_screens()
kwin = self._kwin("DP-2")
widget = self.overlay(follow_pointer=True)
with mock.patch.object(overlay_module, "_kwin", kwin), \
mock.patch.object(QApplication, "screens", return_value=screens):
widget.show_recording()
kwin.call.reset_mock()
self._ticks_on(widget, screens, kwin)
self.assertEqual(kwin.call.call_count, 1)
def test_a_warning_and_an_error_both_show(self):
widget = self.overlay()
widget.show_warning("cleanup failed")
@@ -1220,6 +1336,28 @@ class LocalModels(DikteTest):
self.assertIn("10", box.program_label.text())
self.assertIn("20", box.status.text())
def test_a_download_says_something_before_the_first_byte(self):
# Opening the connection takes ten or twenty seconds, and the byte
# counts only start after it. The line underneath still read "has not
# been downloaded yet" beside a button that now said Stop, so a
# download that had started looked like a click that had not landed.
box = self.window(cfg.Config()).local_llm
box.load("", "ggml-org/SmolLM3-3B-GGUF")
box.repo.blockSignals(True)
box.repo.setCurrentText("ggml-org/SmolLM3-3B-GGUF")
box.repo.blockSignals(False)
box._on_listed([("models", [self._item("SmolLM3-Q4_K_M.gguf")],
"ggml-org/SmolLM3-3B-GGUF")], "")
with mock.patch.object(settings_ui.threading, "Thread"):
box._download()
self.assertIn("Starting", box.status.text())
# And the same again for the stop, which is read between blocks and so
# not read at all while the connection is still being opened.
with mock.patch.object(settings_ui.threading, "Thread"):
box._download()
self.assertTrue(box._stop)
self.assertIn("Stopping", box.status.text())
def test_a_long_model_name_is_not_cut_in_half(self):
# The list under a combo box takes the box's width and elides what does
# not fit, in the middle: "ggml-org/Qwen....7B-Base-GGUF".