mirror of
https://github.com/yusufipk/dikte.git
synced 2026-09-11 19:06:11 +00:00
Not every model behind /audio/transcriptions marks segments the way whisper does. microsoft/mai-transcribe-2 answers a fourteen minute video with three of them, one per paragraph, and to_srt turns each into a cue that stays up for minutes. The model is worth keeping for what it hears, so the times are taken from somewhere else instead: the same request now asks for word timestamps too, and where the segments come back too long to be cues, the cues are cut out of the words. A cue ends where a sentence does, and failing that where it has grown too long to read or to leave up. A full stop too early in a cue is not the end of a sentence but a list marker or a shortened word, and one that ends up short anyway is held on screen until the next needs the space. Whisper still answers with its own segments and nothing on that path changes; the local server is not asked for words it was never asked for, and a hosted model that refuses the field falls back to the request it used to answer. A cue is short enough now that two can begin in the same second, so to_srt hands out every timing a second holds rather than the first.