On silence after speech, the server finalizes the segment (commits the
partial) automatically.
Restores punctuation on the committed transcript at VAD endpoints /
finalizes (CPU model; the trailing mark is held back until a true end).
Labels each finalized segment with its dominant speaker (S1–S4,
streaming Sortformer on CPU). Works for transcribe and translate.
Requires VAD — enabling this turns VAD on automatically.
off
Restrict output to the selected language's script (needs exactly one
language). E.g. English then never emits Korean/Arabic/Japanese.
Audio Input Device (mic session; e.g. BlackHole)
Input Gain (applies to mic & file, client-side)
0 dB
Audio File (decoded → 16 kHz mono, gain applied, streamed through model)