Frame-Level Speaker Encoder for Streaming Diarization
FASTDIAR is a streaming diarization architecture with a confidence-gated online clustering at a fixed 960 ms delay, which is the most accurate streaming diarizer on low-overlap benchmarks, degrades far less than cache-based systems beyond four speakers, and needs no diarization corpus in training.
Diarization error rate (DER, %) with a collar of 0 and overlapped speech excluded, split by the number of speakers. Streaming Sortformer is capped at four speakers, so the three corpora that cross that boundary are also split at it. RTF is measured on one CPU thread. Bold marks the lowest DER in each column.
| Model | #Params | Latency (ms) |
RTF | NOTSOFAR | VoxConverse | DIHARD3 | AMI | RAMC | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ≤4 | ≥5 | all | ≤4 | ≥5 | all | ≤4 | ≥5 | all | 3–4 | 2 | ||||
| diart | 5.8M | 1000 | 0.16 | 42.56 | 54.16 | 48.01 | 12.81 | 17.17 | 16.12 | 17.40 | 46.14 | 22.04 | 28.79 | 25.87 |
| Streaming Sortformer | 117.7M | 1040 | 1.54 | 14.48 | 33.45 | 23.39 | 5.41 | 23.55 | 19.15 | 14.08 | 34.25 | 17.34 | 25.56 | 25.27 |
| FASTDIAR | 10.4M | 960 | 0.19 | 19.23 | 24.25 | 21.58 | 10.70 | 12.88 | 12.35 | 24.73 | 32.95 | 26.06 | 22.15 | 21.61 |
The confidence score collapses at every speaker change and recovers once the encoder has gathered enough context from the new speaker. Its dips line up with the block borders of the similarity matrix, so the score marks where one cluster ends and the next begins, without needing any centroid or label. Frames below the threshold (dashed line) are still labeled, but only confident frames update a speaker centroid, so embeddings from a speaker change never blur two clusters together.
Both encoders see the same two seconds of past audio and emit an embedding every 80 ms. The cluster borders of the baseline are blurred, while the proposed encoder forms sharp, well-separated clusters, which improves diarization quality. With the same VAD and clustering, DER drops from 14.79% to 12.35% on VoxConverse and from 27.97% to 21.61% on RAMC, and the system runs 25× faster.
@article{torgashov2026fastdiar,
title={{FASTDIAR}: Frame-Level Speaker Encoder for Streaming Diarization},
author={Torgashov, Nikita and K{\"o}p{\"u}kl{\"u}, Okan},
journal={arXiv preprint arXiv:2610.02941},
year={2026}
}