FASTDIAR

Frame-Level Speaker Encoder for Streaming Diarization

View the Project on GitHub microsoft/fastdiar

FASTDIAR: Frame-Level Speaker Encoder for Streaming Diarization

arXiv paper GitHub code HuggingFace demo

1Department of Speech, Music and Hearing, KTH Royal Institute of Technology, Stockholm, Sweden
2Microsoft, Munich, Germany
*Work done during an internship at Microsoft.
FASTDIAR online clustering running frame by frame on a recording

Overview

FASTDIAR is a streaming diarization architecture with a confidence-gated online clustering at a fixed 960 ms delay, which is the most accurate streaming diarizer on low-overlap benchmarks, degrades far less than cache-based systems beyond four speakers, and needs no diarization corpus in training.

  • Frame-level streaming encoder: A state-of-the-art speaker recognition model, ReDimNet2, made causal and stripped of temporal pooling. It reads the audio stream once and emits a speaker embedding every 80 ms from a bounded two-second window of past audio, with no chunking and no recomputation.
  • Confidence-gated clustering: The similarity between each embedding and the one 800 ms earlier decides which frames may update a speaker centroid, so frames that straddle a speaker change never pull two clusters together. Every label leaves the system at a fixed 960 ms delay.
  • Fast on CPU: With 10.4M parameters, the full system (encoder, VAD and clustering) runs 5× faster than real time on a single CPU thread.

Results

Diarization error rate (DER, %) with a collar of 0 and overlapped speech excluded, split by the number of speakers. Streaming Sortformer is capped at four speakers, so the three corpora that cross that boundary are also split at it. RTF is measured on one CPU thread. Bold marks the lowest DER in each column.

Model #Params Latency
(ms)
RTF NOTSOFAR VoxConverse DIHARD3 AMI RAMC
≤4≥5all ≤4≥5all ≤4≥5all 3–4 2
diart 5.8M10000.16 42.5654.1648.01 12.8117.1716.12 17.4046.1422.04 28.79 25.87
Streaming Sortformer 117.7M10401.54 14.4833.4523.39 5.4123.5519.15 14.0834.2517.34 25.56 25.27
FASTDIAR 10.4M9600.19 19.2324.2521.58 10.7012.8812.35 24.7332.9526.06 22.15 21.61

Ablations

Confidence score detects cluster borders

Left: confidence score over an 18-second excerpt, dropping at each speaker change. Right: frame-to-frame cosine similarity matrix of the same excerpt.
Left: the confidence score, the cosine similarity between the current embedding and the one 10 frames (800 ms) earlier. Right: cosine similarity between all pairs of frames of the same excerpt. Reference speakers are shown on top, and dotted lines mark reference speaker changes.

The confidence score collapses at every speaker change and recovers once the encoder has gathered enough context from the new speaker. Its dips line up with the block borders of the similarity matrix, so the score marks where one cluster ends and the next begins, without needing any centroid or label. Frames below the threshold (dashed line) are still labeled, but only confident frames update a speaker centroid, so embeddings from a speaker change never blur two clusters together.

Frame-level encoder forms sharper clusters

Frame-to-frame cosine similarity matrices of the baseline chunk-based encoder (left) and the proposed frame-level encoder (right).
Frame-to-frame cosine similarity of the same excerpt. Left: the baseline, offline ReDimNet2-B6 run on 2-second chunks with an 80 ms shift to simulate streaming. Right: the proposed frame-level encoder.

Both encoders see the same two seconds of past audio and emit an embedding every 80 ms. The cluster borders of the baseline are blurred, while the proposed encoder forms sharp, well-separated clusters, which improves diarization quality. With the same VAD and clustering, DER drops from 14.79% to 12.35% on VoxConverse and from 27.97% to 21.61% on RAMC, and the system runs 25× faster.

Citation

@article{torgashov2026fastdiar,
  title={{FASTDIAR}: Frame-Level Speaker Encoder for Streaming Diarization},
  author={Torgashov, Nikita and K{\"o}p{\"u}kl{\"u}, Okan},
  journal={arXiv preprint arXiv:2610.02941},
  year={2026}
}