The usual pipeline
Mix it down, then guess it back
- Mix every participant into one audio track
- Run diarisation to find the speaker turns
- Cluster the turns into voices
- Guess which voice is which person
Speaker 1, Speaker 2, and a confidence
score you never see. Similar voices, a cross-talking pair or
somebody joining late is where it fails — and it fails
quietly, which is the part that matters.