System Detail
microsoft/ unispeech-sat-base-plus-sv
unispeech_sat_base_plus
This page shows the global rank, scenario ranks, and trial-set results for this submission.
Global Position
#18/18
Public ranking position among reviewed full-core submissions.
Scenario-Macro EER
25.6711
Scenario-Macro minDCF
0.890736
Coverage
26/26
0 scenarios ranked #1
0 scenarios in top 3
Public Listing
Ranked Publicly
This full-core result is approved and included in the public ranking.
Model Provenance
Source and architecture
- Displayed name:
microsoft/unispeech-sat-base-plus-sv - Source: Microsoft
- Architecture: UniSpeech-SAT Base+
- Submitted alias:
unispeech-sat-base-plus-sv - Created at:
2026-05-31T02:57:39.101571+00:00
Training And Links
Training data and references
- Training data: Pretrained on 60k h Libri-Light, 10k h GigaSpeech, and 24k h VoxPopuli; SV head fine-tuned on VoxCeleb1
- Training setup: UniSpeech-SAT Base+ speaker-aware SSL model fine-tuned for speaker verification with an x-vector head and additive-margin softmax.
- Submission mode:
full-core - Paper: UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
Higher-Ranked Scenarios
Scenario ranks above the global position
Show 3 supporting trial sets
- VoxCeleb Short: #12 on its trial-set ranking.
- VoxCeleb1-E: #13 on its trial-set ranking.
- VoxCeleb1-O: #13 on its trial-set ranking.
Lower-Ranked Scenarios
Scenario ranks below the global position
Show 3 supporting trial sets
- Whisper40 Whisper: #18 on its trial-set ranking.
- CN-Celeb: #18 on its trial-set ranking.
- VoxKnesset: #18 on its trial-set ranking.
Scenario Rankings
Full ranking across scenarios
This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.
| Scenario | Rank | Lens | Evidence | Score |
|---|---|---|---|---|
|
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
|
#16/18
Needs work
|
Noise and reverberation
18.4000 EER from leader
|
20.9800
0.741000 minDCF
|
|
|
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
|
#17/18
Needs work
|
Short-duration speech
13.6334 EER from leader
|
26.1478
0.927872 minDCF
|
|
|
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
|
#17/18
Needs work
|
Distance mismatch
19.5817 EER from leader
|
28.8455
0.982128 minDCF
|
|
|
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
|
#17/18
Needs work
|
Accent and dialect variation
12.1707 EER from leader
|
32.6058
0.912056 minDCF
|
|
|
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
|
#17/18
Needs work
|
Source-genre variation
22.7618 EER from leader
|
37.0343
0.985940 minDCF
|
|
|
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
|
#18/18
Needs work
|
Open-domain media speech
12.5012 EER from leader
|
17.6846
0.670598 minDCF
|
|
|
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
|
#18/18
Needs work
|
Speaking style shift
15.4003 EER from leader
|
18.1898
0.805772 minDCF
|
|
|
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
|
#18/18
Needs work
|
Speaker aging and time gaps
17.3082 EER from leader
|
19.3778
0.858414 minDCF
|
|
|
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
|
#18/18
Needs work
|
Language mismatch
18.7612 EER from leader
|
23.2288
0.970581 minDCF
|
|
|
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
|
#18/18
Needs work
|
Overlapping speakers
16.3639 EER from leader
|
26.5022
0.944448 minDCF
|
|
|
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
|
#18/18
Needs work
|
Device and channel variation
25.4754 EER from leader
|
31.7857
0.999283 minDCF
|
Trial-Set Results
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
Show
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
| Trial Set | Rank | Scenarios | Trials | Score |
|---|---|---|---|---|
|
TidyVoiceX2-ASV
TidyVoiceX2-ASV
|
#18/18
|
Cross-lingual
|
200000 |
23.2288
0.970581 minDCF
|
|
HI-MIA
HI-MIA
|
#17/18
|
Short-duration
|
660000 |
23.7600
0.944420 minDCF
|
|
GSC Short
Google Speech Commands
|
#17/18
|
Short-duration
|
220000 |
24.4600
0.966790 minDCF
|
|
GLOBE
GLOBE
|
#14/18
|
Accent/dialect
|
56848 |
26.9350
0.824245 minDCF
|
|
CN-Celeb
CN-Celeb
|
#18/18
|
In-the-wild
|
3484292 |
29.8022
0.945410 minDCF
|
|
CN-Celeb Genre
CN-Celeb
|
#17/18
|
Genre shift
|
440000 |
37.0343
0.985940 minDCF
|
|
CN-Celeb Short
CN-Celeb
|
#17/18
|
Short-duration
|
546964 |
34.0117
0.998459 minDCF
|
|
VoxCeleb1-O
VoxCeleb1
|
#13/18
|
In-the-wild
|
37611 |
5.6037
0.466737 minDCF
|
|
VoxCeleb1-E
VoxCeleb1
|
#13/18
|
In-the-wild
|
579818 |
3.6382
0.261858 minDCF
|
|
VoxCeleb1-H
VoxCeleb1
|
#13/18
|
In-the-wild
|
550894 |
7.4593
0.458766 minDCF
|
|
VoxCeleb Short
VoxCeleb1
|
#12/18
|
Short-duration
|
394724 |
22.3595
0.801817 minDCF
|
|
3D-Speaker Device
3D-Speaker
|
#17/18
|
Channel/device
|
180000 |
36.8213
0.999900 minDCF
|
|
3D-Speaker Distance
3D-Speaker
|
#17/18
|
Distance
|
175163 |
37.8673
0.999921 minDCF
|
|
3D-Speaker Dialect
3D-Speaker
|
#17/18
|
Accent/dialect
|
180000 |
38.2767
0.999867 minDCF
|
|
FFSVC 2022 Cross-Channel
FFSVC 2022
|
#18/18
|
Channel/device
|
72000 |
26.7500
0.998667 minDCF
|
|
FFSVC 2022 Cross-Domain
FFSVC 2022
|
#18/18
|
Distance
|
66546 |
26.9545
0.999067 minDCF
|
|
Whisper40 Whisper
Whisper40
|
#18/18
|
Speaking style
|
17600 |
33.6875
0.997938 minDCF
|
|
Lombard Grid Lombard
Lombard Grid
|
#17/18
|
Speaking style
|
29524 |
7.9993
0.495715 minDCF
|
|
VOiCES Noise/Reverb
VOiCES
|
#16/18
|
Noise/reverb
|
55000 |
20.9800
0.741000 minDCF
|
|
CHiME-6 Domestic Far-Field
CHiME-6
|
#17/18
|
Distance
|
36487 |
26.7802
0.970516 minDCF
|
|
CHiME-6 Overlap
CHiME-6
|
#16/18
|
Overlap
|
39600 |
29.4111
0.923250 minDCF
|
|
ESD
ESD
|
#16/18
|
Speaking style
|
437408 |
12.8825
0.923664 minDCF
|
|
AliMeeting Near/Far
AliMeeting
|
#17/18
|
Distance
|
220000 |
23.7800
0.959010 minDCF
|
|
AliMeeting Overlap
AliMeeting
|
#18/18
|
Overlap
|
165000 |
23.5933
0.965647 minDCF
|
|
VoxKnesset
VoxKnesset
|
#18/18
|
Aging
|
158312 |
27.7733
0.972844 minDCF
|
|
VoxPopuli Aging
VoxPopuli
|
#17/18
|
Aging
|
146575 |
10.9824
0.743985 minDCF
|