OpenSVBench Scenario-driven SV leaderboard
System Detail

microsoft/ wavlm-base-plus-sv

wavlm_base

ranked review: approved Microsoft WavLM-Base+

This page shows the global rank, scenario ranks, and trial-set results for this submission.

Global Position

#17/18

Public ranking position among reviewed full-core submissions.

Scenario-Macro EER
25.1577
Scenario-Macro minDCF
0.875090
Coverage

26/26

0 scenarios ranked #1 0 scenarios in top 3
Public Listing

Ranked Publicly

This full-core result is approved and included in the public ranking.

Model Provenance

Source and architecture

  • Displayed name: microsoft/wavlm-base-plus-sv
  • Source: Microsoft
  • Architecture: WavLM-Base+
  • Submitted alias: wavlm-base-plus-sv
  • Created at: 2026-05-29T14:44:54.637184+00:00
Training And Links

Training data and references

Higher-Ranked Scenarios

Scenario ranks above the global position

Show 3 supporting trial sets
  • VoxCeleb1-E: #12 on its trial-set ranking.
  • VoxCeleb1-O: #12 on its trial-set ranking.
  • VoxCeleb1-H: #12 on its trial-set ranking.
Lower-Ranked Scenarios

Scenario ranks below the global position

Show 3 supporting trial sets
  • CN-Celeb Genre: #18 on its trial-set ranking.
  • CHiME-6 Overlap: #18 on its trial-set ranking.
  • CHiME-6 Domestic Far-Field: #18 on its trial-set ranking.
Scenario Rankings

Full ranking across scenarios

This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.

Scenario Rank Lens Evidence Score
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
#16/18 Needs work
Speaker aging and time gaps 14.3863 EER from leader
16.4560
0.801948 minDCF
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
#16/18 Needs work
Open-domain media speech 11.3223 EER from leader
16.5057
0.632600 minDCF
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
#16/18 Needs work
Accent and dialect variation 10.9217 EER from leader
31.3569
0.895666 minDCF
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
#17/18 Needs work
Speaking style shift 15.2840 EER from leader
18.0734
0.805194 minDCF
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
#17/18 Needs work
Noise and reverberation 18.4800 EER from leader
21.0600
0.742060 minDCF
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
#17/18 Needs work
Language mismatch 17.2476 EER from leader
21.7152
0.966351 minDCF
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
#17/18 Needs work
Overlapping speakers 15.9522 EER from leader
26.0906
0.926016 minDCF
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
#17/18 Needs work
Device and channel variation 24.9714 EER from leader
31.2817
0.997779 minDCF
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
#18/18 Needs work
Short-duration speech 14.1955 EER from leader
26.7099
0.906048 minDCF
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
#18/18 Needs work
Distance mismatch 19.8418 EER from leader
29.1056
0.976812 minDCF
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
#18/18 Needs work
Source-genre variation 24.1075 EER from leader
38.3800
0.975515 minDCF
Trial-Set Results

Raw ranking by trial set

Expand this section to inspect the detailed trial-set rankings.

Show
Trial Set Rank Scenarios Trials Score
TidyVoiceX2-ASV
TidyVoiceX2-ASV
#17/18
Cross-lingual
200000
21.7152
0.966351 minDCF
HI-MIA
HI-MIA
#18/18
Short-duration
660000
25.0650
0.910383 minDCF
GSC Short
Google Speech Commands
#18/18
Short-duration
220000
25.7750
0.967470 minDCF
GLOBE
GLOBE
#15/18
Accent/dialect
56848
27.0704
0.791331 minDCF
CN-Celeb
CN-Celeb
#17/18
In-the-wild
3484292
27.9358
0.916549 minDCF
CN-Celeb Genre
CN-Celeb
#18/18
Genre shift
440000
38.3800
0.975515 minDCF
CN-Celeb Short
CN-Celeb
#16/18
Short-duration
546964
33.2598
0.989844 minDCF
VoxCeleb1-O
VoxCeleb1
#12/18
In-the-wild
37611
5.5047
0.436697 minDCF
VoxCeleb1-E
VoxCeleb1
#12/18
In-the-wild
579818
3.2098
0.213272 minDCF
VoxCeleb1-H
VoxCeleb1
#12/18
In-the-wild
550894
6.5121
0.395984 minDCF
VoxCeleb Short
VoxCeleb1
#13/18
Short-duration
394724
22.7399
0.756493 minDCF
3D-Speaker Device
3D-Speaker
#16/18
Channel/device
180000
36.5133
0.999780 minDCF
3D-Speaker Distance
3D-Speaker
#16/18
Distance
175163
37.3027
0.999801 minDCF
3D-Speaker Dialect
3D-Speaker
#16/18
Accent/dialect
180000
35.6433
1.000000 minDCF
FFSVC 2022 Cross-Channel
FFSVC 2022
#17/18
Channel/device
72000
26.0500
0.995778 minDCF
FFSVC 2022 Cross-Domain
FFSVC 2022
#17/18
Distance
66546
26.2121
0.996206 minDCF
Whisper40 Whisper
Whisper40
#17/18
Speaking style
17600
32.6563
0.989187 minDCF
Lombard Grid Lombard
Lombard Grid
#18/18
Speaking style
29524
8.1967
0.544113 minDCF
VOiCES Noise/Reverb
VOiCES
#17/18
Noise/reverb
55000
21.0600
0.742060 minDCF
CHiME-6 Domestic Far-Field
CHiME-6
#18/18
Distance
36487
29.0021
0.989720 minDCF
CHiME-6 Overlap
CHiME-6
#18/18
Overlap
39600
30.3611
0.924111 minDCF
ESD
ESD
#18/18
Speaking style
437408
13.3674
0.882281 minDCF
AliMeeting Near/Far
AliMeeting
#18/18
Distance
220000
23.9055
0.921520 minDCF
AliMeeting Overlap
AliMeeting
#17/18
Overlap
165000
21.8200
0.927920 minDCF
VoxKnesset
VoxKnesset
#17/18
Aging
158312
23.1333
0.940091 minDCF
VoxPopuli Aging
VoxPopuli
#15/18
Aging
146575
9.7786
0.663805 minDCF