OpenSVBench Scenario-driven SV leaderboard
System Detail

wespeaker/ res152-voxceleb

wespeaker_resnet152

ranked review: approved WeSpeaker ResNet

This page shows the global rank, scenario ranks, and trial-set results for this submission.

Global Position

#4/18

Public ranking position among reviewed full-core submissions.

Scenario-Macro EER
11.2275
Scenario-Macro minDCF
0.442748
Coverage

26/26

0 scenarios ranked #1 1 scenario in top 3
Public Listing

Ranked Publicly

This full-core result is approved and included in the public ranking.

Model Provenance

Source and architecture

  • Displayed name: wespeaker/res152-voxceleb
  • Source: WeSpeaker
  • Architecture: ResNet
  • Submitted alias: wespeaker-voxceleb-resnet152-LM
  • Created at: 2026-05-29T14:53:46.991718+00:00
Training And Links

Training data and references

  • Training data: VoxCeleb2 dev, 5,994 speakers
  • Training setup: ResNet152-TSTP-emb256 r-vector, ArcMargin, 150-epoch VoxCeleb2 recipe with speed perturbation and MUSAN/RIRS augmentation.
  • Submission mode: full-core
  • Paper: Deep Residual Learning for Image Recognition
Higher-Ranked Scenarios

Scenario ranks above the global position

Show 3 supporting trial sets
  • Lombard Grid Lombard: #3 on its trial-set ranking.
  • ESD: #3 on its trial-set ranking.
  • Whisper40 Whisper: #3 on its trial-set ranking.
Lower-Ranked Scenarios

Scenario ranks below the global position

Show 3 supporting trial sets
  • 3D-Speaker Device: #8 on its trial-set ranking.
  • AliMeeting Overlap: #8 on its trial-set ranking.
  • CN-Celeb Genre: #7 on its trial-set ranking.
Scenario Rankings

Full ranking across scenarios

This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.

Scenario Rank Lens Evidence Score
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
#2/18 Top 3
Speaking style shift 1.3434 EER from leader
4.1329
0.268357 minDCF
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
#4/18 Competitive
Open-domain media speech 0.9177 EER from leader
6.1011
0.224787 minDCF
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
#4/18 Competitive
Accent and dialect variation 1.0902 EER from leader
21.5254
0.675759 minDCF
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
#5/18 Competitive
Speaker aging and time gaps 0.6780 EER from leader
2.7476
0.111015 minDCF
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
#5/18 Competitive
Language mismatch 0.3592 EER from leader
4.8269
0.305505 minDCF
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
#5/18 Competitive
Noise and reverberation 2.6800 EER from leader
5.2600
0.154640 minDCF
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
#5/18 Competitive
Overlapping speakers 2.8169 EER from leader
12.9553
0.565796 minDCF
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
#5/18 Competitive
Short-duration speech 0.9914 EER from leader
13.5058
0.545781 minDCF
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
#7/18 Competitive
Distance mismatch 2.6914 EER from leader
11.9552
0.523450 minDCF
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
#7/18 Competitive
Device and channel variation 7.9353 EER from leader
14.2456
0.688961 minDCF
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
#7/18 Competitive
Source-genre variation 11.9747 EER from leader
26.2472
0.806172 minDCF
Variant Compare

Compare ResNet variants

Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.

Show
Variant Source Training Data Training Setup Global Rank Macro EER Macro minDCF Open
wespeaker/res152-voxceleb
Current
WeSpeaker VoxCeleb2 dev, 5,994 speakers ResNet152-TSTP-emb256 r-vector, ArcMargin, 150-epoch VoxCeleb2 recipe with speed perturbation and MUSAN/RIRS augmentation. #4
ranked
11.2275 0.442748 Open
wespeaker/res293-voxceleb WeSpeaker VoxCeleb2 dev, 5,994 speakers ResNet293-TSTP-emb256 r-vector, large-margin fine-tuned; official card reports 28.62M parameters and 28.10G FLOPs. #5
ranked
11.3130 0.439607 Open
wespeaker/res34-voxceleb WeSpeaker VoxCeleb2 dev, 5,994 speakers ResNet34-TSTP-emb256 r-vector, large-margin fine-tuned; official card reports 6.63M parameters and 4.55G FLOPs. #7
ranked
12.0516 0.477054 Open
wespeaker/res34-cnceleb WeSpeaker CN-Celeb train ResNet34 r-vector with TSTP pooling and large-margin fine-tuning on the CN-Celeb WeSpeaker recipe. #14
ranked
14.4910 0.607964 Open
Trial-Set Results

Raw ranking by trial set

Expand this section to inspect the detailed trial-set rankings.

Show
Trial Set Rank Scenarios Trials Score
TidyVoiceX2-ASV
TidyVoiceX2-ASV
#5/18
Cross-lingual
200000
4.8269
0.305505 minDCF
HI-MIA
HI-MIA
#4/18
Short-duration
660000
5.8453
0.284905 minDCF
GSC Short
Google Speech Commands
#5/18
Short-duration
220000
10.8860
0.601260 minDCF
GLOBE
GLOBE
#5/18
Accent/dialect
56848
25.0967
0.533224 minDCF
CN-Celeb
CN-Celeb
#6/18
In-the-wild
3484292
11.3672
0.399395 minDCF
CN-Celeb Genre
CN-Celeb
#7/18
Genre shift
440000
26.2472
0.806172 minDCF
CN-Celeb Short
CN-Celeb
#5/18
Short-duration
546964
24.9437
0.851112 minDCF
VoxCeleb1-O
VoxCeleb1
#4/18
In-the-wild
37611
0.4998
0.035524 minDCF
VoxCeleb1-E
VoxCeleb1
#6/18
In-the-wild
579818
0.7174
0.043120 minDCF
VoxCeleb1-H
VoxCeleb1
#5/18
In-the-wild
550894
1.2879
0.071894 minDCF
VoxCeleb Short
VoxCeleb1
#4/18
Short-duration
394724
12.3481
0.445848 minDCF
3D-Speaker Device
3D-Speaker
#8/18
Channel/device
180000
18.9800
0.832393 minDCF
3D-Speaker Distance
3D-Speaker
#4/18
Distance
175163
17.5655
0.786467 minDCF
3D-Speaker Dialect
3D-Speaker
#5/18
Accent/dialect
180000
17.9540
0.818293 minDCF
FFSVC 2022 Cross-Channel
FFSVC 2022
#7/18
Channel/device
72000
9.5111
0.545528 minDCF
FFSVC 2022 Cross-Domain
FFSVC 2022
#7/18
Distance
66546
9.6636
0.556279 minDCF
Whisper40 Whisper
Whisper40
#3/18
Speaking style
17600
9.2500
0.567375 minDCF
Lombard Grid Lombard
Lombard Grid
#3/18
Speaking style
29524
0.0745
0.011736 minDCF
VOiCES Noise/Reverb
VOiCES
#5/18
Noise/reverb
55000
5.2600
0.154640 minDCF
CHiME-6 Domestic Far-Field
CHiME-6
#5/18
Distance
36487
12.0320
0.480796 minDCF
CHiME-6 Overlap
CHiME-6
#5/18
Overlap
39600
16.5639
0.642111 minDCF
ESD
ESD
#3/18
Speaking style
437408
3.0742
0.225959 minDCF
AliMeeting Near/Far
AliMeeting
#4/18
Distance
220000
8.5600
0.270260 minDCF
AliMeeting Overlap
AliMeeting
#8/18
Overlap
165000
9.3467
0.489480 minDCF
VoxKnesset
VoxKnesset
#5/18
Aging
158312
3.7467
0.148597 minDCF
VoxPopuli Aging
VoxPopuli
#5/18
Aging
146575
1.7486
0.073433 minDCF