System Detail
wespeaker/ res34-cnceleb
wespeaker_r34
This page shows the global rank, scenario ranks, and trial-set results for this submission.
Global Position
#14/18
Public ranking position among reviewed full-core submissions.
Scenario-Macro EER
14.4910
Scenario-Macro minDCF
0.607964
Coverage
26/26
0 scenarios ranked #1
1 scenario in top 3
Public Listing
Ranked Publicly
This full-core result is approved and included in the public ranking.
Model Provenance
Source and architecture
- Displayed name:
wespeaker/res34-cnceleb - Source: WeSpeaker
- Architecture: ResNet
- Submitted alias:
wespeaker-cnceleb-resnet34-LM - Created at:
2026-05-29T14:26:58.126430+00:00
Training And Links
Training data and references
- Training data: CN-Celeb train
- Training setup: ResNet34 r-vector with TSTP pooling and large-margin fine-tuning on the CN-Celeb WeSpeaker recipe.
- Submission mode:
full-core - Paper: Deep Residual Learning for Image Recognition
Higher-Ranked Scenarios
Scenario ranks above the global position
Show 3 supporting trial sets
- CN-Celeb Short: #1 on its trial-set ranking.
- CN-Celeb Genre: #2 on its trial-set ranking.
- CN-Celeb: #3 on its trial-set ranking.
Lower-Ranked Scenarios
Scenario ranks below the global position
Show 3 supporting trial sets
- VoxPopuli Aging: #16 on its trial-set ranking.
- Lombard Grid Lombard: #16 on its trial-set ranking.
- 3D-Speaker Device: #15 on its trial-set ranking.
Scenario Rankings
Full ranking across scenarios
This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.
| Scenario | Rank | Lens | Evidence | Score |
|---|---|---|---|---|
|
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
|
#2/18
Top 3
|
Source-genre variation
1.4230 EER from leader
|
15.6955
0.591628 minDCF
|
|
|
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
|
#7/18
Competitive
|
Short-duration speech
1.7754 EER from leader
|
14.2898
0.579413 minDCF
|
|
|
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
|
#10/18
Needs work
|
Open-domain media speech
2.6600 EER from leader
|
7.8434
0.404734 minDCF
|
|
|
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
|
#11/18
Needs work
|
Speaking style shift
4.2849 EER from leader
|
7.0743
0.410279 minDCF
|
|
|
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
|
#13/18
Needs work
|
Language mismatch
3.1094 EER from leader
|
7.5770
0.460771 minDCF
|
|
|
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
|
#13/18
Needs work
|
Speaker aging and time gaps
10.0589 EER from leader
|
12.1285
0.635341 minDCF
|
|
|
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
|
#14/18
Needs work
|
Overlapping speakers
6.2739 EER from leader
|
16.4122
0.671276 minDCF
|
|
|
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
|
#14/18
Needs work
|
Noise and reverberation
14.0600 EER from leader
|
16.6400
0.531160 minDCF
|
|
|
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
|
#14/18
Needs work
|
Accent and dialect variation
5.3952 EER from leader
|
25.8303
0.825746 minDCF
|
|
|
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
|
#15/18
Needs work
|
Distance mismatch
7.7286 EER from leader
|
16.9924
0.739294 minDCF
|
|
|
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
|
#15/18
Needs work
|
Device and channel variation
12.6075 EER from leader
|
18.9178
0.837962 minDCF
|
Variant Compare
Compare ResNet variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
Show
Compare ResNet variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
| Variant | Source | Training Data | Training Setup | Global Rank | Macro EER | Macro minDCF | Open |
|---|---|---|---|---|---|---|---|
| wespeaker/res152-voxceleb | WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ResNet152-TSTP-emb256 r-vector, ArcMargin, 150-epoch VoxCeleb2 recipe with speed perturbation and MUSAN/RIRS augmentation. |
#4
ranked
|
11.2275 | 0.442748 | Open |
| wespeaker/res293-voxceleb | WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ResNet293-TSTP-emb256 r-vector, large-margin fine-tuned; official card reports 28.62M parameters and 28.10G FLOPs. |
#5
ranked
|
11.3130 | 0.439607 | Open |
| wespeaker/res34-voxceleb | WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ResNet34-TSTP-emb256 r-vector, large-margin fine-tuned; official card reports 6.63M parameters and 4.55G FLOPs. |
#7
ranked
|
12.0516 | 0.477054 | Open |
|
wespeaker/res34-cnceleb
Current
|
WeSpeaker | CN-Celeb train | ResNet34 r-vector with TSTP pooling and large-margin fine-tuning on the CN-Celeb WeSpeaker recipe. |
#14
ranked
|
14.4910 | 0.607964 | Open |
Trial-Set Results
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
Show
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
| Trial Set | Rank | Scenarios | Trials | Score |
|---|---|---|---|---|
|
TidyVoiceX2-ASV
TidyVoiceX2-ASV
|
#13/18
|
Cross-lingual
|
200000 |
7.5770
0.460771 minDCF
|
|
HI-MIA
HI-MIA
|
#12/18
|
Short-duration
|
660000 |
7.7542
0.426640 minDCF
|
|
GSC Short
Google Speech Commands
|
#4/18
|
Short-duration
|
220000 |
8.3650
0.496165 minDCF
|
|
GLOBE
GLOBE
|
#12/18
|
Accent/dialect
|
56848 |
26.3293
0.658746 minDCF
|
|
CN-Celeb
CN-Celeb
|
#3/18
|
In-the-wild
|
3484292 |
7.4729
0.345407 minDCF
|
|
CN-Celeb Genre
CN-Celeb
|
#2/18
|
Genre shift
|
440000 |
15.6955
0.591628 minDCF
|
|
CN-Celeb Short
CN-Celeb
|
#1/18
|
Short-duration
|
546964 |
17.6675
0.600899 minDCF
|
|
VoxCeleb1-O
VoxCeleb1
|
#15/18
|
In-the-wild
|
37611 |
7.0897
0.433848 minDCF
|
|
VoxCeleb1-E
VoxCeleb1
|
#15/18
|
In-the-wild
|
579818 |
7.1685
0.432123 minDCF
|
|
VoxCeleb1-H
VoxCeleb1
|
#15/18
|
In-the-wild
|
550894 |
10.3836
0.526212 minDCF
|
|
VoxCeleb Short
VoxCeleb1
|
#14/18
|
Short-duration
|
394724 |
23.3725
0.793947 minDCF
|
|
3D-Speaker Device
3D-Speaker
|
#15/18
|
Channel/device
|
180000 |
26.8967
0.998813 minDCF
|
|
3D-Speaker Distance
3D-Speaker
|
#14/18
|
Distance
|
175163 |
25.1560
0.990655 minDCF
|
|
3D-Speaker Dialect
3D-Speaker
|
#14/18
|
Accent/dialect
|
180000 |
25.3313
0.992747 minDCF
|
|
FFSVC 2022 Cross-Channel
FFSVC 2022
|
#9/18
|
Channel/device
|
72000 |
10.9389
0.677111 minDCF
|
|
FFSVC 2022 Cross-Domain
FFSVC 2022
|
#9/18
|
Distance
|
66546 |
11.0514
0.689180 minDCF
|
|
Whisper40 Whisper
Whisper40
|
#5/18
|
Speaking style
|
17600 |
9.8125
0.615313 minDCF
|
|
Lombard Grid Lombard
Lombard Grid
|
#16/18
|
Speaking style
|
29524 |
3.3905
0.196237 minDCF
|
|
VOiCES Noise/Reverb
VOiCES
|
#14/18
|
Noise/reverb
|
55000 |
16.6400
0.531160 minDCF
|
|
CHiME-6 Domestic Far-Field
CHiME-6
|
#14/18
|
Distance
|
36487 |
22.9123
0.914742 minDCF
|
|
CHiME-6 Overlap
CHiME-6
|
#13/18
|
Overlap
|
39600 |
23.6778
0.833472 minDCF
|
|
ESD
ESD
|
#13/18
|
Speaking style
|
437408 |
8.0200
0.419287 minDCF
|
|
AliMeeting Near/Far
AliMeeting
|
#7/18
|
Distance
|
220000 |
8.8500
0.362600 minDCF
|
|
AliMeeting Overlap
AliMeeting
|
#7/18
|
Overlap
|
165000 |
9.1467
0.509080 minDCF
|
|
VoxKnesset
VoxKnesset
|
#13/18
|
Aging
|
158312 |
14.3733
0.632640 minDCF
|
|
VoxPopuli Aging
VoxPopuli
|
#16/18
|
Aging
|
146575 |
9.8837
0.638041 minDCF
|