System Detail
wespeaker/ ecapa1024-voxceleb
wespeaker_ecapa1024
This page shows the global rank, scenario ranks, and trial-set results for this submission.
Global Position
#11/18
Public ranking position among reviewed full-core submissions.
Scenario-Macro EER
14.0835
Scenario-Macro minDCF
0.534111
Coverage
26/26
0 scenarios ranked #1
0 scenarios in top 3
Public Listing
Ranked Publicly
This full-core result is approved and included in the public ranking.
Model Provenance
Source and architecture
- Displayed name:
wespeaker/ecapa1024-voxceleb - Source: WeSpeaker
- Architecture: ECAPA-TDNN
- Submitted alias:
wespeaker-voxceleb-ecapa-tdnn1024-LM - Created at:
2026-05-29T14:49:19.687079+00:00
Training And Links
Training data and references
- Training data: VoxCeleb2 dev, 5,994 speakers
- Training setup: ECAPA-TDNN GLOB c1024 with ASTP pooling, 192-d embedding, ArcMargin, 150-epoch training, speed perturbation, and MUSAN/RIRS augmentation.
- Submission mode:
full-core - Paper: ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Higher-Ranked Scenarios
Scenario ranks above the global position
Show 3 supporting trial sets
- CHiME-6 Overlap: #8 on its trial-set ranking.
- VoxCeleb1-O: #9 on its trial-set ranking.
- ESD: #9 on its trial-set ranking.
Lower-Ranked Scenarios
Scenario ranks below the global position
Show 3 supporting trial sets
- CN-Celeb Genre: #15 on its trial-set ranking.
- CN-Celeb Short: #15 on its trial-set ranking.
- GSC Short: #15 on its trial-set ranking.
Scenario Rankings
Full ranking across scenarios
This table shows where the model sits on each scenario, using source-balanced scenario scores and the linked trial sets as evidence.
| Scenario | Rank | Lens | Evidence | Score |
|---|---|---|---|---|
|
Aging Robustness
How stable identity representations remain across age-derived and longitudinal recording gaps.
|
#9/18
Competitive
|
Speaker aging and time gaps
1.9627 EER from leader
|
4.0323
0.181572 minDCF
|
|
|
Speaking-Style Robustness
Whether the model can preserve identity across emotion-driven change, whispered speech, and noise-induced Lombard speaking style.
|
#10/18
Needs work
|
Speaking style shift
3.7796 EER from leader
|
6.5691
0.361235 minDCF
|
|
|
Noise / Reverb Robustness
How reliably the model preserves identity when room acoustics, distractor noise, and microphone placement deviate from an easier in-corpus reference condition.
|
#10/18
Needs work
|
Noise and reverberation
4.0260 EER from leader
|
6.6060
0.214720 minDCF
|
|
|
Overlap Robustness
How well speaker identity survives light, mid, and heavy overlap in both meeting and domestic recordings.
|
#10/18
Needs work
|
Overlapping speakers
4.5028 EER from leader
|
14.6411
0.618143 minDCF
|
|
|
Cross-Lingual Robustness
How well speaker identity survives enrollment-test language mismatch.
|
#12/18
Needs work
|
Language mismatch
2.5171 EER from leader
|
6.9847
0.436029 minDCF
|
|
|
Distance Robustness
How much performance changes across meeting and domestic distance mismatch.
|
#12/18
Needs work
|
Distance mismatch
6.1802 EER from leader
|
15.4440
0.640785 minDCF
|
|
|
Accent / Dialect Robustness
How reliably the model tracks identity across accent and dialect mismatch.
|
#12/18
Needs work
|
Accent and dialect variation
3.8812 EER from leader
|
24.3163
0.734606 minDCF
|
|
|
In-The-Wild Robustness
How strong the model is on unconstrained celebrity and media speech across official CN-Celeb and VoxCeleb protocols.
|
#13/18
Needs work
|
Open-domain media speech
3.3370 EER from leader
|
8.5204
0.308148 minDCF
|
|
|
Channel / Device Robustness
How resilient the model is to device and channel mismatch.
|
#13/18
Needs work
|
Device and channel variation
12.1464 EER from leader
|
18.4567
0.824736 minDCF
|
|
|
Short-Duration Robustness
How much performance holds up when speech evidence is limited by duration.
|
#15/18
Needs work
|
Short-duration speech
5.7306 EER from leader
|
18.2450
0.644928 minDCF
|
|
|
Genre-Shift Robustness
Whether performance holds when CN-Celeb enrollment and test speech come from different source genres.
|
#15/18
Needs work
|
Source-genre variation
16.8300 EER from leader
|
31.1025
0.910323 minDCF
|
Variant Compare
Compare ECAPA-TDNN variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
Show
Compare ECAPA-TDNN variants
Expand this section to compare checkpoints in the same architecture group by source, training data, training setup, and leaderboard result.
| Variant | Source | Training Data | Training Setup | Global Rank | Macro EER | Macro minDCF | Open |
|---|---|---|---|---|---|---|---|
| speechbrain/ecapa | SpeechBrain | VoxCeleb1 + VoxCeleb2 training data | SpeechBrain ECAPA-TDNN release with attentive statistical pooling and Additive Margin Softmax loss. |
#10
ranked
|
13.5461 | 0.543660 | Open |
|
wespeaker/ecapa1024-voxceleb
Current
|
WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ECAPA-TDNN GLOB c1024 with ASTP pooling, 192-d embedding, ArcMargin, 150-epoch training, speed perturbation, and MUSAN/RIRS augmentation. |
#11
ranked
|
14.0835 | 0.534111 | Open |
| wespeaker/ecapa512-voxceleb | WeSpeaker | VoxCeleb2 dev, 5,994 speakers | ECAPA-TDNN GLOB c512 with ASTP pooling, 192-d embedding, ArcMargin, large-margin fine-tuning, speed perturbation, and MUSAN/RIRS augmentation. |
#13
ranked
|
14.3120 | 0.541141 | Open |
Trial-Set Results
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
Show
Raw ranking by trial set
Expand this section to inspect the detailed trial-set rankings.
| Trial Set | Rank | Scenarios | Trials | Score |
|---|---|---|---|---|
|
TidyVoiceX2-ASV
TidyVoiceX2-ASV
|
#12/18
|
Cross-lingual
|
200000 |
6.9847
0.436029 minDCF
|
|
HI-MIA
HI-MIA
|
#15/18
|
Short-duration
|
660000 |
8.6250
0.381742 minDCF
|
|
GSC Short
Google Speech Commands
|
#15/18
|
Short-duration
|
220000 |
17.4110
0.776190 minDCF
|
|
GLOBE
GLOBE
|
#11/18
|
Accent/dialect
|
56848 |
25.7740
0.541118 minDCF
|
|
CN-Celeb
CN-Celeb
|
#14/18
|
In-the-wild
|
3484292 |
15.8209
0.536618 minDCF
|
|
CN-Celeb Genre
CN-Celeb
|
#15/18
|
Genre shift
|
440000 |
31.1025
0.910323 minDCF
|
|
CN-Celeb Short
CN-Celeb
|
#15/18
|
Short-duration
|
546964 |
30.3869
0.901376 minDCF
|
|
VoxCeleb1-O
VoxCeleb1
|
#9/18
|
In-the-wild
|
37611 |
0.8188
0.065623 minDCF
|
|
VoxCeleb1-E
VoxCeleb1
|
#10/18
|
In-the-wild
|
579818 |
0.9899
0.062012 minDCF
|
|
VoxCeleb1-H
VoxCeleb1
|
#10/18
|
In-the-wild
|
550894 |
1.8513
0.111398 minDCF
|
|
VoxCeleb Short
VoxCeleb1
|
#11/18
|
Short-duration
|
394724 |
16.5570
0.520405 minDCF
|
|
3D-Speaker Device
3D-Speaker
|
#13/18
|
Channel/device
|
180000 |
23.8800
0.940667 minDCF
|
|
3D-Speaker Distance
3D-Speaker
|
#13/18
|
Distance
|
175163 |
23.6657
0.920046 minDCF
|
|
3D-Speaker Dialect
3D-Speaker
|
#12/18
|
Accent/dialect
|
180000 |
22.8587
0.928093 minDCF
|
|
FFSVC 2022 Cross-Channel
FFSVC 2022
|
#15/18
|
Channel/device
|
72000 |
13.0333
0.708806 minDCF
|
|
FFSVC 2022 Cross-Domain
FFSVC 2022
|
#15/18
|
Distance
|
66546 |
13.1585
0.719381 minDCF
|
|
Whisper40 Whisper
Whisper40
|
#12/18
|
Speaking style
|
17600 |
14.8000
0.736062 minDCF
|
|
Lombard Grid Lombard
Lombard Grid
|
#10/18
|
Speaking style
|
29524 |
0.3055
0.017288 minDCF
|
|
VOiCES Noise/Reverb
VOiCES
|
#10/18
|
Noise/reverb
|
55000 |
6.6060
0.214720 minDCF
|
|
CHiME-6 Domestic Far-Field
CHiME-6
|
#9/18
|
Distance
|
36487 |
13.6569
0.577058 minDCF
|
|
CHiME-6 Overlap
CHiME-6
|
#8/18
|
Overlap
|
39600 |
17.8889
0.675139 minDCF
|
|
ESD
ESD
|
#9/18
|
Speaking style
|
437408 |
4.6018
0.330355 minDCF
|
|
AliMeeting Near/Far
AliMeeting
|
#14/18
|
Distance
|
220000 |
11.2950
0.346655 minDCF
|
|
AliMeeting Overlap
AliMeeting
|
#13/18
|
Overlap
|
165000 |
11.3933
0.561147 minDCF
|
|
VoxKnesset
VoxKnesset
|
#9/18
|
Aging
|
158312 |
5.6932
0.252750 minDCF
|
|
VoxPopuli Aging
VoxPopuli
|
#10/18
|
Aging
|
146575 |
2.3715
0.110394 minDCF
|