20 results for “Automatic speaker verification”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper explores the use of hybrid feature sets to improve the performance of Automatic Speaker Verification (ASV) systems under noisy conditions, using Mel-Frequency Cepstral Coefficients (MFCC),…
This paper investigates speaker verification in a multi-utterance, multimodal setting and shows that aggregating information across anonymized speech reduces error rates.
This paper proposes a Cross-Feature Knowledge Distillation (CFKD) framework to improve automatic speaker verification (ASV) performance of neural audio codecs (NACs) by guiding the student model to mi…
This paper evaluates the use of large audio language models for speaker verification systems against conventional pipelines and finds that task-specific adaptation improves performance.
The paper introduces Kiwano, an open-source toolkit for speaker verification research using PyTorch, offering standardized recipes, pretrained models, and integration of various architectures.
Minjae Lee, Hee-Soo Heo, Youngki Kwon, Han-Gyu Kim +2 more
The paper introduces Target Speaker Tagging (TST), a task that combines speaker diarization, verification, and identification into a single workflow for multi-speaker conversations. It presents TST-Be…
This paper introduces DG^VoiC, a voice clustering framework for customer verification and speaker linking in call-centre audio using sensitive information-aligned anonymisation, speech-focused preproc…
Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang +2 more
A multimodal speaker identification system is proposed using Adaptive Modality Routing (AMR) for the POLY-SIM 2026 Grand Challenge, achieving high accuracy in various conditions.
This paper proposes the Identity Sensitivity Score (ISS) as a diagnostic tool for speaker-dependent failure analysis in audio deepfake detection, which shows significant correlation between ISS and ac…
Xuanji He, Gaoyang Dong, Xiaoxiao Li, Minchuan Chen +1 more
The paper presents the tttAI system for time-stamped speaker-attributed speech recognition in smart-glasses recordings, achieving a tcpCER of 7.10% on Track 1 and 34.04% on Track 2.
This paper proposes ZP-KWS, a lightweight framework for user-defined keyword spotting that achieves speaker-invariant zero-shot detection using a phoneme-supervised audio encoder and a compact speaker…
Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou +3 more
This paper introduces Audio BERT (AuB) and SpInv, methods for recovering embeddings from speech tokens and performing speaker inversion attacks using only three seconds of frontend output.
The study tests the semantic independence of Resemble AI's deepfake audio detector, DETECT-3B-Omni, using 10,240 audio samples from various speakers and AI voice-cloning systems, and shows that the ac…
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
The paper presents MaskedFOP, a system for speaker identification in Urdu language without face modality, using a modality-dropout network, two complementary audio embeddings, and a cascaded inference…
This paper investigates the correlation between dimensions of self-supervised speech features and speech characteristics, finding that content dimensions primarily capture intensity, formants, and voi…
Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu +5 more
This paper introduces DramaSR-532K, a large-scale benchmark for speaker recognition in long-form TV dramas, and proposes DramaSR-LRM, a robust approach for speaker recognition using a large reasoning…
Echo is a joint-embedding predictive architecture that uses a single, pretrained ViT encoder to simultaneously perform speaker diarization, speech recognition, and dynamic source separation in a share…
Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu +2 more
The paper proposes TRIDENT, a framework to restore a source speaker's identity from converted audio using a three-pronged architecture.
Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more
This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…