ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Automatic speaker verification”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.HCEmpiricalRecentJul 22, 2026

Improving the performance of an ASV system using hybrid speech features

Stanisław Ciszkiewicz, Artur Janicki

This paper explores the use of hybrid feature sets to improve the performance of Automatic Speaker Verification (ASV) systems under noisy conditions, using Mel-Frequency Cepstral Coefficients (MFCC),…

View →
eess.ASEmpiricalRecentJul 22, 2026

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews

This paper investigates speaker verification in a multi-utterance, multimodal setting and shows that aggregating information across anonymized speech reduces error rates.

View →
eess.ASEmpiricalRecentJul 8, 2026

Text-Independent Speaker Verification Using Discrete Audio Tokens

Zheng Liang, Junjie Li, Kong Aik Lee

This paper proposes a Cross-Feature Knowledge Distillation (CFKD) framework to improve automatic speaker verification (ASV) performance of neural audio codecs (NACs) by guiding the student model to mi…

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

Large Audio Language Models for Spoofing-Aware Speaker Verification

Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak +2 more

This paper evaluates the use of large audio language models for speaker verification systems against conventional pipelines and finds that task-specific adaptation improves performance.

View →
cs.SDcs.LGEmpiricalRecentJun 21, 2026

Kiwano: A Cutting-Edge Open-Source Toolkit for Speaker Verification

Mickael Rouvier, Pierre Michel Bousquet

The paper introduces Kiwano, an open-source toolkit for speaker verification research using PyTorch, offering standardized recipes, pretrained models, and integration of various architectures.

View →
eess.ASEmpiricalRecentJun 12, 2026

Who Spoke When in Multi-Conversation: Target Speaker Tagging Task and Benchmark

Minjae Lee, Hee-Soo Heo, Youngki Kwon, Han-Gyu Kim +2 more

The paper introduces Target Speaker Tagging (TST), a task that combines speaker diarization, verification, and identification into a single workflow for multi-speaker conversations. It presents TST-Be…

View →
cs.SDcs.AIcs.CLEmpiricalRecentJun 26, 2026

DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions

Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth +1 more

This paper introduces DG^VoiC, a voice clustering framework for customer verification and speaker linking in call-centre audio using sensitive information-aligned anonymisation, speech-focused preproc…

View →
cs.LGcs.AIcs.SDEmpiricalRecentJun 28, 2026

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang +2 more

A multimodal speaker identification system is proposed using Adaptive Modality Routing (AMR) for the POLY-SIM 2026 Grand Challenge, achieving high accuracy in various conditions.

View →
cs.SDcs.AIcs.CREmpiricalRecentJul 23, 2026

Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

Daniyal Kabir Dar, Arun Ross

This paper proposes the Identity Sensitivity Score (ISS) as a diagnostic tool for speaker-dependent failure analysis in audio deepfake detection, which shows significant correlation between ISS and ac…

View →
eess.ASEmpiricalRecentJul 20, 2026

The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

Xuanji He, Gaoyang Dong, Xiaoxiao Li, Minchuan Chen +1 more

The paper presents the tttAI system for time-stamped speaker-attributed speech recognition in smart-glasses recordings, achieving a tcpCER of 7.10% on Track 1 and 34.04% on Track 2.

View →
eess.AScs.SDEmpiricalRecentJun 18, 2026

Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang, Hung-Shin Lee +1 more

This paper proposes ZP-KWS, a lightweight framework for user-defined keyword spotting that achieves speaker-invariant zero-shot detection using a phoneme-supervised audio encoder and a compact speaker…

View →
cs.SDcs.AIcs.CREmpiricalRecentJul 18, 2026

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou +3 more

This paper introduces Audio BERT (AuB) and SpInv, methods for recovering embeddings from speech tokens and performing speaker inversion attacks using only three seconds of frontend output.

View →
cs.SDcs.AIEmpiricalRecentJul 3, 2026

DETECT-3B-Omni is Agnostic of Content and Demographics

Nicolas M. Müller, Aditya Tirumala Bukkapatnam, Dominik Schnieders, Zohaib Ahmed

The study tests the semantic independence of Resemble AI's deepfake audio detector, DETECT-3B-Omni, using 10,240 audio samples from various speakers and AI voice-cloning systems, and shows that the ac…

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
cs.SDcs.MMEmpiricalRecentJun 12, 2026

MaskedFOP: Polyglot Speaker Identification under Missing Visual Modality via Cascaded Graph Label Propagation

Ayoub Elkhouzari, Youssef Iraqi, Loubna Mekouar

The paper presents MaskedFOP, a system for speaker identification in Urdu language without face modality, using a modality-dropout network, two complementary audio embeddings, and a cascaded inference…

View →
eess.ASEmpiricalRecentJun 18, 2026

Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces

Kyle Janse van Rensburg, Herman Kamper

This paper investigates the correlation between dimensions of self-supervised speech features and speech characteristics, finding that content dimensions primarily capture intensity, formants, and voi…

View →
cs.CLcs.AIcs.CVEmpiricalRecentJul 2, 2026

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu +5 more

This paper introduces DramaSR-532K, a large-scale benchmark for speaker recognition in long-form TV dramas, and proposes DramaSR-LRM, a robust approach for speaker recognition using a large reasoning…

View →
cs.SDcs.AIeess.ASRecentJun 1, 2026

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

Louis Mouchon

Echo is a joint-embedding predictive architecture that uses a single, pretrained ViT encoder to simultaneously perform speaker diarization, speech recognition, and dynamic source separation in a share…

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu +2 more

The paper proposes TRIDENT, a framework to restore a source speaker's identity from converted audio using a three-pronged architecture.

View →
eess.AScs.SDEmpiricalRecentJul 16, 2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han +8 more

This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol,…

View →