ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “speech tokens, voiceprints, speaker inversion attack, Audio BERT (AuB), SpInv”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.AIcs.CREmpiricalRecentJul 18, 2026

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou +3 more

This paper introduces Audio BERT (AuB) and SpInv, methods for recovering embeddings from speech tokens and performing speaker inversion attacks using only three seconds of frontend output.

View →
cs.SDcs.AIcs.CRRecentJun 4, 2026

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun +2 more

The paper introduces a novel Clean-Referenced Feature-Vocoder Attack, a black-box adversarial attack that perturbs high-level SSL feature representations instead of raw audio waveforms, achieving supe…

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

Large Audio Language Models for Spoofing-Aware Speaker Verification

Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak +2 more

This paper evaluates the use of large audio language models for speaker verification systems against conventional pipelines and finds that task-specific adaptation improves performance.

View →
cs.SDEmpiricalRecentJul 20, 2026

SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

Kuan-Lin Chu, Jun-Cheng Chen, Chun-Shien Lu

This paper introduces SSTMark, a training-free speech watermarking framework that encodes watermark information into the semantic content of generated speech.

View →
cs.LGcs.AIcs.CREmpiricalRecentJun 26, 2026

What Was That Again? Certified Robustness for Automatic Speech Recognition

Andrew C. Cullen, Neil Marchant, Jiani Xie, Paul Montague +1 more

This paper proposes a method to decrease Word Error Rate (WER) and increase recall in Automatic Speech Recognition systems using a certification-inspired mechanism, a Two-Sided Atomic Audit, and a Ran…

View →
cs.SDcs.CRRecentMay 2, 2026

MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech

Yutong Jin, Qi Li, Lingshuang Liu, Jianbing Ni

MelShield is a robust, in-generation audio watermarking framework that embeds identifiable signals into AI-generated speech in the Mel-spectrogram domain for reliable copyright protection and attribut…

View →
cs.SDcs.AIcs.CLRecentMay 28, 2026

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

Bo-Han Feng, Yu-Hsuan Li Liang, Chien-Feng Liu, You-Hsuan Chang +1 more

This paper provides a unified taxonomy and controlled empirical evaluation of jailbreak attacks and defenses for Large Audio Language Models (LALMs), demonstrating that safety evaluation must consider…

View →
eess.ASEmpiricalRecentJul 22, 2026

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews

This paper investigates speaker verification in a multi-utterance, multimodal setting and shows that aggregating information across anonymized speech reduces error rates.

View →
cs.SDcs.CREmpiricalRecentJul 19, 2026

Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning

Henriette Flore Kenne, Raphael Anaadumba, Mohammad Arif Ul Alam

This paper proposes a framework to protect speaker privacy in dementia detection speech recordings by introducing Cumulative Signal Attack (CSA) at the signal level and Gradient Reversal Layer (GRL) w…

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
cs.SDcs.LGEmpiricalRecentJun 21, 2026

Kiwano: A Cutting-Edge Open-Source Toolkit for Speaker Verification

Mickael Rouvier, Pierre Michel Bousquet

The paper introduces Kiwano, an open-source toolkit for speaker verification research using PyTorch, offering standardized recipes, pretrained models, and integration of various architectures.

View →
eess.AScs.SDEmpiricalRecentJun 18, 2026

Personalized Keyword Spotting for User-Defined Keywords Leveraging Text-Independent Speaker Verification

Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang, Hung-Shin Lee +1 more

This paper proposes ZP-KWS, a lightweight framework for user-defined keyword spotting that achieves speaker-invariant zero-shot detection using a phoneme-supervised audio encoder and a compact speaker…

View →
eess.AScs.LGEmpiricalRecentJul 3, 2026

Open-Set Source Tracing as Compositional Factors via Structured Prototypes

Santiago Rubio, Antonio Almudévar, Antonio Miguel, Eduardo Lleida +1 more

This paper proposes a new definition of source in source tracing as a compositional tuple of Architecture, Training Data, and other training factors, and introduces a framework using Structured Orthon…

View →
cs.SDcs.CLeess.ASEmpiricalRecentJul 23, 2026

VibeVoice-ASR-BitNet Technical Report

Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng +6 more

The paper introduces VibeVoice-ASR-BitNet, a compressed real-time speech recognition model optimized for edge CPUs using heterogeneous quantization and custom SIMD kernels.

View →
eess.ASEmpiricalRecentJul 2, 2026

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

Jesuraj Bandekar, Prasanta Kumar Ghosh

This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…

View →
eess.AScs.SDEmpiricalRecentJul 9, 2026

Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection

Anna Taylor, Michele Panariello, Massimiliano Todisco, Chiara Galdi +2 more

The paper introduces a phoneme-level analysis framework for interpreting speech deepfake detection results using Gradient-weighted Class Activation Mapping and speech recognition.

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu +2 more

The paper proposes TRIDENT, a framework to restore a source speaker's identity from converted audio using a three-pronged architecture.

View →
cs.CRcs.AIcs.SDRecentApr 16, 2026

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection

Meng Chen, Kun Wang, Li Lu, Jiaheng Zhang +1 more

The paper introduces AudioHijack, a framework that successfully demonstrates context-agnostic and imperceptible auditory prompt injection attacks, showing that commercial Large Audio-Language Models c…

View →