20 results for “speech tokens, voiceprints, speaker inversion attack, Audio BERT (AuB), SpInv”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou +3 more
This paper introduces Audio BERT (AuB) and SpInv, methods for recovering embeddings from speech tokens and performing speaker inversion attacks using only three seconds of frontend output.
Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun +2 more
The paper introduces a novel Clean-Referenced Feature-Vocoder Attack, a black-box adversarial attack that perturbs high-level SSL feature representations instead of raw audio waveforms, achieving supe…
This paper evaluates the use of large audio language models for speaker verification systems against conventional pipelines and finds that task-specific adaptation improves performance.
This paper introduces SSTMark, a training-free speech watermarking framework that encodes watermark information into the semantic content of generated speech.
Andrew C. Cullen, Neil Marchant, Jiani Xie, Paul Montague +1 more
This paper proposes a method to decrease Word Error Rate (WER) and increase recall in Automatic Speech Recognition systems using a certification-inspired mechanism, a Two-Sided Atomic Audit, and a Ran…
MelShield is a robust, in-generation audio watermarking framework that embeds identifiable signals into AI-generated speech in the Mel-spectrogram domain for reliable copyright protection and attribut…
This paper provides a unified taxonomy and controlled empirical evaluation of jailbreak attacks and defenses for Large Audio Language Models (LALMs), demonstrating that safety evaluation must consider…
This paper investigates speaker verification in a multi-utterance, multimodal setting and shows that aggregating information across anonymized speech reduces error rates.
This paper proposes a framework to protect speaker privacy in dementia detection speech recordings by introducing Cumulative Signal Attack (CSA) at the signal level and Gradient Reversal Layer (GRL) w…
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
The paper introduces Kiwano, an open-source toolkit for speaker verification research using PyTorch, offering standardized recipes, pretrained models, and integration of various architectures.
This paper proposes ZP-KWS, a lightweight framework for user-defined keyword spotting that achieves speaker-invariant zero-shot detection using a phoneme-supervised audio encoder and a compact speaker…
This paper proposes a new definition of source in source tracing as a compositional tuple of Architecture, Training Data, and other training factors, and introduces a framework using Structured Orthon…
Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng +6 more
The paper introduces VibeVoice-ASR-BitNet, a compressed real-time speech recognition model optimized for edge CPUs using heterogeneous quantization and custom SIMD kernels.
This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…
The paper introduces a phoneme-level analysis framework for interpreting speech deepfake detection results using Gradient-weighted Class Activation Mapping and speech recognition.
Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu +2 more
The paper proposes TRIDENT, a framework to restore a source speaker's identity from converted audio using a three-pronged architecture.
Meng Chen, Kun Wang, Li Lu, Jiaheng Zhang +1 more
The paper introduces AudioHijack, a framework that successfully demonstrates context-agnostic and imperceptible auditory prompt injection attacks, showing that commercial Large Audio-Language Models c…