ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

18 results for “audio intelligence”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.AIEmpiricalRecentJul 21, 2026

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar

Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle ge…

View →
cs.CRcs.AIcs.SDRecentApr 16, 2026

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection

Meng Chen, Kun Wang, Li Lu, Jiaheng Zhang +1 more

The paper introduces AudioHijack, a framework that successfully demonstrates context-agnostic and imperceptible auditory prompt injection attacks, showing that commercial Large Audio-Language Models c…

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more

The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…

View →
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View →
cs.SDEmpiricalRecentJul 22, 2026

Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R

Yixuan Xiao, Ngoc Thang Vu

This paper proposes a novel layer-wise decision fusion method for fake audio detection using deep speech models, achieving the best cross-dataset performance.

View →
cs.SDcs.AIEmpiricalRecentJul 3, 2026

DETECT-3B-Omni is Agnostic of Content and Demographics

Nicolas M. Müller, Aditya Tirumala Bukkapatnam, Dominik Schnieders, Zohaib Ahmed

The study tests the semantic independence of Resemble AI's deepfake audio detector, DETECT-3B-Omni, using 10,240 audio samples from various speakers and AI voice-cloning systems, and shows that the ac…

View →
cs.SDcs.AIRecentJun 1, 2026

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark

Seonghyeon Go, Yumin Kim

The paper introduces HAIM, a new benchmark dataset designed to move AI music detection beyond simple binary classification by tracking specific stages and types of AI integration in music production.

View →
eess.AScs.PLcs.SDEmpiricalRecentJun 19, 2026

Compiling Differentiable Audio Graphs to Real-Time DSP

Facundo Franchino, Sebastian J. Schlecht

This paper presents ADAC, a compiler that converts trained differentiable audio models into efficient FAUST code for real-time audio effects.

View →
eess.AScs.SDEmpiricalRecentJun 20, 2026

Learning from Audio-Dependency Errors: Data Curation Strategies Based on Model Confusion Patterns in Audio Question Answering

Hyeonuk Nam

The authors identify confusion patterns in a large audio-language model and use them to curate diagnostic data for fine-tuning, achieving higher accuracy than the baseline.

View →
cs.CLcs.SDRecentMay 29, 2026

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

Yuhan Song, Linhao Zhang, Aiwei Liu, Chuhan Wu +5 more

UniAudio-Token is a framework that enhances existing semantic speech tokenizers with general audio perception, allowing them to handle diverse audio types while maintaining high-fidelity speech capabi…

View →
eess.AScs.SDEmpiricalRecentJun 23, 2026

Evaluation of Headrest-Integrated Loudspeakers for Enhanced Spatial Audio Immersion in Automotive Cabins

Martin Wolters, Jacobo Giralt, Harald Mundt, Arijit Biswas

This paper conducts subjective assessments to evaluate the preference and spatial audio attributes of headrest-integrated speakers for immersive audio scenarios in automobiles.

View →
cs.SDcs.AReess.ASRecentJun 2, 2026

Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing Aid

Feyisayo Olalere, Umut Altin, Kiki van der Heijden, Marcel van Gerven

This paper characterizes the gap between current DNN-based speech enhancement systems and hearing aid constraints, and proposes a lightweight architecture to meet these constraints.

View →
cs.CRcs.SDRecentMay 18, 2026

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

Yanyun Wang, Yu Huang, Zi Liang, Xixin Wu +1 more

The paper introduces Acoustic Interference Attack (AIA), a novel jailbreak method that bypasses Large Audio Language Model (LALM) safety alignments by manipulating the underlying acoustic latent seman…

View →
cs.SDeess.ASEmpiricalRecentJul 26, 2026

Automatic Audio Equalization with Semantic Embeddings

Eloi Moliner, Vesa Välimäki, Konstantinos Drossos, Matti S. Hämäläinen

This paper proposes a data-driven method for automatic blind audio equalization using a deep neural network and semantic embeddings.

View →