ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

18 results for “Audio-Dependent Question Answering”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

eess.ASEmpiricalRecentJul 21, 2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

Haolin He, Renhe Sun, Zheqi Dai, Xingjian Du +15 more

This paper introduces Audio-Dependency Filtering (ADF) pipeline for Audio-Dependent Question Answering (ADQA) task in DCASE~2026, achieving top overall and sub-10B accuracy.

View →
eess.AScs.SDEmpiricalRecentJun 20, 2026

Learning from Audio-Dependency Errors: Data Curation Strategies Based on Model Confusion Patterns in Audio Question Answering

Hyeonuk Nam

The authors identify confusion patterns in a large audio-language model and use them to curate diagnostic data for fine-tuning, achieving higher accuracy than the baseline.

View →
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View →
cs.SDcs.CLRecentJun 3, 2026

Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

Yichen Gao, Yiqun Zhang, Zijing Wang, Yujia Li +6 more

The paper demonstrates that audio-language models often ignore conflicting audio evidence in favor of text, and proposes a training-free decoding rule, GACL, that significantly improves faithfulness b…

View →
eess.ASEmpiricalRecentJul 18, 2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker +2 more

The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…

View →
cs.SDcs.AIcs.MMRecentMay 27, 2026

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

Yuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen +1 more

The paper introduces PlanAudio, a unified LLM-based framework that directly synthesizes natural, composite audio containing speech and sounds from unconstrained free-form text prompts, outperforming e…

View →
cs.SDcs.AIEmpiricalRecentJul 16, 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer +10 more

The paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition…

View →
cs.CLRecentMay 30, 2026

LaSR: Context-Aware Speech Recognition via Latent Reasoning

Heyang Liu, Ziyang Cheng, Jiayi Huang, Wenyang Xiao +4 more

The paper proposes LaSR, a context-aware training paradigm that uses latent reasoning to significantly improve speech recognition, especially for specialized terminology, without adding latency.

View →
cs.CLcs.AIeess.ASEmpiricalRecentJul 6, 2026

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar +4 more

The paper introduces SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models using a multidimensional protocol.

View →
cs.CLcs.IREmpiricalRecentJun 10, 2026

uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking

Simon Lupart, Kidist Amde Mekonnen, Zahra Abbasiantaeb, Mohammad Aliannejadi

This paper proposes a multi-turn retrieval-augmented generation pipeline for conversational systems across four domains.

View →
cs.CLcs.AIcs.SDRecentMay 29, 2026

Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus

Máté Gedeon, Piroska Zsófia Barta, Péter Mihajlik, Katalin Mády

The paper introduces BEA-Dialogue+, an expanded 200-hour corpus for Hungarian conversational ASR, demonstrating that while larger data is challenging, specialized fine-tuning techniques significantly…

View →
cs.HCcs.AIcs.SDEmpiricalRecentJun 19, 2026

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

Youngwon Choi, Hyeonyu Kim, Taeyoun Kwon, Donghyuk Jung +1 more

CORTIS is a text-only adaptation framework that fine-tunes spoken language models for task-oriented voice agents using text-form task supervision.

View →
cs.LGEmpiricalRecentJul 23, 2026

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang +5 more

This paper proposes X$^3$-OPD, a framework for transferring reasoning capabilities from text-based models to audio-language models using on-policy distillation.

View →
eess.AScs.AIcs.SDRecentMay 27, 2026

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

Zhisheng Zhang, Xiang Li, Yixuan Zhou, Jing Peng +2 more

LoSATok proposes a low-dimensional semantic-acoustic tokenizer that efficiently compresses high-dimensional audio features into a compact latent space, significantly improving the performance and effi…

View →
cs.CLcs.AIRecentMay 27, 2026

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

Haechan Kim, Seungjun Chung, Inkyu Park, Jihoo Lee +1 more

The paper introduces three new Korean speech benchmarks (KVoiceBench, KOpenAudioBench, and KMMAU) to evaluate SpeechLMs, demonstrating that English-centric evaluation fails to capture performance gaps…

View →