ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

17 results for “Familiarity with audio-visual data”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

eess.AScs.CVEmpiricalRecentJul 17, 2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more

The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…

View →
eess.AScs.SDEmpiricalRecentJun 20, 2026

Learning from Audio-Dependency Errors: Data Curation Strategies Based on Model Confusion Patterns in Audio Question Answering

Hyeonuk Nam

The authors identify confusion patterns in a large audio-language model and use them to curate diagnostic data for fine-tuning, achieving higher accuracy than the baseline.

View →
cs.IRcs.DLEmpiricalRecentJul 22, 2026

Using Hierarchical Controlled Vocabularies to Understand CLIP Retrieval Failures in Historical Photo Collections

Ratan Sebastian, Anett Hoppe, Christoph Rippe, Ralph Ewerth

This paper investigates how the structural properties of controlled vocabularies like the Art and Architecture Thesaurus impact the performance of vision-language models like CLIP for content-based im…

View →
cs.SDcs.IRcs.LGEmpiricalRecentJul 3, 2026

Taste-aware music retrieval from audio embeddings

Matteo Spanio, Antonio Rodà

This paper establishes a content-based music information retrieval benchmark for predicting taste from audio using ten frozen audio encoders and compares their performance.

View →
cs.SDcs.AIRecentMay 30, 2026

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

Zhou Yang, Yueyi Yang

This paper investigates if upper-face affective cues enhance audiovisual sentence recognition, especially when audio is degraded, finding that while mouth cues are crucial for robustness, upper-face c…

View →
cs.CLRecentMay 28, 2026

Your Multimodal Speech Model Says I Have a Face for Radio

Maya K. Nachesa, Vlad Niculae, Vagrant Gautam

This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…

View →
eess.ASEmpiricalRecentJul 3, 2026

CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds

Harshit Rajgarhia, Shuubham Ojha, Akhil Pothanapalli, Rachuri Lokesh +3 more

The paper introduces CaReCoS, a benchmark for multimodal reasoning over medical acoustic spectrograms, and evaluates the performance of vision and omni models, finding a maximum accuracy of 51.2%.

View →
cs.CVEmpiricalRecentJun 26, 2026

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao +13 more

Introduce PerceptionRubrics, a rubric-based evaluation framework for addressing real-world brittleness of models using 1,038 images and over 12,000 instance-specific rubrics.

View →
cs.CLRecentMay 29, 2026

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

Yeil Jeong, Youngjin Yoo, Seobin Sohn, Hyejin Han +3 more

The paper introduces TeachObs, a comprehensive, human-validated benchmark for multimodal teaching observation, and evaluates frontier LLMs, finding that no single model consistently outperforms others…

View →
cs.SDcs.AIDatasetRecentJul 8, 2026

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu +7 more

The paper introduces MADB, a large-scale dataset and benchmark for music aesthetic assessment with 9,999 tracks annotated by 30 trained annotators across 10 perceptual dimensions.

View →
cs.SDcs.AIEmpiricalRecentJul 21, 2026

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar

Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle ge…

View →
cs.SDEmpiricalRecentJul 7, 2026

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music

Tomáš Sourada, Katia Vendrame, Jan Hajič

The paper introduces MusICA-MetaBench, a framework for deriving on-demand music perception benchmarks from user-provided data, ensuring statistically reliable model comparisons.

View →
cs.CVcs.AIRecentJun 1, 2026

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang +8 more

The paper introduces Moment-Video, a new benchmark that diagnoses the ability of video MLLMs to understand brief, critical visual events, revealing that current models struggle significantly with temp…

View →
cs.SDcs.AIcs.CLRecentMay 28, 2026

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

Yonggang Zhu, Liting Gao, Aidong Men, Wenwu Wang

The paper introduces COMET, a novel PLS-SVD framework, to analyze the audio-text modality gap in CLAP models, showing that shared concepts are captured by a small subset of axes, and proposes a spectr…

View →