17 results for “Familiarity with audio-visual data”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more
The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…
The authors identify confusion patterns in a large audio-language model and use them to curate diagnostic data for fine-tuning, achieving higher accuracy than the baseline.
This paper investigates how the structural properties of controlled vocabularies like the Art and Architecture Thesaurus impact the performance of vision-language models like CLIP for content-based im…
This paper establishes a content-based music information retrieval benchmark for predicting taste from audio using ten frozen audio encoders and compares their performance.
This paper investigates if upper-face affective cues enhance audiovisual sentence recognition, especially when audio is degraded, finding that while mouth cues are crucial for robustness, upper-face c…
This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…
The paper introduces CaReCoS, a benchmark for multimodal reasoning over medical acoustic spectrograms, and evaluates the performance of vision and omni models, finding a maximum accuracy of 51.2%.
Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao +13 more
Introduce PerceptionRubrics, a rubric-based evaluation framework for addressing real-world brittleness of models using 1,038 images and over 12,000 instance-specific rubrics.
Yeil Jeong, Youngjin Yoo, Seobin Sohn, Hyejin Han +3 more
The paper introduces TeachObs, a comprehensive, human-validated benchmark for multimodal teaching observation, and evaluates frontier LLMs, finding that no single model consistently outperforms others…
Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu +7 more
The paper introduces MADB, a large-scale dataset and benchmark for music aesthetic assessment with 9,999 tracks annotated by 30 trained annotators across 10 perceptual dimensions.
Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle ge…
The paper introduces MusICA-MetaBench, a framework for deriving on-demand music perception benchmarks from user-provided data, ensuring statistically reliable model comparisons.
Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang +8 more
The paper introduces Moment-Video, a new benchmark that diagnoses the ability of video MLLMs to understand brief, critical visual events, revealing that current models struggle significantly with temp…
The paper introduces COMET, a novel PLS-SVD framework, to analyze the audio-text modality gap in CLAP models, showing that shared concepts are captured by a small subset of axes, and proposes a spectr…