CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds
The paper introduces CaReCoS, a benchmark for multimodal reasoning over medical acoustic spectrograms, and evaluates the performance of vision and omni models, finding a maximum accuracy of 51.2%.
Introduced CaReCoS benchmark for multimodal reasoning over medical acoustic spectrograms.
Before reading this…
Applications
- →Medical diagnosis
To understand this paper, make sure you know these concepts first:
- Understanding of medical acoustic signals and spectrogramsfind papers →
- Familiarity with multimodal reasoning and benchmarksfind papers →
Abstract
More Like ThisMedical acoustic signals such as respiratory sounds, cardiac auscultations, and cough audio carry rich diagnostic information, yet no existing benchmark evaluates multimodal reasoning over their spectrogram representations. We address both gaps with CaReCoS, a benchmark pairing clinically grounded questions with mel-spectrogram images derived from seven medical audio datasets. Evaluating 9 state-of-the-art vision and omni models, we find that all struggle with fine-grained acoustic features encoded in spectrograms: no model reliably combines visual pattern recognition with medical knowledge, achieving a maximum accuracy of 51.2%, underscoring the need for training on medical sound visualizations.