20 results for “multimodal inputs”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces Partial Information Decomposition (PID) to quantitatively separate unique, redundant, and synergistic contributions of different modalities (e.g., vision, language) in multimodal…
This paper evaluates biases in multimodal speech recognition by testing how pairing different faces with the same audio affects transcription accuracy, finding significant quality-of-service drops acr…
The paper proposes a novel multimodal framework for session-based music recommendation that jointly models audio, lyric, and semantic content signals within a unified LLM-based sequential reasoning sy…
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more
The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…
The paper introduces COMET, a novel PLS-SVD framework, to analyze the audio-text modality gap in CLAP models, showing that shared concepts are captured by a small subset of axes, and proposes a spectr…
The paper introduces MLLM-Microscope, a system that analyzes the internal structure of multimodal large language models (MLLMs), finding that modality fusion significantly impacts the linearity and di…
Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang +2 more
A multimodal speaker identification system is proposed using Adaptive Modality Routing (AMR) for the POLY-SIM 2026 Grand Challenge, achieving high accuracy in various conditions.
Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat +4 more
The paper introduces a new inclusive, multimodal Hindi ASR benchmark with real-world recordings and diverse demographic groups, enabling more robust and realistic evaluation.
The paper introduces CaReCoS, a benchmark for multimodal reasoning over medical acoustic spectrograms, and evaluates the performance of vision and omni models, finding a maximum accuracy of 51.2%.
This paper identifies modality-order sensitivity as a failure in vision-language models and introduces a test-time training method to mitigate it, resulting in improved performance.
Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang +2 more
This paper introduces MMEE, a multilingual and multi-emotion corpus for emphasis detection, and evaluates two state-of-the-art models under various settings.
Zixin Zhang, Fan Qi, Shuai Li, Xiaoshan Yang +1 more
The paper proposes FedMChain, a novel federated learning framework that structures multimodal training into sequential phases to mitigate modality competition and improve model performance while reduc…
The paper systematically compares multimodal transformer and LLM approaches for document type classification, finding that specialized multimodal Transformers outperform LLM-based models, especially w…
This paper proposes a multimodal framework for jointly improving Automatic Speech Recognition (ASR) and Dialect Identification (DID) in Indian languages using a Bottleneck Encoder, RoBERTa encoder, ga…
Sarmistha Das, Vaibhav Vishal, Shreyas Guha, Amaan Ali +2 more
This paper introduces a Hybrid Mixture-of-Experts (HybridMoE) framework and a specialized corpus (Varnika) to significantly improve language models' ability to understand and retain figurative, cultur…
Zhiwei Chen, Yijie Li, Yimo Zhang, Shiyun Shao +8 more
GaMi is a multimodal material identification system that uses mmWave and acoustic sensing with a cross-modal subtractive disentanglement framework to achieve high accuracy (95.2%) for material identif…
The paper introduces a Conflict-aware Penalty (CP) and Statistical Loss (SL) framework to stabilize and balance the training of multimodal sentiment analysis models, achieving state-of-the-art perform…
Xiaoyang Jiang, Yanlai Yang, Kenneth A. Norman, Brenden Lake +1 more
The paper introduces BabyCL, a continual multimodal learning framework that processes egocentric video data in a single chronological pass, demonstrating that meaningful word-referent mappings can be…