20 results for “multimodal reasoning”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun +4 more
The paper proposes CSMR, a cognitive scheduling framework that allows a language model to dynamically decide when to acquire task-relevant visual evidence, significantly improving multimodal reasoning…
The paper introduces a new benchmark (BGTD) and a multimodal framework (mmTraffic) that enables explainable, evidence-grounded interpretation of encrypted network traffic using LLMs.
Jiawei Kong, Hao Fang, Shunxiang Liao, Jinyu Li +4 more
The paper proposes Reasoning-Conditioned Direct Preference Optimization (RC-DPO) to effectively mitigate hallucinations in multimodal large reasoning models by explicitly conditioning the preference o…
Seojeong Park, Jiho Choi, Junyong Kang, Seonho Lee +2 more
The paper addresses Perceptual Judgment Bias in multimodal LLM judges by introducing a new dataset and a unified training framework that forces models to prioritize visual evidence over plausible text…
The paper introduces Partial Information Decomposition (PID) to quantitatively separate unique, redundant, and synergistic contributions of different modalities (e.g., vision, language) in multimodal…
Hongxing Li, Xiufeng Huang, Dingming Li, Wenjing Jiang +10 more
This paper proposes Perceive-to-Reason (P2R), a framework for fine-grained visual reasoning that decouples perception from reasoning and introduces a new reinforcement learning strategy.
Ke Xu, Yuhao Wang, Ziyang Cheng, Hongcheng Liu +2 more
The paper introduces MOV-Bench, a challenging benchmark for multi-hop audio-visual reasoning, and proposes AOP-Agent, an agentic framework that significantly improves open-source Omni-LLMs' ability to…
Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang +5 more
This paper proposes X$^3$-OPD, a framework for transferring reasoning capabilities from text-based models to audio-language models using on-policy distillation.
Pengchao Feng, Chao-Hong Tan, Qian Chen, Wen Wang +2 more
This paper proposes Efficient Chain-of-Modality Reasoning (ECoM Reasoning), a framework to improve reasoning ability in spoken language models (SLMs) for mathematical question answering tasks by compr…
Lianyu Hu, Shengqian Qin, Zeqin Liao, Qing Guo +3 more
This paper proposes CoLT, a framework that enables multi-modal models to reason through a chain of latent thought representations instead of text tokens, improving performance and reducing inference t…
Garvin Guo, Yu Chen, Xiang Wang, Shuai Li +3 more
The paper deconstructs latent visual reasoning tokens into components and finds that the performance gains are primarily due to boundary markers and attention patterns, not the tokens' ability to enco…
Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu +5 more
This paper introduces DramaSR-532K, a large-scale benchmark for speaker recognition in long-form TV dramas, and proposes DramaSR-LRM, a robust approach for speaker recognition using a large reasoning…
Yinsong Xu, Wei Jing, Liuxin Zhang, Wanjun Lv +1 more
The paper proposes a unified framework that decouples long-video reasoning into semantic and visual evidence, significantly improving performance on the HD-EPIC VQA Challenge.
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more
The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…
Yuhan Wang, Shuochen Chang, Yalin Feng, Dongsheng Ma +7 more
The paper proposes EAGLE, a novel evidence-aligned multi-agent framework, demonstrating that requiring shared visual evidence among agents is crucial for achieving reliable and trustworthy consensus i…
Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton +2 more
This paper introduces a new evaluation framework, SpatialUncertain, demonstrating that current Vision-Language Models (VLMs) are prone to overconfident and incorrect answers to spatial questions when…