20 results for “multimodal fusion”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces Partial Information Decomposition (PID) to quantitatively separate unique, redundant, and synergistic contributions of different modalities (e.g., vision, language) in multimodal…
Zihao He, Hongjie Fang, Shirun Tang, Cewu Lu +1 more
The paper proposes LAG-Fusion, a framework for asynchronous multimodal diffusion policy composition with latency-aware guidance fusion.
The paper proposes UF-AMA, a unified framework that achieves state-of-the-art cross-domain emotion recognition by adaptively aligning and fusing multimodal physiological signals like EEG and eye-track…
The paper proposes a novel multimodal framework for session-based music recommendation that jointly models audio, lyric, and semantic content signals within a unified LLM-based sequential reasoning sy…
The paper introduces MLLM-Microscope, a system that analyzes the internal structure of multimodal large language models (MLLMs), finding that modality fusion significantly impacts the linearity and di…
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more
The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…
This paper presents Fusion Embedding family, which adds audio to a pre-trained vision-language embedding base using a connector and deep adapters.
Zixin Zhang, Fan Qi, Shuai Li, Xiaoshan Yang +1 more
The paper proposes FedMChain, a novel federated learning framework that structures multimodal training into sequential phases to mitigate modality competition and improve model performance while reduc…
The paper proposes xModel-KD, a cross-modal knowledge distillation framework, to improve 3D point cloud segmentation by effectively transferring rich appearance cues from 2D images to sparse 3D geomet…
This paper proposes a multi-level consistency-driven framework, $C^3$ASD, for robust active speaker detection in video, addressing the limitations of recent audio-visual fusion methods.
This paper introduces Align4D, a framework for generating coherent video-3D pairs using any-modal input, achieving state-of-the-art quality and consistency in X-to-4D generation.
This paper proposes DPNeXt, a streamlined multi-scale feature fusion decoder for Multi-Task Learning (MTL) in robotics perception systems, improving frozen VFM utilization and mitigating negative indu…
The paper demonstrates a coordinated, cross-modal spoofing attack that successfully deceives state-of-the-art multi-sensor fusion systems in autonomous vehicles by making multiple sensors agree on a f…
This paper proposes a multimodal framework for jointly improving Automatic Speech Recognition (ASR) and Dialect Identification (DID) in Indian languages using a Bottleneck Encoder, RoBERTa encoder, ga…
The paper introduces COMET, a novel PLS-SVD framework, to analyze the audio-text modality gap in CLAP models, showing that shared concepts are captured by a small subset of axes, and proposes a spectr…
Chong Bao, Shichen Liu, Lijun Yu, David Futschik +8 more
The paper introduces Archon, a unified, fully pretrained multimodal model that addresses the challenge of generating holistic digital humans by integrating seven modalities (including text, audio, mot…
The paper systematically compares multimodal transformer and LLM approaches for document type classification, finding that specialized multimodal Transformers outperform LLM-based models, especially w…
This paper explores the use of Large Language Models (LLMs) in data fusion tasks for tabular data and shows their superiority over traditional methods.
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…