Weihua Chen
3 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
Lumos-Nexus is a training-efficient framework that enhances video generation quality by progressively bridging generation from a lightweight model to a high-fidelity generator in a shared latent space, without sacrificing reasoning capabilities.
The paper proposes a novel render-free framework that conditions video diffusion models directly on compressed 3D human mesh tokens, enabling robust 3D-aware human motion control without relying on rendered 2D guidance.
This paper introduces ClinFusion, a vision-centric multimodal large language model designed for holistic medical understanding, featuring a Cascade Spatial-Aware Locality Fusion operator and a vision-grounded evaluation framework.
Papers
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu +20 more
This paper introduces ClinFusion, a vision-centric multimodal large language model designed for holistic medical understanding, featuring a Cascade Spatial-Aware Locality Fusion operator and a vision-…