Shijian Lu
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper introduces Camera-Centric VLA, a new model for Vision-Language-Action policies that predicts camera-centric actions and hand-eye matrix, allowing the policy to figure out camera geometry on its own.
The paper introduces VLM-IE3D, a framework that enhances 2D vision-language models with implicit and explicit 3D geometries learned from RGB videos, achieving superior performance on various 3D tasks.
Papers
3D-Aware VLMs with Implicit and Explicit Geometries
Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao +3 more
The paper introduces VLM-IE3D, a framework that enhances 2D vision-language models with implicit and explicit 3D geometries learned from RGB videos, achieving superior performance on various 3D tasks.