~ similar to 2607.06555· 20 results
Xuanyi Liu, Deyi Ji, Liqun Liu, Lanyun Zhu +7 more
CamGeo is a novel framework that improves sparse camera-conditioned image-to-video generation by distilling rich 3D geometric priors into the diffusion backbone, resulting in geometrically consistent…
This paper introduces Align4D, a framework for generating coherent video-3D pairs using any-modal input, achieving state-of-the-art quality and consistency in X-to-4D generation.
Jiahao He, Yihua Shao, Zhengkai Zhao, Pan Gao +5 more
The paper presents GrainGS, a dynamic Gaussian framework for scene reconstruction that balances fine-grained motion modeling, structural stability, and compact representation.
The paper introduces PIXIE, a zero-shot framework for estimating 6D pose of an object from an RGB image using only an untextured 3D model.
The paper proposes a novel cross-axis feature fusion architecture and an auxiliary joint-difference prediction task to significantly improve text-based 3D human motion editing by better understanding…
Panfei Cheng, Hongshan Yu, Wenrui Chen, Xiaojun Tang +2 more
The paper proposes a novel symmetry-aware, category-level method for 9D object pose estimation that accurately estimates translation and size first, followed by rotation, achieving state-of-the-art re…
This paper introduces FutureSurf, a benchmark and dataset for evaluating dynamic-scene reconstruction methods' ability to predict future surface geometry.
Inhee Lee, Sangwon Baik, Sungjoo Kim, Hyeonwoo Kim +2 more
SimuScene introduces a novel compositional 3D reconstruction pipeline that integrates physics simulation directly into the shape and layout estimation process to generate stable, simulation-ready 3D s…
Erich Liang, Caleb Kha-Uong, Chinmaya Saran, Sreemanti Dey +4 more
This paper introduces InFlux++, a large-scale synthetic and real-world dataset for training and evaluating per-frame camera intrinsics prediction models.
T2Mo is a novel framework that generates controllable dynamic 3D object shapes by combining explicit 3D trajectories for spatial guidance with natural language text semantics.
Jingyun Liang, Min Wei, Shikai Li, Yizeng Han +4 more
The paper proposes a novel render-free framework that conditions video diffusion models directly on compressed 3D human mesh tokens, enabling robust 3D-aware human motion control without relying on re…
Kai Luo, Fei Teng, Mengfei Duan, Wanjun Jia +5 more
This paper introduces PS-Track, a hierarchical pipeline for point-supervised Multi-Object Tracking (PS-MOT), which addresses spatial ambiguity and identity drift through Temporal-Feedback Prompting (T…
Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat +2 more
This paper introduces Flex-Forcing, a framework for video generation that enables a model to operate under both bidirectional and autoregressive generation regimes, achieving better video quality and…
Seoyoon Kim, Kanghyun Kim, Dongwoo Ko, Yeong Jin Heo +1 more
This paper introduces Spatially Conditioned Diffusion Policy (SCDP), a single-camera manipulation system that uses end-effector trajectories as visual attention anchors.
Jiayi Wu, Haoming Cai, Cornelia Fermuller, Christopher Metzler +1 more
Real2SAM2Real introduces a framework that uses explicit 3D caches, derived from 3D lifting models, to provide robust geometric guidance to Video Diffusion Models, significantly improving spatiotempora…