ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.06555· 20 results

cs.CERecentMay 29, 2026

CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry Priors

Xuanyi Liu, Deyi Ji, Liqun Liu, Lanyun Zhu +7 more

CamGeo is a novel framework that improves sparse camera-conditioned image-to-video generation by distilling rich 3D geometric priors into the diffusion backbone, resulting in geometrically consistent…

View →
cs.CVEmpiricalRecentJul 2, 2026

Alignment Is All You Need For X-to-4D Generation

Qiaowei Miao, Kehan Li, Yawei Luo, Yi Yang

This paper introduces Align4D, a framework for generating coherent video-3D pairs using any-modal input, achieving state-of-the-art quality and consistency in X-to-4D generation.

View →
cs.CVEmpiricalRecentJul 23, 2026

GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis

Jiahao He, Yihua Shao, Zhengkai Zhao, Pan Gao +5 more

The paper presents GrainGS, a dynamic Gaussian framework for scene reconstruction that balances fine-grained motion modeling, structural stability, and compact representation.

View →
cs.CVcs.ROEmpiricalRecentJul 17, 2026

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

Leon Jungemeyer, Alejandro Magaña, Gautham Mohan, Matthias Karl +1 more

The paper introduces PIXIE, a zero-shot framework for estimating 6D pose of an object from an RGB image using only an untextured 3D model.

View →
cs.CVcs.AIRecentMay 31, 2026

Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing

Gyojin Han, Junmo Kim

The paper proposes a novel cross-axis feature fusion architecture and an auxiliary joint-difference prediction task to significantly improve text-based 3D human motion editing by better understanding…

View →
cs.CVRecentJun 1, 2026

Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution

Panfei Cheng, Hongshan Yu, Wenrui Chen, Xiaojun Tang +2 more

The paper proposes a novel symmetry-aware, category-level method for 9D object pose estimation that accurately estimates translation and size first, followed by rotation, achieving state-of-the-art re…

View →
cs.CVEmpiricalRecentJul 23, 2026

Future Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window

Yukun Shi, Minglun Gong

This paper introduces FutureSurf, a benchmark and dataset for evaluating dynamic-scene reconstruction methods' ability to predict future surface geometry.

View →
cs.CVcs.RORecentJun 2, 2026

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image

Inhee Lee, Sangwon Baik, Sungjoo Kim, Hyeonwoo Kim +2 more

SimuScene introduces a novel compositional 3D reconstruction pipeline that integrates physics simulation directly into the shape and layout estimation process to generate stable, simulation-ready 3D s…

View →
cs.CVDatasetRecentJul 6, 2026

InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics

Erich Liang, Caleb Kha-Uong, Chinmaya Saran, Sreemanti Dey +4 more

This paper introduces InFlux++, a large-scale synthetic and real-world dataset for training and evaluating per-frame camera intrinsics prediction models.

View →
cs.CVRecentJun 3, 2026

Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text

Jaeyeong Kim, Ines Kim, Jahyeok Koo, Seungryong Kim

T2Mo is a novel framework that generates controllable dynamic 3D object shapes by combining explicit 3D trajectories for spatial guidance with natural language text semantics.

View →
cs.CVcs.AIeess.IVRecentJun 1, 2026

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

Jingyun Liang, Min Wei, Shikai Li, Yizeng Han +4 more

The paper proposes a novel render-free framework that conditions video diffusion models directly on compressed 3D human mesh tokens, enabling robust 3D-aware human motion control without relying on re…

View →
cs.CVcs.ROeess.IVEmpiricalRecentJun 29, 2026

PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking

Kai Luo, Fei Teng, Mengfei Duan, Wanjun Jia +5 more

This paper introduces PS-Track, a hierarchical pipeline for point-supervised Multi-Object Tracking (PS-MOT), which addresses spatial ambiguity and identity drift through Temporal-Feedback Prompting (T…

View →
cs.CVEmpiricalRecentJul 3, 2026

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat +2 more

This paper introduces Flex-Forcing, a framework for video generation that enables a model to operate under both bidirectional and autoregressive generation regimes, achieving better video quality and…

View →
cs.ROEmpiricalRecentJun 12, 2026

Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera

Seoyoon Kim, Kanghyun Kim, Dongwoo Ko, Yeong Jin Heo +1 more

This paper introduces Spatially Conditioned Diffusion Policy (SCDP), a single-camera manipulation system that uses end-effector trajectories as visual attention anchors.

View →
cs.CVcs.AIRecentMay 29, 2026

Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion

Jiayi Wu, Haoming Cai, Cornelia Fermuller, Christopher Metzler +1 more

Real2SAM2Real introduces a framework that uses explicit 3D caches, derived from 3D lifting models, to provide robust geometric guidance to Video Diffusion Models, significantly improving spatiotempora…

View →