ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “pose estimation”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVRecentJun 1, 2026

Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution

Panfei Cheng, Hongshan Yu, Wenrui Chen, Xiaojun Tang +2 more

The paper proposes a novel symmetry-aware, category-level method for 9D object pose estimation that accurately estimates translation and size first, followed by rotation, achieving state-of-the-art re…

View →
cs.CVcs.ROEmpiricalRecentJul 17, 2026

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

Leon Jungemeyer, Alejandro Magaña, Gautham Mohan, Matthias Karl +1 more

The paper introduces PIXIE, a zero-shot framework for estimating 6D pose of an object from an RGB image using only an untextured 3D model.

View →
cs.CVcs.RORecentJun 3, 2026

CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation

Yurim Jeon, Dongseong Seo, Seung-Woo Seo

CIPER proposes a unified transformer framework to simultaneously perform cross-view image retrieval and precise 3-DoF pose estimation, overcoming the limitations of cascaded, separate methods.

View →
cs.ROEmpiricalRecentJul 23, 2026

Factorized Spatio-Temporal Convolutions for Human Pose Estimation from Planar Lidar

Simone Arreghini, Mirko Nava, Nicholas Carlotti, Antonio Paolillo +1 more

This paper presents a lightweight network for omnidirectional human detection and relative 2D pose estimation from planar LiDAR sequences using Space-Time Blocks.

View →
cs.CVEmpiricalRecentJul 7, 2026

ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

Ruihang Zhang, Felix Taubner, Pooja Ravi, Kiriakos N. Kutulakos +1 more

The paper introduces ProxyPose, a method for six-degree-of-freedom (6-DoF) pose tracking using a video diffusion model and a single marked pixel in the first frame.

View →
cs.ROEmpiricalRecentJul 17, 2026

BayesContact: Uncertain Pose Estimation via Visuo-Tactile Proposals and Simulation-based Inference

Aditya Kamireddypalli, Matias Mattamala, Joao Moura, Russell Buchanan +2 more

The paper introduces BayesContact, a framework for visuo-tactile pose estimation using simulation-based inference, improving pose observability and insertion success by 30%.

View →
cs.CVRecentJun 1, 2026

PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation

Xiaohang Yu, Ti Wang, Mackenzie Weygandt Mathis

PRIMA is a framework that significantly improves 3D quadruped mesh recovery by integrating biological knowledge and a test-time adaptation strategy, achieving state-of-the-art results on diverse and c…

View →
cs.ROEmpiricalRecentJul 1, 2026

Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization

Suraj Borate, Aarav Shah, Madhu Vadali

This paper addresses robot localization in GPS-denied indoor environments using a semantic reasoning approach with a vision-language model, achieving high accuracy with a composite loss and curriculum…

View →
cs.CVcs.MAcs.MMEmpiricalRecentJul 17, 2026

Toward Semantic Communication for Real-time Mobile 3D Reconstruction

Fangzhou Zhao, Yao Sun, Xuesong Liu, Runze Cheng +2 more

This paper proposes a SemCom framework for real-time mobile 3D reconstruction, which includes a semantic transceiver and a confidence-guided geometric estimation method.

View →
cs.ROcs.AIRecentMay 28, 2026

V2I Work Zone Geometry Reconstruction with Pose-Conditioned UWB Range Denoising

Jiaxi Liu, Hangyu Li, Yang Cheng, Rui Gana +6 more

The paper proposes a pose-conditioned, permutation-equivariant denoiser to accurately reconstruct work zone geometry using noisy Ultra-Wideband (UWB) range data from connected and autonomous vehicles…

View →
cs.CVEmpiricalRecentJul 7, 2026

Vision as Unified Multimodal Generation

Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more

The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…

View →
cs.ROEmpiricalRecentJul 19, 2026

VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model

Diyari Mohammed Salih, Lingxiang Hu, Naima AitOufroukh-Mammar, Fabien Bonardi

This paper introduces VIDAR, a framework for metric dense monocular reconstruction using visual-inertial odometry and Depth Anything 3.

View →
cs.CVRecentJun 1, 2026

Honey, I Shrunk the Arc de Triomphe!

Yuanbo Xiangli, Hanyu Chen, Xueqing Tsang, Noah Snavely

The paper introduces MetricScenes, a new large-scale, in-the-wild dataset, and demonstrates that fine-tuning existing geometry models on this dataset significantly mitigates the scale-collapse problem…

View →
cs.CVcs.AIcs.LGRecentMay 30, 2026

MoEIoU: Rethinking Bounding-Box Regression as a Mixture of Experts

Vinay Edula, Priyanka Bagade

The paper proposes MoEIoU, a novel mixture-of-experts based regression loss that adaptively models bounding-box localization errors, achieving superior convergence and accuracy in object detection.

View →
cs.CVcs.AIcs.LGRecentJun 1, 2026

Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association

Matvei Shelukhan, Timur Mamedov, Aleksandr Chukhrov, Karina Kvanchiani

The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…

View →
cs.CERecentMay 29, 2026

CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry Priors

Xuanyi Liu, Deyi Ji, Liqun Liu, Lanyun Zhu +7 more

CamGeo is a novel framework that improves sparse camera-conditioned image-to-video generation by distilling rich 3D geometric priors into the diffusion backbone, resulting in geometrically consistent…

View →
cs.CVcs.AIcs.LGRecentMay 29, 2026

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

Mohammed Asad Karim, Vinay Kumar Verma

The paper introduces a novel two-stage framework to achieve robust, category-agnostic object localization in-context (ICL) by optimizing attention and minimizing localization error using reinforcement…

View →