ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Monocular 3D object detection”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVEmpiricalRecentJul 9, 2026

Wat3R: Underwater 3D Geometry Learning without Annotations

Jiangwei Ren, Xingyu Jiang, Zijie Song, Wei Xu +3 more

This paper proposes Wat3R, a cross-domain semi-supervised learning framework for adapting 3D reconstruction models from air to underwater scenes using unlabeled real underwater video footage and a tea…

View →
cs.CVcs.AIRecentMay 28, 2026

VLM3: Vision Language Models Are Native 3D Learners

Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more

The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…

View →
cs.CVcs.AIcs.LGRecentMay 30, 2026

MoEIoU: Rethinking Bounding-Box Regression as a Mixture of Experts

Vinay Edula, Priyanka Bagade

The paper proposes MoEIoU, a novel mixture-of-experts based regression loss that adaptively models bounding-box localization errors, achieving superior convergence and accuracy in object detection.

View →
cs.CVEmpiricalRecentJul 9, 2026

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

Fabio Tosi, Luca Bartolomei, Matteo Poggi, Stefano Mattoccia

The paper introduces ZipDepth, a compact monocular depth network that achieves high zero-shot accuracy with low computational demands by combining an efficient encoder-decoder and large-scale knowledg…

View →
cs.CVRecentJun 1, 2026

Honey, I Shrunk the Arc de Triomphe!

Yuanbo Xiangli, Hanyu Chen, Xueqing Tsang, Noah Snavely

The paper introduces MetricScenes, a new large-scale, in-the-wild dataset, and demonstrates that fine-tuning existing geometry models on this dataset significantly mitigates the scale-collapse problem…

View →
cs.CVcs.AIcs.ROEmpiricalRecentJul 17, 2026

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim

This paper proposes DPNeXt, a streamlined multi-scale feature fusion decoder for Multi-Task Learning (MTL) in robotics perception systems, improving frozen VFM utilization and mitigating negative indu…

View →
cs.CVEmpiricalRecentJul 1, 2026

Relation-Centric Open-Vocabulary 3D Gaussian Segmentation

Eunsung Cha, Hyunjoon Lee, Jaesik Park

This paper proposes PairGS, a framework for open-vocabulary 3D Gaussian segmentation that models pairwise relations between Gaussians using rich signals from 3D Gaussian representations.

View →
cs.ROEmpiricalRecentJul 23, 2026

Factorized Spatio-Temporal Convolutions for Human Pose Estimation from Planar Lidar

Simone Arreghini, Mirko Nava, Nicholas Carlotti, Antonio Paolillo +1 more

This paper presents a lightweight network for omnidirectional human detection and relative 2D pose estimation from planar LiDAR sequences using Space-Time Blocks.

View →
cs.ROcs.LGcs.SEEmpiricalRecentJul 13, 2026

Self-Healing Visual Recovery for Autonomous Ground Vehicles Using Camera-Only Visual Odometry

Jakob Solberg Berntzen, Safia Fatima, Leon Moonen

This paper presents a two-stage recovery approach for camera-only low-cost unmanned ground vehicles to restore guideline tracking when lines are lost.

View →
cs.CVRecentJun 1, 2026

Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution

Panfei Cheng, Hongshan Yu, Wenrui Chen, Xiaojun Tang +2 more

The paper proposes a novel symmetry-aware, category-level method for 9D object pose estimation that accurately estimates translation and size first, followed by rotation, achieving state-of-the-art re…

View →
cs.CVcs.AIRecentMay 28, 2026

xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR

Thenukan Pathmanathan, Kanchan Keisham, Thangarajah Akilan

The paper proposes xModel-KD, a cross-modal knowledge distillation framework, to improve 3D point cloud segmentation by effectively transferring rich appearance cues from 2D images to sparse 3D geomet…

View →
cs.CVcs.AIcs.LGEmpiricalRecentJul 23, 2026

3D-Aware VLMs with Implicit and Explicit Geometries

Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao +3 more

The paper introduces VLM-IE3D, a framework that enhances 2D vision-language models with implicit and explicit 3D geometries learned from RGB videos, achieving superior performance on various 3D tasks.

View →
cs.ROEmpiricalRecentJul 19, 2026

VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model

Diyari Mohammed Salih, Lingxiang Hu, Naima AitOufroukh-Mammar, Fabien Bonardi

This paper introduces VIDAR, a framework for metric dense monocular reconstruction using visual-inertial odometry and Depth Anything 3.

View →
cs.CVcs.AIRecentMay 27, 2026

SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs

Jiawei Li, Ziyi Liu, Weijie Shi, Long Chen +2 more

SSR3D-LLM introduces a structured spatial reasoning interface for unified 3D-LLMs, allowing fine-grained object grounding by generating and processing sequential latent spatial steps.

View →
cs.CVcs.AIRecentJun 1, 2026

Modeling Depth Ambiguity: A Mixture-Density Representation for Flying-Point-Free Depth Estimation

Siyuan Bian, Congrong Xu, Jun Gao

The paper introduces a Mixture-Density Representation (MDA) to model depth ambiguity, effectively eliminating 'flying-point' artifacts at object boundaries by allowing pixels to predict multiple possi…

View →