20 results for “Monocular 3D object detection”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Jiangwei Ren, Xingyu Jiang, Zijie Song, Wei Xu +3 more
This paper proposes Wat3R, a cross-domain semi-supervised learning framework for adapting 3D reconstruction models from air to underwater scenes using unlabeled real underwater video footage and a tea…
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…
The paper proposes MoEIoU, a novel mixture-of-experts based regression loss that adaptively models bounding-box localization errors, achieving superior convergence and accuracy in object detection.
The paper introduces ZipDepth, a compact monocular depth network that achieves high zero-shot accuracy with low computational demands by combining an efficient encoder-decoder and large-scale knowledg…
The paper introduces MetricScenes, a new large-scale, in-the-wild dataset, and demonstrates that fine-tuning existing geometry models on this dataset significantly mitigates the scale-collapse problem…
This paper proposes DPNeXt, a streamlined multi-scale feature fusion decoder for Multi-Task Learning (MTL) in robotics perception systems, improving frozen VFM utilization and mitigating negative indu…
This paper proposes PairGS, a framework for open-vocabulary 3D Gaussian segmentation that models pairwise relations between Gaussians using rich signals from 3D Gaussian representations.
This paper presents a lightweight network for omnidirectional human detection and relative 2D pose estimation from planar LiDAR sequences using Space-Time Blocks.
This paper presents a two-stage recovery approach for camera-only low-cost unmanned ground vehicles to restore guideline tracking when lines are lost.
Panfei Cheng, Hongshan Yu, Wenrui Chen, Xiaojun Tang +2 more
The paper proposes a novel symmetry-aware, category-level method for 9D object pose estimation that accurately estimates translation and size first, followed by rotation, achieving state-of-the-art re…
The paper proposes xModel-KD, a cross-modal knowledge distillation framework, to improve 3D point cloud segmentation by effectively transferring rich appearance cues from 2D images to sparse 3D geomet…
Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao +3 more
The paper introduces VLM-IE3D, a framework that enhances 2D vision-language models with implicit and explicit 3D geometries learned from RGB videos, achieving superior performance on various 3D tasks.
This paper introduces VIDAR, a framework for metric dense monocular reconstruction using visual-inertial odometry and Depth Anything 3.
Jiawei Li, Ziyi Liu, Weijie Shi, Long Chen +2 more
SSR3D-LLM introduces a structured spatial reasoning interface for unified 3D-LLMs, allowing fine-grained object grounding by generating and processing sequential latent spatial steps.
The paper introduces a Mixture-Density Representation (MDA) to model depth ambiguity, effectively eliminating 'flying-point' artifacts at object boundaries by allowing pixels to predict multiple possi…