20 results for “Monocular depth estimation”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper introduces ZipDepth, a compact monocular depth network that achieves high zero-shot accuracy with low computational demands by combining an efficient encoder-decoder and large-scale knowledg…
The paper introduces MetricScenes, a new large-scale, in-the-wild dataset, and demonstrates that fine-tuning existing geometry models on this dataset significantly mitigates the scale-collapse problem…
This paper proposes Metric-DROID, an end-to-end recurrent architecture for precise metric depth estimation in monocular autonomous robot navigation, using proprioceptive odometry and a LSTM Update Ope…
The paper introduces a Mixture-Density Representation (MDA) to model depth ambiguity, effectively eliminating 'flying-point' artifacts at object boundaries by allowing pixels to predict multiple possi…
This paper introduces VIDAR, a framework for metric dense monocular reconstruction using visual-inertial odometry and Depth Anything 3.
PixVOD proposes a fully parallelizable, pixel-distributed framework for visual odometry and depth estimation that performs computations directly on the sensor using Gaussian Belief Propagation.
This paper proposes a novel self-training pipeline using unpaired real all-weather data for self-supervised depth estimation in adverse conditions, employing uncertainty modeling and robust radar fusi…
Jiangwei Ren, Xingyu Jiang, Zijie Song, Wei Xu +3 more
This paper proposes Wat3R, a cross-domain semi-supervised learning framework for adapting 3D reconstruction models from air to underwater scenes using unlabeled real underwater video footage and a tea…
This paper introduces LiteMatch, a lightweight stereo matching framework that achieves strong zero-shot generalization through cost volume stabilization without expensive 3D convolutions.
This paper proposes DPNeXt, a streamlined multi-scale feature fusion decoder for Multi-Task Learning (MTL) in robotics perception systems, improving frozen VFM utilization and mitigating negative indu…
GLAM-SLAM is a real-time, decoupled Gaussian-splatting SLAM system for large-scale outdoor scenes with a robust feature-based frontend and structured sparse mapping representation.
Kangrui Wang, Linjie Li, Zhengyuan Yang, Shiqi Chen +6 more
The paper addresses the challenge of multi-turn view planning for VLMs by proposing an iterative framework that uses self-exploration and view graph distillation, significantly improving planning perf…
RayDer introduces a unified, feed-forward transformer that simplifies self-supervised novel view synthesis (NVS) by consolidating camera estimation, scene reconstruction, and rendering into a single,…
TROPHIES introduces a unified framework to jointly reconstruct dynamic humans, static scenes, and camera poses from multi-view videos, achieving globally consistent and physically plausible 4D reconst…
This paper presents a two-stage recovery approach for camera-only low-cost unmanned ground vehicles to restore guideline tracking when lines are lost.
This paper presents MotionForesight, a method for predicting future 3D trajectories of objects in human-object interaction videos using existing video prediction models.