ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Monocular depth estimation”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVEmpiricalRecentJul 9, 2026

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

Fabio Tosi, Luca Bartolomei, Matteo Poggi, Stefano Mattoccia

The paper introduces ZipDepth, a compact monocular depth network that achieves high zero-shot accuracy with low computational demands by combining an efficient encoder-decoder and large-scale knowledg…

View →
cs.CVRecentJun 1, 2026

Honey, I Shrunk the Arc de Triomphe!

Yuanbo Xiangli, Hanyu Chen, Xueqing Tsang, Noah Snavely

The paper introduces MetricScenes, a new large-scale, in-the-wild dataset, and demonstrates that fine-tuning existing geometry models on this dataset significantly mitigates the scale-collapse problem…

View →
cs.ROcs.CVEmpiricalRecentJul 19, 2026

DROID-ANCHOR: Odometry-Anchored Recurrent Metric Depth Estimation

Yuxuan Chen, Brook Du

This paper proposes Metric-DROID, an end-to-end recurrent architecture for precise metric depth estimation in monocular autonomous robot navigation, using proprioceptive odometry and a LSTM Update Ope…

View →
cs.CVcs.AIRecentJun 1, 2026

Modeling Depth Ambiguity: A Mixture-Density Representation for Flying-Point-Free Depth Estimation

Siyuan Bian, Congrong Xu, Jun Gao

The paper introduces a Mixture-Density Representation (MDA) to model depth ambiguity, effectively eliminating 'flying-point' artifacts at object boundaries by allowing pixels to predict multiple possi…

View →
cs.ROEmpiricalRecentJul 19, 2026

VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model

Diyari Mohammed Salih, Lingxiang Hu, Naima AitOufroukh-Mammar, Fabien Bonardi

This paper introduces VIDAR, a framework for metric dense monocular reconstruction using visual-inertial odometry and Depth Anything 3.

View →
cs.CVRecentJun 2, 2026

PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation

Shinjeong Kim, Ignacio Alzugaray, Callum Rhodes, Paul H. J. Kelly +1 more

PixVOD proposes a fully parallelizable, pixel-distributed framework for visual odometry and depth estimation that performs computations directly on the sensor using Gaussian Belief Propagation.

View →
cs.CVEmpiricalRecentJul 23, 2026

Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving

Mengshi Qi, Xiaoyang Bi, Xianlin Zhang, Huadong Ma

This paper proposes a novel self-training pipeline using unpaired real all-weather data for self-supervised depth estimation in adverse conditions, employing uncertainty modeling and robust radar fusi…

View →
cs.CVEmpiricalRecentJul 9, 2026

Wat3R: Underwater 3D Geometry Learning without Annotations

Jiangwei Ren, Xingyu Jiang, Zijie Song, Wei Xu +3 more

This paper proposes Wat3R, a cross-domain semi-supervised learning framework for adapting 3D reconstruction models from air to underwater scenes using unlabeled real underwater video footage and a tea…

View →
cs.CVEmpiricalRecentJun 30, 2026

LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization

Md Raqib Khan, Santosh Kumar Vipparthi, Subrahmanyam Murala

This paper introduces LiteMatch, a lightweight stereo matching framework that achieves strong zero-shot generalization through cost volume stabilization without expensive 3D convolutions.

View →
cs.CVcs.AIcs.ROEmpiricalRecentJul 17, 2026

DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction

Jehun Kang, Jungha Wang, Youngjun Hwang, David Hyunchul Shim

This paper proposes DPNeXt, a streamlined multi-scale feature fusion decoder for Multi-Task Learning (MTL) in robotics perception systems, improving frozen VFM utilization and mitigating negative indu…

View →
cs.ROcs.CVEmpiricalRecentJul 23, 2026

GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial Decomposition

Panagiotis Mermigkas, Argyris Manetas, Petros Maragos

GLAM-SLAM is a real-time, decoupled Gaussian-splatting SLAM system for large-scale outdoor scenes with a robust feature-based frontend and structured sparse mapping representation.

View →
cs.AIcs.CVcs.RORecentMay 28, 2026

Planning with the Views via Scene Self-Exploration

Kangrui Wang, Linjie Li, Zhengyuan Yang, Shiqi Chen +6 more

The paper addresses the challenge of multi-turn view planning for VLMs by proposing an iterative framework that uses self-exploration and view graph distillation, significantly improving planning perf…

View →
cs.CVcs.AIcs.LGRecentMay 29, 2026

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

Ulrich Prestel, Stefan Andreas Baumann, Nick Stracke, Björn Ommer

RayDer introduces a unified, feed-forward transformer that simplifies self-supervised novel view synthesis (NVS) by consolidating camera estimation, scene reconstruction, and rendering into a single,…

View →
cs.CVRecentJun 1, 2026

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

Jinpeng Liu, Yukang Xu, Yutong Li, Xingyu Liu

TROPHIES introduces a unified framework to jointly reconstruct dynamic humans, static scenes, and camera poses from multi-view videos, achieving globally consistent and physically plausible 4D reconst…

View →
cs.ROcs.LGcs.SEEmpiricalRecentJul 13, 2026

Self-Healing Visual Recovery for Autonomous Ground Vehicles Using Camera-Only Visual Odometry

Jakob Solberg Berntzen, Safia Fatima, Leon Moonen

This paper presents a two-stage recovery approach for camera-only low-cost unmanned ground vehicles to restore guideline tracking when lines are lost.

View →
cs.CVEmpiricalRecentJul 17, 2026

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction

Homanga Bharadhwaj, Yash Jangir

This paper presents MotionForesight, a method for predicting future 3D trajectories of objects in human-object interaction videos using existing video prediction models.

View →