20 results for “scene description accuracy”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper introduces the Complex Social Behavior (CSB) dataset and evaluates the progress of scene description accuracy in vision language models (VLMs) from 2017 to 2025. The authors find that MLLMs…
The paper argues that benchmarking Vision-Language Models (VLMs) for urban perception must treat human disagreement and non-response as key measurement outcomes, rather than assuming perfect consensus…
Places in the Wild introduces a massive, high-resolution RAW photograph dataset of 67,574 images captured in situ across 810 locations, providing unprecedented detail for ecologically valid vision res…
The paper introduces the Image Reconstruction Game, a benchmark showing that the quality of the descriptive model is the primary determinant of image reconstruction success, while the generator's role…
The paper introduces AdvScene, a novel scene-grounded framework that measures the real-world 'scene robustness' of adversarial patches by characterizing their operational envelope across varying viewp…
This paper addresses robot localization in GPS-denied indoor environments using a semantic reasoning approach with a vision-language model, achieving high accuracy with a composite loss and curriculum…
This paper introduces FutureSurf, a benchmark and dataset for evaluating dynamic-scene reconstruction methods' ability to predict future surface geometry.
Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang +10 more
The paper introduces SceneActBench, a benchmark for evaluating vision-language model agents' ability to perform actions on multi-object 3D scenes.
Lianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang +2 more
This paper formalizes the binding problem using information theory and develops a probing method to measure binding information in deep learning representations, demonstrating that binding is crucial…
The paper reframes industrial visual sim-to-real transfer as a domain-gap problem categorized by the availability of explicit object geometry (CAD), arguing that the required prior evidence dictates t…
Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton +2 more
This paper introduces a new evaluation framework, SpatialUncertain, demonstrating that current Vision-Language Models (VLMs) are prone to overconfident and incorrect answers to spatial questions when…
The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…
FLORO is a multimodal geospatial foundation model that learns transferable remote sensing representations from a small, diverse corpus, achieving strong performance across various sensor types and res…
This paper introduces Vision-Language-Motion Maps (VLMM), an open-vocabulary, natural-language-queryable 3D map with fused motion attributes and per-element uncertainty, which outperforms semantic-onl…
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more
The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…
Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao +13 more
Introduce PerceptionRubrics, a rubric-based evaluation framework for addressing real-world brittleness of models using 1,038 images and over 12,000 instance-specific rubrics.
The paper introduces SynCity 3000, a framework for generating large, coherent 3D scenes using a convolutional generator, addressing the scarcity of 3D scene data for training.
The paper proposes a unified framework to systematically redefine instance matching for Panoptic Quality evaluation, moving beyond the standard One-to-One matching to accommodate complex scenarios lik…
This pilot study evaluates curator-guided multilingual art description using a small, on-premise VLM (Qwen2.5-VL-3B-Instruct) for German, Romanian, and Serbian, finding that language-specific adapters…