20 results for “Understanding of computer vision concepts”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Lianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang +2 more
This paper formalizes the binding problem using information theory and develops a probing method to measure binding information in deep learning representations, demonstrating that binding is crucial…
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more
The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…
This paper addresses the vulnerability of DNNs used in robotic semantic segmentation to adversarial attacks by proposing specialized detection strategies to enhance safety in robotic perception system…
The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…
This paper demonstrates that standard image classifiers can be interpreted as multiple-instance learning models, allowing for the recovery of spatial class evidence from image-level logits.
The paper reframes industrial visual sim-to-real transfer as a domain-gap problem categorized by the availability of explicit object geometry (CAD), arguing that the required prior evidence dictates t…
This book provides a compact, derivation-oriented mathematical primer that connects major families of generative AI models, showing their underlying structural relationships.
This paper systematically explores the convex polygon reconstruction problem with specified sets of features, contributing new testing algorithms and hardness results.
Fangzhou Lin, Peiran Li, Lingyu Xu, Wenjing Chen +11 more
The paper introduces CV-Arena, a large-scale open benchmark for instructional computer vision, demonstrating that professional-grade image editing requires advanced capabilities in physical reasoning…
Ziying Chen, Yang Cao, He Sun, Beining Yang +1 more
The paper proposes a novel geometric embedding hashing method to recover object correspondences (vector links) between two embedding clouds generated by different black-box encoders using only a small…
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…
Zhishan Tao, Ruoyu Wang, Yucheng Wu, Enjun Du +5 more
The paper proposes CARA, an interpretable framework for collision anticipation in autonomous driving using domain-grounded risk concepts, aligning them with video frames, and organizing them into evol…
This paper presents an open-source computer vision pipeline for classifying vehicle body types from naturalistic roadway video.
Hongxing Li, Xiufeng Huang, Dingming Li, Wenjing Jiang +10 more
This paper proposes Perceive-to-Reason (P2R), a framework for fine-grained visual reasoning that decouples perception from reasoning and introduces a new reinforcement learning strategy.
Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao +3 more
The paper introduces VLM-IE3D, a framework that enhances 2D vision-language models with implicit and explicit 3D geometries learned from RGB videos, achieving superior performance on various 3D tasks.
Yue Feng, Jingjing Li, Qijia Lu, Wei Ji +8 more
This paper addresses the challenge of detecting and explaining AI-manipulated segments within long, untrimmed videos by proposing a new benchmark and a coarse-to-fine forensic detection framework.