20 results for “Familiarity with computer vision and natural language processing concepts”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper evaluates the performance of a Large Language Model (LLM) in a high-stakes context by comparing it to human experts and measuring variance and error magnitude.
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more
The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…
This paper introduces the Complex Social Behavior (CSB) dataset and evaluates the progress of scene description accuracy in vision language models (VLMs) from 2017 to 2025. The authors find that MLLMs…
Fangzhou Lin, Peiran Li, Lingyu Xu, Wenjing Chen +11 more
The paper introduces CV-Arena, a large-scale open benchmark for instructional computer vision, demonstrating that professional-grade image editing requires advanced capabilities in physical reasoning…
This paper investigates how the structural properties of controlled vocabularies like the Art and Architecture Thesaurus impact the performance of vision-language models like CLIP for content-based im…
The paper proposes two novel CAPTCHA types—ASCII art and overlapping audio—and demonstrates that current frontier LLMs struggle significantly to solve them, suggesting they are highly effective anti-b…
Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo +21 more
The paper introduces Dr. DocBench, a difficulty-aware, comprehensive benchmark designed to rigorously test expert-level and challenging document parsing capabilities for VLMs, demonstrating that curre…
KidsNanny is a two-stage multimodal content moderation pipeline that achieves high accuracy and efficiency in detecting child safety threats, particularly excelling in text-embedded content.
The paper introduces a structured benchmark (TGAD) showing that current text-guided anomaly detection models often overstate their language conditioning, as performance significantly degrades when the…
Yuxiang Xie, Qi Lv, Jianming Xing, Zijian Hong +3 more
The paper presents RoboSpatialBrain, a system that combines two mechanisms to improve spatial reasoning in vision-language models for embodied tasks, achieving first place in the RoboSpatial Challenge…
The paper identifies a fundamental mismatch between standard pairwise ranking metrics (like AP and FPR-95) and the true assignment objective in multi-view object association, proposing a Sinkhorn-base…
Aakash Pant, Kavya Shah, Apoorv Agnihotri, Sneha Nikam +2 more
The paper critiques current AI benchmarking practices for low-resource settings, arguing that evaluation must shift focus from isolated model performance to the holistic performance of the deployed sy…
Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao +13 more
Introduce PerceptionRubrics, a rubric-based evaluation framework for addressing real-world brittleness of models using 1,038 images and over 12,000 instance-specific rubrics.
The paper proposes CYKNN, a novel recurrent neural network architecture that directly encodes the CYK parsing algorithm, demonstrating superior performance over large language models on syntactic pars…
This paper identifies modality-order sensitivity as a failure in vision-language models and introduces a test-time training method to mitigate it, resulting in improved performance.
Meng Chen, Anya Ji, Tsung-Han Wu, Tobias Maringgele +3 more
The paper introduces DigitalCoach, a dataset of human expert-novice computer use coaching sessions, and evaluates the ability of state-of-the-art models to teach humans how to use computers.
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…
Zijie Zhou, Dandan Zhu, Hangxiangpan Wang, Heng Zhang +2 more
The paper proposes AsyMoE, a novel Mixture of Experts architecture for Large Vision-Language Models that explicitly models the inherent asymmetry between visual and linguistic modalities, achieving si…
This paper addresses the vulnerability of DNNs used in robotic semantic segmentation to adversarial attacks by proposing specialized detection strategies to enhance safety in robotic perception system…
Lianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang +2 more
This paper formalizes the binding problem using information theory and develops a probing method to measure binding information in deep learning representations, demonstrating that binding is crucial…