20 results for “Vision-language-action models”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The paper proposes Continuous Reasoning for Vision-Language-Action (VLA) models, arguing that effective reasoning must be a shared, verifiable internal latent space rather than discrete text tokens, l…
Shiyuan Yang, Borong Zhang, Jizheng Zhang, Zhijia Tao +4 more
The paper introduces FabriVLA, a lightweight Vision-Language-Action model that achieves strong performance on the Meta-World MT50 benchmark using a compact 1B scale VLM backbone and a flow-matching ac…
Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang +10 more
The paper introduces SceneActBench, a benchmark for evaluating vision-language model agents' ability to perform actions on multi-object 3D scenes.
Yuefeng Peng, Mingzhe Li, Kejing Xia, Renhao Zhang +1 more
This paper presents the first systematic study of membership inference attacks (MIAs) against Vision-Language-Action (VLA) models, demonstrating that these models are highly vulnerable to privacy brea…
Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye +36 more
Qwen-VLA introduces a unified embodied foundation model that extends vision-language understanding to continuous action generation, enabling robust, multi-task generalization across diverse robotic ta…
This paper introduces Vision-Language-Motion Maps (VLMM), an open-vocabulary, natural-language-queryable 3D map with fused motion attributes and per-element uncertainty, which outperforms semantic-onl…
The paper addresses the difficulty of using general vision-language models (VLMs) for fine-grained driver behavior recognition by creating a new, richly described dataset and demonstrating that fine-t…
Jingtao He, Hongliang Lu, Xiaoyun Qiu, Yixuan Wang +1 more
The paper introduces a structured multi-level visual perturbation framework to systematically analyze how dependent VLA-based driving behavior is on visual information, revealing uneven visual groundi…
Dong Jing, Tianqi Zhang, Jiaqi Liu, Jinman Zhao +4 more
This paper proposes a two-stage training framework to pretrain action modules with motion priors before Vision-Language-Action (VLA) alignment, improving VLA performance and reducing optimization chal…
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…
Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang +8 more
VLA-Trace is a diagnostic framework that analyzes Vision-Language-Action (VLA) models by tracing their internal representations and external behaviors, revealing that while these models are good at vi…
Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham +4 more
Proposed G$^3$VLA, a camera-aware geometric module for pretrained vision-language-action models, injecting calibrated structure into visual tokens using intrinsic-conditioned ray embeddings, projectiv…
Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen +13 more
The paper presents SenseNova-Vision, a unified multimodal model for computer vision tasks using natural language instructions and optional visual prompts, trained primarily on a new corpus and requiri…
The paper evaluates the performance of Vision-Language Models (VLMs) in a collaborative dialogue task requiring spatial reconstruction, finding that while detailed text representations improve results…
Mingjian Gao, Wenqiao Zhang, Yuqian Yuan, Yang Dai +8 more
VISUALTHINK-VLA introduces a visual intermediate-reasoning framework that guides action prediction using compact visual evidence, achieving high accuracy and significantly low latency for real-time Vi…
Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang +25 more
This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchron…
The paper argues that large language models (LLMs) are a special case of world models and proposes a continuous spectrum between token prediction and latent-space architectures.
The paper introduces pause-and-think-T, a reasoning-centric dataset and benchmark that enables compact Vision-Language Models to perform visually grounded, context-aware action suggestion, matching la…