20 results for “video-action models”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Sizhe Lester Li, Evan Kim, Xingjian Bai, Tong Zhao +3 more
The paper proposes VERA, a decoupled policy that uses an action-free video world model combined with an embodiment-specific Inverse Dynamics Model (IDM) to achieve generalizable, zero-shot robot contr…
Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang +25 more
This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchron…
Boyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang +2 more
The paper proposes an Understanding-Enhanced Model Collaboration Method (UE-MCM) to accurately detect subtle and rare mistakes in egocentric videos by combining coarse-grained workflow understanding w…
Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang +23 more
This paper introduces LingBot-Video, a video pretraining paradigm for embodied intelligence using a DiT-based approach, Mixture-of-Experts framework, and extensive robot-oriented data.
Dong Jing, Tianqi Zhang, Jiaqi Liu, Jinman Zhao +4 more
This paper proposes a two-stage training framework to pretrain action modules with motion priors before Vision-Language-Action (VLA) alignment, improving VLA performance and reducing optimization chal…
This paper proposes robot-factored world models for action-conditioned video prediction in robotics, which factor out action realization and robot rendering to avoid learning the realization process a…
Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang +10 more
The paper introduces SceneActBench, a benchmark for evaluating vision-language model agents' ability to perform actions on multi-object 3D scenes.
The paper introduces pause-and-think-T, a reasoning-centric dataset and benchmark that enables compact Vision-Language Models to perform visually grounded, context-aware action suggestion, matching la…
Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat +2 more
This paper introduces Flex-Forcing, a framework for video generation that enables a model to operate under both bidirectional and autoregressive generation regimes, achieving better video quality and…
Shiyuan Yang, Borong Zhang, Jizheng Zhang, Zhijia Tao +4 more
The paper introduces FabriVLA, a lightweight Vision-Language-Action model that achieves strong performance on the Meta-World MT50 benchmark using a compact 1B scale VLM backbone and a flow-matching ac…
Wenhao Li, Xueying Jiang, Quanhao Qian, Deli Zhao +3 more
This paper introduces Camera-Centric VLA, a new model for Vision-Language-Action policies that predicts camera-centric actions and hand-eye matrix, allowing the policy to figure out camera geometry on…
Yue Peng, Yongzhe Zhao, Artur Habuda, Khuyen Pham +4 more
Proposed G$^3$VLA, a camera-aware geometric module for pretrained vision-language-action models, injecting calibrated structure into visual tokens using intrinsic-conditioned ray embeddings, projectiv…
The paper proposes Continuous Reasoning for Vision-Language-Action (VLA) models, arguing that effective reasoning must be a shared, verifiable internal latent space rather than discrete text tokens, l…
Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang +8 more
The paper introduces Moment-Video, a new benchmark that diagnoses the ability of video MLLMs to understand brief, critical visual events, revealing that current models struggle significantly with temp…
The paper introduces ConTrans, a novel local-global multi-scale encoder that combines convolutional and transformer features to significantly improve zero-shot temporal action localization by capturin…
Yuefeng Peng, Mingzhe Li, Kejing Xia, Renhao Zhang +1 more
This paper presents the first systematic study of membership inference attacks (MIAs) against Vision-Language-Action (VLA) models, demonstrating that these models are highly vulnerable to privacy brea…
Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang +9 more
The paper introduces WorldDirector, a framework for creating controllable video worlds with persistent dynamic object memory and exact visual identities.