20 results for “video pretraining”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang +23 more
This paper introduces LingBot-Video, a video pretraining paradigm for embodied intelligence using a DiT-based approach, Mixture-of-Experts framework, and extensive robot-oriented data.
Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang +25 more
This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchron…
The paper introduces Knowledge-Intensive Video Generation (KIVI) as a challenging benchmark for evaluating video models on factuality and practical usefulness, showing that current state-of-the-art sy…
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more
The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…
This paper introduces ElasticTTT, a framework to preserve the generative prior in Test-Time Tuning (TTT) of pretrained diffusion models for video editing, preventing Prior Collapse.
Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin +3 more
This paper proposes LongE2V, a method for high-quality video recovery from sparse event streams using pre-trained video diffusion priors.
Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya +5 more
The paper introduces AnyGroundBench, a domain-adaptation benchmark for Spatio-Temporal Video Grounding, to evaluate the ability of Vision-Language Models to adapt to specialized domains.
The paper introduces LongVQUBench, a comprehensive benchmark for long-term video quality understanding with 1200 diverse videos and 1500 questions.
Yeil Jeong, Youngjin Yoo, Seobin Sohn, Hyejin Han +3 more
The paper introduces TeachObs, a comprehensive, human-validated benchmark for multimodal teaching observation, and evaluates frontier LLMs, finding that no single model consistently outperforms others…
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai +1 more
The paper proposes a training-free framework, Visual Representation-Guided Video-LLM Reasoning, to perform composed video retrieval by using visual examples and text instructions, achieving strong per…
Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li +2 more
This paper presents video-SALMONN-R$^3$, an end-to-end video-LLM that enables re-watch through reinforcement learning, improving question answering performance with a two-stage paradigm.
VidPrism introduces a novel heterogeneous Mixture-of-Experts framework that specializes temporal processing by dividing labor among experts, achieving state-of-the-art performance in image-to-video tr…
Peng Yu, Yuankai Fan, Yang Qiu, Tian Li +3 more
This paper introduces Arachne, a framework for efficient Text-to-Video model training at scale, reducing iteration time by up to 65% over leading frameworks.
Yinsong Xu, Wei Jing, Liuxin Zhang, Wanjun Lv +1 more
The paper proposes a unified framework that decouples long-video reasoning into semantic and visual evidence, significantly improving performance on the HD-EPIC VQA Challenge.
Xiaoyang Jiang, Yanlai Yang, Kenneth A. Norman, Brenden Lake +1 more
The paper introduces BabyCL, a continual multimodal learning framework that processes egocentric video data in a single chronological pass, demonstrating that meaningful word-referent mappings can be…
This paper presents MotionForesight, a method for predicting future 3D trajectories of objects in human-object interaction videos using existing video prediction models.
Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola +5 more
This paper proposes VideoTreeSearch (VTS), a framework for grounded long-video question answering that uses iterative self-correcting search over an adaptive temporal tree.
Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He +5 more
This paper proposes OPSD-V, an on-policy self-distillation method for reducing long-horizon degradation in few-step autoregressive video diffusion models by introducing real long-video data as tempora…
Xinyu Mao, Yuhui Zeng, Xiaokun Liu, Wenyu Qin +5 more
This paper introduces CineCap, a framework for cinematographic captioning using structured reasoning, spatio-temporal anchors, and reinforcement learning.