20 results for “Understanding of video pretraining”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang +23 more
This paper introduces LingBot-Video, a video pretraining paradigm for embodied intelligence using a DiT-based approach, Mixture-of-Experts framework, and extensive robot-oriented data.
The paper introduces LongVQUBench, a comprehensive benchmark for long-term video quality understanding with 1200 diverse videos and 1500 questions.
Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more
The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…
Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li +2 more
This paper presents video-SALMONN-R$^3$, an end-to-end video-LLM that enables re-watch through reinforcement learning, improving question answering performance with a two-stage paradigm.
Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya +5 more
The paper introduces AnyGroundBench, a domain-adaptation benchmark for Spatio-Temporal Video Grounding, to evaluate the ability of Vision-Language Models to adapt to specialized domains.
This paper introduces ElasticTTT, a framework to preserve the generative prior in Test-Time Tuning (TTT) of pretrained diffusion models for video editing, preventing Prior Collapse.
VidPrism introduces a novel heterogeneous Mixture-of-Experts framework that specializes temporal processing by dividing labor among experts, achieving state-of-the-art performance in image-to-video tr…
Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin +3 more
This paper proposes LongE2V, a method for high-quality video recovery from sparse event streams using pre-trained video diffusion priors.
The paper introduces Knowledge-Intensive Video Generation (KIVI) as a challenging benchmark for evaluating video models on factuality and practical usefulness, showing that current state-of-the-art sy…
Xiaoyang Jiang, Yanlai Yang, Kenneth A. Norman, Brenden Lake +1 more
The paper introduces BabyCL, a continual multimodal learning framework that processes egocentric video data in a single chronological pass, demonstrating that meaningful word-referent mappings can be…
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai +1 more
The paper proposes a training-free framework, Visual Representation-Guided Video-LLM Reasoning, to perform composed video retrieval by using visual examples and text instructions, achieving strong per…
This paper presents MotionForesight, a method for predicting future 3D trajectories of objects in human-object interaction videos using existing video prediction models.
Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He +5 more
This paper proposes OPSD-V, an on-policy self-distillation method for reducing long-horizon degradation in few-step autoregressive video diffusion models by introducing real long-video data as tempora…
VideoMLA introduces a novel Multi-Head Latent Attention (MLA) mechanism that replaces per-head KV caches with a shared low-rank content latent, significantly reducing memory and improving throughput f…
Yinsong Xu, Wei Jing, Liuxin Zhang, Wanjun Lv +1 more
The paper proposes a unified framework that decouples long-video reasoning into semantic and visual evidence, significantly improving performance on the HD-EPIC VQA Challenge.
Boyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang +2 more
The paper proposes an Understanding-Enhanced Model Collaboration Method (UE-MCM) to accurately detect subtle and rare mistakes in egocentric videos by combining coarse-grained workflow understanding w…
Yeil Jeong, Youngjin Yoo, Seobin Sohn, Hyejin Han +3 more
The paper introduces TeachObs, a comprehensive, human-validated benchmark for multimodal teaching observation, and evaluates frontier LLMs, finding that no single model consistently outperforms others…
This paper identifies modality-order sensitivity as a failure in vision-language models and introduces a test-time training method to mitigate it, resulting in improved performance.