ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Understanding of video pretraining”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVEmpiricalRecentJul 8, 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang +23 more

This paper introduces LingBot-Video, a video pretraining paradigm for embodied intelligence using a DiT-based approach, Mixture-of-Experts framework, and extensive robot-oriented data.

View →
cs.CVcs.AIEmpiricalRecentJul 1, 2026

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

Arpita Nema, Hanwei Zhu, Xi Zhang, Weisi Lin

The paper introduces LongVQUBench, a comprehensive benchmark for long-term video quality understanding with 1200 diverse videos and 1500 questions.

View →
eess.AScs.CVEmpiricalRecentJul 17, 2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more

The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…

View →
cs.CVcs.AIcs.SDEmpiricalRecentJun 23, 2026

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li +2 more

This paper presents video-SALMONN-R$^3$, an end-to-end video-LLM that enables re-watch through reinforcement learning, improving question answering performance with a two-stage paradigm.

View →
cs.CVcs.AIEmpiricalRecentJul 2, 2026

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya +5 more

The paper introduces AnyGroundBench, a domain-adaptation benchmark for Spatio-Temporal Video Grounding, to evaluate the ability of Vision-Language Models to adapt to specialized domains.

View →
cs.CVcs.AIEmpiricalRecentJul 23, 2026

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu

This paper introduces ElasticTTT, a framework to preserve the generative prior in Test-Time Tuning (TTT) of pretrained diffusion models for video editing, preventing Prior Collapse.

View →
cs.CVcs.AIRecentMay 27, 2026

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer

Rui Lin, Chuanming Wang, Huadong Ma

VidPrism introduces a novel heterogeneous Mixture-of-Experts framework that specializes temporal processing by dividing labor among experts, achieving state-of-the-art performance in image-to-video tr…

View →
cs.CVEmpiricalRecentJul 9, 2026

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin +3 more

This paper proposes LongE2V, a method for high-quality video recovery from sparse event streams using pre-trained video diffusion priors.

View →
cs.CVcs.AIRecentMay 31, 2026

Knowledge-Intensive Video Generation

Chenxu Wang, Mingda Chen

The paper introduces Knowledge-Intensive Video Generation (KIVI) as a challenging benchmark for evaluating video models on factuality and practical usefulness, showing that current state-of-the-art sy…

View →
cs.CVcs.AIcs.CLRecentJun 3, 2026

Continual Visual and Verbal Learning Through a Child's Egocentric Input

Xiaoyang Jiang, Yanlai Yang, Kenneth A. Norman, Brenden Lake +1 more

The paper introduces BabyCL, a continual multimodal learning framework that processes egocentric video data in a single chronological pass, demonstrating that meaningful word-referent mappings can be…

View →
cs.CVRecentJun 1, 2026

Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning

Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai +1 more

The paper proposes a training-free framework, Visual Representation-Guided Video-LLM Reasoning, to perform composed video retrieval by using visual examples and text instructions, achieving strong per…

View →
cs.CVEmpiricalRecentJul 17, 2026

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction

Homanga Bharadhwaj, Yash Jangir

This paper presents MotionForesight, a method for predicting future 3D trajectories of objects in human-object interaction videos using existing video prediction models.

View →
cs.CVEmpiricalRecentJul 9, 2026

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He +5 more

This paper proposes OPSD-V, an on-policy self-distillation method for reducing long-horizon degradation in few-step autoregressive video diffusion models by introducing real long-video data as tempora…

View →
cs.CVcs.AIRecentMay 28, 2026

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral, Adil Kaan Akan +3 more

VideoMLA introduces a novel Multi-Head Latent Attention (MLA) mechanism that replaces per-head KV caches with a shared low-rank content latent, significantly reducing memory and improving throughput f…

View →
cs.CVcs.AIRecentMay 28, 2026

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

Yinsong Xu, Wei Jing, Liuxin Zhang, Wanjun Lv +1 more

The paper proposes a unified framework that decouples long-video reasoning into semantic and visual evidence, significantly improving performance on the HD-EPIC VQA Challenge.

View →
cs.CVcs.AIcs.LGRecentJun 1, 2026

Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection

Boyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang +2 more

The paper proposes an Understanding-Enhanced Model Collaboration Method (UE-MCM) to accurately detect subtle and rare mistakes in egocentric videos by combining coarse-grained workflow understanding w…

View →
cs.CLRecentMay 29, 2026

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

Yeil Jeong, Youngjin Yoo, Seobin Sohn, Hyejin Han +3 more

The paper introduces TeachObs, a comprehensive, human-validated benchmark for multimodal teaching observation, and evaluates frontier LLMs, finding that no single model consistently outperforms others…

View →
cs.CVcs.CLEmpiricalRecentJul 22, 2026

Test-Time Training for Modality Order Consistency in Vision-Language Models

Aditi Gupta, Yossi Gandelsman

This paper identifies modality-order sensitivity as a failure in vision-language models and introduces a test-time training method to mitigate it, resulting in improved performance.

View →