ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “video pretraining”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVEmpiricalRecentJul 8, 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang +23 more

This paper introduces LingBot-Video, a video pretraining paradigm for embodied intelligence using a DiT-based approach, Mixture-of-Experts framework, and extensive robot-oriented data.

View →
cs.ROcs.CVEmpiricalRecentJul 9, 2026

Native Video-Action Pretraining for Generalizable Robot Control

Qihang Zhang, Lin Li, Luyao Zhang, Shuai Yang +25 more

This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchron…

View →
cs.CVcs.AIRecentMay 31, 2026

Knowledge-Intensive Video Generation

Chenxu Wang, Mingda Chen

The paper introduces Knowledge-Intensive Video Generation (KIVI) as a challenging benchmark for evaluating video models on factuality and practical usefulness, showing that current state-of-the-art sy…

View →
eess.AScs.CVEmpiricalRecentJul 17, 2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar, Lasha Koroshinadze +18 more

The paper introduces Audio-Visual Flamingo (AV-Flamingo), an open-source audio-visual large language model designed for understanding and reasoning over long and complex real-world audio-visual videos…

View →
cs.CVcs.AIEmpiricalRecentJul 23, 2026

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Yueyi Liu, Chi Zhang, Sen Cui, Miao Liu

This paper introduces ElasticTTT, a framework to preserve the generative prior in Test-Time Tuning (TTT) of pretrained diffusion models for video editing, preventing Prior Collapse.

View →
cs.CVEmpiricalRecentJul 9, 2026

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin +3 more

This paper proposes LongE2V, a method for high-quality video recovery from sparse event streams using pre-trained video diffusion priors.

View →
cs.CVcs.AIEmpiricalRecentJul 2, 2026

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya +5 more

The paper introduces AnyGroundBench, a domain-adaptation benchmark for Spatio-Temporal Video Grounding, to evaluate the ability of Vision-Language Models to adapt to specialized domains.

View →
cs.CVcs.AIEmpiricalRecentJul 1, 2026

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

Arpita Nema, Hanwei Zhu, Xi Zhang, Weisi Lin

The paper introduces LongVQUBench, a comprehensive benchmark for long-term video quality understanding with 1200 diverse videos and 1500 questions.

View →
cs.CLRecentMay 29, 2026

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

Yeil Jeong, Youngjin Yoo, Seobin Sohn, Hyejin Han +3 more

The paper introduces TeachObs, a comprehensive, human-validated benchmark for multimodal teaching observation, and evaluates frontier LLMs, finding that no single model consistently outperforms others…

View →
cs.CVRecentJun 1, 2026

Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning

Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai +1 more

The paper proposes a training-free framework, Visual Representation-Guided Video-LLM Reasoning, to perform composed video retrieval by using visual examples and text instructions, achieving strong per…

View →
cs.CVcs.AIcs.SDEmpiricalRecentJun 23, 2026

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li +2 more

This paper presents video-SALMONN-R$^3$, an end-to-end video-LLM that enables re-watch through reinforcement learning, improving question answering performance with a two-stage paradigm.

View →
cs.CVcs.AIRecentMay 27, 2026

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer

Rui Lin, Chuanming Wang, Huadong Ma

VidPrism introduces a novel heterogeneous Mixture-of-Experts framework that specializes temporal processing by dividing labor among experts, achieving state-of-the-art performance in image-to-video tr…

View →
cs.DCEmpiricalRecentJul 2, 2026

Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training

Peng Yu, Yuankai Fan, Yang Qiu, Tian Li +3 more

This paper introduces Arachne, a framework for efficient Text-to-Video model training at scale, reducing iteration time by up to 65% over leading frameworks.

View →
cs.CVcs.AIRecentMay 28, 2026

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

Yinsong Xu, Wei Jing, Liuxin Zhang, Wanjun Lv +1 more

The paper proposes a unified framework that decouples long-video reasoning into semantic and visual evidence, significantly improving performance on the HD-EPIC VQA Challenge.

View →
cs.CVcs.AIcs.CLRecentJun 3, 2026

Continual Visual and Verbal Learning Through a Child's Egocentric Input

Xiaoyang Jiang, Yanlai Yang, Kenneth A. Norman, Brenden Lake +1 more

The paper introduces BabyCL, a continual multimodal learning framework that processes egocentric video data in a single chronological pass, demonstrating that meaningful word-referent mappings can be…

View →
cs.CVEmpiricalRecentJul 17, 2026

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction

Homanga Bharadhwaj, Yash Jangir

This paper presents MotionForesight, a method for predicting future 3D trajectories of objects in human-object interaction videos using existing video prediction models.

View →
cs.CVEmpiricalRecentJul 17, 2026

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola +5 more

This paper proposes VideoTreeSearch (VTS), a framework for grounded long-video question answering that uses iterative self-correcting search over an adaptive temporal tree.

View →
cs.CVEmpiricalRecentJul 9, 2026

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He +5 more

This paper proposes OPSD-V, an on-policy self-distillation method for reducing long-horizon degradation in few-step autoregressive video diffusion models by introducing real long-video data as tempora…

View →
cs.AIEmpiricalRecentJun 23, 2026

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

Xinyu Mao, Yuhui Zeng, Xiaokun Liu, Wenyu Qin +5 more

This paper introduces CineCap, a framework for cinematographic captioning using structured reasoning, spatio-temporal anchors, and reinforcement learning.

View →