~ similar to 2607.08770· 17 results
The paper introduces LongVQUBench, a comprehensive benchmark for long-term video quality understanding with 1200 diverse videos and 1500 questions.
Qixin Hu, Shuai Yang, Wei Huang, Song Han +1 more
LongLive-RAG proposes a novel Retrieval-Augmented Generation (RAG) framework to stabilize and improve the quality of long-horizon video generation by treating the entire generated history as a searcha…
Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He +5 more
This paper proposes OPSD-V, an on-policy self-distillation method for reducing long-horizon degradation in few-step autoregressive video diffusion models by introducing real long-video data as tempora…
Ruotong Liao, Guowen Huang, Qing Cheng, Guangyao Zhai +5 more
TunerDiT introduces a training-free progressive steering method to enhance multi-event video generation using Diffusion Transformers, achieving state-of-the-art performance by explicitly managing even…
Haoxuan Wu, Lai Man Po, Mengyang Liu, Kun Li +2 more
The paper introduces PRISM, a method for decoding preference signals from noisy latents using a lightweight Query-based Aggregation head and a frozen video diffusion backbone, achieving state-of-the-a…
VideoMLA introduces a novel Multi-Head Latent Attention (MLA) mechanism that replaces per-head KV caches with a shared low-rank content latent, significantly reducing memory and improving throughput f…
Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat +2 more
This paper introduces Flex-Forcing, a framework for video generation that enables a model to operate under both bidirectional and autoregressive generation regimes, achieving better video quality and…
Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza +7 more
The paper proposes a method for online novel view synthesis from multi-view streaming videos, decoupling memory update and application frequencies, using cross-view attention, Memory Loss, and Memory…
This paper introduces Align4D, a framework for generating coherent video-3D pairs using any-modal input, achieving state-of-the-art quality and consistency in X-to-4D generation.
Jiayi Wu, Haoming Cai, Cornelia Fermuller, Christopher Metzler +1 more
Real2SAM2Real introduces a framework that uses explicit 3D caches, derived from 3D lifting models, to provide robust geometric guidance to Video Diffusion Models, significantly improving spatiotempora…
The paper proposes VISTA, a multi-level event semantics mining framework, to accurately predict complex events in long videos, addressing the limitations of current LLMs in this domain.
Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu +5 more
SANA-Streaming introduces a novel, efficient framework that enables real-time, high-resolution streaming video-to-video editing by combining a hybrid diffusion transformer with specialized training an…
Minseok Joo, Dogyun Park, Taehoon Lee, Kyujin Lee +1 more
The paper proposes COVRAG, a depth-based memory retrieval framework that maximizes the coverage of target-view regions to significantly improve long-term geometric consistency in autoregressive long v…
This paper introduces ElasticTTT, a framework to preserve the generative prior in Test-Time Tuning (TTT) of pretrained diffusion models for video editing, preventing Prior Collapse.
Xuanyi Liu, Deyi Ji, Liqun Liu, Lanyun Zhu +7 more
CamGeo is a novel framework that improves sparse camera-conditioned image-to-video generation by distilling rich 3D geometric priors into the diffusion backbone, resulting in geometrically consistent…
Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang +8 more
The paper introduces Moment-Video, a new benchmark that diagnoses the ability of video MLLMs to understand brief, critical visual events, revealing that current models struggle significantly with temp…
RayDer introduces a unified, feed-forward transformer that simplifies self-supervised novel view synthesis (NVS) by consolidating camera estimation, scene reconstruction, and rendering into a single,…