ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “long-form streaming”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CVcs.AIRecentMay 28, 2026

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu +5 more

SANA-Streaming introduces a novel, efficient framework that enables real-time, high-resolution streaming video-to-video editing by combining a hybrid diffusion transformer with specialized training an…

View →
cs.SDEmpiricalRecentJul 22, 2026

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more

This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.

View →
cs.CVRecentJun 1, 2026

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

Peiwen Sun, Xudong Lu, Huadai Liu, Yang Bo +8 more

The paper introduces X-Stream, a new benchmark for multi-stream video understanding, and finds that current state-of-the-art MLLMs perform poorly when required to process multiple concurrent video str…

View →
cs.CVRecentJun 1, 2026

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

Qixin Hu, Shuai Yang, Wei Huang, Song Han +1 more

LongLive-RAG proposes a novel Retrieval-Augmented Generation (RAG) framework to stabilize and improve the quality of long-horizon video generation by treating the entire generated history as a searcha…

View →
cs.LGcs.AIeess.ASRecentMay 31, 2026

MURMUR: An Efficient Inference System for Long-Form ASR

Wei-Tzu Lee, Keisuke Kamahori, Baris Kasikci

Murmur is an efficient inference system for long-form ASR that resolves the accuracy-latency trade-off by optimizing both inter-chunk processing and intra-chunk attention mechanisms.

View →
cs.CVEmpiricalRecentJul 9, 2026

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin +3 more

This paper proposes LongE2V, a method for high-quality video recovery from sparse event streams using pre-trained video diffusion priors.

View →
eess.AScs.CVEmpiricalRecentJul 16, 2026

WanSong v1.0 Technical Report

Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…

View →
cs.LGcs.AIcs.NITheoreticalRecentJul 27, 2026

Adaptive Data Admission and Retention for Streaming Federated Learning

Zhuoyi Zhao, Ben Liang

This paper proposes a framework for federated learning with limited client memory and time-varying sampling costs, and develops an Active-Constraint Drift-Plus-Penalty (ACDPP) policy to minimize the c…

View →
cs.CLcs.DSRecentMay 29, 2026

Incremental BPE Tokenization

Shenghu Jiang, Ruihao Gong

The paper introduces an efficient, novel algorithm for incremental Byte Pair Encoding (BPE) tokenization that processes input text prefix by prefix, achieving significant speedups and enabling streami…

View →
cs.SDcs.LGEmpiricalRecentJul 22, 2026

Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting

Mahesh Godavarti

This paper proposes cumsum-composable phase transport, a streaming-native temporal layer for keyword spotting using unitary transport, prefix differences, and gated residual updates.

View →
cs.CRcs.SERecentMar 18, 2026

Circumventing Platform Defenses at Scale: Automated Content Replication from YouTube to Blockchain-Based Decentralized Storage

Muhammad Zeeshan Akram

The paper introduces YouTube-Synch, a robust system that successfully replicates content from thousands of YouTube channels to decentralized storage by continuously adapting to and bypassing YouTube's…

View →
cs.DScs.CCTheoreticalRecentJul 16, 2026

Semi-Streaming Matching in a Single Pass II: Greedy is Optimal

Sepehr Assadi, Max Jiang, Mars Xiang

This paper proves that no single-pass semi-streaming algorithm can achieve a better-than-half approximation to the maximum matching problem, implying the optimality of the naive greedy algorithm.

View →
cs.AIcs.LGRecentMay 30, 2026

SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition

Jayanta Dey, Shikhar Srivastava, Itamar Lerner, Christopher Kanan +1 more

SHARP proposes a novel sleep-based hierarchical replay framework to efficiently learn long-range non-stationary temporal patterns in streaming data, achieving improved context retention and predictive…

View →
cs.DScs.CCTheoreticalRecentJul 16, 2026

Semi-Streaming Matching in a Single Pass I: A New Framework for Lower Bounds via Blueprints

Sepehr Assadi, Max Jiang, Mars Xiang

The paper develops a new framework for proving lower bounds for the maximum matching problem in the semi-streaming model, improving upon the previous best known bounds.

View →
cs.SDcs.AIcs.CVRecentJun 1, 2026

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

Jiashuo Yu, Yao Yao, Boyu Chen, Alex Wang

JenBridge is a novel, adaptive framework that generates high-fidelity, long-form video soundtracks, significantly improving narrative coherence and naturalness across scene transitions.

View →
cs.SDcs.AIcs.HCEmpiricalRecentJun 23, 2026

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

Baisen Wang, Chenxi Bao, Qisong Han

This paper proposes a framework for creating low-latency, interactive generative music AI using distillation in a streaming autoregressive latent space and music-aware consistency objectives.

View →