ArXivCSExplorer
β˜†β˜†BookmarksπŸ†RSSHow to UseFAQ
Built with and by Teycir Ben Soltaneβ€’
How to Useβ€’FAQβ€’GitHubβ€’arXiv.orgβ€’
Share:

18 results for β€œlong-form audio”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.β“˜

Want pure semantic search? Try claim verification β†’

cs.SDcs.AIcs.CVRecentJun 1, 2026

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

Jiashuo Yu, Yao Yao, Boyu Chen, Alex Wang

JenBridge is a novel, adaptive framework that generates high-fidelity, long-form video soundtracks, significantly improving narrative coherence and naturalness across scene transitions.

View β†’
eess.AScs.CVEmpiricalRecentJul 16, 2026

WanSong v1.0 Technical Report

Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…

View β†’
eess.AScs.PLcs.SDEmpiricalRecentJun 19, 2026

Compiling Differentiable Audio Graphs to Real-Time DSP

Facundo Franchino, Sebastian J. Schlecht

This paper presents ADAC, a compiler that converts trained differentiable audio models into efficient FAUST code for real-time audio effects.

View β†’
cs.LGcs.AIeess.ASRecentMay 31, 2026

MURMUR: An Efficient Inference System for Long-Form ASR

Wei-Tzu Lee, Keisuke Kamahori, Baris Kasikci

Murmur is an efficient inference system for long-form ASR that resolves the accuracy-latency trade-off by optimizing both inter-chunk processing and intra-chunk attention mechanisms.

View β†’
cs.SDcs.AIRecentJun 1, 2026

MOSS-Audio Technical Report

Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more

MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.

View β†’
eess.AScs.AIcs.SDRecentMay 27, 2026

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

Zhisheng Zhang, Xiang Li, Yixuan Zhou, Jing Peng +2 more

LoSATok proposes a low-dimensional semantic-acoustic tokenizer that efficiently compresses high-dimensional audio features into a compact latent space, significantly improving the performance and effi…

View β†’
cs.SDcs.AIcs.MMRecentMay 27, 2026

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

Yuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen +1 more

The paper introduces PlanAudio, a unified LLM-based framework that directly synthesizes natural, composite audio containing speech and sounds from unconstrained free-form text prompts, outperforming e…

View β†’
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View β†’
cs.CLcs.AIcs.SDEmpiricalRecentJul 6, 2026

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu +1 more

This paper investigates timestamp drift in modern autoregressive ASR systems and proposes REDDIT, a two-stage post-training framework to correct timestamps while avoiding forgetting.

View β†’
eess.AScs.AIcs.SDDatasetRecentJul 18, 2026

RealDESED: A Real-World Domestic Sound Event Detection Benchmark

Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi +2 more

This paper introduces RealDESED, a new benchmark for domestic sound event detection with 5,710 recordings, precise annotations, and multi-annotator labeling.

View β†’
eess.ASEmpiricalRecentJul 18, 2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker +2 more

The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…

View β†’
eess.ASDatasetRecentJul 15, 2026

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

Ilya Shigabeev, Ilya Latyshev

The paper introduces Dialogs, a new Russian conversational speech corpus with high-quality recordings, segmented utterances, and expressive prosody labels.

View β†’
eess.ASEmpiricalRecentJul 19, 2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more

The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mul…

View β†’