18 results for βlong-form audioβ
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.β
Want pure semantic search? Try claim verification β
JenBridge is a novel, adaptive framework that generates high-fidelity, long-form video soundtracks, significantly improving narrative coherence and naturalness across scene transitions.
This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,β¦
This paper presents ADAC, a compiler that converts trained differentiable audio models into efficient FAUST code for real-time audio effects.
Murmur is an efficient inference system for long-form ASR that resolves the accuracy-latency trade-off by optimizing both inter-chunk processing and intra-chunk attention mechanisms.
Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu +21 more
MOSS-Audio is a unified audio-language model designed for comprehensive understanding of speech, environmental sounds, and music, achieving strong performance across various audio-grounded tasks.
Zhisheng Zhang, Xiang Li, Yixuan Zhou, Jing Peng +2 more
LoSATok proposes a low-dimensional semantic-acoustic tokenizer that efficiently compresses high-dimensional audio features into a compact latent space, significantly improving the performance and effiβ¦
Yuyue Wang, Xihua Wang, Xin Cheng, Yijing Chen +1 more
The paper introduces PlanAudio, a unified LLM-based framework that directly synthesizes natural, composite audio containing speech and sounds from unconstrained free-form text prompts, outperforming eβ¦
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu +1 more
This paper investigates timestamp drift in modern autoregressive ASR systems and proposes REDDIT, a two-stage post-training framework to correct timestamps while avoiding forgetting.
Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi +2 more
This paper introduces RealDESED, a new benchmark for domestic sound event detection with 5,710 recordings, precise annotations, and multi-annotator labeling.
The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referencβ¦
The paper introduces Dialogs, a new Russian conversational speech corpus with high-quality recordings, segmented utterances, and expressive prosody labels.
Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more
The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mulβ¦