~ similar to 2607.08545· 11 results
This paper introduces a sinusoidal recurrence mechanism for harmonic spectral enrichment in implicit neural representations, validated against various models and tasks.
Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more
The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mul…
Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu +2 more
This paper proposes RPPNet, a two-stage deep learning architecture for music generation with variable structural boundaries, which automatically derives grouping of Rhythm-Pitch Primitive sequences fr…
The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…
A deep neural timbre trait predictor is introduced to evaluate neural audio synthesizers' performance using human judgments and correlate with average human ratings.
Shengkui Zhao, Zexu Pan, Haoxu Wang, Biao Tian +2 more
This paper proposes TF-MossFormer, a time-frequency transformer for monaural speech separation that combines local and global attention using a content-aware sliding-window mechanism.
Zi Hu, Houmin Sun, Linxi Li, Yechen Wang +3 more
This paper proposes a neural audio watermarking method that embeds a message into the continuous latent representation of a codec-like speech autoencoder for improved codec robustness.
This paper explains how Rotary Position Embeddings (RoPE) frequencies in transformer models correspond to the relative-distance structure of training data, and the implications for long-context genera…