ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.08545· 11 results

cs.CVEmpiricalRecentJul 23, 2026

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

Hyunmin Cho, Jaejun Yoo, Kyong Hwan Jin

This paper introduces a sinusoidal recurrence mechanism for harmonic spectral enrichment in implicit neural representations, validated against various models and tasks.

View →
eess.ASEmpiricalRecentJul 19, 2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more

The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mul…

View →
cs.SDcs.AIEmpiricalRecentJul 22, 2026

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu +2 more

This paper proposes RPPNet, a two-stage deep learning architecture for music generation with variable structural boundaries, which automatically derives grouping of Rhythm-Pitch Primitive sequences fr…

View →
eess.ASEmpiricalRecentJul 18, 2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker +2 more

The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…

View →
cs.SDEmpiricalRecentJun 29, 2026

Predicting Timbre Traits for Interpretable Assessment of Musical Sound Synthesizers

Théo Chasle Cauchy, Modan Tailleur, Lindsey Reymore, Fanny Roche +1 more

A deep neural timbre trait predictor is introduced to evaluate neural audio synthesizers' performance using human judgments and correlate with average human ratings.

View →
cs.SDEmpiricalRecentJul 23, 2026

TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation

Shengkui Zhao, Zexu Pan, Haoxu Wang, Biao Tian +2 more

This paper proposes TF-MossFormer, a time-frequency transformer for monaural speech separation that combines local and global attention using a content-aware sliding-window mechanism.

View →
cs.SDEmpiricalRecentJul 23, 2026

Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness

Zi Hu, Houmin Sun, Linxi Li, Yechen Wang +3 more

This paper proposes a neural audio watermarking method that embeds a message into the continuous latent representation of a codec-like speech autoencoder for improved codec robustness.

View →
cs.LGEmpiricalRecentJul 8, 2026

How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization

Xinyi Wu, Siyuan Liu, Ali Jadbabaie

This paper explains how Rotary Position Embeddings (RoPE) frequencies in transformer models correspond to the relative-distance structure of training data, and the implications for long-context genera…

View →