ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “music-aware consistency objectives”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDEmpiricalRecentJul 7, 2026

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music

Tomáš Sourada, Katia Vendrame, Jan Hajič

The paper introduces MusICA-MetaBench, a framework for deriving on-demand music perception benchmarks from user-provided data, ensuring statistically reliable model comparisons.

View →
cs.SDcs.IRcs.LGEmpiricalRecentJul 24, 2026

Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition

Austin Rockman

The paper introduces Reflector, an interactive audio workstation that adapts pitch-class retrieval as compositions evolve, using a learned embedding space based on a hand-designed oracle.

View →
cs.SDcs.AIcs.ETEmpiricalRecentJul 7, 2026

Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation

Josef Pavlicek

This paper introduces a maintainable hybrid architecture for generating harmonies from melodies using quantum-inspired exploration and rule-based optimization.

View →
cs.CLcs.AIcs.SDRecentMay 28, 2026

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo +5 more

The paper introduces MusTBENCH, a new benchmark, and MusT, an optimization recipe, to rigorously test and improve the ability of Large Audio-Language Models (LALMs) to accurately ground their musical…

View →
eess.AScs.SDEmpiricalRecentJun 29, 2026

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling

Yoonjeong Park, Jaekwon Im, Juhan Nam

This paper introduces MeloDISinger, a text-based singing voice editing model that preserves melody and duration while enabling melody-aware duration control.

View →
cs.SDcs.AIcs.HCEmpiricalRecentJun 23, 2026

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

Baisen Wang, Chenxi Bao, Qisong Han

This paper proposes a framework for creating low-latency, interactive generative music AI using distillation in a streaming autoregressive latent space and music-aware consistency objectives.

View →
cs.IRcs.AIcs.LGRecentMay 28, 2026

Multimodal Music Recommendation System using LLMs

Srikar Prabhas Kandagatla, Sreehitha R. Narayana, Chandana Magapu, Swetha Mohan +5 more

The paper proposes a novel multimodal framework for session-based music recommendation that jointly models audio, lyric, and semantic content signals within a unified LLM-based sequential reasoning sy…

View →
cs.SDcs.PFEmpiricalRecentJul 9, 2026

A Quantized Native Runtime for On-Device Semantic Audio Generation

Matteo Spanio, Antonio Rodà

The paper presents aria, a dependency-free runtime for generating text-to-music using Stable Audio 3 on commodity hardware, with a focus on quantization for memory savings and activation steering.

View →
eess.AScs.CVEmpiricalRecentJul 16, 2026

WanSong v1.0 Technical Report

Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

This paper introduces WanSong, a diffusion-based model for long-form, commercial-grade song generation that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems,…

View →
cs.SDcs.AIEmpiricalRecentJul 22, 2026

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu +2 more

This paper proposes RPPNet, a two-stage deep learning architecture for music generation with variable structural boundaries, which automatically derives grouping of Rhythm-Pitch Primitive sequences fr…

View →
cs.SDcs.AIDatasetRecentJul 8, 2026

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu +7 more

The paper introduces MADB, a large-scale dataset and benchmark for music aesthetic assessment with 9,999 tracks annotated by 30 trained annotators across 10 perceptual dimensions.

View →
cs.SDEmpiricalRecentJul 24, 2026

CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following

Yining Yang, Ruogu Chen, Jie Han

This paper introduces CODA, a real-time score following system that exploits the cascaded structure of music scores for prediction consistency and enables recovery from score discontinuities.

View →
cs.SDEmpiricalRecentJul 24, 2026

Music-JEPA: Learning a World Model of Sound from Action

Ziyu Wang, Kun Fang, Yann LeCun

This paper proposes a method for learning a world model of piano sound using Joint Embedding Predictive Architectures (JEPA), treating music as an action-conditioned system.

View →
cs.SDmath.COmath.HOTheoreticalRecentJul 25, 2026

Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions

Clifton Callender

This paper presents the structure of infinite melodic lines with self-similarity under all rational and irrational tempo ratios, allowing for infinite solutions in table canons.

View →
cs.DScs.LGstat.MLRecentJun 3, 2026

A General Framework for Dynamic Consistent Submodular Maximization

Paul Dütting, Federico Fusco, Silvio Lattanzi, Ashkan Norouzi-Fard +2 more

The paper develops a general framework for dynamic consistent submodular maximization, achieving constant-factor approximations with sublinear consistency for both cardinality and rank-$k$ matroid con…

View →
eess.ASEmpiricalRecentJul 18, 2026

NABEATs: Noise-Aware Audio Representation Learning

Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker +2 more

The paper introduces Noise-Aware BEATs (NABEATs), a noise-aware audio self-supervised learning framework that estimates clean BEATs representations from noisy audio signals using an auxiliary referenc…

View →