Automatic Audio Equalization with Semantic Embeddings
This paper proposes a data-driven method for automatic blind audio equalization using a deep neural network and semantic embeddings.
Proposes a data-driven approach for blind audio equalization using a pre-trained deep neural network
Before reading this…
Applications
- →Real-world audio enhancement applications
To understand this paper, make sure you know these concepts first:
- Understanding of deep learning conceptsfind papers →
- Familiarity with audio processingfind papers →
Abstract
More Like ThisThis paper presents a data-driven approach to automatic blind equalization of audio by predicting log-mel spectral features and deriving an inverse filter. The method uses a deep neural network, where a pre-trained model provides semantic embeddings as a backbone, and only a lightweight head is trained. This design is intended to enhance training efficiency and generalization. Trained on both music and speech, the model is robust to noise and reverberation. Objective evaluations confirm its effectiveness, and subjective tests show performance comparable to that of an oracle that uses true log-mel spectral features, indicating that the model accurately estimates the desired characteristics, with remaining limitations attributed to the filtering stage. Overall, the results highlight the potential of the method for real-world audio enhancement applications.