20 results for “bird vocalizations”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Yuzhu Wang, Kalle Lahtinen, Patrik Lauha, Shiqi Zhang +3 more
This paper proposes an ensemble of two source separators, FTRNN and TF-Locoformer, trained with mixture invariant training (MixIT), and introduces mixture-constrained max pooling (MCM) to improve bird…
Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more
The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.
This paper extends BirdCallNet for joint species and call-type classification on the long-tailed WiWa dataset and investigates the interaction of task-loss balancing with pretrained representations an…
The paper introduces Dialogs, a new Russian conversational speech corpus with high-quality recordings, segmented utterances, and expressive prosody labels.
Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more
The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.
This paper proposes a framework for measuring stadium noise levels with spatially distributed acoustic measurements, criticizing the lack of standardization in current record-breaking claims.
CARE-DPP is a batch active-learning method for eco-acoustic monitoring that combines predictive uncertainty and embedding-space novelty with a determinantal point process objective.
Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li +12 more
This paper introduces a unified framework for generating high-quality full-length music from lyrics, text descriptions, and musical attributes, consisting of a semantic-aware tokenizer, hybird-LM, Ful…
This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…
This paper introduces DG^VoiC, a voice clustering framework for customer verification and speaker linking in call-centre audio using sensitive information-aligned anonymisation, speech-focused preproc…
ImmersiveTTS is an environment-aware text-to-speech model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions, achieving s…
The paper introduces SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models using a multidimensional protocol.
A Wave-U-Net model is trained to extract a fundamental waveform from input speech signals for accurate and robust instantaneous pitch estimation.