ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “bird vocalizations”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

eess.AScs.SDEmpiricalRecentJul 3, 2026

Mixture-Constrained Max Pooling Improves Separation-Based Bird Species Classification

Yuzhu Wang, Kalle Lahtinen, Patrik Lauha, Shiqi Zhang +3 more

This paper proposes an ensemble of two source separators, FTRNN and TF-Locoformer, trained with mixture invariant training (MixIT), and introduces mixture-constrained max pooling (MCM) to improve bird…

View →
eess.AScs.SDDatasetRecentJul 23, 2026

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more

The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.

View →
cs.SDcs.LGeess.ASEmpiricalRecentJul 3, 2026

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

Paria Vali Zadeh, Sven Tomforde

This paper extends BirdCallNet for joint species and call-type classification on the long-tailed WiWa dataset and investigates the interaction of task-loss balancing with pretrained representations an…

View →
eess.ASDatasetRecentJul 15, 2026

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

Ilya Shigabeev, Ilya Latyshev

The paper introduces Dialogs, a new Russian conversational speech corpus with high-quality recordings, segmented utterances, and expressive prosody labels.

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
eess.ASEmpiricalRecentJul 21, 2026

Towards a reproducible cross-venue method for quantifying crowd noise in stadiums

Alejandro Osses, Bente Ackermans, Helmer Nuijens, Rick Scholte

This paper proposes a framework for measuring stadium noise levels with spatially distributed acoustic measurements, criticizing the lack of standardization in current record-breaking claims.

View →
cs.SDcs.LGEmpiricalRecentJul 7, 2026

Determinantal point process sampling for bioacoustic active learning

Hugo Magaldi, Gabriel Dubus

CARE-DPP is a batch active-learning method for eco-acoustic monitoring that combines predictive uncertainty and embedding-space novelty with a determinantal point process objective.

View →
cs.SDcs.AIeess.ASEmpiricalRecentJul 22, 2026

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li +12 more

This paper introduces a unified framework for generating high-quality full-length music from lyrics, text descriptions, and musical attributes, consisting of a semantic-aware tokenizer, hybird-LM, Ful…

View →
eess.ASEmpiricalRecentJul 2, 2026

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

Jesuraj Bandekar, Prasanta Kumar Ghosh

This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…

View →
cs.SDcs.AIcs.CLEmpiricalRecentJun 26, 2026

DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions

Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth +1 more

This paper introduces DG^VoiC, a voice clustering framework for customer verification and speaker linking in call-centre audio using sensitive information-aligned anonymisation, speech-focused preproc…

View →
eess.AScs.AIcs.CLRecentMay 29, 2026

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee

ImmersiveTTS is an environment-aware text-to-speech model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions, achieving s…

View →
cs.CLcs.AIeess.ASEmpiricalRecentJul 6, 2026

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar +4 more

The paper introduces SPEARBench, a benchmark for evaluating naturalness in speech-to-speech language models using a multidimensional protocol.

View →
cs.SDEmpiricalRecentJun 12, 2026

Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement

Junya Koguchi, Tomoki Koriyama

A Wave-U-Net model is trained to extract a fundamental waveform from input speech signals for accurate and robust instantaneous pitch estimation.

View →