ArXivCSExplorer
β˜†β˜†BookmarksπŸ†RSSHow to UseFAQ
Built with and by Teycir Ben Soltaneβ€’
How to Useβ€’FAQβ€’GitHubβ€’arXiv.orgβ€’
Share:

20 results for β€œpitch-accent”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.β“˜

Want pure semantic search? Try claim verification β†’

eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu

The paper proposes Pitch-Accent-focused Speech Quality Assessment (PASQA) to explicitly target pitch-accent correctness in speech quality assessment, using a controlled Japanese accent-error dataset a…

View β†’
eess.ASEmpiricalRecentJul 25, 2026

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

Ivan Kukanov, Zheng Xin Chai

This paper fine-tunes two zero-shot text-to-speech models, Chatterbox and CosyVoice 3, on Singlish speakers from the IMDA National Speech Corpus to improve accent similarity and naturalness.

View β†’
cs.SDEmpiricalRecentJun 12, 2026

Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement

Junya Koguchi, Tomoki Koriyama

A Wave-U-Net model is trained to extract a fundamental waveform from input speech signals for accurate and robust instantaneous pitch estimation.

View β†’
eess.AScs.CLRecentMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…

View β†’
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View β†’
cs.CLcs.AIeess.ASEmpiricalRecentJun 16, 2026

Perceptual compensation for tonal context in self-supervised speech models

James Kirby, Ioana Krehan, Michele Gubian

This paper examines the absence of phonological context compensation in wav2vec2.0 architecture using Mandarin Chinese tones, contrasting self-supervised pre-training with fine-tuning for ASR.

View β†’
eess.ASEmpiricalRecentJun 18, 2026

Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces

Kyle Janse van Rensburg, Herman Kamper

This paper investigates the correlation between dimensions of self-supervised speech features and speech characteristics, finding that content dimensions primarily capture intensity, formants, and voi…

View β†’
cs.AIcs.CLcs.HCRecentMay 27, 2026

Mind Your Tone: Does Tone Alter LLM Performance?

Om Dobariya, Akhil Kumar

This study demonstrates that the tone of a prompt significantly affects the accuracy of various LLMs, requiring users to exercise caution regarding tone-robust reliability.

View β†’
eess.ASEmpiricalRecentJun 12, 2026

Unsupervised Approaches for Global Prosodic Embedding Extraction

Martin Meza, Luciana Ferrer, Pablo Riera

The paper proposes methods for generating global prosodic embeddings using auto-encoder models of pitch and energy, demonstrating competitive or superior performance under challenging conditions.

View β†’
cs.CLcs.AIcs.LGEmpiricalRecentJun 26, 2026

Do Speech Emphasis Models Generalize across Languages and Emotions?

Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang +2 more

This paper introduces MMEE, a multilingual and multi-emotion corpus for emphasis detection, and evaluates two state-of-the-art models under various settings.

View β†’
eess.AScs.SDDatasetRecentJul 23, 2026

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

Seolhee Lee, Minsu Kang, Yangsun Lee, Woosun Min +2 more

The paper introduces the Designed Vocalizations Dataset for AI-based voice conversion research on non-human vocalizations and effects, providing a standardized test set and benchmark results.

View β†’
eess.ASEmpiricalRecentJun 18, 2026

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian

Ruchi Pandey, Tomi Kinnunen

This paper systematically evaluates acoustic-to-articulatory inversion under domain shifts on FROST-EMA, a Finnish-Russian bilingual EMA corpus, and establishes benchmarks for articulatory targets, ac…

View β†’
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

Masato Takagi, Masaya Kawamura, Reo Shimizu, Yuma Shirahata

This paper investigates the ability of mean opinion score (MOS) prediction models to capture quality differences in text-to-speech beyond acoustic fidelity through controlled perturbations on speech.

View β†’
eess.ASEmpiricalRecentJul 2, 2026

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

Jesuraj Bandekar, Prasanta Kumar Ghosh

This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…

View β†’
cs.CLcs.AIcs.CYEmpiricalRecentJul 22, 2026

Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction

Mohamed Aziz Khadraoui, Adel Ammar, Bilel Benjdira, Zahid Khan +2 more

A regression-based approach is presented for Arabic dialect geolocation using a hierarchical neural architecture and spherical geodesic loss, achieving a median localization error of 481.2 km.

View β†’
eess.AScs.SDEmpiricalRecentJun 29, 2026

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling

Yoonjeong Park, Jaekwon Im, Juhan Nam

This paper introduces MeloDISinger, a text-based singing voice editing model that preserves melody and duration while enabling melody-aware duration control.

View β†’
cs.SDEmpiricalRecentJul 8, 2026

MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

Wenhao Feng, Yuxun Tang, Jiatong Shi, Qin Jin

The paper introduces MMGenre, a benchmark for multi-genre singing voice synthesis diagnosis, revealing limited genre discrimination and proposing lightweight genre-specific continued training.

View β†’