ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2607.19751· 15 results

cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →
cs.CLcs.AIEmpiricalRecentJul 8, 2026

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Jordan Painter, Dipankar Srirag, Adarsh Kappiyath, Diptesh Kanojia +2 more

The paper introduces DiaLLM, a method for continually pretraining language models on the International Corpus of English to generate dialectal English, and compares the effectiveness of different alig…

View →
cs.CLcs.CYcs.HCRecentJun 1, 2026

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

Victor Tolulope Olufemi, Oreoluwa Babatunde, Ramsey Njema, Bolarinwa Gbotemi +27 more

This paper demonstrates that compact, domain-specialized Automatic Speech Recognition (ASR) models significantly outperform large, general-purpose foundation models for conversational speech across 19…

View →
cs.CLRecentJun 1, 2026

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

Sarah Almeida Carneiro, Christos Xypolopoulos, Xiao Fei, Yang Zhang +1 more

The paper introduces CARTE, a new benchmark designed to test how well large language models understand fine-grained, regionally differentiated knowledge across the 13 metropolitan regions of France, r…

View →
eess.ASEmpiricalRecentJul 25, 2026

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

Ivan Kukanov, Zheng Xin Chai

This paper fine-tunes two zero-shot text-to-speech models, Chatterbox and CosyVoice 3, on Singlish speakers from the IMDA National Speech Corpus to improve accent similarity and naturalness.

View →
cs.CLcs.SDeess.ASEmpiricalRecentJun 18, 2026

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury

A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.

View →
eess.AScs.CLcs.LGEmpiricalRecentJun 18, 2026

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors

Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu

The paper proposes Pitch-Accent-focused Speech Quality Assessment (PASQA) to explicitly target pitch-accent correctness in speech quality assessment, using a controlled Japanese accent-error dataset a…

View →
eess.AScs.CLRecentMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…

View →
cs.LGcs.AIcs.SDEmpiricalRecentJun 28, 2026

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

Chuxiao Zuo, Yao Zhu, Minqiang Xu, Manhong Wang +2 more

A multimodal speaker identification system is proposed using Adaptive Modality Routing (AMR) for the POLY-SIM 2026 Grand Challenge, achieving high accuracy in various conditions.

View →
eess.AScs.CLcs.SDRecentMay 30, 2026

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

Xinwei Cao, Mengxuan Lu, Torbjørn Svendsen, Giampiero Salvi

The paper proposes a Lagrangian sub-flow (LSF) framework and geometric diagnostic signals to improve out-of-distribution detection using Continuous Normalizing Flows, overcoming the likelihood paradox…

View →
eess.ASEmpiricalRecentJul 2, 2026

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

Jesuraj Bandekar, Prasanta Kumar Ghosh

This paper proposes a novel pretraining method for Acoustic-to-Articulatory Inversion (AAI) using Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels, improving performance an…

View →
cs.SDcs.AIEmpiricalRecentJul 9, 2026

PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction

Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo +2 more

The paper introduces PS4, a framework for training target speaker extraction models using a large-scale corpus and proxy-supervised joint training strategy.

View →
eess.ASEmpiricalRecentJun 16, 2026

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more

This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…

View →
cs.NEEmpiricalRecentJul 16, 2026

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models

Ryona Noda

This paper proposes a method for quantizing large language models with cross-layer error compensation and finite-sample feature-statistics matching.

View →