ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “speech classification”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDeess.ASEmpiricalRecentJul 26, 2026

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai, Tatsuya Kawahara

The study investigates the limitations of Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription using Grapheme-to-Phoneme (G2P) labels and proposes a new approach based on continuous Artic…

View →
eess.AScs.CLRecentMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…

View →
eess.AScs.AIcs.LGEmpiricalRecentJun 18, 2026

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

Paban Sapkota, Hemant Kumar Kathania, Mikko Kurimo, Sudarsana Reddy Kadiri +1 more

This paper investigates the use of various acoustic features for recognizing dysarthric speech using a Factorized Time Delay Neural Network (F-TDNN) model, achieving a relative improvement of 4.65% in…

View →
eess.AScs.AIcs.LGEmpiricalRecentJun 18, 2026

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

Rostislav Makarov, Timo Gerkmann

The paper repurposes a pre-trained speech classifier as the backbone for diffusion generation, reducing the need for two separately trained models.

View →
eess.AScs.AIcs.SDRecentMay 29, 2026

A Unified and Reproducible Experimentation Framework for Speech Understanding

Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li +20 more

The paper introduces SURE, a unified framework designed to standardize and improve the comparability and reproducibility of evaluations for advanced speech understanding models.

View →
eess.ASEmpiricalRecentJul 7, 2026

Few-Shot Class-Incremental Audio Classification Using Pseudo-Incrementally Trained Embedding Learner and Continually Updated Stochastic Classifier

Yanxiong Li, Wenchang Cao, Jiaxin Tan, Qianqian Li +1 more

This paper proposes a model for few-shot class-incremental audio classification, which consists of a pseudo-incrementally trained embedding learner and a continually updated stochastic classifier.

View →
eess.AScs.AIcs.CLEmpiricalRecentJul 10, 2026

Phone Segmentation and Recognition through Phonological Activation Mapping

Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li +7 more

The authors propose a method for phone segmentation and recognition using self-supervised speech models, requiring minimal phonetic transcriptions and generalizing to unseen phones.

View →
eess.ASEmpiricalRecentJun 16, 2026

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

Sujith Pulikodan, Agneedh Basu, Pavan Kumar, Pranav Bhat +3 more

This paper investigates the effectiveness of incorporating synthetic speech data in Automatic Speech Recognition (ASR) Systems for three Indic languages by analyzing performance gains, script sources,…

View →
cs.CLcs.AIeess.ASRecentMay 31, 2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu +3 more

PolySpeech-100 introduces a massive, multi-lingual benchmark covering 110 linguistic variants to rigorously test Speech-LLMs, demonstrating that open-source models struggle with low-resource languages…

View →
eess.ASEmpiricalRecentJun 27, 2026

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Yujie Tu, Yifan Yang, Tianrui Wang, Yanqiao Zhu +32 more

The paper introduces GigaSpeechBench, a comprehensive multilingual and multidimensional ASR & AST benchmark with 680 hours of human-annotated speech, featuring 12 low-resource languages, 6 Chinese dia…

View →
eess.ASEmpiricalRecentJun 18, 2026

Stuttering Classification and Segmentation with Attention-Based Multiple Instance Learning

Petar Sušac, Sebastian P. Bayerl, Hrvoje Džapo

The paper proposes a multiple instance neural network architecture using fine-tuned wav2vec 2.0, WavLM and Whisper encoders for stuttering detection and classification, achieving improvements in both…

View →
eess.AScs.SDEmpiricalRecentJul 7, 2026

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao +1 more

The paper proposes SpeechKV, a method to compress speech sequences inside large language models using a learned pooling, maintaining performance and delivering decoding speedup.

View →
eess.ASEmpiricalRecentJun 18, 2026

Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces

Kyle Janse van Rensburg, Herman Kamper

This paper investigates the correlation between dimensions of self-supervised speech features and speech characteristics, finding that content dimensions primarily capture intensity, formants, and voi…

View →
cs.SDcs.AIEmpiricalRecentJun 23, 2026

ZONOS2 Technical Report

Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close +1 more

The authors present ZONOS2 8B, a TTS model with improved naturalness, prosody, and voice cloning fidelity, achieved through scaling, data expansion, and simplification.

View →
eess.AScs.SDEmpiricalRecentJul 3, 2026

Mixture-Constrained Max Pooling Improves Separation-Based Bird Species Classification

Yuzhu Wang, Kalle Lahtinen, Patrik Lauha, Shiqi Zhang +3 more

This paper proposes an ensemble of two source separators, FTRNN and TF-Locoformer, trained with mixture invariant training (MixIT), and introduces mixture-constrained max pooling (MCM) to improve bird…

View →
eess.AScs.AIcs.SDEmpiricalRecentJul 9, 2026

On the Role of Conversational Timing in Synthetic Training Data for ASR

Máté Gedeon, Péter Mihajlik

This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.

View →
eess.AScs.SDNEWEmpiricalJul 28, 2026

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

Behrad TaghiBeyglou, Fatemeh Bagheri, Ervin Sejdic

This paper proposes a graph framework using pretrained SSL embeddings for speech analysis in Amyotrophic Lateral Sclerosis (ALS) patients, achieving better results than validation baselines on the SAN…

View →
eess.ASEmpiricalRecentJul 19, 2026

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang +9 more

The paper proposes SALMONN-2, an ALLM built on a unified SSL encoder, and presents a multi-layer feature fusion adapter to better exploit hierarchical SSL encoder representations. It also explores mul…

View →
cs.CLeess.ASEmpiricalRecentJul 3, 2026

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

Saurabh Kumar, Amartyaveer, Prasanta Kumar Ghosh

This paper proposes a multimodal framework for jointly improving Automatic Speech Recognition (ASR) and Dialect Identification (DID) in Indian languages using a Bottleneck Encoder, RoBERTa encoder, ga…

View →