20 results for βProsodic cuesβ
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.β
Want pure semantic search? Try claim verification β
The paper proposes methods for generating global prosodic embeddings using auto-encoder models of pitch and energy, demonstrating competitive or superior performance under challenging conditions.
Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang +2 more
This paper introduces MMEE, a multilingual and multi-emotion corpus for emphasis detection, and evaluates two state-of-the-art models under various settings.
This paper investigates the ability of mean opinion score (MOS) prediction models to capture quality differences in text-to-speech beyond acoustic fidelity through controlled perturbations on speech.
This paper explores the effect of conversational timing properties on automatic speech recognition (ASR) systems by controlling and optimizing pause and overlap timing distributions.
A lightweight framework for automated pronunciation assessment using native speech resources and unsupervised or lightly calibrated methods.
This paper investigates if upper-face affective cues enhance audiovisual sentence recognition, especially when audio is degraded, finding that while mouth cues are crucial for robustness, upper-face cβ¦
SooHwan Eom, Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +1 more
The paper proposes RTFree-F5, a method to make flow-matching TTS models like F5-TTS independent of reference transcripts, improving performance and naturalness for dysarthric speakers.
This paper proposes a new training criterion to reduce a classifier's reliance on shortcuts in language proficiency assessment systems, improving their correlation with human references.
This paper examines the absence of phonological context compensation in wav2vec2.0 architecture using Mandarin Chinese tones, contrasting self-supervised pre-training with fine-tuning for ASR.
The paper introduces Dialogs, a new Russian conversational speech corpus with high-quality recordings, segmented utterances, and expressive prosody labels.
The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featuβ¦
This study demonstrates that the tone of a prompt significantly affects the accuracy of various LLMs, requiring users to exercise caution regarding tone-robust reliability.
This paper compares hard-label and distributional objectives for speech emotion recognition using a WavLM-Base multitask model and evaluates their alignment with human vote distributions.
Jing Peng, Junhao Du, Chenghao Wang, Hanqi Li +20 more
The paper introduces SURE, a unified framework designed to standardize and improve the comparability and reproducibility of evaluations for advanced speech understanding models.
Yujian Ma, Jinqiu Sang, Ruizhe Li, Jiaao Yu +1 more
This paper examines how fine-tuning large audio-language models affects the semantics, decoder accessibility, and temporal output alignment of native audio-token states using temporal audio grounding.
This paper investigates how the final prompt in conversational AI-search evaluations differs from the conversation history, using two corpora of commercial and PRISM conversations.
Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle geβ¦