ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “diarization”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SDcs.AIEmpiricalRecentJul 21, 2026

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar

Caption Studio is a transparency-first speech and audio intelligence platform that provides automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle ge…

View →
eess.ASEmpiricalRecentJun 12, 2026

Who Spoke When in Multi-Conversation: Target Speaker Tagging Task and Benchmark

Minjae Lee, Hee-Soo Heo, Youngki Kwon, Han-Gyu Kim +2 more

The paper introduces Target Speaker Tagging (TST), a task that combines speaker diarization, verification, and identification into a single workflow for multi-speaker conversations. It presents TST-Be…

View →
cs.SDcs.AIeess.ASRecentJun 1, 2026

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

Louis Mouchon

Echo is a joint-embedding predictive architecture that uses a single, pretrained ViT encoder to simultaneously perform speaker diarization, speech recognition, and dynamic source separation in a share…

View →
eess.AScs.CLRecentMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

The paper proposes a novel workflow to extract fine-grained regional accent features in Brazilian Portuguese using only acoustic labels and a phoneme-based forced aligner, showing that localized featu…

View →
cs.CLRecentJun 1, 2026

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

Baris Karacan, Vaibhav Bhargava, Barbara Di Eugenio, Natalie Parde +20 more

The paper introduces a supervised fine-tuning pipeline using large language models to accurately categorize sentence-level clinical provenance across multi-disciplinary hospital notes, demonstrating t…

View →
cs.CRcs.CLRecentApr 28, 2026

A Quantitative Confirmation of the Currier Language Distinction

Christophe Parisel

The paper quantitatively confirms the Currier A/B language distinction in the Voynich Manuscript, demonstrating it is governed by a higher-dimensional, context-dependent boolean switch rather than a s…

View →
cs.CLcs.AIEmpiricalRecentJul 23, 2026

Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it

Federico Boggia

This paper measures and analyzes the overuse of epanorthosis, a rhetorical figure, in large language models and proposes techniques to mitigate it.

View →
cs.CRcs.CLcs.IRRecentApr 11, 2026

Hijacking Text Heritage: Hiding the Human Signature through Homoglyphic Substitution

Robert Dilworth

This paper proposes using homoglyphic substitution, replacing characters with visually similar alternatives, as a method to degrade and prevent the extraction of personal information via adversarial s…

View →
cs.CLEmpiricalRecentJul 16, 2026

Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry

Paul-Andrei Pogăcean, Sanda-Maria Avram

This paper proposes a method for language identification using compositional vectors and the centered log-ratio transformation, achieving robust accuracy and strong performance for longer sequences.

View →
cs.CLcs.AIcs.LGRecentMay 28, 2026

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

Yaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang +4 more

The paper introduces LLMSurgeon, a framework that estimates the domain-level data mixture of a Large Language Model (LLM) using only generated text, thereby providing a post-hoc method to audit the mo…

View →
cs.CRcs.CLcs.LGRecentMay 28, 2026

Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content

Bing Liu, Shunping Wang, Yufan Zhu, Xinyi Yu +4 more

This paper introduces 'implicit identity' as a unifying framework to survey and categorize LLM fingerprinting and watermarking techniques for verifying ownership and provenance across datasets, models…

View →
cs.LGcs.AIRecentMay 31, 2026

A Fiber Criterion for Representation Identifiability in Supervised Learning

Vasileios Sevetlidis

The paper formalizes the problem of representation identifiability in supervised learning, showing that a representation property is identifiable if and only if it is constant across all possible fact…

View →
cs.CLcs.AIcs.CYEmpiricalRecentJul 9, 2026

Validity of LLMs as data annotators: AMALIA on authority

Manuel Pita

The study tests the validity of Portugal's AMALIA language model by examining its ability to follow a construct's theory and not just rely on surface correlates.

View →
stat.MLcs.LGEmpiricalRecentJul 21, 2026

Algebraic Signatures for Structural Learning in Probability Tensors

Akihiro Maeda, Shohei Hidaka, Satoshi Aoki

This paper presents a method for identifying probabilistic structures from empirical probability tensors using algebraic statistics and Kronecker-stack class of configuration matrices.

View →
cs.CLcs.AIcs.CYEmpiricalRecentJul 22, 2026

Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction

Mohamed Aziz Khadraoui, Adel Ammar, Bilel Benjdira, Zahid Khan +2 more

A regression-based approach is presented for Arabic dialect geolocation using a hierarchical neural architecture and spherical geodesic loss, achieving a median localization error of 481.2 km.

View →
cs.CLEmpiricalRecentJul 7, 2026

On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?

Ramon Ferrer-i-Cancho, Catherine Hobaiter, Thore Bergman, Morgan Gustison

The paper applies network science to demonstrate the feasibility of unsupervised dependency parsing in non-human primates due to their sequence length distribution.

View →
cs.CLcs.AIcs.SDEmpiricalRecentJun 12, 2026

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter

This paper uses continual learning with explicit disfluency tokens to improve Automatic Speech Recognition (ASR) systems on disfluent speech, addressing the information loss and hallucinations caused…

View →
cs.CLcs.LGRecentMay 30, 2026

French parsing enhanced with a word clustering method based on a syntactic lexicon

Anthony Sigogne, Matthieu Constant, Eric Laporte

The paper enhances French parsing accuracy by integrating data from a syntactic lexicon and applying word clustering methods to verbs within a Probabilistic Context-Free Grammar framework.

View →