SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
This paper introduces the REAL-TSE Challenge, a satellite challenge on target speaker extraction from real conversational recordings, and describes its task definition, datasets, evaluation protocol, and submitted systems.
First real-world TSE benchmark with natural overlap, reverberation, noise, channel mismatch, and conversational dynamics
Before reading this…
Applications
- →Speech recognition, Speaker diarization, Speech separation
To understand this paper, make sure you know these concepts first:
- Understanding of Speech and Language Technology, Target Speaker Extractionfind papers →
Abstract
More Like ThisWe introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.