Zhizheng Wu
3 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geometry and multi-view acoustic context.
This paper introduces Speaker Anonymization (SA) as a novel perturbation mechanism for zero-shot voice conversion, balancing timbre leakage and prosodic utility while enabling strictly causal, zero-lookahead networks.
This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.
Papers
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more
This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.