Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Zheng Wu

Zheng Wu

6 indexed papers

Recent (6 mo)
6
With code
0
Influential cites
0
Benchmarked
0

Publications per year

6
26

Top categories

Sound×3AI×2Robotics×1HCI×1Signal Processing×1NLP×1Multimedia×1Crypto×1

Frequent co-authors

Zhizheng Wu3×
Rongshen He1×
Xinyu Liang1×
Dekun Chen1×
Jiaqi Li1×
Mingjie Chen1×

Research Timeline

2026
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis

This paper provides the first comprehensive security analysis of the Agent Skills framework, identifying severe structural vulnerabilities that require fundamental architectural changes rather than simple mitigations.

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geometry and multi-view acoustic context.

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

The paper introduces MineExplorer, a new benchmark in Minecraft, to evaluate the sustained open-world exploration capabilities of MLLM agents, finding that long-horizon coordination remains a significant challenge.

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

This paper introduces Speaker Anonymization (SA) as a novel perturbation mechanism for zero-shot voice conversion, balancing timbre leakage and prosodic utility while enabling strictly causal, zero-lookahead networks.

How defensive driving enhances driving safety: A driving simulator study on drivers' defensive driving behaviors

This study investigates defensive driving behaviors, their impact on driving safety, and underlying mechanisms through driving simulator experiments.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.

Highlighted terms show continued research focus across papers

Papers

cs.SDEmpiricalRecentJul 22, 2026

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li +2 more

This paper introduces a training recipe for sentence-level and long-form streaming speech-to-speech translation using only 2k hours of paired cross-lingual data and auxiliary supervision.

View →
cs.ROcs.HCeess.SPEmpirical
Recent
Jul 21, 2026

How defensive driving enhances driving safety: A driving simulator study on drivers' defensive driving behaviors

Xinzheng Wu, Junyi Chen, Shaolingfeng Ye, Yong Shen

This study investigates defensive driving behaviors, their impact on driving safety, and underlying mechanisms through driving simulator experiments.

View →
cs.SDEmpiricalRecentJun 18, 2026

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

Yudong Li, Zihao Fang, Junwen Qiu, Ruihai Jing +3 more

This paper introduces Speaker Anonymization (SA) as a novel perturbation mechanism for zero-shot voice conversion, balancing timbre leakage and prosodic utility while enabling strictly causal, zero-lo…

View →
cs.CLRecentMay 29, 2026

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

Tianjie Ju, Yueqing Sun, Zheng Wu, Wei Zhang +6 more

The paper introduces MineExplorer, a new benchmark in Minecraft, to evaluate the sustained open-world exploration capabilities of MLLM agents, finding that long-horizon coordination remains a signific…

View →
cs.SDcs.AIcs.MMRecentMay 27, 2026

EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

Chong Jing, Zitong Lan, Junan Zhang, Zhizheng Wu

EigeNet introduces a geometry-informed multi-modal Transformer framework to achieve state-of-the-art few-shot novel view Room Impulse Response (RIR) prediction by effectively integrating spatial geome…

View →
cs.CRcs.AIRecentApr 3, 2026

Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis

Zhiyuan Li, Jingzheng Wu, Xiang Ling, Xing Cui +1 more

This paper provides the first comprehensive security analysis of the Agent Skills framework, identifying severe structural vulnerabilities that require fundamental architectural changes rather than si…

View →