Ming Lu
6 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper proposes VerFU, a client-verifiable federated unlearning framework for low-altitude wireless networks that allows devices to ensure the server accurately removes their historical data contributions without revealing the original data.
The paper proposes RPM-Net, a novel framework using a reciprocal point mechanism and adversarial margin constraints to achieve superior detection of unknown network security threats in imbalanced multi-class environments.
This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy distillation and GRPO.
This paper proposes Perceive-to-Reason (P2R), a framework for fine-grained visual reasoning that decouples perception from reasoning and introduces a new reinforcement learning strategy.
This paper introduces LingBot-VA 2.0, a video-action foundation model designed for embodiment, with semantic visual-action tokenization, causal pretraining, sparse MoE backbone, and enhanced asynchronous inference.
The paper introduces Relay On-Policy Distillation (Relay-OPD), a method for on-policy distillation that constructs relay trajectories to address prefix failure and improve performance.
Papers
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni +4 more
The paper introduces Relay On-Policy Distillation (Relay-OPD), a method for on-policy distillation that constructs relay trajectories to address prefix failure and improve performance.