Jiang Liu
4 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces DiffErase, a black-box attack that effectively removes inaudible audio watermarks while preserving perceptual quality by utilizing diffusion models.
The paper proposes Predictive Routing Replay (PR2) to stabilize reinforcement learning on Mixture of Experts (MoE) LLMs by predicting and incorporating short-horizon router evolution during training and rollout.
The paper introduces Lookahead Group Reward (&) to combat Supervision Fidelity Decay (SFD) in on-policy distillation, significantly improving student model performance on long reasoning tasks.
This paper introduces LLM-as-a-Verifier, a framework for fine-grained verification of LLMs using continuous scores, achieving state-of-the-art performance on various benchmarks.
Papers
LLM-as-a-Verifier: A General-Purpose Verification Framework
Jacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu +5 more
This paper introduces LLM-as-a-Verifier, a framework for fine-grained verification of LLMs using continuous scores, achieving state-of-the-art performance on various benchmarks.