Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Jiang Liu

Jiang Liu

4 indexed papers

Recent (6 mo)
4
With code
0
Influential cites
0
Benchmarked
0

Publications per year

4
26

Top categories

AI×3NLP×2ML×2Multiagent×1Robotics×1Crypto×1Sound×1

Frequent co-authors

Jacky Kwok1×
Shulu Li1×
Pranav Atreya1×
Yuejiang Liu1×
Yixing Jiang1×
Chelsea Finn1×

Research Timeline

2026
Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors

The paper introduces DiffErase, a black-box attack that effectively removes inaudible audio watermarks while preserving perceptual quality by utilizing diffusion models.

PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning

The paper proposes Predictive Routing Replay (PR2) to stabilize reinforcement learning on Mixture of Experts (MoE) LLMs by predicting and incorporating short-horizon router evolution during training and rollout.

Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

The paper introduces Lookahead Group Reward (&) to combat Supervision Fidelity Decay (SFD) in on-policy distillation, significantly improving student model performance on long reasoning tasks.

LLM-as-a-Verifier: A General-Purpose Verification Framework

This paper introduces LLM-as-a-Verifier, a framework for fine-grained verification of LLMs using continuous scores, achieving state-of-the-art performance on various benchmarks.

Highlighted terms show continued research focus across papers

Papers

cs.AIcs.CLcs.LGEmpiricalRecentJul 6, 2026

LLM-as-a-Verifier: A General-Purpose Verification Framework

Jacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu +5 more

This paper introduces LLM-as-a-Verifier, a framework for fine-grained verification of LLMs using continuous scores, achieving state-of-the-art performance on various benchmarks.

View →
cs.LGcs.AIRecent
May 29, 2026

PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning

Daize Dong, Junlin Chen, Haolong Jia, Jiawei Wu +8 more

The paper proposes Predictive Routing Replay (PR2) to stabilize reinforcement learning on Mixture of Experts (MoE) LLMs by predicting and incorporating short-horizon router evolution during training a…

View →
cs.CLcs.AIRecentMay 29, 2026

Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

Yanjiang Liu, Jie Lou, Xinyan Guan, Yuqiu Ji +6 more

The paper introduces Lookahead Group Reward (&) to combat Supervision Fidelity Decay (SFD) in on-policy distillation, significantly improving student model performance on long reasoning tasks.

View →
cs.CRcs.SDRecentMay 28, 2026

Audio Pirates: Black-box Audio Watermark Removal via Diffusion Priors

Lingfeng Yao, Xincong Zhong, Chenpei Huang, Xuandong Zhao +5 more

The paper introduces DiffErase, a black-box attack that effectively removes inaudible audio watermarks while preserving perceptual quality by utilizing diffusion models.

View →