20 results for “Group Relative Policy Optimization”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Azal Ahmad Khan, Ammar Ahmed, Zeshan Fayyaz, Sheng Di +2 more
The paper introduces Straggler-Aware Group Control (SAGC), a dynamic group-size controller that optimizes synchronous on-policy RL training by adapting group size to minimize delays caused by slow rol…
Yiming Ren, Yiran Xu, Zicheng Lin, Chufan Shi +7 more
The paper proposes S2L-PO, a framework that uses smaller, naturally diverse models as structured explorers to enhance the policy-level diversity and performance of larger language models during traini…
The paper introduces AdaPrefix-GRPO, a method that adjusts the amount of reference solution assistance during training to improve the success rate and accuracy of Group Relative Policy Optimization (G…
The paper introduces Safe Equilibrium Policy Optimization (σepo{}) to train language models for multi-agent strategic tasks, achieving improved safety and robustness across various game domains.
Chao Wang, Hongtao Tian, Tao Yang, Yunsheng Shi +2 more
This paper introduces PASS (Process Advantage Signal Shaping), a method to address three pathologies in GRPO (Group Relative Policy Optimization) for process-supervised reinforcement learning of LLM r…
This paper provides a mathematical framework for studying different policy learning problems and shows reductions between them.
The paper proposes a scalable, distributed approach for constrained Multi-Agent Reinforcement Learning by using local consensus over dual variables to ensure global constraint satisfaction without cen…
The paper introduces the Markov decision contest, a new framework for reinforcement learning using pairwise preferences, and proves that stationary Markov policies are optimal and solvable efficiently…
This paper trains sparse sensor policies for Rayleigh-Bénard convection control using multi-agent reinforce learning and grouped regularization.
This paper introduces mean field reinforcement learning through Markov decision processes in large-population stochastic control, developing the necessary framework for representative-agent learning,…
The paper introduces Prompted Policy Optimization (PromptPO), an LLM-based method that successfully optimizes policies for various sequential RL tasks, demonstrating that LLMs can replace classical RL…
This paper introduces SAFE, a new framework for multi-agent reinforce learning with continuous action spaces using a counterfactual baseline conditioned on a self-evolving default action.
The paper introduces a learned 'rerooter' mechanism to improve subgoal-based policy tree search, allowing scalable search in complex environments without the overhead of explicit subgoal generation.
This paper develops a new algorithm for policy optimization in online episodic tabular Markov decision processes with unknown transition kernels, providing data-dependent regret bounds and best-of-bot…
This paper proposes Primitive-Guided Tree Search (PGTS), a hybrid framework for computing Nash equilibrium policies in multi-agent Pursuit-Evasion games by integrating offline exact computation with o…
Johanna Menn, Miriam Kober, Paul Brunzema, David Stenger +1 more
The paper introduces local Preferential Bayesian Optimization (PBO) methods that adapt high-dimensional Bayesian Optimization techniques, such as trust-region and derivative-informed local search, to…