ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Understanding of policy learning, mathematical analysis”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

stat.MLcs.LGTheoreticalRecentJul 3, 2026

A Hierarchy of Policy Learning Problems

Hamsa Bastani, Osbert Bastani, Shihan Chen

This paper provides a mathematical framework for studying different policy learning problems and shows reductions between them.

View →
math.OCcs.LGcs.MATheoreticalRecentJul 1, 2026

Mean Field Reinforcement Learning

René Carmona, Mathieu Laurière

This paper introduces mean field reinforcement learning through Markov decision processes in large-population stochastic control, developing the necessary framework for representative-agent learning,…

View →
stat.MLcs.LGTheoreticalRecentJul 19, 2026

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning

Joseph Lazzaro, Alessio Russo, Aldo Pacchiano

This paper provides non-asytotic sample complexity guarantees for the Navigate and Stop algorithm in online tabular Reinforcement Learning, identifying additional attributes that affect the overall sa…

View →
cs.AIRecentMay 30, 2026

AI Sovereignty as National Learning Capacity: A Human-Centered Learning Mechanics Viewpoint on France, the United States, and China

Kim Phuc Tran

The paper proposes viewing national AI development, specifically in France, as a 'national AI learning system' governed by a controlled balance between information injection and entropy dissipation, a…

View →
cs.HCcs.AIRecentMay 27, 2026

Learning to Assign Prediction Tasks to Agents with Capacity Constraints

Shang Wu, Saatvik Kher, Padhraic Smyth

This paper develops a policy-learning framework to optimally assign prediction tasks to multiple agents, considering individual agent expertise and capacity constraints, achieving systematic performan…

View →
stat.MLcs.LGmath.PRTheoreticalRecentJun 16, 2026

A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise

M. Forzo, E. Monzio Compagnoni, A. Russo, A. Pacchiano

This paper introduces a stochastic differential equation approximation for linear Temporal Difference (TD) learning under Markovian noise, explaining the constant-stepsize error floor.

View →
cs.LGstat.MLRecentJun 1, 2026

Minimax-Optimal Policy Regret in Partially Observable Markov Games

Raman Arora

The paper develops an optimistic maximum-likelihood algorithm that achieves $ ilde{O}(\sqrt{T})$ policy regret for sequential decision-making in partially observable Markov games against adaptive oppo…

View →
cs.LGcs.AIRecentMay 28, 2026

On Distributional Reinforcement Learning in Chaotic Dynamical Systems

James Rudd-Jones, Mirco Musolesi, María Pérez-Ortiz

The paper proposes using distributional Reinforcement Learning (RL) to stabilize learning in chaotic dynamical systems by optimizing the smooth evolution of the return distribution rather than individ…

View →
cs.LGstat.MLTheoreticalRecentJun 30, 2026

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

Mingyi Li, Taira Tsuchiya, Kenji Yamanishi

This paper develops a new algorithm for policy optimization in online episodic tabular Markov decision processes with unknown transition kernels, providing data-dependent regret bounds and best-of-bot…

View →
cs.AIRecentMay 27, 2026

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

Junyu Zhang, Feihong Yang, Jian Wang, Chao Wang +1 more

The paper introduces Global PSRO, a novel deep reinforcement learning framework that efficiently approximates Nash equilibria in large two-player zero-sum games by intelligently expanding the strategy…

View →
cs.LGcs.AIstat.MLRecentMay 29, 2026

Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning

Yike Zhao, Onno Eberhard, Malek Khammassi, Ali H. Sayed +1 more

This paper theoretically justifies the strong performance of linear recurrent neural networks as memory units in partially observable reinforcement learning by constructing specific linear filters tha…

View →
cs.LGcs.AIRecentMay 29, 2026

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

Jonathan Colaço Carr, Prakash Panangaden, Doina Precup, Benjamin Van Roy

The paper introduces the Markov decision contest, a new framework for reinforcement learning using pairwise preferences, and proves that stationary Markov policies are optimal and solvable efficiently…

View →
cs.ROcs.LGEmpiricalRecentJul 26, 2026

Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

Zahra Abdalla Elashaal, Afef Hfaiedh, Nahla Khraief, Issmail Ellabib +1 more

The paper proposes a Hierarchical Reinforcement Learning framework with two levels for handling high-level strategic planning and low-level continuous-control using Soft Actor-Critic and entropy-regul…

View →
cs.ROcs.AIcs.CVEmpiricalRecentJun 29, 2026

Learning from Mistakes: Rollout-Retrieval Lifelong Policy Learning for Autonomous Driving

Cheng Gong, Haoyang Wang, Chao Lu, Zirui Li +1 more

This paper proposes Rollout-Retrieval Lifelong Policy Learning (R$^2$LPL), a framework for continual policy improvement in autonomous driving by retrieving corrective targets from recoverable mistakes…

View →
cs.AIcs.LGRecentMay 30, 2026

Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief

Hongqiang Lin, Pengfei Wang, Nenggan Zheng

The paper introduces Posterior Hybrid Bayesian Belief (PhyB), a novel framework that reformulates policy optimization in Bayesian Offline RL by approximating expectations as a convex combination over…

View →
cs.NESurveyRecentJul 17, 2026

From Optimal Policies to Individual Differences: Rethinking Reinforcement Learning for Biology

Patrick Govoni, Palina Bartashevich, Clémence Bergerot, Valerii Chirkov +2 more

This paper explores approaches to generating behavioral diversity in reinforcement learning models to bridge the gap between simulation and biology.

View →