20 results for “Understanding of reinforcement learning, generative modeling, and diffusion processes.”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
Renhao Zhang, Haotian Fu, Mingxi Jia, George Konidaris +2 more
The Parameterized Diffusion Policy (PDP) framework transforms diffusion models from general stochastic generators into precise, steerable tools for learning and adapting complex robotic behaviors by e…
This paper introduces mean field reinforcement learning through Markov decision processes in large-population stochastic control, developing the necessary framework for representative-agent learning,…
A model-driven approach is proposed for generating families of reinforcement learning training environments using a hybrid genetic algorithm and model transformations.
This paper introduces DADiff, a diffusion-based framework for domain adaptation in reinforcement learning, which estimates dynamics mismatch based on generative trajectory deviation.
This book provides a compact, derivation-oriented mathematical primer that connects major families of generative AI models, showing their underlying structural relationships.
The paper proposes DIBS, a decoupled behavioral cloning approach that stabilizes inductive generalization in RL by separating task-specific policy learning from the evolution function, leading to impr…
This paper proposes two strategies to improve feedback efficiency of reinforcement learning from human feedback (RLHF) in diffusion models.
This paper provides a mathematical framework for studying different policy learning problems and shows reductions between them.
The paper introduces the Terminal Representation (TR), a novel, lower-dimensional, and structurally distinct formulation for encoding reward-weighted trajectories in RL that bypasses the need for eige…
Yushi Huang, Xiangxin Zhou, Jun Zhang, Liefeng Bo +1 more
This paper introduces MeanFlowNFT, a method that applies reinforcement learning to MeanFlow generators to improve performance and consistency.
This paper introduces a stochastic differential equation approximation for linear Temporal Difference (TD) learning under Markovian noise, explaining the constant-stepsize error floor.
The paper proposes using distributional Reinforcement Learning (RL) to stabilize learning in chaotic dynamical systems by optimizing the smooth evolution of the return distribution rather than individ…
The paper introduces the Markov decision contest, a new framework for reinforcement learning using pairwise preferences, and proves that stationary Markov policies are optimal and solvable efficiently…
This paper proposes a new imitation learning algorithm called DistIL that uses distributional feedback to improve policy improvement and regret guarantees.
Yiming Ren, Yiran Xu, Zicheng Lin, Chufan Shi +7 more
The paper proposes S2L-PO, a framework that uses smaller, naturally diverse models as structured explorers to enhance the policy-level diversity and performance of larger language models during traini…
Jiakang Li, Guanyu Zhu, Can Jin, Chenxi Huang +7 more
The paper introduces Latent Reward Steering (LRS), an adaptive inference-time framework that implicitly improves the reasoning ability of LLMs by guiding the model's internal latent states based on a…
This paper explores approaches to generating behavioral diversity in reinforcement learning models to bridge the gap between simulation and biology.
Zihao He, Hongjie Fang, Shirun Tang, Cewu Lu +1 more
The paper proposes LAG-Fusion, a framework for asynchronous multimodal diffusion policy composition with latency-aware guidance fusion.