Soumik Mukhopadhyay
1 indexed paper
Recent (6 mo)
1With code
0Influential cites
0Benchmarked
0Publications per year
126
Top categories
ML×1AI×1Vision×1
Frequent co-authors
Research Timeline
2026
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
This paper proposes two strategies to improve feedback efficiency of reinforcement learning from human feedback (RLHF) in diffusion models.
Highlighted terms show continued research focus across papers