Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design
Researchers conducted a systematic analysis of reinforcement learning design space for diffusion models, focusing on likelihood estimation. They found that using an evidence lower bound (ELBO) based model likelihood estimator enables effective, efficient, and stable RL optimization, improving performance on multiple tasks.
Save an API key to vote.