Skip to content
r/MachineLearning · Communities

Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]

I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because of limited compute I cannot try these papers out and learn them by implementing them myself. If someone here has worked w