X · @teortaxesTex
· X / Twitter
RT Taiqiang Wu: How to maximize OPD performance? 🧐 One important thing is warm-up. Then the student-sampled sequence is well defined in the teacher…
RT Taiqiang WuHow to maximize OPD performance? 🧐One important thing is warm-up. Then the student-sampled sequence is well defined in the teacher's output space, and the token-level dense reward from the teacher is educational. 🤜In this paper, we demystify the warm-up process:🧵