Skip to content
X · @teortaxesTex · X / Twitter

> 1. Draft with a quantized copy of the model itself – it re-syncs for free every training step yeah… that's no DSpark… but every acceleration count…

> 1. Draft with a quantized copy of the model itself - it re-syncs for free every training stepyeah… that's no DSpark… but every acceleration countsI wonder how stale drafter acceptance lengths actually change between rollouts though, my gut says < 5% per 100 RL stepsVuk Rosić 武克: RL training for LLMs spends nearly 70%