X · @teortaxesTex
· X / Twitter
Interesting
InterestingXidulu: 6/ Of course, models trained this way can perform- (Soft, free) Latent feedback decoding- (fused) Double prefill + soft decodingThe extra 30% FLOPs spent at training time translate to free-form generation performance comparable with models trained with 2x the FLOPs