X · @teortaxesTex
· X / Twitter
RT wesley hsieh: The method is quite simple: For each data point x, they sample a timestep t, create K noised latents zt with K different noises, and …
RT wesley hsiehThe method is quite simple:For each data point x, they sample a timestep t, create K noised latents zt with K different noises, and only run backward propagation on the sample with minimum loss.If they skip the process of selecting the best sample, then this algorithm reduces to a variant of multiplicity