X · @swyx
· X / Twitter
RT Latent.Space: The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www…
RT Latent.SpaceThe Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent.space/p/inference-eng@Baseten @philipkiely and @waterloo_intern explain what actually happens after a model is trained, why turning weights into a fast and reliable