Skip to content
X · @swyx · X / Twitter

RT Latent.Space: The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www…

RT Latent.SpaceThe Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI https://www.latent.space/p/inference-eng@Baseten @philipkiely and @waterloo_intern explain what actually happens after a model is trained, why turning weights into a fast and reliable