X · @soumithchintala
· X / Twitter
Modal trained a DFlash speculator that's much faster than MTP, making it a great boost for inference speeds!
Modal trained a DFlash speculator that's much faster than MTP, making it a great boost for inference speeds!Modal: Inkling by @thinkymachines is now available on Modal, backed by a custom DFlash speculator for 67% higher throughput and interactivity.Running on Modal Auto Endpoints with SGLang today.