Skip to content
X · @ylecun · X / Twitter

RT Steeve Morin: We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It s…

RT Steeve MorinWe're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack.It ships with 5 architectures out of the box: NVIDIA, AMD, Metal, Intel and TPU. All transparent.It supports DFlash, continuous batching, prefix caching, the whole deal.Oh, and