X · @ylecun
· X / Twitter
RT Steeve Morin: We're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack. It s…
RT Steeve MorinWe're releasing ZML/LLMD, our homegrown LLM server built on top of our homegrown high performance heterogeneous inference stack.It ships with 5 architectures out of the box: NVIDIA, AMD, Metal, Intel and TPU. All transparent.It supports DFlash, continuous batching, prefix caching, the whole deal.Oh, and