Skip to content
X · @huggingface · X / Twitter

RT clem 🤗: Big unlock for open-source AI inference: Hugging Face Transformers models can now run in vLLM at native speed, often matching or beating…

RT clem 🤗Big unlock for open-source AI inference: Hugging Face Transformers models can now run in vLLM at native speed, often matching or beating hand-written implementations.Until now, every new architecture often needed to be built twice:- Once in Transformers for training and research- Again in vLLM for fast product