X · @togethercompute
· X / Twitter
This is what latency optimization looks like below the API 👇 Together ATLAS, NVIDIA Blackwell, CUDA, TensorRT-LLM, Dynamo, and custom kernels all w…
This is what latency optimization looks like below the API 👇Together ATLAS, NVIDIA Blackwell, CUDA, TensorRT-LLM, Dynamo, and custom kernels all working together to make inference faster for users.Proud of @realDanFu and our team!NVIDIA AI Infrastructure: AI that responds in under 100ms doesn't happen by accident.@toge