Skip to content
X · @teortaxesTex · X / Twitter

RT Nathan Barry: Turns out you can make LLM inference fully deterministic across devices, with no loss to quality or speed. This weekend at the @Space…

RT Nathan BarryTurns out you can make LLM inference fully deterministic across devices, with no loss to quality or speed.This weekend at the @SpaceXAI hackathon I got Qwen3-0.6B to produce identical hashed logits from a 512-token generation across 2 GPUs and 3 CPUs: an A100, an H100, an Apple M5 Max, an AMD EPYC, and a