X · @teortaxesTex
· X / Twitter
RT Nathan Barry: Turns out you can make LLM inference fully deterministic across devices, with no loss to quality or speed. This weekend at the @Space…
RT Nathan BarryTurns out you can make LLM inference fully deterministic across devices, with no loss to quality or speed.This weekend at the @SpaceXAI hackathon I got Qwen3-0.6B to produce identical hashed logits from a 512-token generation across 2 GPUs and 3 CPUs: an A100, an H100, an Apple M5 Max, an AMD EPYC, and a