X · @natolambert
· X / Twitter
nice rl experiment on train-inference mismatch
nice rl experiment on train-inference mismatchYichuan Wang: Zero Train–Inference Mismatch — now for linear attention, and under async RL 🎯We got bitwise-exact trainer/generator parity for Gated DeltaNet (Qwen3.5-9B / 35B-A3B) on TorchTitan RL + vLLM, then asked the question nobody had actually tested in open source: do