X · @jeremyphoward
· X / Twitter
RT Sara Dragutinovic: Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older op…
RT Sara DragutinovicIs Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.