Skip to content
X · @jeremyphoward · X / Twitter

RT Sara Dragutinovic: Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older op…

RT Sara DragutinovicIs Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.