r/LocalLLaMA
· Communities
20B Looping model (paper) matches or beats Qwen3 Coder 30B at 10% of pre-training tokens
No weights yet. I feel sad for them, that training run cost maybe 100s of thousands of dollars and they didn't even beat GPT-OSS 20B in every regard But the ability to train a model from scratch on 3.5 trillion tokens instead of 35 trillion sure gives me hope. They only spent 100s of thousands of $ instead of millions,