Skip to content
X · @teortaxesTex · X / Twitter

…oh, right of course they instantly did that @TheZvi hardest hit

…oh, rightof course they instantly did that@TheZvi hardest hitGoodfire: Our team spent months developing RLFR, our method which uses probes on a model's internals as reward signals for RL.Silico reproduced it in 2 days, reducing hallucinations in Qwen3-8B by 37% without capability loss. (3/6)