X · @teortaxesTex
· X / Twitter
RT _horror: Stutton's Dwarkesh appearance didn't really convince me. But this idea of using a primitive grounding reward and let the model learn its o…
RT _horrorStutton's Dwarkesh appearance didn't really convince me. But this idea of using a primitive grounding reward and let the model learn its own internal utility schedule feels inherently right by unprincipled brute intuition.https://oaklab.ai/posts/learning-from-experience-instead-of-curated-datasets