Skip to content
X · @swyx · X / Twitter

RT Daniel Han: My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but …

RT Daniel HanMy 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out!1. Closed vs open models2. Throughput maxxing but accuracy minimizing3. Benchmaxxing & cheating4. Distillation & RL5. Stopping reward hacking6. @UnslothAI Dynamic QuantsDetails:1. If reasoning wasn't discovered via o1-previe