X · @swyx
· X / Twitter
RT Daniel Han: My 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out! 1. Closed vs open models 2. Throughput maxxing but …
RT Daniel HanMy 2hr workshop on Open vs Closed models, reward hacking, benchmaxxing & RL is out!1. Closed vs open models2. Throughput maxxing but accuracy minimizing3. Benchmaxxing & cheating4. Distillation & RL5. Stopping reward hacking6. @UnslothAI Dynamic QuantsDetails:1. If reasoning wasn't discovered via o1-previe