X · @swyx
· X / Twitter
RT Nathan Lambert: My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now…
RT Nathan LambertMy book, Reinforcement Learning from Human Feedback is done!This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024. Transferring as mu