X · @natolambert
· X / Twitter
Number 1 on day 1, is a good way to start.
Number 1 on day 1, is a good way to start.Nathan Lambert: My book, Reinforcement Learning from Human Feedback is done!This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights an