Skip to content
X · @natolambert · X / Twitter

Number 1 on day 1, is a good way to start.

Number 1 on day 1, is a good way to start.Nathan Lambert: My book, Reinforcement Learning from Human Feedback is done!This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights an