X · @natolambert
· X / Twitter
New lecture! This one is a recap of a bunch of history of preferences, the nature of rewards, how RLHF is formulated, which were once seen as central …
New lecture! This one is a recap of a bunch of history of preferences, the nature of rewards, how RLHF is formulated, which were once seen as central problems in the field. How much as changed. Still... super interesting to understand our optimization tools today. Books coming soon :D 00:00 Intro & context07:34 A short