Skip to content
X · @dwarkesh_sp · X / Twitter

RT λux: i think dwarkesh’s strongest point here is that deployed agents are already collecting off-policy experience across thousands of real domain…

RT λuxi think dwarkesh’s strongest point here is that deployed agents are already collecting off-policy experience across thousands of real domains but almost none of it becomes gradient signal. the agent is mostly treated as an inference endpoint and not as a learner (deployment is not automatically learning). imo the