Skip to content
X · @cwolferesearch · X / Twitter

RT Viv: great write up, evals are hard! here are 2 broad buckets we use to evaluate agents: 1. Measure the State of the World 2. Agent as a Judge on t…

RT Vivgreat write up, evals are hard! here are 2 broad buckets we use to evaluate agents:1. Measure the State of the World2. Agent as a Judge on the Trajectory1. Measure the state of the environment before and after the Agent does the Task.@harborframework and containerized Evals make this easier. The agent thinks in o