r/LocalLLaMA
· Communities
eval-harness: A solution for generating personal evaluations that I have put together to evaluate agentic-cli harnesses
Hello LocalLLaMA, I wanted to build out my own personal list of evaluations, early on into putting this together I realised I wanted a way to not just evaluate the model but also the agentic harness that the model is running within, as I find the majority of my use of LLMs is more and more inside of a suite of CLI agen