What are agent evals?

Evals measure whether an agent behaves correctly on a known set of cases, so a prompt, model or tool change can be checked before it ships.

Enterprise

How AgentOven does it

Test suites run a set of inputs against an agent on demand or on a schedule, and results are kept with each run. LLM-judge scoring is on the roadmap.

Test suitesScheduled runs