02 December 2025 16:30 - 17:00
Panel | Is it doing the right thing? The new job of checking your agents' work
An agent can complete a task successfully and still produce the wrong outcome. As agents take on longer workflows and greater autonomy, teams need to evaluate more than final outputs, they must determine whether the agent followed the right process, used tools appropriately and acted within its intended boundaries.
This panel will examine how engineering teams can evaluate agent behaviour at scale. We'll explore what “correct” looks like for non-deterministic systems, where human review still belongs and how continuous evaluation can help teams trust agents without slowing down deployment.
Key takeaways:
→ Define success for agents when there may be more than one valid path or outcome.
→ Evaluate both final outputs and the actions taken to produce them.
→ Combine automated evaluation with targeted human review.
→Build continuous checks that keep pace as agents, tools and workflows evolve.