26 August 2026 13:30 - 14:00
Simulating human emotion at scale: How swarm AI predicts what polls and sentiment tools cannot
Agentic AI introduces a new set of evaluation challenges that go well beyond traditional LLM benchmarking.
When a model plans, calls tools, recovers from errors, and operates over long horizons, the question shifts from "is this answer correct?" to "did this trajectory accomplish the goal - safely, efficiently, and for the right reasons?"
This talk surveys the practical challenges of evaluating agents in production: why static benchmarks saturate and mispredict real-world behavior, how path-dependence and multiple valid solutions complicate scoring, and why trajectory-level metrics (steps, tool calls, cost, latency) often matter as much as final-task success.