Sign In
Register

Partnership opportunities

Secure your seat

Call to action
Your text goes here. Insert your content, thoughts, or information in this space.
Button

Back to speakers

Maheshwar
Kuchana
Senior Machine Learning Engineer
Deliveroo
Maheshwar Kuchana is a Senior Machine Learning Engineer specialising in productionising AI agents and building scalable ML systems across retail, healthcare, finance, and food‑tech. With experience at Deliveroo, Faculty, and ASOS, he has led high‑impact projects including multi‑agent platforms, KYC automation, diagnostic imaging systems, and enterprise‑grade MLOps tooling. He brings deep expertise in cloud technologies, LLMs, multi‑agent orchestration, and end‑to‑end ML pipelines, consistently delivering reliable, real‑world AI products.
Button
01 December 2026 16:30 - 17:00
Panel | The new evaluation stack: measuring reasoning, workflow quality and real-world performance
How do you know your AI system is actually working? Not in a test environment, not on a benchmark, but in production, across multi-step tasks, where failures do not always throw an error and drift does not always announce itself. Most existing tooling was not built for this problem. This session gets into how engineering teams are building eval infrastructure for workflow-based generative AI: catching silent failures, measuring reasoning quality across decision chains, and closing the gap between controlled evals and real-world performance. What this session will cover: - How to measure reasoning quality across multi-step workflows, not just final output accuracy - The patterns that signal silent failure or drift before they surface as visible errors - What a production-grade eval stack looks like when you are evaluating a workflow, not a single model - Where controlled benchmark evals and real-world performance diverge, and why that gap matters more as workflows get longer