Partnership opportunities

Save $200 on your pass

Call to action
Your text goes here. Insert your content, thoughts, or information in this space.
Button

Back to speakers

Subrata
Das
Principal Data Scientist
Humana
Subrata Das is Principal AI and Data Scientist and part-time Gen AI professor, with experience across machine learning, deep learning, reinforcement learning, NLP and generative AI. He takes a practical, stakeholder-focused approach grounded in strong theory, delivering AI and data solutions across finance, healthcare, law, ecommerce and defense. Previously, he held senior roles including Chief Data Scientist and AI Lab Manager at Xerox and worked as a consultant at MIT Lincoln Laboratory. He has led research programs in data fusion and analytics funded by DARPA and NASA. He currently teaches graduate NLP and Gen AI courses at Northeastern University and is the author of multiple books on AI and analytics, contributing through teaching, writing and seminars.
Button
29 October 2026 15:00 - 15:30
Panel | The new evaluation stack: Measuring reasoning, workflow quality, and real-world performance
How do you know your AI system is actually working? Not in a test environment. Not on a benchmark. In production, across multi-step tasks, where failures don't always throw an error and drift doesn't always announce itself. Most existing tooling wasn't built for this problem. This session gets into how engineering teams are building eval infrastructure for workflow-based AI: catching silent failures, measuring reasoning quality across decision chains, and closing the gap between controlled evals and real-world performance. Key takeaways: - How to measure reasoning quality across multi-step workflows, not just final output accuracy - The patterns that signal silent failure or drift before they surface as visible errors - What a production-grade eval stack looks like when you're evaluating a workflow, not a single model