29 October 2026 12:00 - 12:30
Inference at scale: The real cost of running GenAI in production
What happens after the demo?
The prototype works, the stakeholders are excited, and then you push to production and everything gets harder. Latency spikes. Costs balloon. Reliability drops in ways that are difficult to debug and even harder to explain. This session is for the engineers who've already hit that wall, or who can see it coming.
Covering the decisions that actually determine whether GenAI survives contact with real traffic: when to cache, when to distill, how to architect for throughput without destroying your unit economics, and what teams who've cracked this are doing differently. Practical and numbers-grounded, with no hand-waving about scale.
If you're running GenAI at any meaningful volume, or building toward it, this session will change how you think about your infrastructure roadmap.