01 December 2026 11:30 - 12:00
Multimodal generation in production: the video, image and voice models that actually ship
Multimodal models are no longer a research demo category, they are a deployed one, sitting alongside generative SaaS platforms and AI copilots as the shape most enterprise adoption now takes. Text still accounts for the largest share of enterprise generative AI usage by data modality, but the gap to image, voice and video is closing fast as more teams put multimodal generation directly into customer-facing products.
This session looks at what is actually shipping today across video, image and voice generation, not what is promised on a roadmap, and what changes technically once a multimodal model moves from a creative tool into a production pipeline.
What this session will cover:
- Which multimodal use cases across image, voice and video have moved from novelty to reliable production tooling in the last year
- What changes in latency, cost and quality control once generation output faces real customers instead of internal reviewers
- How teams are combining multiple specialized models rather than betting on one model to do everything
- What to evaluate before committing a multimodal model to a customer-facing workflow