Google Research has unveiled a significant breakthrough in video generation technology, introducing a multi-agent framework called "AI video co-director" that automatically creates long-form video narratives while maintaining visual and narrative consistency. The framework, built on top of Google's Gemini and Veo models, addresses fundamental challenges that have plagued current AI video pipelines: semantic drift (where characters and scenery subtly change between shots), cascading failures, and content collapse. By modeling video generation as a global optimization and world-state tracking problem rather than a linear chain of independent processes, the system can now generate minute-long videos with significantly improved character persistence and visual continuity. The architecture employs a hierarchical multi-agent orchestration system that decouples creative intent from consistency enforcement. An Orchestrator Agent uses a multi-armed bandit algorithm to navigate creative choices across three dimensions—narrative strategy, story structure, and aesthetic tone—while specialized sub-agents handle keyframes, video synthesis, and audio. This top-down steering approach ensures the entire pipeline operates under a unified creative vision rather than accumulating errors from independently crafted prompts. The framework natively inherits Google's SynthID watermarking and other safety mechanisms, with additional classifiers available for production use. The research, which will be presented at COLM 2026 and EMNLP 2026 conferences, represents a crucial step toward making AI-assisted video creation more practical for creative professionals. By automating the burdensome orchestration tasks that currently require extensive manual intervention, the system frees creators to focus on storytelling while the AI handles the technical complexity of maintaining visual continuity across a multi-shot narrative.