Google researchers have developed Retrieve-for-Train, a new framework that dramatically reduces computational costs in AI-powered search systems by replacing expensive inference-time reasoning with offline training. The innovation, detailed in a paper presented at ICML 2026, addresses a fundamental limitation of large language models: their tendency to generate redundant search query variants and require substantial chain-of-thought reasoning tokens that create unacceptable latency in production search systems.
The research identifies two critical problems with standard LLMs for query decomposition: "paraphrastic collapse," where models generate near-synonymous queries instead of exploring complementary facets, and autoregressive latency bottlenecks that make sub-second search responses impossible at scale. When a user searches for "camping gear," current systems struggle to return diverse results covering tents, sleeping bags, stoves, and headlamps—instead producing semantic variations of the same item.
The solution uses reinforcement learning to train a fan-out model offline, then distills its optimized behaviors into a lightweight 53.9M-parameter diffusion model. This compact model can generate a complete, coherent slate of search results in a single non-autoregressive pass, eliminating expensive reasoning tokens while maintaining sophisticated set-level properties like diversity and coverage. Google's framework enables production-grade search performance meeting sub-second latency requirements without sacrificing result quality.
Key Points
Retrieve-for-Train replaces expensive inference-time reasoning with offline reinforcement learning to accelerate search query decomposition
Addresses 'paraphrastic collapse' where LLMs generate redundant queries instead of diverse, complementary results
Distills learned behaviors into a 53.9M-parameter diffusion model that executes in a single non-autoregressive pass
Eliminates chain-of-thought reasoning tokens while maintaining set-level diversity and coverage properties