Google Research has introduced TimesFM-3, a state-of-the-art foundation model that advances time series forecasting by enabling accurate multivariate predictions in a single forward pass. The 330-million parameter model, pre-trained on more than one trillion time points, represents a significant upgrade from TimesFM-2.5, which was limited to univariate forecasting—analyzing only single time series in isolation.
Unlike its predecessors, TimesFM-3 is natively designed to handle the complexity of real-world forecasting problems where multiple related time series and external factors jointly influence predictions. The model supports three types of data inputs: multiple target series to forecast simultaneously, historically known past covariates, and past-future covariates such as known promotional campaigns or weather forecasts. This capability is particularly valuable for enterprise use cases like retail inventory planning, where ice cream sales depend not just on historical patterns but also on foot traffic, competitor offerings, and scheduled promotions.
Technically, TimesFM-3 employs an alternating attention architecture that blends temporal patterns with cross-series relationships. The model uses causal temporal attention to analyze trends within individual time series while leveraging full variate attention to capture how one series influences another. The use of contiguous patch masking enables non-autoregressive decoding—generating entire forecasting horizons in a single pass rather than iteratively, reducing latency and computational cost. The model outputs probabilistic forecasts across nine quantiles, providing a comprehensive view of forecast uncertainty.
Key Points
TimesFM-3 introduces native multivariate forecasting, enabling simultaneous prediction of multiple related time series with captured cross-dependencies
The model supports three covariate types: multiple targets, past covariates, and past-future covariates like planned promotions or weather forecasts
Uses alternating attention architecture combining causal temporal attention with full variate attention for efficient cross-series correlation learning
Generates complete forecasts in a single forward pass using contiguous patch masking, reducing latency and computational overhead
Pre-trained on 1 trillion time points, the 330M-parameter model achieves zero-shot generalization across retail, finance, manufacturing, and healthcare domains