For decades, modern weather forecasting has followed a very precise pipeline: collecting observations from satellites, radar, stations, and sounding balloons, converting them into a coherent representation of the state of the atmosphere, feeding them into massive physical models, and letting extremely powerful supercomputers solve, step by step, the equations governing the evolution of air, moisture, and energy. Artificial intelligence has begun disrupting this paradigm in recent years, demonstrating that a model trained on decades of data can produce competitive forecasts with a fraction of the compute. With WeatherNext 3, however, Google is attempting to take a step further: moving beyond simply mimicking the output of traditional systems to bring the machine directly closer to real-world raw data.

The new model developed by Google DeepMind and Google Research ingests low-latency geostationary satellite imagery, can be initialized hourly, and outputs forecasts with resolutions of up to roughly five kilometers for select surface variables. Compared to WeatherNext 2, which operated primarily on a 25-kilometer grid with new initializations every six hours, the leap is substantial. Google has already started integrating WeatherNext 3 into Search, Maps, Gemini, Google Maps Platform, and Cloud, making the project far more than a mere research paper: it is a system directly entering products used daily by millions of people and, via APIs, into enterprise and developer services.

The real breakthrough is in the model's input data

The most compelling aspect of WeatherNext 3 is less intuitive than its resolution or update frequency. Many of the leading AI weather models do not directly observe the atmosphere: they are trained and initialized using so-called analyses—reconstructions of the atmospheric state produced by numerical prediction systems through a complex data assimilation process. In practice, the AI learns from a world that has already been interpreted by a physical model.

This approach has yielded extraordinary results, but it creates a structural dependency. If the traditional analysis carries latency, bias, or insufficient resolution, the AI model inherits those limitations, at least in part. Google is attempting to reduce this dependency by directly feeding the system hourly updated geostationary satellite mosaics alongside conventional analysis data. The WeatherNext 3 paper highlights this shift as one of the central pillars of the project: moving AI beyond simply emulating the separate stages of assimilation, forecasting, and post-processing.

This does not mean Google has eliminated physical meteorology from the process, nor that the model can operate in a vacuum. WeatherNext 3 continues to rely on data generated by traditional systems, and the research itself fits into an ecosystem where observations, reanalyses, and numerical models remain fundamental. What is new is that satellite imagery is no longer merely a source that must first be digested by another machine: it becomes a direct input for the AI model, accelerating the link between what is unfolding above Earth and the resulting forecast.

Hourly forecasts change more than just convenience

When it comes to weather forecasting, six hours can be an enormous window. Fronts, storm cells, and precipitation systems can shift structure in much shorter timeframes—especially when the challenge is not knowing whether it will be hot in five days, but pinpointing where a storm system will hit over the coming hours. WeatherNext 3 can generate fresh initializations every hour, whereas the primary cycles of traditional global models are generally updated at wider intervals.

The outcome is particularly compelling because Google has not only increased frequency. Forecasts for temperature and dew point can reach roughly five-kilometer resolution, many surface variables hover around ten kilometers, and three-dimensional atmospheric fields sit at roughly twenty-five kilometers. Resolution does not automatically equate to accuracy, but it allows for a much better representation of coastlines, valleys, topography, and local gradients that tend to be smoothed out on coarser grids.

Google also claims to have achieved particularly significant improvements in precipitation forecasting—historically one of the toughest challenges for global models. WeatherNext 3 was trained partly on NASA's IMERG satellite dataset and its own proprietary precipitation reanalysis built on satellite and radar data. In benchmarks published by the company, the model demonstrates substantial reductions in probabilistic error compared to numerical baselines, with gains varying depending on the metric and dataset used. Consequently, when Google touts precipitation forecasts that are up to 50% more accurate for those planning at least a day in advance, it is referring to specific benchmarks rather than a universal rule applicable to every location, weather event, and lead time.

Meteorology has become one of AI's most serious fields

To understand the importance of WeatherNext 3, one must dismiss the idea that Google is playing alone. Meteorology is currently one of the scientific fields where machine learning is closest to a true operational transformation. The European Centre for Medium-Range Weather Forecasts, ECMWF, already put its Artificial Intelligence Forecasting System, AIFS, into production in February 2025, deploying it alongside the traditional Integrated Forecasting System. Among the advantages, ECMWF highlighted vastly lower energy consumption during the forecasting phase and competitive or superior results across several metrics, while still maintaining physical systems as a core part of its infrastructure.

This context is crucial because it prevents framing the story as a simplistic clash between “old supercomputers” and “new AI”. Major meteorological centres are incorporating machine learning, while AI labs continue to rely on the datasets and expertise generated by numerical meteorology. The most likely future is hybrid: physical systems and learned systems feeding into each other, with AI applied where it can deliver greater speed, more ensembles, and higher resolution, and physical models utilized for constraints, interpretability, assimilation, and verification.

WeatherNext 3 is particularly interesting because it shifts the boundary of this collaboration. Instead of merely tasking AI with predicting the future from an already reconstructed snapshot of the atmosphere, it brings it closer to direct observations. It is a step that could progressively narrow the divide between data collection and forecasting.

Weather is also an economic issue

Weather forecasting is often perceived as a service to help decide whether to carry an umbrella, but a massive portion of the economy relies on weather-dependent decisions. An airline adjusts flight paths and fuel loads, a farmer decides when to irrigate or treat a field, a logistics firm evaluates delays and risks, and an electrical grid must estimate how much power will be generated by solar and wind installations. Improving a forecast even marginally, when multiplied across millions of decisions, can yield substantial economic benefits.

Google developed WeatherNext 3 with the energy sector explicitly in mind. The model generates wind forecasts at an altitude of one hundred metres—closer to the hub height of wind turbines—as well as cloud cover and solar radiation fields useful for estimating photovoltaic output. This type of data is becoming increasingly critical with the expansion of renewable energy sources, as wind and solar are inherently variable and the power grid must continuously balance supply and demand.

The relationship between weather AI and energy is virtually circular. Artificial intelligence models consume infrastructure and electrical power, yet they can also help make an energy system with a rising share of intermittent production far more predictable. A more accurate forecast of wind or cloud cover can reduce the required safety margins, optimize plant scheduling, and make the deployment of battery storage and backup generation more efficient.

The greatest benefit could be in countries with less infrastructure

One of the most compelling arguments put forward by Google concerns parts of the world that lack the dense networks of radars, weather stations, and computing power found in Europe and North America. Building and operating high-resolution regional models is expensive, as is maintaining dense observational networks. A global model capable of directly utilizing satellite observations to produce high-resolution outputs could help bridge at least part of this divide.

This is a significant prospect, especially for regions vulnerable to extreme weather events that have limited meteorological infrastructure, but it must be approached with caution. The quality of a global model does not eliminate the need for local observations, national expertise, radar networks, and public meteorological services. Operational forecasting and early warnings require ground-level knowledge and institutional accountability. Google itself, within the WeatherNext 3 documentation, explicitly advises consulting local meteorological services for official warnings and safety guidance.

The real opportunity does not lie in replacing those institutions, but in providing them with a new source of global forecasting that can be run and queried at a fraction of the cost of developing an entire numerical infrastructure independently.

The risk of mistaking a good benchmark for absolute truth

WeatherNext 3 is presented by Google as its most accurate global weather model to date, with the company citing independent live evaluations from Operational WeatherBench, the Brightband project that benchmarks different systems under operational conditions. The presence of an external benchmark is positive, particularly in a domain where comparative performance can vary considerably depending on the variable, region, forecast lead time, and metric selected.

However, it remains essential to steer clear of absolute headlines such as “AI definitively beats traditional meteorology”. No model is superior across every parameter. A forecast may perform exceptionally well on average temperatures while proving less effective on local extremes; it may accurately capture the broader trajectory of a weather system yet miss a localized thunderstorm; it may demonstrate superiority on a global average but fall short within a specific territory. Recent literature on AIFS also demonstrates that comparisons between AI-driven and physics-based models can yield differing conclusions depending on the verification methodology and post-processing techniques applied.

Furthermore, the WeatherNext 3 paper is currently available as a preprint, meaning the scientific work is public and open to analysis, but should not be treated as though it has already passed through every stage of academic peer review. The fact that the model is already integrated into Google products makes it even more critical to distinguish between claimed performance, independent benchmarks, and the responsibilities of official forecasting agencies.

Google is not just building a model: it is building a meteorological platform

The decision to deploy WeatherNext 3 across Search, Maps, and Gemini is only the most visible part. For businesses and researchers, Google provides data access through BigQuery, Earth Engine, and Cloud Storage, while the technical catalog outlines global forecasts of up to fifteen days for main cycles and hourly intermediate updates with shorter horizons.

This means WeatherNext 3 can become a component within other products: software for agriculture, energy management, insurance, logistics, maritime shipping, aviation, or risk analysis. When a weather forecast can be queried like a cloud service, the value no longer lies solely in the model, but in the ecosystem of applications that can be built on top of it.

This is where Google holds a distinct advantage. DeepMind can develop the model, Google Cloud can distribute the data, Maps can integrate it into geographic space, Search can deliver it to the public, and Gemini can turn it into a conversational interface. Very few organizations possess research, cloud infrastructure, global mapping, and consumer products simultaneously at this scale.

The real revolution could be making forecasting nearly continuous

Traditional meteorology has been built around cycles: data is gathered, an analysis is assembled, the model is run, the forecast is generated, and then the process starts over a few hours later. WeatherNext 3 suggests a different path, where the boundary between one cycle and the next becomes less distinct because fresh satellite data feeds the system every hour.

We are not yet facing a global forecast that updates instant by instant, and there is no reason to believe atmospheric uncertainty can ever be eliminated. Chaos remains a physical property of the system, imposing fundamental limits on predictability. However, the faster a model can incorporate what is actually happening, the more it can reduce at least that portion of error stemming from an already outdated initial snapshot.

This, more than the five-kilometer resolution or the percentage accuracy on a single benchmark, is the breakthrough to watch. Weather AI is shifting from the phase of proving it can reproduce a forecast to the phase of actively redesigning the entire pipeline through which that forecast is born.

If this transition continues, in a few years the question will no longer be whether weather forecasts are made “with AI” or “with physics.” It will be far harder to separate the two, because learned models will use real-world observations, physical systems, historical data, and local corrections within a single infrastructure. WeatherNext 3 does not yet represent the end of the old paradigm, but it is one of the clearest signs that the new one is beginning to take shape.

Sources