Meridia Insight Science Breakthroughs Knowledge

AI Matches Swiss Weather Forecasts at 1 km Resolution

A new AI weather model matches or beats Switzerland’s official high-resolution forecasts across the Alps—without solving a single physics equation.

25% reduction in temperature forecast error after just 2 years of fine-tuning data.

In the Swiss Alps, where a sudden wind gust can derail a train and an hour of unforecast rain can trigger a landslide, weather prediction isn’t just science—it’s infrastructure. Now, a new artificial intelligence system called Varda-single-1.0 has achieved something remarkable: it forecasts the weather across this complex mountain terrain at a resolution of 1 kilometer, every hour, for up to five days—and it matches or beats the performance of Switzerland’s official high-resolution numerical weather models.

That’s not just a technical milestone. It means that for the first time, a data-driven model has crossed the threshold from experimental curiosity to operational parity in one of the world’s most challenging forecasting environments. Over a full year of testing, Varda-single-1.0 matched the skill of MeteoSwiss’ flagship 1 km ICON-CH1-EPS model out to 33 hours and outperformed its 2 km ICON-CH2-EPS counterpart all the way to 120 hours (Smith et al., 2026). For a country where avalanches, flash floods, and windstorms are routine threats, this leap in forecast capability could save lives, protect infrastructure, and stabilize energy grids.

And it did so without solving a single differential equation.

The Science

Varda-single-1.0, developed by researchers at MeteoSwiss and partners using the open-source Anemoi framework (Lang et al., 2024a), is a purely data-driven forecasting system. Unlike traditional numerical weather prediction (NWP) models, which simulate the atmosphere by solving the physical laws of fluid dynamics and thermodynamics on supercomputers, Varda learns patterns directly from decades of historical weather data.

The system operates on a "stretched-grid" design: a global mesh at 31 km resolution that zooms in to 1 km over the Alpine region

Figure 3: 
Input meshes of Varda-single. Blue nodes are obtained from the native ERA5 N320 reduced Gaussian grid and represent the spatial boundary conditions necessary to model the regional domain. Red nodes are the cell centres of the REA-L-CH1 grid at 1 km\mathrm{km} spatial resolution, our regional domain of interest. During pre-training, only global nodes are considered to learn global weather dynamics. In the stretched-grid training step, the regional domain is embedded within the global grid, and regional dynamics are learned conditional on the global dynamics. We display a close-up of a 24 km\mathrm{km} ×\times 24 km\mathrm{km} cutout to depict the complex topography of the Alpine region, which spans a substantial portion of the regional domain.
Figure 3: Input meshes of Varda-single. Blue nodes are obtained from the native ERA5 N320 reduced Gaussian grid and represent the spatial boundary conditions necessary to model the regional domain. Red nodes are the cell centres of the REA-L-CH1 grid at 1 km\mathrm{km} spatial resolution, our regional domain of interest. During pre-training, only global nodes are considered to learn global weather dynamics. In the stretched-grid training step, the regional domain is embedded within the global grid, and regional dynamics are learned conditional on the global dynamics. We display a close-up of a 24 km\mathrm{km} ×\times 24 km\mathrm{km} cutout to depict the complex topography of the Alpine region, which spans a substantial portion of the regional domain. Source: Alberto Pennino, Francesco Zanetta

. This avoids the need for lateral boundary conditions—external forecasts that feed into regional models—which simplifies deployment and reduces error sources. The model is built using Graph Transformers, a type of neural network that treats the Earth’s surface as a graph, with weather stations and grid points as nodes connected by edges that encode spatial relationships.

Varda-single consists of two independently trained models:

  1. A 6-hourly autoregressive forecaster that predicts the state of the atmosphere every 6 hours.
  2. A temporal downscaler that fills in the missing hourly steps between those 6-hour forecasts.

This two-stage architecture is key. Running an AI model autoregressively at 1-hour intervals over 120 hours would accumulate too much error. By forecasting every 6 hours and then "filling in the blanks" with a separate model conditioned on both endpoints, Varda avoids compounding mistakes while still delivering hourly output—exactly what sectors like aviation, energy, and emergency management need.

The training process was equally sophisticated. The team used a four-stage curriculum:

  • Global pre-training on ERA5 reanalysis data (1979–2023) to learn broad atmospheric dynamics.
  • Stretched-grid training on a 20-year, 1 km regional reanalysis (REA-L-CH1) combined with global data.
  • Rollout training to improve multi-step forecast stability.
  • Fine-tuning on real-time operational analyses (KENDA-CH1) from 2024–2025.

The final fine-tuning stage was surprisingly decisive. Despite using only two years of data, it produced the largest skill jump in the entire pipeline—cutting the RMSE of 2-meter temperature by 25% at +60 hours. This suggests that even high-quality reanalyses like REA-L-CH1 don’t perfectly match real-time operational conditions. The model already knew the physics; it just needed to adapt to the "real world" data it would actually ingest.

What They Found

When tested over a full year (April 2025 to April 2026) against MeteoSwiss’ operational analyses and surface observations, Varda-single-1.0 didn’t just keep pace—it led.

For 2-meter temperature and 10-meter wind speed, two of the most critical variables for public safety and energy planning, Varda matched or exceeded the skill of both ICON-CH1-CTRL (1 km) and ICON-CH2-CTRL (2 km) across most lead times. At +6 hours, Varda outperformed ICON-CH1-CTRL in over 60% of the Alpine domain for temperature and nearly 50% for wind speed

Spatial Skill Advantage at +6h Lead Time

Fraction of Alpine domain where Varda-single outperforms baseline models in 2m temperature at +6h lead time.

Spatial Skill Advantage at +6h Lead Time
LabelValue
Varda-single vs ICON-CH1-CTRL0.62
Varda-single vs ICON-CH2-CTRL0.48
ICON-CH1-CTRL vs ICON-CH2-CTRL0.55

.

The gains were especially pronounced in wind forecasts. While both models struggled with extreme gusts in narrow valleys—a known challenge in complex terrain—Varda showed a more consistent bias pattern and lower random error. This is critical: underestimating wind speed in the Alps isn’t just an academic misstep; it can mean failing to close mountain passes or issue avalanche warnings in time.

Precipitation was more mixed. Varda captured large-scale winter storms well, but during summer convection, it produced smoother, more diffuse rainfall fields than observed. This is a known limitation of models trained with mean squared error (MSE) loss, which penalizes large errors heavily and thus "averages out" extremes. As a result, Varda tended to miss the sharp peaks of convective cells, though it got the overall storm location and timing right.

To quantify this, the team used the SAL (Structure, Amplitude, Location) metric, which breaks down precipitation errors into three components. Across 69 precipitation events, Varda and ICON-CH1-CTRL performed similarly in storm location and structure, but Varda consistently underestimated amplitude—by about 15% on average (Figure 11). In practical terms, this means the model might forecast 10 mm of rain when 12 mm actually falls: enough to matter for flood risk.

Yet even here, the story isn’t all downside. The temporal downscaler successfully reconstructed the diurnal cycle of convection—something many AI models miss. In the afternoon heat of July, when thunderstorms pop up over the Alps, Varda captured the timing and spatial evolution of cloud formation with striking fidelity

Figure 10: 6 h\mathrm{h} accumulated precipitation for a well-forecast winter case (top; initialised at 00 UTC on 7 December 2025, lead time +24 h\mathrm{h}) and a poorly forecast summer case (bottom; initialised at 12 UTC on 2 July 2025, lead time +6 h\mathrm{h}). Left: Varda-single. Right: the verifying KENDA-CH1 analysis. Inset boxes give the SAL components for the displayed 6 h\mathrm{h} window. Note that these are single-window values, whereas the values in Figure 11 are per-initialisation averages over the 6–30 h\mathrm{h} windows and are therefore numerically different.
Figure 10: 6 h\mathrm{h} accumulated precipitation for a well-forecast winter case (top; initialised at 00 UTC on 7 December 2025, lead time +24 h\mathrm{h}) and a poorly forecast summer case (bottom; initialised at 12 UTC on 2 July 2025, lead time +6 h\mathrm{h}). Left: Varda-single. Right: the verifying KENDA-CH1 analysis. Inset boxes give the SAL components for the displayed 6 h\mathrm{h} window. Note that these are single-window values, whereas the values in Figure 11 are per-initialisation averages over the 6–30 h\mathrm{h} windows and are therefore numerically different. Source: Alberto Pennino, Francesco Zanetta

.

Why This Changes Things

The success of Varda-single-1.0 signals a turning point: data-driven models are no longer just global curiosities. They can now operate at the high resolutions and complex terrains where national meteorological services live and die.

Consider the implications. Traditional NWP models like ICON require massive supercomputing resources—MeteoSwiss runs its 1 km ensemble on one of Europe’s most powerful machines. These systems are expensive, energy-intensive, and slow to update. Varda, by contrast, runs at a fraction of the computational cost. While the paper doesn’t give exact numbers, prior work suggests AI models can be 100–1,000 times faster at inference (Keisler, 2022; Bi et al., 2023). That speed opens doors: faster update cycles, higher spatial resolution, and the ability to run ensembles for probabilistic forecasting—all without new hardware.

But the real revolution is in accessibility. Varda was built using Anemoi, an open framework developed by ECMWF and European meteorological services. The model’s architecture, training pipeline, and even its pre-trained weights are designed to be shared. This means a small country with limited computing power could fine-tune Varda for its own terrain—say, the Himalayas or the Andes—without starting from scratch.

And the benefits extend beyond accuracy. Because Varda learns from data, not equations, it can implicitly capture processes that are too small or complex to resolve in traditional models—like the way cold air drains down a valley at night or how a ridge triggers a rotor cloud. These "sub-grid" effects are usually parameterized in NWP, a major source of error. Varda, trained on 20 years of high-resolution reanalysis, learns them directly.

Still, challenges remain. The model’s smoothing of convective precipitation is a reminder that AI isn’t magic. It reflects the limitations of its training objective. Future versions could use quantile loss or generative models to better capture extremes. And while Varda is deterministic, the next step—probabilistic forecasting—is already in motion, with teams exploring diffusion models and ensemble distillation (Price et al., 2024; Bonev et al., 2025).

What’s Next

Varda-single-1.0 is not a replacement for physics-based models. It’s a complement. The future likely lies in hybrid systems: AI for speed and pattern recognition, physics for consistency and extreme events. Some teams are already coupling AI models to physical solvers, using neural networks to correct biases or accelerate subroutines (Xu et al., 2024).

For MeteoSwiss, the path forward is clear. Varda will be integrated into operational workflows, providing backup forecasts and specialized products—like wind gusts for wind farms or fog for airports. The team is already working on Varda-ensemble, a probabilistic version that could provide uncertainty estimates for high-impact events.

Globally, the implications are profound. If a 1 km AI model can work in the Alps, it can work in the Rockies, the Alps, the Japanese Alps, or the Ethiopian Highlands. And because AI models are cheaper to run, they could democratize high-resolution forecasting for low- and middle-income countries, where weather warnings are often delayed or absent.

The era of AI weather forecasting is no longer coming. It’s here. And it’s already saving lives—one kilometer at a time.

Forecast Skill at +24h Lead Time

Correlation skill of Varda-single against KENDA-CH1 analysis at +24h lead time.

Forecast Skill at +24h Lead Time
LabelValue
Temperature (2m)0.85
Wind Speed (10m)0.78
Dew Point (2m)0.82
MSL Pressure0.91

The outsized effect of this stage is itself a symptom of how much the REA-L-CH1 reanalysis differs from the operational KENDA-CH1 analysis.

Comments (0)

No comments yet. Be the first to share your thoughts.