A New AI Framework Can Rewrite the Past — And Improve Weather Forecasts

40% — that’s the improvement in forecast accuracy DAWIS achieves over traditional filters in chaotic, high-dimensional systems like the atmosphere, according to new research (Wikingsson et al., 2026). But the real shock isn’t the number. It’s how it works: by revising past estimates when new data arrives, like a detective who rethinks yesterday’s alibi after today’s fingerprint match. For decades, weather models and other forecasting systems have been locked into their earlier guesses. Once a state is estimated, it’s treated as fixed history. But what if that history was wrong? Errors compound, forecasts drift, and by day five, even the best models can be off by hundreds of miles. DAWIS — short for Data Assimilation with Windowed Inverse Sampling — breaks this chain. It doesn’t just predict the future. It rewrites the past, dynamically refining earlier states as new observations come in. And it does so within a single, unified framework that handles filtering, smoothing, and forecasting — without needing separate models for each task.
This isn’t just incremental progress. It’s a conceptual shift in how machines reason about time, uncertainty, and evidence. And its implications stretch far beyond meteorology, into climate modeling, oceanography, neuroscience, and any field where we must infer hidden states from noisy, incomplete data.
The Science
At the heart of modern forecasting lies data assimilation (DA) — the process of blending prior knowledge (like a physics-based weather model) with real-world observations (satellite readings, ground sensors) to estimate the true state of a system. In weather prediction, this happens every six hours: models ingest millions of data points to correct their trajectory.
But traditional DA methods face a fundamental limitation: they’re causal. They move forward in time, updating the present based on the past and current data, but never revisiting old estimates. Filters like the Ensemble Kalman Filter (EnKF) or particle filters condition on a fixed history. Once a state is estimated, it’s locked in. This creates a ratchet effect: errors accumulate, and no future observation can undo them.
Smoothing methods — which use future data to refine past states — exist, but they’re often separate, computationally expensive, and not integrated into real-time forecasting. Hybrid approaches like 4D-Var optimize over a fixed window, but rely on linear approximations and struggle with non-Gaussian uncertainty.
Enter generative models. Recently, flow- and diffusion-based models have emerged as powerful tools for sampling from complex, high-dimensional distributions. These models learn a “prior” — a statistical representation of how states typically evolve — and can generate new samples by gradually denoising from randomness to structure.
The key innovation in DAWIS is to apply this idea not to single states, but to windows of consecutive states. Instead of learning a prior over individual snapshots, DAWIS learns a multitask stochastic interpolant over sequences of $w+1$ states. Each state in the window has its own flow time — a parameter controlling how much it’s been denoised. This allows independent control over how strongly each state is held fixed, revised, or regenerated from scratch.
The method builds on inverse sampling: first, the current estimate is partially “noised up” to an intermediate latent representation (the turning point), then regenerated under guidance from new observations. This inversion step is what enables revision: by not starting from pure noise, DAWIS retains temporal consistency; by not starting from the raw previous estimate, it allows correction.
Critically, DAWIS unifies three DA tasks — filtering, fixed-lag smoothing, and block smoothing — under one framework. The same trained model can switch between modes simply by adjusting the turning points $\bm{\tau}_{\min}$. And in its DAWIS-Joint variant, it can even generate forecasts within the assimilation cycle, eliminating the need for a separate forecasting model.
What They Found
The authors tested DAWIS on two challenging nonlinear dynamical systems: the Stochastic Quasigeostrophic (SQG) model, a simplified representation of atmospheric turbulence, and the Lorenz-96 system, a classic benchmark for chaotic dynamics. Observations were sparse (only 10–30% of state variables observed), noisy, and nonlinearly related to the true state — mimicking real-world conditions.
The results were striking. Across multiple metrics and configurations, DAWIS outperformed state-of-the-art baselines.
In filtering tasks, DAWIS reduced the Continuous Ranked Probability Score (CRPS) — a measure of probabilistic forecast accuracy — by up to 40% compared to DAISI, a leading flow-based filter (Andrae et al., 2026).
CRPS Reduction in Filtering Tasks
DAWIS achieves up to 40% lower CRPS than DAISI across multiple configurations in the SQG model.
| Label | Value |
|---|---|
| DAISI | 0.85 |
| DAWIS Filter | 0.51 |
This means not only better point estimates, but more reliable uncertainty quantification: DAWIS knows when it’s uncertain.
For smoothing, DAWIS achieved a 32% lower Root Mean Square Error (RMSE) than the Smoothing Diffusion Assimilator (SDA), a recent diffusion-based smoother (Rozet & Louppe, 2023).
RMSE Reduction in Smoothing Tasks
DAWIS Block Smoother reduces RMSE by 32% compared to SDA in the SQG Sparse setting.
| Label | Value |
|---|---|
| SDA | 1.42 |
| DAWIS Block Smoother | 0.96 |
Even more impressively, the DAWIS Lagged Smoother — which produces smoothed estimates as a byproduct of filtering — matched or exceeded dedicated smoothing methods, despite operating in real time.
and
show side-by-side comparisons of filtering and smoothing performance, with DAWIS consistently recovering finer spatial structures and sharper gradients in the SQG model, even under extreme noise and sparsity.
One of the most compelling demonstrations was in the block smoothing regime. Here, DAWIS iteratively refined long trajectories by updating overlapping windows of states. Unlike traditional block Gibbs samplers that resample interiors from scratch, DAWIS inverts and regenerates, preserving more information from previous cycles. This led to faster convergence and higher final accuracy. After just five iterations, DAWIS reduced RMSE by 28% over the initial filtered trajectory — a level of refinement previously unattainable in high-dimensional systems.
Perhaps most surprisingly, DAWIS-Joint — the variant that generates forecasts internally — performed on par with models using external simulators. This suggests that the multitask interpolant doesn’t just learn a static prior; it captures the dynamics of the system well enough to replace a dedicated forecast model.
Why This Changes Things
The implications of DAWIS go far beyond beating benchmarks. It represents a new paradigm for how AI systems can reason about time and evidence.
Consider weather prediction. Today’s operational models use a “cold start” every six hours: they ingest data, run a 4D-Var or EnKF cycle, and launch a new forecast. But the analysis at hour zero is still constrained by the model’s prior trajectory. If the model drifted off course in the previous cycle, that error persists. DAWIS could change that. By allowing past states to be revised, it could correct for model bias, sensor drift, or unexpected events — like a sudden storm front — even after the fact.
In climate modeling, where we seek to reconstruct past states from proxy data (ice cores, tree rings), DAWIS-style block smoothing could provide more accurate paleoclimate reconstructions. Unlike traditional methods that assume fixed priors, DAWIS can iteratively refine centuries-long trajectories, propagating information from sparse, noisy observations across time.
But the impact isn’t limited to geophysics. In neuroscience, researchers use DA to infer brain states from EEG or fMRI data. These signals are indirect, noisy, and nonlinearly related to neural activity. A method that can revise past estimates as new data arrives could lead to more accurate brain-computer interfaces or epilepsy prediction systems.
In robotics, autonomous systems must maintain a belief about their environment. Current SLAM (Simultaneous Localization and Mapping) algorithms struggle when sensor data contradicts earlier assumptions. DAWIS offers a path to lifelong correction — a robot that can realize it misread a hallway three minutes ago and update its entire map accordingly.
The unified nature of DAWIS is also transformative. Today, teams maintain separate models for forecasting, filtering, and smoothing — each with its own training data, hyperparameters, and failure modes. DAWIS collapses these into one. This isn’t just more efficient; it’s more coherent. The same statistical assumptions govern all tasks, reducing inconsistencies and improving trust.
And by eliminating the need for an external forecast model in DAWIS-Joint, the framework opens the door to model-free DA. In domains where physics-based simulators are inaccurate or unavailable — like financial markets or social dynamics — DAWIS could still provide high-quality state estimates purely from data.
What’s Next
DAWIS is not a panacea. The method assumes the multitask interpolant is well-trained and that the guidance approximation (here, MMPS) is accurate. In highly chaotic systems, small errors in the score function could amplify during inversion and regeneration. The authors note that performance degrades when the assimilation window $w$ is too short to capture system memory, or too long to train effectively.
Scaling to operational weather models — with billions of state variables — remains a challenge. While flow-based models are more efficient than particle filters, training a multitask interpolant on global atmospheric data would require massive compute and careful architecture design. The current experiments used state dimensions in the thousands; real weather models operate in the millions.
Another open question is how to set the turning points $\bm{\tau}_{\min}$ adaptively. In the paper, they’re fixed or tuned offline. But in practice, some states may need more revision than others — a hurricane eye versus open ocean. Learning to control the “plasticity” of each state slot could further improve performance.
There’s also the question of uncertainty calibration. While DAWIS produces ensemble forecasts, it’s unclear how well the spread matches true error — a critical issue for decision-making in high-stakes domains like disaster response.
Still, the direction is clear. DAWIS demonstrates that AI systems don’t have to be slaves to causality. They can be retrodictive as well as predictive. They can learn from their mistakes, even after the fact.
As we build more complex models of the world — from digital twins of cities to AI-driven climate interventions — the ability to revise the past may become as important as predicting the future. DAWIS isn’t just a better filter. It’s a step toward machines that think more like scientists: always questioning, always updating, always learning.