← News
Tech for Good Tech for Good Frontiers

Can AI Learn What 6,000 Years of Ocean History Teaches Us?

Can AI Learn What 6,000 Years of Ocean History Teaches Us?
6,000 years Ocean history studied
Reproduces Spatial Patterns AI capability
Misses Deep Ocean Dynamics Limitation

The Science

Climate models are hungry beasts. Simulating the ocean's behavior over centuries requires supercomputers churning for weeks, vast institutional expertise, and energy budgets that make a small country's power grid sweat. But a new generation of AI models — climate emulators — promises something different: run the same simulation in minutes on a laptop, no physics PhD required.

The problem is trust. Weather forecasting AI has clear goals — predict whether it'll rain in London next Tuesday — and clear metrics to measure success. Climate emulators face murkier terrain. They need to capture not just tomorrow's weather, but how the ocean's behavior shifts over decades and centuries when orbital parameters change, when carbon dioxide rises, when ice sheets retreat. These are fundamentally different challenges, and the tools we use to validate weather models don't necessarily tell us whether a climate emulator will hold up when pushed to scenarios it hasn't seen before.

Adam Subel and Laure Zanna at NYU's Courant Institute have developed a clever solution: instead of testing AI ocean emulators on future warming scenarios (which push the models into territory they've never seen), they look to the past. Specifically, they use the midHolocene — a period roughly 6,000 years ago when Earth's orbital parameters were slightly different, causing a subtle but measurable shift in how sunlight hit the planet's surface.

The setup is elegant. They trained neural network emulators on data from a standard climate model (CESM2) running under pre-industrial control conditions — a stable climate with no changing forcings. They then asked a simple question: if you force this trained emulator with the surface conditions from the midHolocene, does it produce the ocean response you'd expect from a model actually run under midHolocene conditions? The answer reveals both the power and the limits of these AI tools.

What They Found

The emulators passed some tests with flying colors. When forced with midHolocene boundary conditions — the pattern of heat flowing into the ocean, the shear of winds on the surface, the incoming sunlight — the AI correctly reproduced the spatial structure of the ocean's response. Temperature anomalies appeared in the right places, followed the right geographic patterns, and correlated strongly with the ground-truth climate model runs (correlations of 0.75-0.95 across major ocean basins). The emulators captured the shifts in seasonal cycles, correctly modeling how the timing and amplitude of ocean temperature swings change when orbital parameters shift. They even reproduced changes in modes of natural variability — the Niño 3.4 index in the tropical Pacific and the Dipole Mode Index in the Indian Ocean both shifted in the right ways.

Figure 2: Comparison of the tropical Pacific seasonal SST anomaly changes between the midHolocene and piControl. We define SST as the uppermost layer (0-5m) and compute anomalies as the difference between the monthly and annual climatologies. We then compute the meridional average of these anomalies between 1.5∘​S−1.5∘​N1.5^{\circ}\mathrm{S}-1.5^{\circ}\mathrm{N}. A: difference between the midHolocene and piControl CESM2 numerical experiments. For B-D, we use the title to indicate the networks and rollout parameters, using abbreviated notation defined in Section 2.5.3. For each rollout, we use initial conditions and insolation from the same experiment listed for the boundary forcings (e.g., where we state 𝝉pi\boldsymbol{\tau}_{\mathrm{pi}}, we also use 𝚽pi[0]\boldsymbol{\Phi}_{\mathrm{pi}}^{[0]} and Ipi\operatorname{I}_{\mathrm{pi}}). B: difference between ℱmH\mathcal{F}_{\mathrm{mH}} rolled out for the midHolocene climate and ℱpi\mathcal{F}_{\mathrm{pi}} rolled out for the piControl climate. C: difference between ℱpi\mathcal{F}_{\mathrm{pi}} rolled out for the midHolocene and piControl climate. D: same difference as C, but for ℱmH\mathcal{F}_{\mathrm{mH}}. We evaluate the climatology over the final 50 years of each 100-year rollout. For each panel, we compute the correlation, ρ\rho, and RMSE with respect to the CESM2 difference in A.
Figure 2: Comparison of the tropical Pacific seasonal SST anomaly changes between the midHolocene and piControl. We define SST as the uppermost layer (0-5m) and compute anomalies as the difference between the monthly and annual climatologies. We then compute the meridional average of these anomalies between 1.5∘​S−1.5∘​N1.5^{\circ}\mathrm{S}-1.5^{\circ}\mathrm{N}. A: difference between the midHolocene and piControl CESM2 numerical experiments. For B-D, we use the title to indicate the networks and rollout parameters, using abbreviated notation defined in Section 2.5.3. For each rollout, we use initial conditions and insolation from the same experiment listed for the boundary forcings (e.g., where we state 𝝉pi\boldsymbol{\tau}_{\mathrm{pi}}, we also use 𝚽pi[0]\boldsymbol{\Phi}_{\mathrm{pi}}^{[0]} and Ipi\operatorname{I}_{\mathrm{pi}}). B: difference between ℱmH\mathcal{F}_{\mathrm{mH}} rolled out for the midHolocene climate and ℱpi\mathcal{F}_{\mathrm{pi}} rolled out for the piControl climate. C: difference between ℱpi\mathcal{F}_{\mathrm{pi}} rolled out for the midHolocene and piControl climate. D: same difference as C, but for ℱmH\mathcal{F}_{\mathrm{mH}}. We evaluate the climatology over the final 50 years of each 100-year rollout. For each panel, we compute the correlation, ρ\rho, and RMSE with respect to the CESM2 difference in A. Source: Adam Subel, Laure Zanna

But the emulators stumbled in crucial ways. They consistently underestimated the amplitude of their responses. The patterns were correct; the magnitudes were too small, by 30-70% depending on the region and metric. And while surface dynamics were handled reasonably well, the emulators completely failed to capture the slow, internally-driven evolution of the ocean interior — the deep circulation changes that unfold over centuries. This isn't a minor gap; it's the difference between understanding next year's weather versus next century's climate.

The researchers also tested simpler baseline models that mapped boundary conditions directly to ocean state — essentially asking: "what if the ocean just responds to what's at the surface, with no internal dynamics?" These baselines recovered some large-scale patterns near the surface, but missed the seasonal changes, the spatial patterns of variability, and almost everything at depth. The conclusion: capturing the full climate response requires the emulator to learn something about how the ocean actually works, not just correlate surface inputs with surface outputs.

Figure 7: Comparison of the depth-averaged potential temperature response over the upper 200m (A-C) and 200-1000m (D-F) when perturbing an emulator with climatological forcings from the midHolocene/piControl CESM2 experiments. A and D: difference between the CESM2 midHolocene and piControl experiments as an estimate of the true response. B and E: responses of the emulator trained on piControl data when perturbed with midHolocene boundary forcings, R⁡(ℱpi)R(\mathcal{F}_{\mathrm{pi}}). C and F: responses of the emulator trained on midHolocene data when perturbed with piControl boundary forcings, R⁡(ℱmH)R(\mathcal{F}_{\mathrm{mH}}). For each emulator response, we include the correlation, ρ\rho, and RMSE with respect to the difference between CESM2 experiments in A and D. Note that for R⁡(ℱmH)R(\mathcal{F}_{\mathrm{mH}}), we flip the sign of the true response before computing the correlation and RMSE. We stipple the CESM2 panels to mark regions where the reference response is not significant relative to internal variability (two-tailed Student’s tt-test at the 95% level, with internal variability estimated from ten 25-year chunks of each experiment).
Figure 7: Comparison of the depth-averaged potential temperature response over the upper 200m (A-C) and 200-1000m (D-F) when perturbing an emulator with climatological forcings from the midHolocene/piControl CESM2 experiments. A and D: difference between the CESM2 midHolocene and piControl experiments as an estimate of the true response. B and E: responses of the emulator trained on piControl data when perturbed with midHolocene boundary forcings, R⁡(ℱpi)R(\mathcal{F}_{\mathrm{pi}}). C and F: responses of the emulator trained on midHolocene data when perturbed with piControl boundary forcings, R⁡(ℱmH)R(\mathcal{F}_{\mathrm{mH}}). For each emulator response, we include the correlation, ρ\rho, and RMSE with respect to the difference between CESM2 experiments in A and D. Note that for R⁡(ℱmH)R(\mathcal{F}_{\mathrm{mH}}), we flip the sign of the true response before computing the correlation and RMSE. We stipple the CESM2 panels to mark regions where the reference response is not significant relative to internal variability (two-tailed Student’s tt-test at the 95% level, with internal variability estimated from ten 25-year chunks of each experiment). Source: Adam Subel, Laure Zanna

Perhaps most provocatively, Subel and Zanna showed that the emulators' responses to individual forcing components — heat flux, zonal wind stress, meridional wind stress — superimpose linearly. Run the emulator with just the heat flux changed, just the wind stress changed, or all components changed together, and the sum of the individual responses almost exactly equals the combined response. This linearity is a double-edged sword. It's useful: it means researchers can isolate which forcing component drives which response. But it's also a limitation — it suggests the emulator isn't learning the full complexity of how forcings interact.

Why This Changes Things

The midHolocene isn't just a convenient test case. It's a template for how we should be building and vetting climate AI. The standard approach has been to validate emulators on historical data and then apply them to future warming scenarios — pushing them far outside the distribution they were trained on. It's like training a self-driving car on California highways and then deploying it in Mumbai during monsoon season. The midHolocene, by contrast, provides an in-distribution challenge: the values are different, but they're not alien. The emulator hasn't seen these exact conditions, but it hasn't been asked to extrapolate beyond the physics it learned.

This matters because climate emulators are already being deployed for practical work. Researchers are using them to rapidly test hypotheses, explore perturbation experiments, and generate low-cost surrogates for expensive simulations. Before we trust these tools with decisions that affect billions of people, we need to understand what they're actually learning — and what they're missing.

Boundary Forcing Comparison: Heat Flux

Surface heat flux variability differs slightly between pre-industrial and midHolocene conditions, reflecting changed orbital parameters.

Boundary Forcing Comparison: Heat Flux
LabelValue
piControl91.17 W/m²
midHolocene93.36 W/m²

The finding that training metrics don't capture dynamical skill is particularly important. The standard loss function — mean squared error, how close the emulator gets to the right answer on average — tells you almost nothing about whether the emulator will respond correctly to a forcing it hasn't seen. Two emulators can score identically on MSE while one correctly captures seasonal phase shifts and the other doesn't. If we're selecting models based only on training metrics, we may be systematically choosing emulators that look good but perform poorly on the tasks we actually care about.

The linear superposition result is both reassuring and concerning. It's reassuring because it means the physics is reasonably well-behaved — you can decompose responses and recombine them. It's concerning because real climate dynamics involve nonlinear interactions: the way warming changes cloud cover, which changes albedo, which changes warming in a feedback loop. Linear superposition works for the forcings this emulator was tested on, but it may not hold for the more extreme perturbations that climate change demands we understand.

What's Next

The failures point toward research priorities. Capturing the ocean interior's slow dynamics may require a different training objective — something that explicitly rewards the deep circulation response, not just surface accuracy. The amplitude underestimation suggests the emulator is dampening too strongly, behaving too much like a passive system rather than one that can amplify and propagate signals. These are fixable problems, but they require moving beyond "does it match the training data" to "does it capture the physics that matters."

The midHolocene itself offers a broader testing ground. Subel and Zanna focused on the pre-industrial control and midHolocene experiments, but the paleoclimate record contains many other test cases: the Last Glacial Maximum 21,000 years ago, the Eocene with its ice-free Antarctica, the Younger Dryas with its abrupt climate swings. Each represents a different set of boundary conditions, a different stress test for the emulator's ability to generalize.

The bigger picture is about building trustworthy AI for climate science. These tools have enormous potential to democratize climate modeling, to let researchers anywhere run experiments that previously required access to national supercomputers. But potential and reliability aren't the same thing. The midHolocene gives us a controlled setting — a known answer key — for checking whether these emulators will hold up when we ask them questions we don't yet know the answers to. That's not just good science. It's the foundation for using AI to understand what happens next.