Aivora
arXivAI ResearchIntermediate

SciExam for ENSO: Evaluating AI Agents' Ability to Build Unsolved Climate Models

AI Agent 能建立氣候模型嗎?評估未知科學領域的「SciExam for ENSO」基準測試

2 min read
SciExam for ENSO: Evaluating AI Agents' Ability to Build Unsolved Climate Models
The 30-second version

Evaluating scientific AI is difficult when there is no pre-existing ground truth. The SciExam for ENSO benchmark challenges AI agents to build low-order stochastic models of the El Niño-Southern Oscillation within a 6-hour window, relying only on self-written frozen diagnostics for feedback. Out of 12 agent systems tested, 6 successfully built models that outperformed an established, human-published model. Intriguingly, these agent-generated models naturally aligned with competing theories in an ongoing scientific debate regarding ENSO's asymmetry.

Key points

01

Evaluation without known answers

Evaluates agents without known answers, using hidden graders to score models on statistical reproduction, unobserved variable recovery, and future forecasting.

02

Outperforming human baselines

Six out of twelve tested agent systems successfully generated climate models that scored higher than the baseline published human model.

03

Shedding light on open debates

The simplified structures of the top-performing agent models map directly onto two competing explanations in an active, unresolved scientific debate.

04

Robustness and anti-memorization

Control runs confirmed that the agents' high scores did not stem from memorizing past climate records, but rather from interactive feedback during the task.

How it works

SciExam for ENSO Workflow
Data inputLock diagnosticsOnly feedback sourceEvaluate forecastReal ObservationsWrite DiagnosticsFreeze DiagnosticsModel Building (6hr)Hidden Graders

Why it matters

This research demonstrates that AI agents are capable of genuine, open-ended scientific discovery rather than just solving problems with known answers. By independently formulating competitive physical models that align with ongoing academic debates, AI agents prove their potential to act as active collaborators in complex scientific fields like climate science and physics.

Who it affects

  • AI Researcher
  • AI Developer
  • Enterprise Leader

How to use it

  1. 1Automated Climate Modeling: Utilizing AI agents to rapidly prototype and build predictive models for specific atmospheric or oceanic phenomena.
  2. 2Scientific Hypothesis Exploration: Helping researchers discover and test underlying physical mechanisms by analyzing agent-generated model architectures.

Limitations & caveats

  • Currently limited to low-order stochastic modeling and has not yet been extended to high-resolution, global-scale general circulation models.
  • Performance is highly dependent on the quality of self-written diagnostics; poor initial diagnostics can lead to compounding failures in model building.

Related