FuguReport

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

Authors Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao
Affiliations Massachusetts Institute of Technology / National University of Singapore / Tsinghua University / The University of Texas at Austin / The Hong Kong University of Science and Technology / The University of Hong Kong / McGill University / Georgia Institute of Technology / Griffith University / NUIST
Categories Evaluation / Scientific Agent Evaluation / Benchmarking scientific agents on Earth systems, Method / Agent Reasoning / Evidence selection and transparent computation, Application / Natural Hazard Modeling / Agent performance on dynamic Earth events
License CC BY 4.0

Abstract Overview

EarthVerse is a benchmark designed to evaluate scientific agents conducting multi-source investigations across dynamic Earth systems and natural hazards. It comprises 405 reproducible tasks grounded in 199 documented historical events spanning 19 hazard families, where agents must navigate multi-file event packages, reconcile conflicting data sources, execute calculations, and document evidence provenance. The evaluation framework pairs fine-grained answer units with process rubrics to assess both answer correctness and the supporting investigative trajectory. Evaluating 25 model and agent configurations under a standardized protocol reveals key failure modes in evidence localization, tool usage, reasoning, and multi-step scientific execution.

Novelty

EarthVerse introduces an end-to-end benchmark that evaluates an agent's ability to discover, align, and maintain claim-evidence chains across heterogeneous, multi-source event packages. Unlike existing Earth science benchmarks that supply curated observations or designated data layers upfront, EarthVerse requires agents to autonomously identify relevant records and navigate multi-source misalignments.

Results

Across 25 evaluated systems, top mean answer-unit accuracy reaches 84.65%, yet the highest Strict@95 reliability is only 34.81%, highlighting a substantial gap between partial step competence and full investigation completion. Controlled interventions identify evidence localization as the primary performance bottleneck, where providing relevant-file and evidence-map oracles yields the only statistically significant Core score gains (+14.72 and +11.34 points, respectively).

Key Points

  1. EarthVerse establishes a reproducible benchmark of 405 package-scoped investigation tasks grounded in 199 real hazard events across 19 hazard categories.
  2. The evaluation framework jointly scores fine-grained scientific answer units and research trajectories to evaluate claim-evidence consistency without mandating a rigid tool path.
  3. Empirical results across 25 systems show that current models frequently solve isolated subtasks but struggle with end-to-end reliability, with evidence localization serving as the primary bottleneck.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.