EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
Abstract Overview
EarthVerse is a benchmark designed to evaluate scientific agents conducting multi-source investigations across dynamic Earth systems and natural hazards. It comprises 405 reproducible tasks grounded in 199 documented historical events spanning 19 hazard families, where agents must navigate multi-file event packages, reconcile conflicting data sources, execute calculations, and document evidence provenance. The evaluation framework pairs fine-grained answer units with process rubrics to assess both answer correctness and the supporting investigative trajectory. Evaluating 25 model and agent configurations under a standardized protocol reveals key failure modes in evidence localization, tool usage, reasoning, and multi-step scientific execution.
Novelty
EarthVerse introduces an end-to-end benchmark that evaluates an agent's ability to discover, align, and maintain claim-evidence chains across heterogeneous, multi-source event packages. Unlike existing Earth science benchmarks that supply curated observations or designated data layers upfront, EarthVerse requires agents to autonomously identify relevant records and navigate multi-source misalignments.
Results
Across 25 evaluated systems, top mean answer-unit accuracy reaches 84.65%, yet the highest Strict@95 reliability is only 34.81%, highlighting a substantial gap between partial step competence and full investigation completion. Controlled interventions identify evidence localization as the primary performance bottleneck, where providing relevant-file and evidence-map oracles yields the only statistically significant Core score gains (+14.72 and +11.34 points, respectively).
Key Points
- EarthVerse establishes a reproducible benchmark of 405 package-scoped investigation tasks grounded in 199 real hazard events across 19 hazard categories.
- The evaluation framework jointly scores fine-grained scientific answer units and research trajectories to evaluate claim-evidence consistency without mandating a rigid tool path.
- Empirical results across 25 systems show that current models frequently solve isolated subtasks but struggle with end-to-end reliability, with evidence localization serving as the primary bottleneck.