Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control
- URL: http://arxiv.org/abs/2512.23292v2
- Date: Tue, 06 Jan 2026 02:29:00 GMT
- Title: Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control
- Authors: Yoonpyo Lee, Kazuma Kobayashi, Sai Puppala, Sajedul Talukder, Seid Koric, Souvik Chakraborty, Syed Bahauddin Alam,
- Abstract summary: Recent benchmarks show that vision-language models achieve only 50-53% accuracy on basic quantitative physics tasks.<n> Perception-centric architectures optimize parameter-space imitation, whereas safety-critical control demands outcome-space guarantees.<n>We present a fundamentally different pathway toward domain-specific foundation models by introducing compact language models operating as Agentic Physical AI.
- Score: 3.9610256846446554
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The prevailing paradigm in AI for physical systems, scaling general-purpose foundation models toward universal multimodal reasoning, confronts a fundamental barrier at the control interface. Recent benchmarks show that even frontier vision-language models achieve only 50-53% accuracy on basic quantitative physics tasks, behaving as approximate guessers that preserve semantic plausibility while violating physical constraints. This input unfaithfulness is not a scaling deficiency but a structural limitation. Perception-centric architectures optimize parameter-space imitation, whereas safety-critical control demands outcome-space guarantees over executed actions. Here, we present a fundamentally different pathway toward domain-specific foundation models by introducing compact language models operating as Agentic Physical AI, in which policy optimization is driven by physics-based validation rather than perceptual inference. We train a 360-million-parameter model on synthetic reactor control scenarios, scaling the dataset from 10^3 to 10^5 examples. This induces a sharp phase transition absent in general-purpose models. Small-scale systems exhibit high-variance imitation with catastrophic tail risk, while large-scale models undergo variance collapse exceeding 500x reduction, stabilizing execution-level behavior. Despite balanced exposure to four actuation families, the model autonomously rejects approximately 70% of the training distribution and concentrates 95% of runtime execution on a single-bank strategy. Learned representations transfer across distinct physics and continuous input modalities without architectural modification.
Related papers
- Hybrid Model Predictive Control with Physics-Informed Neural Network for Satellite Attitude Control [2.7222301668137483]
Reliable spacecraft attitude control depends on accurate prediction of attitude dynamics.<n>For spacecraft with complex dynamics, obtaining accurate physics-based models can be difficult, time-consuming, or computationally heavy.<n>This work explores Physics-Informed Neural Networks (PINNs) for modeling spacecraft attitude dynamics.
arXiv Detail & Related papers (2026-02-17T19:08:48Z) - MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer [55.982504915794514]
Cross-embodiment policies typically rely on shared-private architectures.<n>We introduce MOTIF for efficient few-shot cross-embodiment transfer.<n>We show that MOTIF significantly outperforms strong baselines in few-shot transfer scenarios.
arXiv Detail & Related papers (2026-02-14T13:21:40Z) - THOR: A Versatile Foundation Model for Earth Observation Climate and Society Applications [9.852915112122567]
THOR is a "computeadaptive" foundation model that solves both input heterogeneity and deployment rigidity.<n>We pre-train THOR with a novel randomized patch and input image size strategy.<n>This allows a single set of pre-trained weights to be deployed at inference with any patch size, enabling a dynamic trade-off between computational cost and feature resolution without retraining.
arXiv Detail & Related papers (2026-01-22T14:38:00Z) - Benchmarking neural surrogates on realistic spatiotemporal multiphysics flows [18.240532888032394]
We present REALM (REalistic AI Learning for Multiphysics), a rigorous benchmarking framework designed to test neural surrogates on challenging, application-driven reactive flows.<n>We benchmark over a dozen representative surrogate model families, including spectral operators, convolutional models, Transformers, pointwise operators, and graph/mesh networks.<n>We identify three robust trends: (i) a scaling barrier governed jointly by dimensionality, stiffness, and mesh irregularity, leading to rapidly growing rollout errors; (ii) performance primarily controlled by architectural inductive biases rather than parameter count; and (iii) a persistent gap between nominal accuracy metrics and physically
arXiv Detail & Related papers (2025-12-21T05:04:13Z) - Towards a Science of Scaling Agent Systems [79.64446272302287]
We formalize a definition for agent evaluation and characterize scaling laws as the interplay between agent quantity, coordination structure, modelic, and task properties.<n>We derive a predictive model using coordination metrics, that cross-validated R2=0, enabling prediction on unseen task domains.<n>We identify three effects: (1) a tool-coordination trade-off: under fixed computational budgets, tool-heavy tasks suffer disproportionately from multi-agent overhead, and (2) a capability saturation: coordination yields diminishing or negative returns once single-agent baselines exceed 45%.
arXiv Detail & Related papers (2025-12-09T06:52:21Z) - Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning [70.56067503630486]
We argue that sixth-generation (6G) intelligence is not fluent token prediction but calibrated the capacity to imagine and choose.<n>We show that WM-MS3M cuts mean absolute error (MAE) by 1.69% versus MS3M with 32% fewer parameters and similar latency, and achieves 35-80% lower root mean squared error (RMSE) than attention/hybrid baselines with 2.3-4.1x faster inference.
arXiv Detail & Related papers (2025-11-04T17:22:22Z) - Flow marching for a generative PDE foundation model [0.0]
We propose Flow Marching, an algorithm that bridges neural operator learning with flow matching motivated by an analysis of error accumulation in physical dynamical systems.<n>We also introduce a Physics-Pretrained Variational Autoencoder (P2E) to embed physical trajectories into a compact latent space.<n>We curate a corpus of 2.5M trajectories across 12 distinct PDE families and train suites of P2Es and FMTs at multiple scales.
arXiv Detail & Related papers (2025-09-23T04:00:41Z) - Transition Models: Rethinking the Generative Learning Objective [68.16330673177207]
We introduce a continuous-time dynamics equation that analytically defines state transitions across any finite time interval.<n>This leads to a novel generative paradigm, Transition Models (TiM), which adapt to arbitrary-step transitions.<n>TiM achieves state-of-the-art performance, surpassing leading models such as SD3.5 (8B parameters) and FLUX.1 (12B parameters) across all evaluated step counts.
arXiv Detail & Related papers (2025-09-04T17:05:59Z) - Learning Robust Satellite Attitude Dynamics with Physics-Informed Normalising Flow [2.7222301668137483]
We investigate the benefits of incorporating Physics-Informed Neural Networks into the learning of spacecraft attitude dynamics.<n>We train several models on simulated data generated with the Basilisk simulator.<n>We find that PINN-based models consistently outperform their purely data-driven counterparts in terms of control accuracy and robustness.
arXiv Detail & Related papers (2025-08-11T10:50:49Z) - OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks [52.87238755666243]
We present OmniEAR, a framework for evaluating how language models reason about physical interactions, tool usage, and multi-agent coordination in embodied tasks.<n>We model continuous physical properties and complex spatial relationships across 1,500 scenarios spanning household and industrial domains.<n>Our systematic evaluation reveals severe performance degradation when models must reason from constraints.
arXiv Detail & Related papers (2025-08-07T17:54:15Z) - A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models [9.304845676825584]
We propose a novel adversarial training framework that integrates multiple attack strategies and advanced machine learning techniques.
Experiments conducted on real-world datasets, including CIFAR-10 and CIFAR-100, demonstrate that the proposed method significantly enhances model robustness.
arXiv Detail & Related papers (2024-10-18T23:47:46Z) - Physics-Integrated Variational Autoencoders for Robust and Interpretable
Generative Modeling [86.9726984929758]
We focus on the integration of incomplete physics models into deep generative models.
We propose a VAE architecture in which a part of the latent space is grounded by physics.
We demonstrate generative performance improvements over a set of synthetic and real-world datasets.
arXiv Detail & Related papers (2021-02-25T20:28:52Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.