Fugu-MT 論文翻訳(概要): Automatic Generation of High-Performance RL Environments

論文の概要: Automatic Generation of High-Performance RL Environments

arxiv url: http://arxiv.org/abs/2603.12145v1
Date: Thu, 12 Mar 2026 16:45:47 GMT
ステータス: 翻訳完了
システム内更新日: 2026-03-13 14:46:26.225577
Title: Automatic Generation of High-Performance RL Environments
Title（参考訳）: 高性能RL環境の自動生成
Authors: Seth Karten, Rahul Dev Appapogu, Chi Jin,
Abstract要約: 複雑な強化学習環境を高性能な実装に変換するには、これまで何ヶ月もの専門技術が必要だった。計算コスト10ドルで意味論的に等価な高性能環境を創出する再利用可能なレシピを提案する。
参考スコア（独自算出の注目度）: 13.796920626646964
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Abstract: Translating complex reinforcement learning (RL) environments into high-performance implementations has traditionally required months of specialized engineering. We present a reusable recipe - a generic prompt template, hierarchical verification, and iterative agent-assisted repair - that produces semantically equivalent high-performance environments for <$10 in compute cost. We demonstrate three distinct workflows across five environments. Direct translation (no prior performance implementation exists): EmuRust (1.5x PPO speedup via Rust parallelism for a Game Boy emulator) and PokeJAX, the first GPU-parallel Pokemon battle simulator (500M SPS random action, 15.2M SPS PPO; 22,320x over the TypeScript reference). Translation verified against existing performance implementations: throughput parity with MJX (1.04x) and 5x over Brax at matched GPU batch sizes (HalfCheetah JAX); 42x PPO (Puffer Pong). New environment creation: TCGJax, the first deployable JAX Pokemon TCG engine (717K SPS random action, 153K SPS PPO; 6.6x over the Python reference), synthesized from a web-extracted specification. At 200M parameters, the environment overhead drops below 4% of training time. Hierarchical verification (property, interaction, and rollout tests) confirms semantic equivalence for all five environments; cross-backend policy transfer confirms zero sim-to-sim gap for all five environments. TCGJax, synthesized from a private reference absent from public repositories, serves as a contamination control for agent pretraining data concerns. The paper contains sufficient detail - including representative prompts, verification methodology, and complete results - that a coding agent could reproduce the translations directly from the manuscript.
Abstract（参考訳）: 複雑な強化学習(RL)環境を高性能な実装に変換するには、これまで何ヶ月もの専門技術が必要だった。本稿では, 汎用的なプロンプトテンプレート, 階層的検証, 反復的エージェント支援修復による, セマンティックに等価なハイパフォーマンス環境を計算コスト$10で生成する再利用可能なレシピを提案する。 5つの環境にまたがる3つの異なるワークフローを示します。 EmuRust(ゲームボーイエミュレータのラスト並列化による1.5倍のPPOスピードアップ)と、最初のGPU並列ポケモンバトルシミュレータ(500M SPSランダムアクション、15.2M SPS PPO; 22,320x over the TypeScript参照)であるPokeJAXである。 MJX (1.04x) と 5x over Brax (HalfCheetah JAX) のスループットパリティ、42x PPO (Puffer Pong) である。新しい環境の作成: TCGJaxは、最初のデプロイ可能なJAX Pokemon TCGエンジン(717K SPSランダムアクション、153K SPS PPO; 6.6x over the Python reference)で、Webで抽出された仕様から合成された。 2億のパラメータで、環境オーバーヘッドはトレーニング時間の4%以下になる。階層的検証(プロパティ、インタラクション、ロールアウトテスト)は5つの環境すべてにおいて意味論的等価性を確認する。 TCGJaxは、公開リポジトリにないプライベートレファレンスから合成され、データに関する事前訓練を行うエージェントの汚染制御として機能する。この論文には、代表的プロンプト、検証手法、完全な結果を含む十分な詳細が含まれており、コーディングエージェントは原稿から直接翻訳を再現できる。

論文の概要: Automatic Generation of High-Performance RL Environments

関連論文リスト