論文の概要: Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines
- arxiv url: http://arxiv.org/abs/2607.21173v1
- Date: Thu, 23 Jul 2026 10:59:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-24 18:26:25.376963
- Title: Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines
- Title(参考訳): 実行型因果研究パイプラインの自動合成と逆検証
- Abstract要約: 本稿では人工知能(AI)に基づく疫学研究アシスタント(ARA)フレームワークについて紹介する。
ARAはプロトコルの構築、合成データ生成、対角検証を統合パイプラインに統合する。
我々は、自動因果推論ベンチマーク上でARAを評価し、識別戦略、因果量、処理と結果変数、生成したコードと承認されたプロトコル間の整合性を評価する。
- 参考スコア(独自算出の注目度): 5.193393225545891
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes successfully yet relies on invalid causal assumptions. We present the Artificial Intelligence (AI)-based Epidemiology Research Assistant (ARA), a framework that makes these failures visible by explicitly encoding causal design principles, study-specific assumptions, and methodological constraints. ARA integrates protocol construction, synthetic data generation, and adversarial validation into a unified pipeline. The framework translates natural language research questions into structured causal protocols and executable analysis code by first constructing a protocol and then generating synthetic datasets using Structural Causal Models (SCMs) with known ground-truth effects. This synthetic-data step can also support pipeline development when access to confidential data, such as medical data, is restricted. The generated analysis is then evaluated under controlled violations of identification assumptions. We evaluate ARA on the Automated Causal Reasoning Benchmark, assessing recovery of identification strategies, causal quantities, treatment and outcome variables, and consistency between generated code and approved protocol. Protocol construction and adversarial validation did not consistently improve numerical agreement with benchmark estimates compared with standard LLM-based generation. However, they changed the failure mode: instead of silently returning causal estimates, ARA often surfaced protocol concerns, diagnostic failures, incomplete inference, or downgraded non-causal interpretations. These findings suggest that validity-first automated science systems should be evaluated not only by answer accuracy, but also by whether they indicate when causal claims are unwarranted.
- Abstract(参考訳): 自動的な研究システムは経験的分析を加速することを約束するが、サイレントな失敗をしがちである。
本稿では,人工知能(AI)に基づく疫学研究アシスタント(ARA)について紹介する。このフレームワークは,因果設計原則,研究固有の仮定,方法論的制約を明示的に符号化することで,これらの障害を可視化する。
ARAはプロトコルの構築、合成データ生成、対角検証を統合パイプラインに統合する。
このフレームワークは、自然言語研究の質問を、まずプロトコルを構築し、次に構造因果モデル(Structure Causal Models, SCM)を用いて、既知の地道効果を持つ合成データセットを生成することによって、構造化因果プロトコルと実行可能な解析コードに翻訳する。
この合成データステップは、医療データなどの機密データへのアクセスが制限された場合にもパイプライン開発をサポートすることができる。
生成した解析は、識別前提の制御された違反の下で評価される。
我々は、自動因果推論ベンチマーク上でARAを評価し、識別戦略、因果量、処理と結果変数、生成したコードと承認されたプロトコル間の整合性を評価する。
プロトコル構築と逆検証は、標準LLM生成と比較して、ベンチマーク推定との数値一致を一貫して改善しなかった。
しかし、彼らは失敗モードを変更した: 因果推定を静かに返す代わりに、ARAはしばしばプロトコルの懸念、診断の失敗、不完全な推論、あるいは非因果解釈の低下を表面化した。
これらの結果から, 正当性優先型自動科学システムは, 正解精度だけでなく, 因果的主張が不適切であるかどうかによって評価されるべきであることが示唆された。
関連論文リスト
- CausalArena: Benchmarking Causal Discovery in the Foundation Model Era [47.81267771630766]
CausalArenaは共通のプロトコルの下で因果発見のための統一されたベンチマークである。
あるベンチマークシステムにおける強いパフォーマンスは、他のベンチマークシステムに確実に移行しないことを示す。
論文 参考訳(メタデータ) (2026-09-10T17:53:15Z) - AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation [48.55515767840701]
AIスロープ(AI slop)は、アーティファクト、幻覚的脆弱性、可視だが正しくないパッチ、セマンティックに再パッケージされたバグレポートである。
本稿では,実証的証拠を調査し,統一メカニズムを特定し,信頼に値するトリアージへの道筋をたどる。
論文 参考訳(メタデータ) (2026-08-26T11:48:38Z) - CasualSynth: Generating Structurally Sound Synthetic Data [44.80087038178069]
大言語モデル(LLM)は、現実的な合成データを生成するが、その出力がターゲットドメインを管理する因果的メカニズムを尊重することを保証しない。
本稿では,意味的実現から因果構造の生成を分離するフレームワークCausal Synthを紹介し,因果的妥当性と言語学的にリッチな合成データを生成する。
論文 参考訳(メタデータ) (2026-05-17T16:21:01Z) - Federated Semantic Knowledge Graphs for Laboratory Workflows: A Structured Expert Elicitation Methodology Demonstrated Through Bioanalytical Workflow Twins [0.0]
医薬・生物医学研究の研究室は、かなりの暗黙の知識をコードしている。
本稿では,この知識を収集・クエリするために,構造化された専門知識抽出手法と連合セマンティック知識グラフ(SKG)アーキテクチャを提案する。
このアーキテクチャは、AI研究所のエージェントが人間の判断が不可能な場所で、障害を検出するのではなく、実行資産がマスクされるようなクエリ可能な表現を欠いているセマンティックワールドモデルを提供する。
論文 参考訳(メタデータ) (2026-05-15T18:44:52Z) - LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation [75.05397479715576]
大規模言語モデル(LLM)とエージェントは有望な進歩を示しているが、その真の能力と失敗モードは未だ不明である。
CプログラムのためのLCMおよびエージェントベースの形式仕様生成に関する、最初の体系的および汚染に配慮した研究を提案する。
論文 参考訳(メタデータ) (2026-05-02T11:31:33Z) - Ontology-Aware Design Patterns for Clinical AI Systems: Translating Reification Theory into Software Architecture [0.0]
臨床AIシステムは、ドキュメント、請求インセンティブ、用語の断片化によって構造的に歪められた健康データを定期的に訓練する。
本稿では,Gang-of-Fourパターン言語における7つの設計パターンを提案する。
論文 参考訳(メタデータ) (2026-04-02T06:05:11Z) - Dynamic analysis enhances issue resolution [53.50448142467294]
DAIRA(Dynamic Analysis-enhanced Issue Resolution Agent)は、エージェントの推論サイクルに動的解析を組み込む自動修復フレームワークである。
テストトレース駆動の方法論によって駆動されるDAIRAは、軽量モニタを使用して重要なランタイムデータを抽出する。
Gemini 3 Flash Previewを使用すると、DAIRAは新たな最先端(SOTA)パフォーマンスを確立し、SWE-bench Verifiedデータセットで79.4%の解像度を達成する。
論文 参考訳(メタデータ) (2026-03-23T14:48:54Z) - Procrustean Bed for AI-Driven Retrosynthesis: A Unified Framework for Reproducible Evaluation [0.0]
RetroCastは、異種モデルの出力を共通スキーマに標準化する統合評価スイートである。
我々は、新しい標準ベンチマークスイートを用いて、検索ベースおよびシーケンスベースの主要なアルゴリズムを評価する。
論文 参考訳(メタデータ) (2025-12-08T01:26:39Z) - A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code [49.09545217453401]
Propensity Smelly Score (PSC) は、特定の臭いの種類を生成する確率を推定する計量である。
我々は、生成戦略、モデルサイズ、モデルアーキテクチャ、および生成したコードの構造特性をいかに形成するかを識別する。
PSCは、開発者がモデルの振る舞いを解釈し、コード品質を評価するのに役立つ。
論文 参考訳(メタデータ) (2025-11-19T19:18:28Z) - CALM: A Causal Analysis Language Model for Tabular Data in Complex Systems with Local Scores, Conditional Independence Tests, and Relation Attributes [15.298086464296235]
観測データからの因果発見は生物学のような科学分野に不可欠である。
制約ベースのアプローチやスコアベースのアプローチを含む既存の手法は、重大な制限に直面している。
本稿では,表データに特化して設計された新しい因果解析言語CALMを紹介する。
論文 参考訳(メタデータ) (2025-10-10T20:19:20Z) - Neural Causal Models for Counterfactual Identification and Estimation [62.30444687707919]
本稿では,ニューラルモデルによる反事実文の評価について検討する。
まず、神経因果モデル(NCM)が十分に表現可能であることを示す。
第2に,反事実分布の同時同定と推定を行うアルゴリズムを開発する。
論文 参考訳(メタデータ) (2022-09-30T18:29:09Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。