論文の概要: When LLMs Sit Above Diagnostic Tools: Unrealized Complementarity in Industrial Fault Diagnosis
- arxiv url: http://arxiv.org/abs/2610.05031v1
- Date: Sun, 04 Oct 2026 07:51:31 GMT
- ステータス: 情報取得中
- システム内更新日: 2026-10-06 21:34:58.938524
- Title: When LLMs Sit Above Diagnostic Tools: Unrealized Complementarity in Industrial Fault Diagnosis
- Title(参考訳): LLMが診断ツールの上位に立つとき:産業的故障診断における非現実的な相補性
- Abstract要約: 診断データセットでは、矛盾する外部情報が最初に正しいLCM判定を覆した。
より強力なスタンドアロンソースよりも暗黙的なLLM統合に一貫した優位性を示した者はいない。
統合された出力は295の専門的修正のうち140点(47.5%)を見逃したが、99点のうち15点が当初正しいLCM判定(15.2%)を失った。
- 参考スコア(独自算出の注目度): 19.651323641444726
- License:
- Abstract: Large language models are increasingly used as integration layers above specialized tools, but a stronger component does not necessarily produce a stronger combined system. Across five diagnostic datasets (bearing vibration, process monitoring, semiconductor equipment), we study whether an LLM can reliably use external diagnostic information; paired repeat calls separate advice effects from output instability. In all five, conflicting external information overturned initially correct LLM judgments. Among the four datasets with direct integration comparisons, none showed a consistent advantage for implicit LLM integration over the stronger standalone source. On a Tennessee Eastman confirmation set whose protocol was fixed before evaluation, unaided accuracy was 64.67%, implicit LLM-specialist integration 77.43%, and the specialist alone 83.33%. Specialist information improved the LLM by 12.8 points (95% interval 9.7 to 15.9), yet the integrated output stayed 5.9 points below the specialist (95% interval -12.0 to -0.7). A two-source selector oracle reached 92.76%, indicating complementarity that the integrated output did not fully realize. The integrated output missed 140 of 295 specialist corrections (47.5%) but lost 15 of 99 initially correct LLM judgments (15.2%). The deficit remained under prompt and specialist sensitivity analyses. Among CWRU cases solved under both evidence presentations, task-aligned physical evidence yielded lower estimates of susceptibility to incorrect advice in six of seven models (five intervals excluding zero); higher reasoning effort gave no reliable reduction in five models, and a separate four-model TEP analysis gave no clear evidence that it resolves the integration problem. Source quality and integration quality should be evaluated separately: an integration layer should be compared with its stronger standalone component, not only with the unaided LLM.
- Abstract(参考訳): 大規模な言語モデルは、専門ツール以上の統合レイヤとしてますます使われていますが、より強力なコンポーネントは必ずしもより強力な統合システムを生み出していません。
5つの診断データセット(振動, プロセス監視, 半導体装置)にわたって, LLMが外部診断情報を確実に利用できるかどうかを検討した。
これら5つの中で、矛盾する外部情報が当初正しいLCM判断を覆した。
直接統合比較を行う4つのデータセットのうち、強いスタンドアロンソースよりも暗黙的なLLM統合に一貫した優位性は示されていない。
評価の前にプロトコルが修正されたテネシー・イーストマンの確認セットでは、正確性は64.67%、暗黙のLSM-スペシャリスト統合は77.43%、スペシャリストは83.33%であった。
専門的な情報はLCMを12.8ポイント(95%間隔9.7から15.9まで)改善したが、統合された出力は専門家の5.9ポイント(95%間隔-12.0から-0.7まで)に留まった。
2ソースセレクタのオラクルは92.76%に達し、統合された出力が完全には実現しなかったことの相補性を示している。
統合出力は295の専門的修正点のうち140点(47.5%)を欠いたが、99点のうち15点が当初正しいLCM判定点(15.2%)を失った。
赤字は、迅速かつ専門的な感度分析の下にとどまった。
両エビデンスで解決されたCWRUの事例のうち、タスクアラインな物理的証拠は、7つのモデルのうち6つのモデル(0を除く5つの区間)において誤ったアドバイスに対する感受性の低い推定を導いた。
ソースの品質と統合品質は別々に評価されるべきである。統合レイヤは、未発表のLCMだけでなく、その強力なスタンドアロンコンポーネントと比較されるべきである。
関連論文リスト
- Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies [17.85382282659117]
ORADS(OvarianAd Reportingal and Data Fable System)分類における大規模言語モデル(LLM)の性能と推論戦略を比較した。
機能ベースのハイブリッドアーキテクチャは、ほぼ最高の性能を示し、99.2%の精度(390の387)と参照標準との一致を実現した。
論文 参考訳(メタデータ) (2026-08-24T10:04:17Z) - An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography [1.3023698730561286]
大きな言語モデル(LLM)は、医学的画像解釈において有望であるが、幻覚、限られた精度、実行間不整合に悩まされている。
我々は、眼底写真から緑内障を検出するための特殊なディープラーニングツールとLLMを統合したエージェントAIフレームワークを開発し、検証した。
論文 参考訳(メタデータ) (2026-08-07T17:33:12Z) - LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data [2.1443570696048906]
大規模言語モデル(LLM)は、構造化された臨床データにますます適用される。
Qwen 2.5 7B と XGBoost を比較する。
論文 参考訳(メタデータ) (2026-06-17T18:49:44Z) - The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents [0.0]
2エージェント統合の精度は、詳細が削除されると58%から25%に低下する。
因子的回復実験により、完全な仕様を復元するだけで、単一エージェントの天井が回復することが示された。
このギャップは単に隠された情報の結果ではなく、共有された決定なしに互換性のあるコードを生成することの難しさを反映している。
論文 参考訳(メタデータ) (2026-03-25T13:18:26Z) - PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology [48.732366302949515]
大規模言語モデル(LLM)は、標準化された検査において専門家レベルの性能を達成したが、複数の選択精度は現実の臨床的有用性や安全性を十分に反映していない。
我々は、未確認患者の質問に対して、専門家のルーブリックを作成するための、ループ内人間パイプラインを開発した。
LLM-as-a-judge フレームワークを用いて,22のプロプライエタリおよびオープンソース LLM の評価を行い,臨床完全性,事実精度,Web-search 統合について検討した。
論文 参考訳(メタデータ) (2026-03-02T00:50:39Z) - LLM2: Let Large Language Models Harness System 2 Reasoning [65.89293674479907]
大規模言語モデル(LLM)は、無数のタスクにまたがって印象的な機能を示してきたが、時には望ましくない出力が得られる。
本稿では LLM とプロセスベースの検証器を組み合わせた新しいフレームワーク LLM2 を紹介する。
LLMs2は妥当な候補を生成するのに責任を持ち、検証者は望ましい出力と望ましくない出力を区別するためにタイムリーなプロセスベースのフィードバックを提供する。
論文 参考訳(メタデータ) (2024-12-29T06:32:36Z) - Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios [49.53589774730807]
マルチモーダル大規模言語モデル(MLLM)は近年,視覚的質問応答から映像理解に至るまでのタスクにおいて,最先端のパフォーマンスを実現している。
12件のオープンソースMLLMが, 単一の偽装キューを受けた65%の症例において, 既往の正解を覆した。
論文 参考訳(メタデータ) (2024-11-05T01:11:28Z) - Closing the gap between open-source and commercial large language models for medical evidence summarization [20.60798771155072]
大規模言語モデル(LLM)は、医学的証拠の要約において大きな可能性を秘めている。
最近の研究は、プロプライエタリなLLMの応用に焦点を当てている。
オープンソースのLLMは透明性とカスタマイズを向上するが、そのパフォーマンスはプロプライエタリなものに比べて低下する。
論文 参考訳(メタデータ) (2024-07-25T05:03:01Z) - How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts [54.07541591018305]
提案するMAD-Benchは,既存のオブジェクト,オブジェクト数,空間関係などの5つのカテゴリに分割した1000の試験サンプルを含むベンチマークである。
我々は,GPT-4v,Reka,Gemini-Proから,LLaVA-NeXTやMiniCPM-Llama3といったオープンソースモデルに至るまで,一般的なMLLMを包括的に分析する。
GPT-4oはMAD-Bench上で82.82%の精度を達成するが、実験中の他のモデルの精度は9%から50%である。
論文 参考訳(メタデータ) (2024-02-20T18:31:27Z) - "Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation [90.09260023184932]
Retrieval-Augmented Generation (RAG) は、外部の知識源を活用して、事実の幻覚を減らすことで、Large Language Model (LLM) を出力する。
NoMIRACLは18言語にまたがるRAGにおけるLDM堅牢性を評価するための人為的アノテーション付きデータセットである。
本研究は,<i>Halucination rate</i>,<i>Halucination rate</i>,<i>Halucination rate</i>,<i>Sorucination rate</i>,<i>Sorucination rate</i>,<i>Sorucination rate</i>,<i>Sorucination rate</i>,<i>Sorucination rate</i>,<i>Sorucination rate</i>,<i>Sr。
論文 参考訳(メタデータ) (2023-12-18T17:18:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。