論文の概要: Data and Evaluation Closed-Loop for Model Capability Enhancement
- arxiv url: http://arxiv.org/abs/2606.28471v1
- Date: Fri, 26 Jun 2026 14:45:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-30 18:07:15.579473
- Title: Data and Evaluation Closed-Loop for Model Capability Enhancement
- Title(参考訳): モデル能力向上のためのデータとクローズドループ
- Authors: Zhixuan Li, Jiangan Yuan, Han Xu,
- Abstract要約: 評価は直感的ではなく,日常的,監査可能,実験的に検証可能であることを示す。
評価データ間推論は直感的ではなく,日常的,監査可能,実験的に検証可能であることを示す。
- 参考スコア(独自算出の注目度): 8.245706308161504
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compressing samples, prompts, decoding, and scoring rules into one noisy score. Practical optimization runs this backward: a failure is observed first, and the engineer must infer the corpus fix. The two sides speak incompatible vocabularies -- benchmark names and per-sample correctness versus data sources, domains, and quality labels -- so this inference is usually intuition, not method. We close this gap with the \emph{capability slice}: a group of evaluation samples sharing background condition, task type, solving operation, and output constraint -- precise enough to localize a single weakness yet stable enough to survive aggregation, unlike a benchmark name, too coarse, or a single sample, too noisy. Built around this unit, an evaluation taxonomy, a non-instruction data taxonomy, and mapping rules form a closed loop turning a benchmark-level failure into a targeted, testable data intervention. We test this loop on two case studies pulling in opposite directions. First, the loop rules the data out: continued pre-training drives BBH down by $-46.82\%$, but diagnosis traces this to a single masked \texttt{\textless EOS\textgreater} loss rather than weakened reasoning; restoring it recovers BBH to $66.44$, above the original checkpoint, without changing the data. Second, the loop rules the data in: a persistent math-reasoning weakness is decomposed by solving operation into specific failing combinations, and a weakness-targeted sampling procedure built from it lifts AIME2025/AIME2026 Pass@128 from $6.67$/$0.00$ to $26.67$ each. The same unmodified loop reaches opposite, correct verdicts in both cases, showing the evaluation-to-data inference can be routine, auditable, and experimentally validated rather than intuitive.
- Abstract(参考訳): モデル能力はLLMの事前トレーニングにおいて中心的な変数であるが、直接観察されることはない。データ形状は前方に形成され、評価は振り返りのみを示し、サンプルを圧縮し、プロンプト、デコードし、ルールを1つのノイズスコアにスコアする。
失敗を最初に観察し、エンジニアはコーパスの修正を推論しなければならない。
両者は互換性のない語彙 - ベンチマーク名とサンプルごとの正確さ -- データソース、ドメイン、品質ラベル -- を比較した場合、この推論は通常直感的であり、メソッドではない。我々はこのギャップを \emph{capability slice} – バックグラウンド条件、タスクタイプ、解決操作、出力制約を共有する評価サンプルのグループ – を閉じている。ベンチマーク名、粗すぎる、あるいはノイズの多い単一のサンプルとは異なり、単一の弱点をローカライズするのに十分安定している。
このユニットを中心に構築された評価分類、非命令データ分類、マッピングルールは、ベンチマークレベルの障害をターゲットとするテスト可能なデータ介入に変換するクローズドループを形成する。
このループを2つのケーススタディで検証し、反対方向に引っ張る。
継続事前トレーニング駆動BBHを$-46.82\%$で下げるが、診断は、データを変更せずにBBHを6.44ドルまで回復させるという、弱い推論ではなく、単一のマスク付き \texttt{\textless EOS\textgreater} 損失に辿り着く。
第2に、ループはデータを規定する: 永続的な数学推論の弱点は、特定の故障した組み合わせに演算を分解することで分解され、AIME2025/AIME2026 Pass@128を6.67$/0.00$から26.67$に引き上げる。
同じ修正されていないループが両方のケースで正解に達し、データ間の推測が直感的ではなく、ルーチン、監査可能、実験的に検証可能であることを示す。
関連論文リスト
- The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval [21.243421703047037]
単語境界の破損がLarge Language Modelベンチマークのターゲット情報の検出方法に与える影響について検討する。
単語に空白文字を挿入して断片に分解することで、LLMの精度は挿入率の増加とともにU字型曲線に従う。
論文 参考訳(メタデータ) (2026-05-08T03:26:42Z) - SpatialBench-UC: Uncertainty-Aware Evaluation of Spatial Prompt Following in Text-to-Image Generation [0.0]
SpaceBench-UCは、ペアの空間関係を再現可能な小さなベンチマークである。
ベンチマークパッケージ、バージョン付きプロンプト、ピン付き構成、サンプルごとのチェッカー出力、レポートテーブルをリリースします。
安定拡散1.5, SD 1.5 BoxDiff, SD 1.4 GLIGENの3つのベースラインについて検討した。
論文 参考訳(メタデータ) (2026-01-19T23:37:10Z) - Training on the Benchmark Is Not All You Need [52.01920740114261]
本稿では,複数選択肢の内容に基づいた簡易かつ効果的なデータ漏洩検出手法を提案する。
本手法は,モデルトレーニングデータや重みを使用せずに,グレーボックス条件下で動作可能である。
4つのベンチマークデータセットから35個の主要なオープンソースLCMのデータ漏洩度を評価する。
論文 参考訳(メタデータ) (2024-09-03T11:09:44Z) - Fact Checking Beyond Training Set [64.88575826304024]
本稿では,レトリバーリーダが,あるドメインのラベル付きデータに基づいてトレーニングし,別のドメインで使用する場合,性能劣化に悩まされることを示す。
本稿では,レトリバー成分を分散シフトに対して頑健にするための逆アルゴリズムを提案する。
次に、これらのデータセットから8つの事実チェックシナリオを構築し、モデルと強力なベースラインモデルのセットを比較します。
論文 参考訳(メタデータ) (2024-03-27T15:15:14Z) - Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models [25.022166664832596]
大規模言語モデル(LLM)におけるデータ汚染を簡易かつ効果的に検出する手法を提案する。
我々は、データ汚染検出を複数項目の質問としてフレーム化し、各インスタンスの3つの摂動バージョンが特定のデータセットパーティションからサブサンプリングされたクイズフォーマットを考案した。
生成された摂動は、元のデータセットインスタンスとともにDCQのオプションを形成し、提供されたオプションの選択を調節する余分なオプションを提供する。
論文 参考訳(メタデータ) (2023-11-10T18:48:58Z) - Semi-DETR: Semi-Supervised Object Detection with Detection Transformers [105.45018934087076]
半教師付き物体検出(SSOD)におけるDETRに基づくフレームワークの解析
本報告では,第1次変圧器を用いたエンド・ツー・エンド半教師対象検出器であるSemi-DETRについて述べる。
我々の手法は、最先端の手法をクリアマージンで上回る。
論文 参考訳(メタデータ) (2023-07-16T16:32:14Z) - Bias Mimicking: A Simple Sampling Approach for Bias Mitigation [57.17709477668213]
本稿では,新しいクラス条件サンプリング手法であるBias Mimickingを紹介する。
Bias Mimickingは、4つのベンチマークで3%の精度でサンプリングの精度を向上する。
論文 参考訳(メタデータ) (2022-09-30T17:33:00Z) - Double Perturbation: On the Robustness of Robustness and Counterfactual
Bias Evaluation [109.06060143938052]
テストデータセットを超えたモデル弱点を明らかにするための"ダブル摂動"フレームワークを提案する。
この枠組みを,モデルの頑健さと英語における反事実バイアスの分析に使用される2つの摂動に基づくアプローチに応用する。
論文 参考訳(メタデータ) (2021-04-12T06:57:36Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。