論文の概要: Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget
- arxiv url: http://arxiv.org/abs/2609.01108v1
- Date: Tue, 01 Sep 2026 11:48:16 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.610742
- Title: Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget
- Title(参考訳): TRACEのリプリケート--その閾値と粒子予算に対する実践者ガイド
- Authors: Alex Chadyuk, Alicia Zhang, Roy Kucukates,
- Abstract要約: TRACEは、事前訓練された自己回帰シーケンスモデルから因果グラフを読み出す。
正確な介入真理に対する平均F1は、語彙サイズ1000で0.90-0.91に達する。
- 参考スコア(独自算出の注目度): 0.10923877073891443
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: TRACE (Math & Lienhart, arXiv:2602.01135) reads causal graphs over event types out of a pretrained autoregressive sequence model by thresholding a per-position conditional-mutual-information estimate at a fixed tau. We independently replicate its headline synthetic result: with tau selected on a validation split, mean per-sequence F1 against exact interventional truth reaches 0.90-0.91 at vocabulary size 1000 (paper: 0.91) and 0.86-0.91 from 100 to 2000. First, the optimal threshold is pinned to the truth margin, not to any constant: at every size the errors at tau* straddle the delta = 0.05 margin defining ground truth (missed true edges lie just above it, accepted false ones just below), and the blind optimum lands near delta/2 times the estimator's calibration, confirmed out of sample at 5000. Second, at a single global threshold TRACE mostly recovers a direct, adjacent-influence graph: lag-1 true edges are recalled at 0.97-0.99, while true edges at lag 2 or more read orders of magnitude lower---the reading-scale price of randomizing mediating positions, which an exact test of direct causal effect requires when the truth is unknown. A per-lag threshold family recovers a third to a half of lag-2 truth; on lag-uniform data one validated threshold recalls every lag at 0.40-0.87, 8-26 pp below an atomic-intervention control at lags 3-6. Third, the default lag decay of the paper's synthetic benchmark concentrates about 85% of interventional truth at lag 1 and pushes the rest below the estimator's noise floor, so headline F1 there certifies lag-1 recovery only and conflates the benchmark's skew with the algorithm's own limit; a flatter decay separates the two. Fourth, F1 saturates from N = 2 particles at the selected threshold---a property of the threshold's margin over the noise floor, not of the estimator, which converges as N^(-1/2). We distill five practitioner rules.
- Abstract(参考訳): TRACE (Math & Lienhart, arXiv:2602.01135) は、事前訓練された自己回帰配列モデルからイベントタイプに関する因果グラフを読み出す。
検証分割で選択されたタウ平均F1は,100~2000年の語彙サイズ1000(紙:0.91)で0.90-0.91,0.86-0.91に達する。
あらゆる大きさにおいて、タウ* の誤差は delta = 0.05 の誤差で基底の真理(真辺は直上にあり、真辺は直下にある)を定め、ブラインド最適地は delta/2 の近傍で推定器の校正値の約5000 倍と確認された。
lag-1 true edges are recalled at 0.97-0.99, while true edges at lag 2 or more read order of size down-------- 中間位置のランダム化の読み取りスケール価格--- 真相が不明な場合に直接因果効果の正確なテストが要求される。
ラグ均一データでは、検証されたしきい値1は、ラグ3〜6での原子干渉制御の下0.04〜0.87,8〜26ppで全てのラグをリコールする。
第3に、論文の合成ベンチマークのデフォルトのラグ崩壊は、ラグ1における介入真理の約85%を集中させ、残りを推定器のノイズフロアの下に押し込むため、ヘッドラインF1はlag-1リカバリのみを認定し、ベンチマークのスキューをアルゴリズムの限界と混同し、フラットな崩壊は2つを分離する。
第4に、F1は選択されたしきい値でN = 2粒子から飽和し、これはN^(-1/2)として収束する推定器ではなく、ノイズフロア上のしきい値のマージンの性質である。
我々は5つの実践規則を蒸留する。
関連論文リスト
- The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models [22.461983628330277]
多くの少数ショット適応手法は、ゼロショットテキストプロトタイプとKラベル付き画像特徴の平均の凸の組み合わせで分類される。
正しい比率は何か、検証データなしで見積もることができるのか、パフォーマンスがどこにあるのかを尋ねる。
4800以上の細胞(SigLIPを含む5つのバックボーン、5つのショットカウント、5つのシード、4つのプロンプトレベル)が、理論上最適な比率は間違った量の信頼できる推定値である。
検証不要な線形プローブは、オラクルチューニングされたブレンド(CLAPは+1.9ポイント、LP++は+1.5ポイント)を破る。
論文 参考訳(メタデータ) (2026-08-23T15:02:33Z) - When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification [0.0]
クラス不均衡処理は、単一のベンチマークデータセットで定期的に評価される。
リークフリーなネスト型クロスバリデーションプロトコルの下では、デフォルト0.5しきい値のプレーンランダムフォレストがF1 = 0.861 +/-0.021に達する。
閾値調整の利点は、不均衡比において非単調である。
論文 参考訳(メタデータ) (2026-08-17T05:58:33Z) - Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection [0.0]
ベンチマーク汚染はn-gram重なり、可能性に基づくメンバーシップ推論、カナリア文字列と診断される。
これを行う自然な方法はうまくいかないことを示し、測定を生き残るものを特定し、それを動作させる補正が、テストされるnullよりも多くのばらつきをもたらすことを見つけます。
論文 参考訳(メタデータ) (2026-08-12T23:27:20Z) - Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk [0.0]
クレジットスコアリングは、決定ロジックをパラメータから読み取ることができないモデルに依存している。
共通する提案は、言語モデルとのギャップを埋める: 特徴属性を計算し、それらを LLM に渡し、合理的に記述させる。
このようなシステムをエンド・ツー・エンドに構築し、約束の後半が成立するかどうかをテストします。
論文 参考訳(メタデータ) (2026-08-08T13:22:14Z) - Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction [51.56484100374058]
テキスト検出器は、事前訓練された典型軸を増幅する。
タスク監督前の生エンコーダでは、3つのアーキテクチャでNYT-vs-HC3 AUROC 0.806/0.944/0.834を達成する。
RoBERTaベースでは、生のプロジェクションは微調整を超えるが、RoBERTaベースでは、フル微調整は、試験された流線型人口の双方で生よりも識別を小さくする。
論文 参考訳(メタデータ) (2026-05-20T19:08:38Z) - Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning [40.342912574072024]
大規模言語モデルは、旅行計画やコードソリューションのような構造化されたアウトプットを生成する。
個々の推論ステップは正しく見えるが、アウトプット全体が予算に違反したり、テストケースに失敗したり、あるいは以前の推論に矛盾することがある。
構造化LCM出力の検証のための決定論的解析制約付き学習品質スコアラを提案する。
論文 参考訳(メタデータ) (2026-05-15T17:08:27Z) - Benchmarking IoT Time-Series AD with Event-Level Augmentations [34.864214444544565]
実世界の問題をシミュレートする統合されたイベントレベル拡張による評価プロトコルを提案する。
5つの公開異常データセット上で14の代表的なモデルを評価する。
論文 参考訳(メタデータ) (2026-02-17T09:45:44Z) - Impact of Labeling Inaccuracy and Image Noise on Tooth Segmentation in Panoramic Radiographs using Federated, Centralized and Local Learning [46.232038247686745]
フェデレートラーニング(FL)は、歯科診断AIにおけるプライバシー制約、不均一なデータ品質、一貫性のないラベル付けを緩和する。
複数のデータ破損シナリオを対象としたパノラマX線撮影において,FLと集中学習(CL)と局所学習(LL)を比較した。
論文 参考訳(メタデータ) (2025-09-08T11:07:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。