論文の概要: Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study
- arxiv url: http://arxiv.org/abs/2607.21866v1
- Date: Thu, 23 Jul 2026 23:45:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-27 20:58:57.013491
- Title: Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study
- Title(参考訳): タブラルデータに基づく古典的機械学習のスケーリング法則:ベンチマークによる検討
- Authors: Kaihua Ding,
- Abstract要約: 本稿では,古典的ML学習曲線の分散教室規模レプリケーションを提案する。
127人の学生が3つの割り当てられたデータセット上でそれぞれ固定されたプロトコルを実行した。
電力法則は 77.7% の細胞の R2 > 0.8 に適合し、ツリーアンサンブルは全データで支配的である。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one team, a handful of cells. We present a distributed classroom-scale replication: 127 students each ran a fixed protocol on 3 assigned datasets, drawn from 18 tabular classification and regression datasets and 6 model families (Boosting, Random Forest, SVM, Linear/Logistic, Ridge, Lasso), yielding 11,536 training runs and 1,648 fitted power-law curves of the form error(N) = a N^(-b) + c. Three findings. (1) Power laws fit: R^2 > 0.8 on 77.7% of cells, with tree ensembles dominating at full data (Boosting 50% of datasets, RandomForest 33%; linear models underperform on classification). (2) Approximate shared exponents within a model family: for 5 of 6 families, a single family-level exponent predicts each family's cross-dataset curves nearly as well as per-dataset exponents (R^2 gap < 0.011), though AIC favors the unconstrained fit and curve collapse is partial (32-58% of points within +/-0.5 dex). We frame this as approximate predictive compressibility, not dataset-independent universality; Lasso fails outright (negative control) and Ridge is fragile under leave-one-dataset-out. (3) Replicator-implementation variance: with random_state=42 fixed, independent re-implementations of the same protocol still differ by mean CV(b) = 0.144 on the fitted exponent -- not seed variance, but the spread induced by unconstrained parts of the protocol (preprocessing, encoding, missing-value handling). We release the aggregated curves, per-cell fits, and a practical data-requirement table for N* to reach target error 0.15.
- Abstract(参考訳): 従来のML学習曲線は、木、線形、カーネルのモデルを表形式で表すのに電力法則に適合するが、小さなスケールでは、1つの曲線、1つのチーム、1つのセル、1つのセルがある。
127人の学生が3つのアサインされたデータセットに対して固定されたプロトコルを実行し、18のグラフ分類と回帰データセットと6つのモデルファミリー(Boosting, Random Forest, SVM, Linear/Logistic, Ridge, Lasso)から抽出し、11,536のトレーニングランと1,648のフォームエラー(N) = a N^(-b) + cのパワーロー曲線を適用。
3つの発見。
1) 電力法則が適合する: R^2 > 0.8 on 77.7%, 樹木のアンサンブルが全データで支配する(データセットの50%, ランダムフォレスト33%, 分類による線形モデルの性能)。
2) モデルファミリー内の共用指数は, モデルファミリー内の共用指数: 6つのファミリーの5つのうち, 1つのファミリーレベル指数は, 各ファミリーのクロスデータセット曲線およびデータセット指数(R^2 ギャップ < 0.011)とほぼ同程度に予測するが, AIC は非制約適合を好んで曲線崩壊は部分的である(+/-0.5 dex 内の点の 32-58%)。
私たちはこれをデータセットに依存しない普遍性ではなく、近似的な予測圧縮性として捉えています。
(3)Replicator-implementation variance: if random_state=42 fixed, independent re-implementations of the same protocol are still different by mean CV(b) = 0.144 on the fitted exponent -- not seed variance, but the spread by unconstrained parts of the protocol (preprocessing, encoding, missing-value handling)。
また,N* がターゲット誤差 0.15 に達するための実用的なデータ要求表を作成した。
関連論文リスト
- Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling [0.0]
合成親和性ツールは、独立大規模言語モデル(LLM)エージェントとして各個人を動作させる。
このパラダイムには基本的な障害モードがあることを示し、分散優先の修正をそれに対して設定する。
論文 参考訳(メタデータ) (2026-07-17T15:45:57Z) - Format-Constraint Coupling in Knowledge Graph Construction from Statistical Tables [2.1795865731681903]
オープンデータポータルの共通レイアウトである,国ごとの時系列行列について検討する。
それらの結合効果は、最大+1.180 (2x2因子、6つのデータセット)の独立効果の和を超える。
ブートストラップ95%CIは4/6データセットに対して厳格に陽性であり、幅広いType-II行列に強い証拠がある。
論文 参考訳(メタデータ) (2026-05-21T04:08:42Z) - Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction [51.56484100374058]
テキスト検出器は、事前訓練された典型軸を増幅する。
タスク監督前の生エンコーダでは、3つのアーキテクチャでNYT-vs-HC3 AUROC 0.806/0.944/0.834を達成する。
RoBERTaベースでは、生のプロジェクションは微調整を超えるが、RoBERTaベースでは、フル微調整は、試験された流線型人口の双方で生よりも識別を小さくする。
論文 参考訳(メタデータ) (2026-05-20T19:08:38Z) - Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling [0.0]
16家系の63塩基モデルにおける推論と真理の結合度を測定した。
我々は、家族依存の臨界スケール(N_c$)以下の損失曲線を目に見えない体制変化を発見し、その上、彼らは協力する。
論文 参考訳(メタデータ) (2026-05-13T03:14:09Z) - Sparse Regression under Correlation and Weak Signals: A Reproducible Benchmark of Classical and Bayesian Methods [1.6679662639178268]
合成データに対する6つのスパース回帰法をベンチマークした。
ベイズ法は予測誤差(MSE 72 vs. 108-267)で勝利し、ホースシューは95%近くをカバーしている(94.8%)。
可変選択の場合、F1 0.47のラッソとスパイク・アンド・スラブのネクタイは、後部が不要な場合に事実上のデフォルトとなる。
論文 参考訳(メタデータ) (2026-04-04T15:46:44Z) - Highly Adaptive Ridge [84.38107748875144]
直交可積分な部分微分を持つ右連続函数のクラスにおいて,$n-2/3$自由次元L2収束率を達成する回帰法を提案する。
Harは、飽和ゼロオーダーテンソル積スプライン基底展開に基づいて、特定のデータ適応型カーネルで正確にカーネルリッジレグレッションを行う。
我々は、特に小さなデータセットに対する最先端アルゴリズムよりも経験的性能が優れていることを示す。
論文 参考訳(メタデータ) (2024-10-03T17:06:06Z) - Scaling Laws in Linear Regression: Compute, Parameters, and Data [86.48154162485712]
無限次元線形回帰セットアップにおけるスケーリング法則の理論について検討する。
テストエラーの再現可能な部分は$Theta(-(a-1) + N-(a-1)/a)$であることを示す。
我々の理論は経験的ニューラルスケーリング法則と一致し、数値シミュレーションによって検証される。
論文 参考訳(メタデータ) (2024-06-12T17:53:29Z) - Pseudo-Labeling for Kernel Ridge Regression under Covariate Shift [1.3597551064547502]
対象分布に対する平均2乗誤差が小さい回帰関数を,ラベルなしデータと異なる特徴分布を持つラベル付きデータに基づいて学習する。
ラベル付きデータを2つのサブセットに分割し、カーネルリッジの回帰処理を行い、候補モデルの集合と計算モデルを得る。
モデル選択に擬似ラベルを用いることで性能を著しく損なうことはないことが判明した。
論文 参考訳(メタデータ) (2023-02-20T18:46:12Z) - Few-Shot Non-Parametric Learning with Deep Latent Variable Model [50.746273235463754]
遅延変数を用いた圧縮による非パラメトリック学習(NPC-LV)を提案する。
NPC-LVは、ラベルなしデータが多いがラベル付きデータはほとんどないデータセットの学習フレームワークである。
我々は,NPC-LVが低データ構造における画像分類における3つのデータセットの教師あり手法よりも優れていることを示す。
論文 参考訳(メタデータ) (2022-06-23T09:35:03Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。