論文の概要: Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
- arxiv url: http://arxiv.org/abs/2605.10129v1
- Date: Mon, 11 May 2026 07:40:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-12 23:28:50.612662
- Title: Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
- Title(参考訳): 合成事前学習が言語モデルロバストネスを改善, ノイズの多い事前学習データに
- Authors: Xu Guo, Runyu Peng, Jian Tong, Yunhua Zhou, Haijun Lv, Zhihui Lu, Qipeng Guo,
- Abstract要約: 大規模な言語モデルは、事前学習のためにWebスケールのコーパスに依存している。
データキュレーションは緩和されるが、そのようなノイズを排除できないため、訓練前のコーパスは実際にはノイズが残る。
本研究は,プレトレーニング段階において,軽量なプレトレーニング段階がノイズデータに抵抗するかどうかを考察する。
- 参考スコア(独自算出の注目度): 32.68918472543787
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model performance. Data curation mitigates but cannot eliminate such noise, so pre-training corpora remain noisy in practice. We therefore study whether a lightweight pre-pre-training (PPT) stage based on synthetic data with learnable temporal structure helps resist noisy data during the pre-training (PT) stage. Across various corruption settings, our method consistently improves robustness to noise during PT, with larger relative gains at higher noise levels. For a 1B-parameter model, a synthetic PPT stage with only 65M tokens achieves the same final loss as the baseline while using up to 49\% fewer natural-text PT tokens across different noise levels. Mechanistic analyses suggest PPT does not immediately suppress attention to noisy tokens. Rather, PPT-initialized models gradually downweight attention between corrupted tokens during noisy PT. This indicates that synthetic PPT inhibits noise self-modeling and shapes the subsequent optimization trajectory. Code is available at https://github.com/guox18/formal-language-prepretraining.
- Abstract(参考訳): 大規模言語モデル(LLM)は、事前学習のためのWebスケールコーパスに依存している。
これらのデータセットに固有のノイズは、意味のあるパターンを曖昧にし、最終的にモデルパフォーマンスを低下させる傾向がある。
データキュレーションは緩和されるが、そのようなノイズを排除できないため、訓練前のコーパスは実際にはノイズが残る。
そこで,学習可能な時間構造を持つ合成データに基づいて,PPT(Pre-Turning Pre-Turning Pre-Turning Pre-Turning Pre-Turning Pre-Turning)段階が,事前学習(PT)段階におけるノイズに抵抗するかどうかを検討した。
本手法は, 様々な汚損設定において, PT中の雑音に対する頑健性を常に改善し, 高い雑音レベルにおいて相対的な利得を増大させる。
1Bパラメータモデルの場合、65Mトークンしか持たない合成PTPステージは、ノイズレベルの異なる自然文PTトークンを最大49\%削減しながら、ベースラインと同じ最終損失を達成する。
メカニスティック分析は、PTTはすぐにノイズトークンへの注意を抑えるものではないことを示唆している。
むしろ、PT初期化モデルは、ノイズのあるPTの間、腐敗したトークン間で徐々に重み付けされる。
このことは、合成PTがノイズ自己モデリングを阻害し、その後の最適化軌道を形作ることを示している。
コードはhttps://github.com/guox18/formal- language-prepretrainingで入手できる。
関連論文リスト
- PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection [30.13331191100816]
大規模言語モデル(LLM)における事前学習データを検出するトレーニングフリーでプラグアンドプレイのフレームワークであるPDRを導入する。
PDRはトークンレベルのスコアを明示的に強調し、初期位置からの異なる信号を増幅し、後の位置からのノイズを抑制する。
論文 参考訳(メタデータ) (2026-01-11T09:32:13Z) - Mixture of Noise for Pre-Trained Model-Based Class-Incremental Learning [59.635264288605946]
クラスインクリメンタルラーニング(CIL)は,旧来の知識を維持しつつ,新たなカテゴリを継続的に学習することを目的としている。
バックボーンに軽量な微調整を適用する既存のアプローチは、依然としてドリフトを誘発する。
バックボーン一般化の劣化を軽減し,新しいタスクを適応させることを目的として,Mixture of Noise (Min)を提案する。
論文 参考訳(メタデータ) (2025-09-20T16:07:20Z) - Impact of Noisy Supervision in Foundation Model Learning [91.56591923244943]
本論文は、事前学習データセットにおけるノイズの性質を包括的に理解し分析する最初の研究である。
雑音の悪影響を緩和し、一般化を改善するため、特徴空間に適応するチューニング法(NMTune)を提案する。
論文 参考訳(メタデータ) (2024-03-11T16:22:41Z) - Understanding and Mitigating the Label Noise in Pre-training on
Downstream Tasks [91.15120211190519]
本稿では、事前学習データセットにおけるノイズの性質を理解し、下流タスクへの影響を軽減することを目的とする。
雑音の悪影響を軽減するために特徴空間に適応する軽量ブラックボックスチューニング法(NMTune)を提案する。
論文 参考訳(メタデータ) (2023-09-29T06:18:15Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。