論文の概要: Not All Synthetic Data Is Yours to Learn From
- arxiv url: http://arxiv.org/abs/2605.31126v1
- Date: Fri, 29 May 2026 10:34:11 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-01 20:56:50.543
- Title: Not All Synthetic Data Is Yours to Learn From
- Title(参考訳): すべての合成データから学ぶべきものではない
- Authors: Sina Alemohammad, Li Chen, Richard G. Baraniuk, Zhangyang Wang,
- Abstract要約: 弱い自己学習は、事前訓練モデルにすでに存在する能力を増幅できることを示す。
我々はこれを、プロンプトフリーな無条件自己学習の最小設定で研究する。
- 参考スコア(独自算出の注目度): 68.13136636413601
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Can a language model improve from plain text sampled from itself, with no prompts, no teacher, no verifier, and no reward model? Yes, but only when the synthetic corpus is compatible with the student, a relational property of the source-student pair rather than an intrinsic property of the data. We call this the latent capability resurfacing hypothesis: weak self-training can amplify capabilities already present in the pretrained model, but only under this compatibility condition. We study this in the minimal setting of prompt-free unconditional self-training, where base language models are fine-tuned on text generated from the BOS token alone, with no task specification or external supervision. We report three findings. First, synthetic utility is relational rather than intrinsic: self-generated data is the most effective source, same-lineage transfer outperforms stronger but differently trained sources, and cross-family transfer is substantially weaker. Second, common intrinsic proxies fail: neither benchmark-level semantic similarity nor average per-token likelihood under the student predicts which corpora help. Third, this regime produces a surprising byproduct. In controlled Pythia experiments, capability and verbatim memorization decouple: benchmark utility is preserved or improved while held-out exact-match extraction drops by over 95 percent, with no forget set, privacy objective, or targeted unlearning. Together, these results suggest that prompt-free self-training works by amplifying what the student already knows, not by importing structure from the data. They also reveal a regime in which capability and verbatim memorization can be separated without any explicit unlearning objective.
- Abstract(参考訳): 言語モデルは、プロンプトなし、教師なし、検証なし、報酬なしのプレーンテキストから改善できるのだろうか?
そう、しかし、合成コーパスが生徒と互換性がある場合に限り、データ固有の性質ではなく、ソースと学生のペアの関係性がある。
弱い自己学習は、事前訓練されたモデルにすでに存在する能力を増幅するが、この互換性条件下でのみ有効である。
我々はこれを,BOSトークンから生成されたテキストに基本言語モデルを微調整し,タスク仕様や外部監視を伴わない,プロンプトフリーな非条件自己学習の最小設定で研究する。
我々は3つの発見を報告した。
まず、合成ユーティリティは本質的なものではなく関係性である: 自己生成データは最も効果的な情報源であり、同系転送はより強力だが異なる訓練されたソースより優れ、クロスファミリー転送は実質的に弱い。
第二に、一般的な内在的プロキシは失敗する: ベンチマークレベルのセマンティックな類似性も、学生の平均的な1つの確率も、どのコーパスが役に立つかを予測しない。
第三に、この体制は驚くべき副産物を生み出している。
制御されたPythia実験では、能力と動詞の暗記を分離する: ベンチマークユーティリティは保存または改善され、正確にマッチされた抽出は95%以上減少する。
これらの結果は、データから構造をインポートするのではなく、学生がすでに知っていることを増幅することで、即時学習が機能することを示唆している。
彼らはまた、明確な未学習の目的なしに、能力と動詞の暗記を分離できる体制を明らかにした。
関連論文リスト
- Privacy-Preserving Model Transcription with Differentially Private Synthetic Distillation [67.76456940243294]
プライベートデータセットでトレーニングされたディープラーニングモデルは、プライバシー漏洩のリスクを引き起こす可能性がある。
本稿では,データフリーモデル-モデル変換ソリューションであるエンフェプライシ保存モデル転写について述べる。
論文 参考訳(メタデータ) (2026-01-27T01:51:35Z) - Synthetic bootstrapped pretraining [52.92577542049469]
本稿では,SBP(Synthetic Bootstrapped Pretraining)について述べる。
SBPはまず、事前学習データセットから文書間の関係のモデルを学び、次にそれを利用して巨大な新しいコーパスを合成する。
SBPは高い繰り返しベースラインを継続的に改善し、オラクル上界で達成可能な性能改善のかなりの部分を提供する。
論文 参考訳(メタデータ) (2025-09-17T22:28:27Z) - From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition [64.59093444558549]
我々はFrom Fake to Realと呼ぶシンプルで簡単に実装できる2段階のトレーニングパイプラインを提案する。
実データと合成データを別々にトレーニングすることで、FFRは実データと合成データの統計的差異にモデルを公開しない。
実験の結果,FFRは3つのデータセットに対して,最先端のグループ精度を最大20%向上させることがわかった。
論文 参考訳(メタデータ) (2023-08-08T19:52:28Z) - Synthetic Pre-Training Tasks for Neural Machine Translation [16.6378815054841]
我々のゴールは、合成資源を使用する場合の事前学習モデルの有効性に寄与する要因を理解することである。
本稿では,語彙的および構造的知識のレベルが異なる事前学習型翻訳モデルを提案する。
複数の言語ペアに対する実験により,高レベルの難読化や純粋に合成された並列データであっても,事前学習のメリットが実現できることが明らかになった。
論文 参考訳(メタデータ) (2022-12-19T21:34:00Z) - On the Transferability of Pre-trained Language Models: A Study from
Artificial Datasets [74.11825654535895]
大規模未ラベルテキストデータ上での事前学習言語モデル(LM)により、ダウンストリームのパフォーマンスが極めて容易になる。
我々は,事前学習データに含まれる特定の特徴について,セマンティクス以外では,下流タスクのスクラッチからトレーニングしたデータよりも,事前学習したLMを優れているか検討した。
論文 参考訳(メタデータ) (2021-09-08T10:39:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。