論文の概要: PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics
- arxiv url: http://arxiv.org/abs/2609.25438v1
- Date: Mon, 21 Sep 2026 21:48:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-23 18:04:04.124187
- Title: PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics
- Title(参考訳): PermuFormer: Algebraic Combinatoricsにおける置換表現のためのマルチタスク事前学習
- Abstract要約: 本稿では,280億トークンのマルチタスク,マルチエンコードコーパスをトレーニングした自動回帰変換器PermuFormerを紹介する。
PermuFormerは、事前トレーニング中に見つからない基本的なタスクの微調整に有効な出発点であることを示す。
また、PermuFormerがトレーニングタスクの解き方を学習する内部メカニズムについても分析する。
- 参考スコア(独自算出の注目度): 7.90000685574129
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine-tuning on downstream tasks. While much of the excitement in AI for math has been concentrated in the use of frontier reasoning models to solve well-specified problems through the medium of language, narrow, specialized models remain an important component of the AI for math ecosystem. In contrast to large language models, specialized models are usually trained directly on the mathematical objects themselves (e.g., graphs, sequences of numbers) rather than the textual descriptions that characterize these objects. However, the common practice of training specialists from scratch may limit their ability to develop domain-aware representations that capture the multifaceted nature of mathematics. In this paper, we describe an approach to pretraining for permutation-focused tasks in algebraic combinatorics. We introduce PermuFormer, an autoregressive transformer trained on a 2.8 billion token multi-task, multi-encoding corpus. We show that PermuFormer is an effective starting point for fine-tuning on basic tasks unseen during pretraining and more complex research-level tasks, frequently outperforming the same architecture trained from scratch, baseline MLPs, and a fine-tuned generic language model of comparable size. We also analyze some of the internal mechanisms by which PermuFormer learns to solve training tasks. For example, we show that while some tasks can be linearly decoded directly from the internal representation of the prompt, other tasks require multiple rounds of generation before the answer can be decoded.
- Abstract(参考訳): 逆事前学習は、下流タスクの微調整の出発点となる、再利用可能なドメイン認識表現を学習するための効果的な方法であることが示されている。
数学におけるAIの興奮の大部分は、言語媒体を通じて明確に特定された問題を解決するためのフロンティア推論モデルの使用に集中しているが、数学エコシステムにおけるAIの重要な構成要素は狭義の特殊モデルである。
大きな言語モデルとは対照的に、特殊モデルは通常、これらのオブジェクトを特徴づけるテキスト記述ではなく、数学的なオブジェクト自身(例えば、グラフ、数字の列)で直接訓練される。
しかし、スクラッチからスペシャリストを訓練する一般的な実践は、数学の多面的な性質を捉えたドメイン認識表現を開発する能力を制限する可能性がある。
本稿では,代数的コンビネータにおける置換に着目したタスクの事前学習手法について述べる。
本稿では,280億トークンのマルチタスク,マルチエンコードコーパスをトレーニングした自動回帰変換器PermuFormerを紹介する。
PermuFormerは、事前学習や複雑な研究レベルのタスクでは見つからない基本タスクの微調整に有効な出発点であり、スクラッチ、ベースラインのMLP、および同等の大きさの微調整された汎用言語モデルで訓練された同じアーキテクチャよりも優れていることを示す。
また、PermuFormerがトレーニングタスクの解き方を学習する内部メカニズムについても分析する。
例えば、あるタスクはプロンプトの内部表現から直接線形に復号化できるが、他のタスクは解を復号化する前に複数ラウンドの生成を必要とする。
関連論文リスト
- What do Language Models Learn and When? The Implicit Curriculum Hypothesis [50.89698323917259]
Inlicit Curriculum hypothesis: Pretraining following a compositional and predictable curriculum across models。
410M-13Bパラメータから4つのモデルファミリの出現点を追跡する。
モデルが一定の精度のしきい値に達する際の出現順序は著しく一致していることがわかった。
論文 参考訳(メタデータ) (2026-04-09T17:50:12Z) - Task Addition and Weight Disentanglement in Closed-Vocabulary Models [75.01322212415435]
タスク算術は、事前学習されたテキストオープン語彙モデルを編集するための有望な方法として登場した。
本稿では,クローズドボキャブラリ画像分類モデルにおけるタスク追加について検討する。
事前学習された視覚変換器もタスク演算で編集できることがわかった。
論文 参考訳(メタデータ) (2025-11-18T15:12:21Z) - JiuZhang 2.0: A Unified Chinese Pre-trained Language Model for
Multi-task Mathematical Problem Solving [77.51817534090789]
マルチタスク数学問題の解法を専門とする統一中国語 PLM である textbfJiuZhang2.0 を提案する。
我々の考えは、中規模のモデルを維持し、マルチタスク設定におけるモデル容量を改善するために、Emphcross-taskの知識共有を利用することである。
論文 参考訳(メタデータ) (2023-06-19T15:45:36Z) - Learning Easily Updated General Purpose Text Representations with
Adaptable Task-Specific Prefixes [22.661527526471996]
ダウンストリームタスク毎にトレーニング済みの大きな言語モデルを微調整すると、計算負荷が発生する。
そこで本研究では,ソースタスクを用いてテキストの固定表現を学習するためのプレフィックスベースの手法を提案する。
論文 参考訳(メタデータ) (2023-05-22T21:31:03Z) - Arithmetic-Based Pretraining -- Improving Numeracy of Pretrained
Language Models [67.48894919842576]
最先端の事前訓練された言語モデルは、数式を必要とするタスクにアウト・オブ・ボックスを適用すると、その能力より劣る傾向にある。
本稿では,Arithmetic-Based Pretrainingと呼ばれる拡張事前学習手法を提案する。
本実験は,算数性の向上を必要とする3つのタスクにおいて,算術的事前学習の有効性を示す。
論文 参考訳(メタデータ) (2022-05-13T16:10:13Z) - Unified Multimodal Pre-training and Prompt-based Tuning for
Vision-Language Understanding and Generation [86.26522210882699]
視覚言語理解と生成のための統一型マルチモーダル事前学習を提案する。
提案したUniVLは、理解タスクと生成タスクの両方を扱うことができる。
実験の結果,同じモデルを用いた場合,理解タスクと生成タスクとの間にはトレードオフがあることが判明した。
論文 参考訳(メタデータ) (2021-12-10T14:59:06Z) - Pre-training Text Representations as Meta Learning [113.3361289756749]
本稿では,下流タスクを効果的に学習するために,モデルがテキスト表現を学習する能力を直接最適化する学習アルゴリズムを提案する。
マルチタスク事前学習とモデル非依存型メタラーニングの間には,一連のメタトレインステップによる本質的な関係があることが示されている。
論文 参考訳(メタデータ) (2020-04-12T09:05:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。