論文の概要: Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases
- arxiv url: http://arxiv.org/abs/2609.09462v1
- Date: Tue, 08 Sep 2026 21:26:37 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-10 19:44:08.819455
- Title: Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases
- Title(参考訳): 固定トークンベースを用いた視覚言語モデルの低ランクプロンプト学習
- Abstract要約: クラス毎にほんの数例からトレーニングした高密度プロンプト行列 $mathbfPinmathbbRmtimes d$ について検討する。
トークン側の要素をまったく学習する必要はありません。
CLIPプロンプト設定では、埋め込み側係数は適応を持ち、トークンベースは簡単に固定できる。
- 参考スコア(独自算出の注目度): 11.437689514839333
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Prompt learning adapts CLIP to downstream recognition by replacing hand-written templates with learned continuous context vectors, which in Context Optimization (CoOp) form a dense prompt matrix $\mathbf{P}\in\mathbb{R}^{m\times d}$ trained from only a few examples per class. We study whether this matrix is over-parameterized by factorizing it as $\mathbf{P}=\mathbf{B}\mathbf{A}$, which cuts the trainable prompt parameters from $md$ to $r(m+d)$, and to $rd$ once the token-side factor $\mathbf{B}$ is fixed. Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization. We then find that the token-side factor need not be learned at all: fixing $\mathbf{B}$ to a Gaussian, orthogonal, SVD-derived, or even random basis and training only the embedding-side factor $\mathbf{A}$ stays on par with the fully trainable factorization, and a source-trained $\mathbf{B}$ offers no advantage over a random one. A prompt-factor asymmetry and a local update-space dimension gap show why fixing $\mathbf{B}$ is far less restrictive than fixing $\mathbf{A}$, and a smoothness-only guarantee certifies that optimizing $\mathbf{A}$ over a fixed $\mathbf{B}$ converges. In the CLIP prompt setting, the embedding-side coefficients carry the adaptation while the token basis can simply be fixed.
- Abstract(参考訳): プロンプト学習はCLIPを下流認識に適応させ、手書きテンプレートを学習された連続したコンテキストベクトルに置き換え、コンテキスト最適化(CoOp)では高密度なプロンプト行列 $\mathbf{P}\in\mathbb{R}^{m\times d}$ をクラス毎にいくつかの例からトレーニングする。
この行列は$\mathbf{P}=\mathbf{B}\mathbf{A}$と分解して過度パラメータ化されているかを調べ、これはトレーニング可能なプロンプトパラメータを$md$から$r(m+d)$に、トークン側因子$\mathbf{B}$が固定されたときに$rd$にカットする。
7つの数ショットのベンチマークと2つのCLIPバックボーン、低ランクのプロンプト、より少ないパラメータで高密度のCoOpにマッチまたは改善されている。
ガウス的、直交的、SVD に由来する、あるいはランダムな基底に$\mathbf{B}$を固定し、埋め込み側因子 $\mathbf{A}$ は、完全に訓練可能な因子化と同程度に留まり、ソース学習された $\mathbf{B}$ は、ランダムな要素に対して何の利点も与えない。
プロンプト因子の非対称性と局所更新空間次元ギャップは、$\mathbf{B}$の固定が$\mathbf{A}$の固定よりもはるかに制限的であることを示し、滑らか性のみの保証は、$\mathbf{A}$を固定された$\mathbf{B}$上で最適化することを証明している。
CLIPプロンプト設定では、埋め込み側係数は適応を持ち、トークンベースは簡単に固定できる。
関連論文リスト
- Robust Learning of a Group DRO Neuron [21.632698901872843]
任意のラベルノイズと群レベルの分布シフトの存在下で,標準2乗損失下で原始ニューロンを学習する問題について検討した。
我々のフレームワークは、任意のラベルの破損やグループ固有の分布シフトに直面して、堅牢な学習保証を提供する。
論文 参考訳(メタデータ) (2026-01-26T04:00:53Z) - Sample and Computationally Efficient Robust Learning of Gaussian Single-Index Models [37.42736399673992]
シングルインデックスモデル (SIM) は $sigma(mathbfwast cdot mathbfx)$ という形式の関数であり、$sigma: mathbbR to mathbbR$ は既知のリンク関数であり、$mathbfwast$ は隠れ単位ベクトルである。
適切な学習者が$L2$-error of $O(mathrmOPT)+epsilon$。
論文 参考訳(メタデータ) (2024-11-08T17:10:38Z) - Compressing Large Language Models using Low Rank and Low Precision Decomposition [46.30918750022739]
この研究は、新しい訓練後のLLM圧縮アルゴリズムである$rm CALDERA$を導入している。
重量行列 $mathbfW$ の固有の低ランク構造を利用して、低ランクで低精度な分解によってそれを近似する。
その結果、LlaMa-$2$$7$B/$13B$/$70$BとLlaMa-$3$B $rm CALDERA$は、既存のトレーニング後の圧縮技術より優れていることが示された。
論文 参考訳(メタデータ) (2024-05-29T08:42:30Z) - Provably learning a multi-head attention layer [55.2904547651831]
マルチヘッドアテンション層は、従来のフィードフォワードモデルとは分離したトランスフォーマーアーキテクチャの重要な構成要素の1つである。
本研究では,ランダムな例から多面的注意層を実証的に学習する研究を開始する。
最悪の場合、$m$に対する指数的依存は避けられないことを示す。
論文 参考訳(メタデータ) (2024-02-06T15:39:09Z) - Statistical Learning under Heterogeneous Distribution Shift [71.8393170225794]
ground-truth predictor is additive $mathbbE[mathbfz mid mathbfx,mathbfy] = f_star(mathbfx) +g_star(mathbfy)$.
論文 参考訳(メタデータ) (2023-02-27T16:34:21Z) - Learning a Single Neuron with Adversarial Label Noise via Gradient
Descent [50.659479930171585]
モノトン活性化に対する $mathbfxmapstosigma(mathbfwcdotmathbfx)$ の関数について検討する。
学習者の目標は仮説ベクトル $mathbfw$ that $F(mathbbw)=C, epsilon$ を高い確率で出力することである。
論文 参考訳(メタデータ) (2022-06-17T17:55:43Z) - Model Selection with Near Optimal Rates for Reinforcement Learning with
General Model Classes [27.361399036211694]
有限地平線エピソディック強化学習(RL)問題に対するモデル選択の問題に対処する。
モデル選択フレームワークでは、$mathcalP*$の代わりに、遷移カーネルのネストされたファミリーが$M$を与えられる。
textttARL-GENが$TildemathcalO(d_mathcalE* H2+sqrtd_mathcalE* mathbbM* H2T)$の後悔を得ることを示す。
論文 参考訳(メタデータ) (2021-07-13T05:00:38Z) - Agnostic Learning of a Single Neuron with Gradient Descent [92.7662890047311]
期待される正方形損失から、最も適合した単一ニューロンを学習することの問題点を考察する。
ReLUアクティベーションでは、我々の人口リスク保証は$O(mathsfOPT1/2)+epsilon$である。
ReLUアクティベーションでは、我々の人口リスク保証は$O(mathsfOPT1/2)+epsilon$である。
論文 参考訳(メタデータ) (2020-05-29T07:20:35Z) - Taking a hint: How to leverage loss predictors in contextual bandits? [63.546913998407405]
我々は,損失予測の助けを借りて,文脈的包帯における学習を研究する。
最適な後悔は$mathcalO(minsqrtT, sqrtmathcalETfrac13)$である。
論文 参考訳(メタデータ) (2020-03-04T07:36:38Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。