論文の概要: On the Capability and Limitation of Hard Prompt
- arxiv url: http://arxiv.org/abs/2609.32302v1
- Date: Sat, 26 Sep 2026 06:51:27 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-07 17:59:07.288856
- Title: On the Capability and Limitation of Hard Prompt
- Title(参考訳): ハードプロンプトの容量と限界について
- Abstract要約: 本稿では,変圧器が下流タスクを解くハードプロンプトの存在を判断することはNP完全であることを示す。
ソフトなプロンプトや連続的なプロンプトとは異なり、ハードプロンプトには必須の制限があることを示す。
有限タスク上でのプロンプトの実行に対して、プロンプト長の観点からタスクのサイズを厳密に制限する。
- 参考スコア(独自算出の注目度): 29.442181557518534
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete prompts is largely open. In this paper, we try to fill this gap either by providing a complete solution or by making substantial progress on the three core theoretical questions regarding hard prompts. First, we show that determining the existence of a hard prompt for a transformer to solve a downstream task is NP-complete and that finding an optimal hard prompt is NP-hard, which is the first computational complexity result for hard prompting, as far as we know. Second, we show that, unlike soft or continuous prompts, hard prompts have essential limitations: hard prompts are not complete; short hard prompts do not significantly enhance the ability of transformers; and long hard prompts exhibit the "prompt dominating answer phenomenon," meaning that, with high probability, the same answer is given for all queries of the same length. On the other hand, linear hard prompts do not have the limitations of short or long prompts. Third, we provide a tight bound on the size of the task in terms of the prompt length for the performance of prompts on the finite task to generalize to the entire data distribution, leading to a necessary and sufficient condition for generalizability. This is the first result on generalization for prompting, as far as we know. Our findings not only offer the first theoretical insights into hard prompts but also provide provably reliable practical guidance for real-world LLM usage.
- Abstract(参考訳): プロンプトエンジニアリングは、大きな言語モデル(LLM)を使用するのに欠かせないツールとなり、LLMを重みを変えることなくタスク固有の専門家に変える。
プロンプト工学における顕著な理論的な進歩にもかかわらず、より実践的なハードプロンプトや離散的なプロンプトの理論はおおむねオープンである。
本稿では,このギャップを,完全な解を与えるか,あるいはハードプロンプトに関する3つの中核的な理論的問題に大きく進展させることで埋めようとしている。
まず、変圧器がダウンストリームタスクを解くためのハードプロンプトの存在を決定することはNP完全であり、最適なハードプロンプトを見つけることは、我々が知る限り、最初の計算複雑性の結果であるNPハードであることを示す。
第二に、ハードプロンプトはソフトプロンプトや連続プロンプトと異なり、ハードプロンプトが完全ではないこと、ショートハードプロンプトはトランスフォーマーの能力を大幅に向上しないこと、ロングハードプロンプトは「プロンプト支配的な応答現象」を示すこと、すなわち高い確率で同じ答えが同じ長さの全てのクエリに対して与えられること、などが示される。
一方、線形ハードプロンプトにはショートプロンプトやロングプロンプトの制限がない。
第3に、有限タスク上のプロンプトのパフォーマンスがデータ分布全体へ一般化するためには、タスク長のプロンプトでタスクのサイズを厳密に制限し、一般化に必要かつ十分な条件を導出する。
これは、我々が知る限り、プロンプトの一般化に関する最初の結果である。
我々の発見は、ハードプロンプトに関する最初の理論的知見を提供するだけでなく、実世界のLLM利用のための信頼性の高い実用的なガイダンスを提供する。
関連論文リスト
- Towards Interpretable Soft Prompts [24.304585350085315]
本研究では,2つのデシラタに基づいて,訓練可能なプロンプトの解釈可能性を評価する。
GPT-2を用いた実験は、解釈可能性と訓練可能なプロンプトのタスク性能の基本的なトレードオフを示す。
論文 参考訳(メタデータ) (2025-04-02T21:42:09Z) - Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMs [15.941209553757274]
いくつかのプロンプトが成功し、他が失敗する理由を説明する理論的フレームワークを提供する。
与えられたタスクに対して、最適なプロンプトを見つけ、プロンプト空間のサイズを特徴付ける複雑さを解析する。
私たちの理論は効果的なプロンプト設計の背景にある原則を明らかにし、CoTを使用する自己指導的なプロンプトである"ステップバイステップ"がパフォーマンスを著しく阻害することを示している。
論文 参考訳(メタデータ) (2025-03-13T06:11:10Z) - Why is prompting hard? Understanding prompts on binary sequence predictors [19.855572748273236]
大規模言語モデル(LLM)は多くのタスクを実行するように促すことができる。
良いプロンプトを見つけることは必ずしも容易ではないし、パフォーマンスのプロンプトを理解するのも容易ではない。
論文 参考訳(メタデータ) (2025-02-15T10:55:47Z) - Efficient Prompting Methods for Large Language Models: A Survey [50.82812214830023]
効率的なプロンプティング手法は幅広い注目を集めている。
本稿では,異なるプロンプト成分に対する自動プロンプトエンジニアリングと連続空間および離散空間におけるプロンプト圧縮について論じる。
論文 参考訳(メタデータ) (2024-04-01T12:19:08Z) - Automatic Engineering of Long Prompts [79.66066613717703]
大規模言語モデル(LLM)は、複雑なオープンドメインタスクを解く際、顕著な能力を示した。
本稿では,自動ロングプロンプトエンジニアリングのためのグリージーアルゴリズムと遺伝的アルゴリズムの性能について検討する。
提案アルゴリズムは,Big Bench Hardの8つのタスクにおいて,平均9.2%の精度向上を実現している。
論文 参考訳(メタデータ) (2023-11-16T07:42:46Z) - Demystifying Prompts in Language Models via Perplexity Estimation [109.59105230163041]
プロンプトのパフォーマンスは、モデルが含んでいる言語に精通している範囲と結合している。
プロンプトの難易度が低ければ低いほど、プロンプトがタスクを実行することができることを示す。
論文 参考訳(メタデータ) (2022-12-08T02:21:47Z) - RLPrompt: Optimizing Discrete Text Prompts With Reinforcement Learning [84.75064077323098]
本稿では、強化学習(RL)を用いた離散的高速最適化手法RLPromptを提案する。
RLPromptは、マスク付きジベリッシュ(例:grammaBERT)や左から右へのモデル(例:GPT)など、様々な種類のLMに柔軟に適用可能である。
少数ショット分類と教師なしテキストスタイル転送の実験は、既存のファインタニングやプロンプト手法よりも優れた性能を示す。
論文 参考訳(メタデータ) (2022-05-25T07:50:31Z) - Least-to-Most Prompting Enables Complex Reasoning in Large Language
Models [52.59923418570378]
本稿では, 難解な一般化の課題を克服するために, 最小限のプロンプト戦略を提案する。
最小限のプロンプトは、プロンプトで見られるものよりも難しい問題に一般化可能であることを示す。
SCANの解決を専門とする文献におけるニューラルシンボリックモデルは、15,000以上のサンプルを含むトレーニングセット全体をトレーニングする。
論文 参考訳(メタデータ) (2022-05-21T15:34:53Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。