論文の概要: MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
- arxiv url: http://arxiv.org/abs/2609.37574v2
- Date: Wed, 30 Sep 2026 17:32:02 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:26.29547
- Title: MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
- Title(参考訳): MERGE: ジェネレーティブエンリッチメントによる検索のためのマルチLLMアンサンブル
- Abstract要約: 大規模言語モデル(LLM)は、情報検索(IR)におけるユーザクエリの強化にますます利用されている。
3つの異種 7-8B オープンソース LLM が独立して候補拡張を生成し、より大きな LLM がそれらを単一のクエリに生成する。
5つのBEIRベンチマーク(NQ、SciFact、FiQA、Touche-2020、DBPedia)では、MERGEはBM25 nDCG@10を+2.1から+14.9ポイント改善し、強力なLLMベースのクエリ拡張ベースラインにマッチまたは性能を向上する。
- 参考スコア(独自算出の注目度): 16.653517324875384
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model -- an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce candidate expansions, and a larger LLM generatively synthesizes them into a single query. To make prompt engineering scalable across the ensemble, we integrate a task-grounded Automatic Prompt Optimization (APO) loop into both stages. Unlike APO methods that judge candidates with an LLM evaluator, our loop scores each candidate by its downstream retrieval performance and runs a small tournament between the current champion prompt and optimizer-proposed drafts, terminating once the champion survives two consecutive rounds; a history-augmented variant additionally feeds the recent tournament trajectory back to the optimizer. MERGE is retriever-agnostic and issues a single BM25 pass with no rank fusion, no supervised document expansion, and no re-indexing. On five BEIR benchmarks (NQ, SciFact, FiQA, Touche-2020, DBPedia), MERGE improves BM25 nDCG@10 over the original queries by +2.1 to +14.9 points and matches or outperforms strong LLM-based query-expansion baselines despite using only compact open-source models. Ablations confirm that the Stage-2 ensemble beats any single Stage-1 LLM, and that task-grounded APO converts large seed-prompt regressions into consistent gains without hand-tuning.
- Abstract(参考訳): 大規模言語モデル(LLM)は、情報検索(IR)におけるユーザクエリの強化にますます使われており、BM25のような標準レトリバーは、ターゲットコーパスと語彙ギャップをブリッジすることができる。
しかしながら、トレーニングデータとアーキテクチャバイアスによって制限される単一のLCMは、そのリッチ化動作は、手作りのプロンプトに依存します。
MERGE(Multi-LLM Ensemble for Retrieval via Generative Enrichment)は3つの異種 7-8B オープンソース LLM が独立に候補拡張を生成し、より大きな LLM がそれらを単一のクエリに合成する2段階のフレームワークである。
アンサンブル全体にわたって迅速なエンジニアリングをスケーラブルにするために,タスクグラウンド付き自動プロンプト最適化(APO)ループを両ステージに統合する。
LLM評価器の候補を判定するAPO法とは異なり、ループは各候補をダウンストリーム検索性能でスコア付けし、現在のチャンピオンプロンプトとオプティマイザが提案するドラフトの小さなトーナメントを実行し、チャンピオンが2ラウンド連続しても終了する。
MERGEはレトリバーに依存しないため、1つのBM25パスにランクフュージョンがなく、ドキュメントの拡張が監督されず、再インデックスもできない。
5つのBEIRベンチマーク(NQ、SciFact、FiQA、Touche-2020、DBPedia)では、MERGEはBM25 nDCG@10を+2.1から+14.9ポイント改善し、コンパクトなオープンソースモデルのみを使用しながら強力なLCMベースのクエリ拡張ベースラインにマッチまたは性能を向上した。
アブレーションにより、ステージ2のアンサンブルがステージ1のLLMに勝っていることが確認され、タスクグラウンドのAPOは、大きなシードプロンプトレグレッションを手作業なしで一貫したゲインに変換する。
関連論文リスト
- Rethinking On-policy Optimization for Query Augmentation [49.87723664806526]
本稿では,様々なベンチマークにおいて,プロンプトベースとRLベースのクエリ拡張の最初の体系的比較を示す。
そこで我々は,検索性能を最大化する擬似文書の生成を学習する,新しいハイブリッド手法 On-policy Pseudo-document Query Expansion (OPQE) を提案する。
論文 参考訳(メタデータ) (2025-10-20T04:16:28Z) - Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers [74.17516978246152]
大規模言語モデル(LLM)は、従来の手法を進化させるために情報検索に広く統合されている。
エージェント検索フレームワークであるEXSEARCHを提案する。
4つの知識集約ベンチマークの実験では、EXSEARCHはベースラインを大幅に上回っている。
論文 参考訳(メタデータ) (2025-05-26T15:27:55Z) - ExpandR: Teaching Dense Retrievers Beyond Queries with LLM Guidance [21.777817032607405]
大規模言語モデル(LLM)は、クエリ拡張による高密度検索の強化に有意な可能性を証明している。
本研究では,LLM拡張高密度検索フレームワークExpandRを提案する。
複数のベンチマーク実験の結果、ExpandRは強いベースラインを一貫して上回ることがわかった。
論文 参考訳(メタデータ) (2025-02-24T11:15:41Z) - Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting [65.00288634420812]
Pairwise Ranking Prompting (PRP)は、大規模言語モデル(LLM)の負担を大幅に軽減する手法である。
本研究は,中等級のオープンソースLCMを用いた標準ベンチマークにおいて,最先端のランク付け性能を達成した文献としては初めてである。
論文 参考訳(メタデータ) (2023-06-30T11:32:25Z) - Large Language Models are Strong Zero-Shot Retriever [89.16756291653371]
ゼロショットシナリオにおける大規模検索に大規模言語モデル(LLM)を適用するための簡単な手法を提案する。
我々の手法であるRetriever(LameR)は,LLM以外のニューラルモデルに基づいて構築された言語モデルである。
論文 参考訳(メタデータ) (2023-04-27T14:45:55Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。