論文の概要: KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search
- arxiv url: http://arxiv.org/abs/2608.21365v1
- Date: Wed, 17 Jun 2026 00:51:03 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-28 05:09:55.312565
- Title: KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search
- Title(参考訳): KSE-Web:低リソースKhmerセマンティックサーチのためのハイブリッド検索とLLM支援クエリ拡張の分析
- Abstract要約: KSE-Webは、Khmerセマンティックサーチのためのハイブリッド検索とLLM支援クエリ拡張の分析である。
データセットには、手動でレビューされたユーザスタイルのKhmer検索クエリと、部分的な人間認証を備えた銀関連ラベルが含まれている。
- 参考スコア(独自算出の注目度): 0.5076419064097734
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage. This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search. We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplication, and document-length control. The dataset includes 300 manually reviewed user-style Khmer search queries and silver relevance labels with partial human verification. We evaluate character n-gram BM25, multilingual dense retrieval, hybrid BM25+dense retrieval, and LLM-assisted query expansion using Qwen2.5 models. Experimental results show that BM25 achieves the strongest overall performance, reaching 0.943 Recall and 0.876 nDCG. Hybrid BM25+dense retrieval performs comparably, achieving 0.929 Recall and 0.871 nDCG, while dense retrieval alone performs lower. LLM-assisted query expansion does not outperform non-expanded retrieval; however, Qwen2.5-3B produces substantially stronger expanded-query results than Qwen2.5-0.5B, suggesting that LLM size and expansion quality matter for low-resource Khmer retrieval. Our analysis further shows that direct LLM expansion can introduce topic drift, generic terms, and noisy reformulations, while simple filtering may remove useful semantic cues. These findings highlight both the potential and limitations of LLM-assisted retrieval for Khmer semantic search and provide a foundation for future Khmer retrieval datasets with stronger human-verified annotations and Khmer-aware retrieval models. The dataset and documentation will be made available at github.com/back-kh/KhmerSemantic-Search.
- Abstract(参考訳): 低リソース言語として、Khmerは、限られた注釈付きデータ、曖昧な単語境界、多言語埋め込みモデルの弱いサポート、Khmer- Englishの頻繁な混合使用など、いくつかの検索課題を提示している。
本稿では、Khmerセマンティックサーチのためのハイブリッド検索とLLM支援クエリ拡張の解析であるKSE-Webを提案する。
約17Kの候補Khmerのタイトルからデータセットを構築し,フィルタリング,正規化,復号化,文書長制御後の全文Khmer文書を3Kクリーニングしたまま保持する。
データセットには、手動でレビューされたユーザスタイルのKhmer検索クエリと、部分的な人間認証を備えた銀関連ラベルが含まれている。
Qwen2.5モデルを用いて,文字n-gram BM25,多言語高密度検索,ハイブリッドBM25+dense検索,LLM支援クエリ拡張の評価を行った。
実験の結果,BM25は0.943リコール,0.876nDCGを達成した。
ハイブリッドBM25+dense検索は、0.929リコールと0.871nDCGを達成し、高密度検索だけでは低性能である。
しかし、Qwen2.5-3BはQwen2.5-0.5Bよりも大幅に拡張されたクエリ結果を生成するため、低リソースのKhmer検索においてLLMのサイズと拡張品質が問題となる。
さらに本解析により,LLM拡張により話題のドリフト,一般用語,雑音の修正が実現し,単純なフィルタリングにより意味的手がかりが取り除かれる可能性が示唆された。
これらの知見は、Khmerセマンティック検索のためのLLM支援検索の可能性と限界を強調し、より強力な人間認証アノテーションとKhmer対応検索モデルを備えた将来のKhmer検索データセットの基礎を提供する。
データセットとドキュメントはgithub.com/back-kh/KhmerSemantic-Searchで公開される。
関連論文リスト
- Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA [0.8594140167290097]
本稿では,Khmer文書画像を用いたオープンMLLMのパイロット診断について述べる。
Khmer-scriptの回答は、英語や数値の分野よりもかなり難しいままである。
論文 参考訳(メタデータ) (2026-08-08T15:52:31Z) - A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset [0.6201364126398703]
CUPは,868のカタログレコードと104のエキスパートアノテートクエリからなる,ギリシャの書籍検索ベンチマークである。
本書検索では,スパース (BM25), 密度 (文変換器), ハイブリッド, LLM を用いた検索手法について検討した。
論文 参考訳(メタデータ) (2026-07-23T12:51:55Z) - Boolean queries are all you need? [10.637635718179075]
TREC 2024 RAGトラックで使用されるMS MARCO V2.1復号セグメントコレクションを検索する。
86のトピックの標準トラックサブセットと100のモデルコール/トピックの予算の下で動作し、エージェントは0.6863のNDCG@10を達成した。
ランク付けは、クエリにマッチするコーパスの密度のみに基づいており、教師付き学習、グローバル統計、用語の重み付けは不要である。
論文 参考訳(メタデータ) (2026-07-13T10:25:46Z) - Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents [0.0]
Average Retrieval Score (L2 distance)、Answer Relevance、Khmer Coverage、Khmer Intersection over Unionの4つの指標を用いてパフォーマンスを評価する。
我々は,300文字のチャンクサイズを持つ文字ベースの再帰的チャンク法において,最適な性能を示す。
論文 参考訳(メタデータ) (2026-05-21T09:06:13Z) - Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval [25.731213365755234]
textitSuperIntelligent Retrieval Agent (SIRA)を紹介する。
SIRAは、複数ラウンド探索探索を単一のコーパス識別検索アクションに圧縮することができる。
解釈可能で、トレーニング不要で、効率的でありながら、より高価なマルチラウンドサーチを超えることができる。
論文 参考訳(メタデータ) (2026-05-07T17:54:29Z) - Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval [56.65147231836708]
SWIM-IRは, 微調整多言語高密度検索のための33言語を含む合成検索訓練データセットである。
SAPは、ターゲット言語における情報クエリの生成において、大きな言語モデル(LLM)を支援する。
我々のモデルはSWIM-Xと呼ばれ、人間に指示された高密度検索モデルと競合する。
論文 参考訳(メタデータ) (2023-11-10T00:17:10Z) - Description-Based Text Similarity [59.552704474862004]
我々は、その内容の抽象的な記述に基づいて、テキストを検索する必要性を特定する。
そこで本研究では,近隣の標準探索で使用する場合の精度を大幅に向上する代替モデルを提案する。
論文 参考訳(メタデータ) (2023-05-21T17:14:31Z) - Large Language Models are Strong Zero-Shot Retriever [89.16756291653371]
ゼロショットシナリオにおける大規模検索に大規模言語モデル(LLM)を適用するための簡単な手法を提案する。
我々の手法であるRetriever(LameR)は,LLM以外のニューラルモデルに基づいて構築された言語モデルである。
論文 参考訳(メタデータ) (2023-04-27T14:45:55Z) - Query2doc: Query Expansion with Large Language Models [69.9707552694766]
提案手法はまず,大言語モデル (LLM) をプロンプトすることで擬似文書を生成する。
query2docは、アドホックIRデータセットでBM25のパフォーマンスを3%から15%向上させる。
また,本手法は,ドメイン内およびドメイン外の両方において,最先端の高密度検索に有効である。
論文 参考訳(メタデータ) (2023-03-14T07:27:30Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。