論文の概要: DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
- arxiv url: http://arxiv.org/abs/2605.04458v1
- Date: Wed, 06 May 2026 03:34:46 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-07 18:41:07.627732
- Title: DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
- Title(参考訳): DoGMaTiQ:レポート評価のための質問応答ナゲットの自動生成
- Abstract要約: DoGMaTiQは高品質なQAベースのナゲットセットを生成するパイプラインである。
生成したレポートの完全自動評価を可能にするために,DoGMaTiQとAutoArgueを統合した。
- 参考スコア(独自算出の注目度): 41.38684270729433
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Evaluation of long-form, citation-backed reports has lately received significant attention due to the wide-scale adoption of retrieval-augmented generation (RAG) systems. Core to many evaluation frameworks is the use of atomic facts, or nuggets, to assess a report's coverage of query-relevant information attested in the underlying collection. While nuggets have traditionally been represented as short statements, recent work has used question-answer (QA) representations, enabling fine-grained evaluations that decouple the information need (i.e. the question) from the potentially diverse content that satisfies it (i.e. its answers). A persistent challenge for nugget-based evaluation is the need to manually curate sets of nuggets for each topic in a test collection -- a laborious process that scales poorly to novel information needs. This challenge is acute in cross-lingual settings, where information is found in multilingual source documents. Accordingly, we introduce DoGMaTiQ, a pipeline for generating high-quality QA-based nugget sets in three stages: (1) document-grounded nugget generation, (2) paraphrase clustering, and (3) nugget subselection based on principled quality criteria. We integrate DoGMaTiQ nuggets with AutoArgue -- a recent nugget-based evaluation framework -- to enable fully automatic evaluation of generated reports. We conduct extensive experiments on two cross-lingual TREC shared tasks, NeuCLIR and RAGTIME, showing strong rank correlations with both human-in-the-loop and fully manual judgments. Finally, detailed analysis of our pipeline reveals that a strong LLM nugget generator is key, and that the system rankings induced by DoGMaTiQ are robust to outlier systems. We facilitate future research in report evaluation by publicly releasing our code and artifacts at https://github.com/manestay/dogmatiq.
- Abstract(参考訳): 近年,RAG (Research-augmented Generation) システムが広く採用されているため,長期にわたる引用型レポートの評価に注目が集まっている。
多くの評価フレームワークの中核となるのは、基礎となるコレクションで証明されたクエリ関連情報に関するレポートのカバレッジを評価するために、アトミックな事実(あるいはナゲット)を使用することである。
ナゲットは伝統的にショートステートメントとして表現されてきたが、最近の研究は質問回答(QA)表現を使用しており、情報の必要性(すなわち質問)とそれを満たす潜在的に多様なコンテンツ(すなわちその答え)を分離するきめ細かい評価を可能にしている。
ナゲットに基づく評価のための永続的な課題は、テストコレクションで各トピックのナゲットセットを手作業でキュレートする必要があることだ。
この課題は、多言語ソースドキュメントに情報がある言語間設定において、急性である。
そこで我々は,(1)文書グラウンドド・ナゲット生成,(2)パラフレーズクラスタリング,(3)原理的品質基準に基づくナゲットサブセレクションという,高品質なQAベースのナゲットセットを生成するパイプラインであるDoGMaTiQを紹介した。
最近のNuggetベースの評価フレームワークであるAutoArgueとDoGMaTiQナゲットを統合して、生成されたレポートの完全な自動評価を可能にします。
我々は,2つの言語間TREC共有タスクであるNeuCLIRとRAGTIMEについて広範囲に実験を行い,人手による判断と人手による判断の相関性を示した。
最後に,我々のパイプラインの詳細な解析から,強力なLLMナゲット生成器が鍵であり,DoGMaTiQによって誘導されるシステムランキングが,システムから外れやすいことを示す。
我々は、コードとアーティファクトをhttps://github.com/manestay/dogmatiq.comで公開することで、レポート評価における将来の研究を促進する。
関連論文リスト
- Incorporating Q&A Nuggets into Retrieval-Augmented Generation [23.32167679162754]
CrucibleはNugget-Augmented Generation Systemであり、取得した文書からQ&Aナゲットの銀行を構築することで、明示的な引用の証明を保存する。
ナゲットの推論は、明確で解釈可能なQ&Aセマンティクスを通じて繰り返し情報を避ける。
我々のシステムは,近年のナゲットベースRAGシステムであるGingerを,ナゲットリコール,密度,励振グラウンドリングで大きく上回っている。
論文 参考訳(メタデータ) (2026-01-19T16:57:33Z) - ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering [54.72902502486611]
ReAG(Reasoning-Augmented Multimodal RAG)は、粗い部分ときめ細かい部分の検索と、無関係な通路をフィルタリングする批評家モデルを組み合わせた手法である。
ReAGは従来の手法よりも優れており、解答精度が向上し、検索された証拠に根ざした解釈可能な推論を提供する。
論文 参考訳(メタデータ) (2025-11-27T19:01:02Z) - Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses [45.2769075498271]
当社のAutoNuggetizerフレームワークを使用して,LMArenaが提供する約7Kの検索アリーナバトルからのデータを分析する。
その結果,ナゲットスコアとヒトの嗜好との間に有意な相関が認められた。
論文 参考訳(メタデータ) (2025-04-28T17:24:36Z) - The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models [53.12387628636912]
本稿では,人間のアノテーションに対して評価を行う自動評価フレームワークを提案する。
この手法は2003年にTREC Question Answering (QA) Trackのために開発された。
完全自動ナゲット評価から得られるスコアと人間に基づく変種とのランニングレベルでの強い一致を観察する。
論文 参考訳(メタデータ) (2025-04-21T12:55:06Z) - Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets [8.734527090842139]
本稿では,RAGシステムによって生成された応答の検索効率と関連性を評価するための新しいリソースを提案する。
我々のデータセットは、TREC iKAT 2024コレクションに拡張され、17の会話と20,575の関連パスアセスメントを含む。
論文 参考訳(メタデータ) (2025-03-12T23:44:10Z) - Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines [17.803396998387665]
Retrieval-augmented Generation (RAG)は、知識集約型視覚質問応答(VQA)タスクに対処するために登場した。
本稿では,知識に基づくVQAタスクに対する従来のRAGモデルの代替としてReAuSEを提案する。
我々のモデルは生成型検索器と正確な回答生成器の両方として機能する。
論文 参考訳(メタデータ) (2025-02-23T16:39:39Z) - Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework [53.12387628636912]
本報告では、TREC 2024 Retrieval-Augmented Generation (RAG) Trackの部分的な結果について概説する。
我々は、情報アクセスの継続的な進歩の障壁としてRAG評価を特定した。
論文 参考訳(メタデータ) (2024-11-14T17:25:43Z) - PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models [72.57329554067195]
ProxyQAは、長文生成を評価するための革新的なフレームワークである。
さまざまなドメインにまたがる詳細なヒューマンキュレートされたメタクエストで構成されており、それぞれに事前にアノテートされた回答を持つ特定のプロキシクエストが伴っている。
プロキシクエリに対処する際の評価器の精度を通じて、生成されたコンテンツの品質を評価する。
論文 参考訳(メタデータ) (2024-01-26T18:12:25Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。