論文の概要: MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
- arxiv url: http://arxiv.org/abs/2608.03397v1
- Date: Tue, 04 Aug 2026 09:51:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-05 15:30:23.11813
- Title: MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
- Title(参考訳): MMLongBench-Doc-V2:MMLongBench-Docの修正
- Authors: Mingtian Zhang,
- Abstract要約: MMLong-Docは135のPDFに対する1,082の質問のQAベンチマークである。
MMLongBench-Doc-V2は106のアノテーションを修正する。
- 参考スコア(独自算出の注目度): 10.169759525075792
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they are meant to capture: the reference metric compares extracted answers, so 1,358,000 loses to 1358000; and a non-trivial share of ground-truth annotations are wrong, ambiguous, or incomplete --- concentrated, because of how they were found, in exactly the questions capable systems answer correctly. MMLongBench-Doc-V2 corrects 106 annotations, each published with the page and arithmetic that settle it, and replaces the string metric with a pinned LLM judge asked whether a response means the reference. Ten questions whose document ships under the wrong filename are removed rather than counted wrong, along with one duplicated question, leaving 1,071 questions over 134 documents. The most reusable contribution is a decision procedure for when an empty set key may be widened and when widening would destroy a deliberate negative sample; applied to all 208 rows, it widened 14. V2 scores are not comparable with published V1 numbers. The corrected corpus, the per-entry correction record and the evaluation harness are available at https://github.com/VectifyAI/MMLongBench-Doc-V2.
- Abstract(参考訳): MMLongBench-Docは135のPDFに対する1,082の質問の長期文書QAベンチマークである。
基準測度は抽出された回答を比較するので、1,358,000 は 1358000 に減少し、非自明な接尾辞アノテーションのシェアは、どのように発見されたか、不明瞭であるか、あるいは不完全である。
MMLongBench-Doc-V2は106のアノテーションを修正し、それぞれのアノテーションをページと演算で公開し、文字列のメトリックをピン留めされたLCM判事に置き換える。
間違ったファイル名の下の文書を削除した10の質問と、重複した1つの質問が134の文書に1,071の質問を残している。
最も再利用可能なコントリビューションは、空のセットキーが拡張され、拡張が故意の負のサンプルを破壊する場合の決定手順であり、全208行に適用され、14行に拡張された。
V2スコアは公表されたV1数字に匹敵するものではない。
修正されたコーパス、エントリー毎の補正レコード、評価ハーネスはhttps://github.com/VectifyAI/MMLongBench-Doc-V2で入手できる。
関連論文リスト
- Inject or Navigate? Token-Efficient Retrieval for LLM Analysis of Transactional Legal Documents [0.0]
フルコーパス注入は検索リコールを最大化するが、トークンのフットプリントは問題ではなくコーパスとスケールする。
NAVINDEXは全18機に1.61倍のトークンフットプリント、56倍の回答コンテキスト、25%のコストで接続されたと判断された。
キャッシュインジェクションは、コーパスが検索ペイロードの約10倍以下にとどまっている間のみ、ドルで安価である。
論文 参考訳(メタデータ) (2026-07-07T02:42:06Z) - CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence [35.06938031991452]
CiteVQAは要素レベルのバウンディングボックスの引用を各回答と一緒に返さなければならないベンチマークである。
CiteVQAは7つのドメインと2つの言語にまたがる711のPDFに1,897の質問があり、1ドキュメントあたり平均40.6ページである。
Strict Attributed Accuracy (SAA) は、回答と引用された領域の両方が正しい場合にのみ、予測を信用する。
論文 参考訳(メタデータ) (2026-05-13T01:54:42Z) - ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge [50.93758649363798]
Impliretは、推論の課題をドキュメント側処理にシフトするベンチマークである。
我々は,この環境下で苦戦している,疎水・密集したレトリバーの幅を評価した。
論文 参考訳(メタデータ) (2025-06-17T11:08:29Z) - MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations [105.10376440302076]
MMLongBench-Doc は 1,062 のエキスパート注釈付き質問を含む長文マルチモーダルベンチマークである。
130の長いPDFフォーマットの文書の上に構築されており、平均49.4ページと20,971のテキストトークンがある。
14個のLVLMの実験により、長いコンテキストのDUが現在のモデルに大きく挑戦することを示した。
論文 参考訳(メタデータ) (2024-07-01T17:59:26Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。