論文の概要: Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
- arxiv url: http://arxiv.org/abs/2606.01393v1
- Date: Sun, 31 May 2026 18:35:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-02 21:34:29.679
- Title: Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
- Title(参考訳): Dr. DocBench: エキスパートレベルで難解な文書解析のための総合ベンチマーク
- Authors: Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo, Zhenting Qi, Konwoo Kim, Longtian Ye, Xiaolong Luo, Jinhe Bi, Henry Zhang, Haris Riaz, Xuan Zhang, Yunze Xiao, Bangya Liu, Tom Tang, Yunfei Zhao, Qunshu Lin, Zihan Wang, Minghao Liu, Michael Lingzhi Li, Yilun Du, Jesse Thomason, Rogerio Feris, Alex Pentland, Zexue He,
- Abstract要約: 我々は、エキスパートレベルの文書解析のための困難を意識したベンチマークであるDocBench博士を紹介する。
Dr. DocBenchは52のBISACドメインにまたがり、障害ベースのサンプリングによってドキュメントを選択する。
約100ページにわたる長いドキュメントから4,514ページの注釈付きページが含まれており、レイアウト、読み込み順序、階層的関係、ドメイン固有のビジュアルコンテンツなど、65kの高品質なアノテーションがある。
本分析では,文書インテリジェンスを診断・進展するための総合的なテストベッドとしてDocBench博士が注目されている。
- 参考スコア(独自算出の注目度): 53.41293908252118
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are increasingly limited in coverage and difficulty: many focus on common document genres or uniformly sampled pages where modern parsers already perform strongly, while offering limited annotation for expert-domain structures such as chemical formula, music notation, complex tables, and cross-page layouts. We introduce Dr. DocBench, a difficulty-aware benchmark for expert-level document parsing. Built from a large-scale multilingual book corpus, Dr. DocBench spans 52 BISAC subject domains and selects challenging documents through parser-failure-based sampling, targeting cases where multiple state-of-the-art systems struggle. It contains 4,514 annotated pages from long documents averaging around 100 pages, with 65k high-quality page- and block-level annotations for layout, reading order, hierarchical relations, and domain-specific visual contents. Evaluations of pipeline-based parsers and general-purpose VLMs show that strong performance on existing benchmarks does not transfer to our expert-level document parsing. Our analysis reveals substantial failures across subjects, content types, and structural attributes, highlighting Dr. DocBench as a comprehensive testbed for diagnosing and advancing document intelligence.
- Abstract(参考訳): 文書解析と認識は視覚言語モデル(VLM)と文書処理システムの基本機能である。
しかし、既存の光学文字認識(OCR)と文書解析ベンチマークは、カバー範囲と難易度にますます制限されている: 一般的な文書ジャンルや、現代のパーサーがすでに強力に機能している一様にサンプリングされたページに焦点を当て、化学式、音楽記法、複雑な表、クロスページレイアウトなどの専門家ドメイン構造に対して限定的なアノテーションを提供している。
我々は、エキスパートレベルの文書解析のための困難を意識したベンチマークであるDocBench博士を紹介する。
DocBenchは、大規模多言語書籍コーパスから構築され、52のBISAC主題ドメインにまたがり、パーサフェイルベースのサンプリングを通じて、複数の最先端システムが苦労するケースをターゲットとして、挑戦的なドキュメントを選択する。
約100ページにわたる長いドキュメントから4,514ページの注釈付きページが含まれており、レイアウト、読み込み順序、階層的関係、ドメイン固有のビジュアルコンテンツのための65kの高品質なページとブロックレベルのアノテーションがある。
パイプラインベースのパーサと汎用VLMの評価は、既存のベンチマークの性能が専門家レベルの文書解析に移行していないことを示している。
本分析では,文書インテリジェンスを診断・進展するための総合的なテストベッドとしてDocBench博士が注目されている。
関連論文リスト
- MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing [74.84107522458798]
MPDocBench-Parseは、現実世界のアプリケーションにおけるマルチページ文書解析のためのベンチマークである。
433の注釈付き文書に3,246ページあり、英語と中国語の15種類の文書を網羅しており、レイアウトは様々である。
論文 参考訳(メタデータ) (2026-05-21T07:36:41Z) - DISCO: Document Intelligence Suite for COmparative Evaluation [1.4425299138308667]
ドキュメントインテリジェンスには、正確なテキスト抽出と、文書コンテンツに対する信頼性の高い推論が必要である。
光文字認識 (OCR) パイプラインと視覚言語モデル (VLM) を個別に評価し, 多様な文書タイプにまたがる解析と質問応答について検討した。
論文 参考訳(メタデータ) (2026-03-04T14:47:34Z) - WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? [64.62909376834601]
本稿では,自然環境における文書理解の評価に特化して設計されたWildDocについて紹介する。
WildDoc上での最先端MLLMの評価は、従来のベンチマークと比べて性能が大幅に低下し、モデルの頑健さが不十分であることを示す。
論文 参考訳(メタデータ) (2025-05-16T09:09:46Z) - OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations [22.336858733121158]
OmniDocBenchは9つのドキュメントソースにまたがる高品質なアノテーションを特徴とする新しいベンチマークです。
パイプラインベースの手法とエンドツーエンドのビジョン言語モデルの両方を徹底的に評価する。
論文 参考訳(メタデータ) (2024-12-10T16:05:56Z) - DocBank: A Benchmark Dataset for Document Layout Analysis [114.81155155508083]
文書レイアウト解析のための詳細なトークンレベルのアノテーションを備えた500Kドキュメントページを含むベンチマークデータセットである textbfDocBank を提示する。
実験の結果,DocBankでトレーニングされたモデルは,さまざまなドキュメントのレイアウト情報を正確に認識することがわかった。
論文 参考訳(メタデータ) (2020-06-01T16:04:30Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。