論文の概要: MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
- arxiv url: http://arxiv.org/abs/2607.11562v1
- Date: Mon, 13 Jul 2026 13:43:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 17:47:21.485432
- Title: MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
- Title(参考訳): MonkeyOCRv2: ドキュメントAIのためのビジュアルテキスト基盤モデル
- Authors: Yuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, Qiang Liu, Jiajun Song, Zidun Guo, Xinhan Wang, Handong Zheng, Yang Liu, Dongliang Luo, Zhiyin Ma, Jiarui Zhang, Xiang Bai,
- Abstract要約: MonkeyOCRv2は、ドキュメントAIのためのビジュアルテキスト事前トレーニングモデルである。
MonkeyDoc v2は17言語にまたがる1億1300万のイメージからなる、最大規模のドキュメントイメージ事前トレーニングコーパスである。
- 参考スコア(独自算出の注目度): 71.06391669976786
- License: http://creativecommons.org/publicdomain/zero/1.0/
- Abstract: Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.
- Abstract(参考訳): 主流のビジュアルエンコーダは、自然画像上で事前訓練されており、文書指向の適応なしには文書画像に効果的に適用できない。
本稿では,文書AIのためのビジュアルテキスト事前学習モデルであるMonkeyOCRv2を提案する。
まず、MonkeyDoc v2を構築し、17言語にまたがる1億1300万のイメージからなる、最大のドキュメントイメージ事前学習コーパスについて知る。
第2に,画像からテキストまでの文書生成と画素レベルの文書再構成を共同で学習する事前学習戦略を提案する。
テキスト認識, 公式認識, テキスト検出, 文書改ざん検出, 重なり合うテキストセグメンテーションを含む5つの代表的な文書解析タスクについて, 広範囲にわたる実験を行った。
MonkeyOCRv2でオリジナルのエンコーダをリプレースすることで,5つのタスクすべてでパフォーマンスが一貫して向上する。
最後に,文書解析と文書理解の課題に対して,マルチモーダルな大規模言語モデルの視覚エンコーダとしての有効性を検証する。
凍結して軽量な言語モデルと組み合わせることで、0.7Bの文書解析モデルが作成され、デジタル生まれで写真化された文書を17言語にまたがる最新のベンチマークであるMDPBenchで、約11$\times$小額の視覚エンコーダで、以前の最高の3Bドットを2.8%上回った。
凍結エンコーダは、同じトレーニング設定下で8つのベンチマークでCLIP、DINO、SAMで構築されたものを上回るドキュメント理解モデルも備えている。
これらの結果は,文書指向の視覚的事前学習が,文書インテリジェンスの基礎となることを示唆している。
関連論文リスト
- Multimodal OCR: Parse Anything from Documents [72.69545534962234]
dots.mocrは、チャート、ダイアグラム、テーブル、アイコンなどのビジュアル要素を第一級解析ターゲットとして扱う。
テキストとグラフィックの両方を構造化出力として再構築し、より忠実なドキュメント再構築を可能にする。
不均一なドキュメント要素に対するエンドツーエンドのトレーニングをサポートする。
論文 参考訳(メタデータ) (2026-03-13T14:42:21Z) - SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models [17.85605201420847]
Visual Document Retrieval (VDR) は通常、文書イメージを直接埋め込むために訓練された特殊なバイエンコーダを使用してテキストから画像の検索を行う。
我々はゼロショット生成・符号化パイプラインを再考し、まず視覚言語モデルを用いて各文書画像の詳細なテキスト記述を生成する。
ViDoRe-v2ベンチマークでは、63.4%のnDCG@5に達し、マルチベクトルビジュアルドキュメントエンコーダで最強である。
論文 参考訳(メタデータ) (2025-09-18T21:11:13Z) - DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents [18.080447065002392]
本稿では,文書内の画像と長文間の相互作用を理解するために,視覚言語事前学習モデルを強制するためのDocumentCLIPを提案する。
我々のモデルは、言語的にも視覚的にもリッチなコンテンツを含む、ニュース記事、雑誌、製品記述などの実世界のマルチモーダル文書理解にとって有益である。
論文 参考訳(メタデータ) (2023-06-09T23:51:11Z) - Unifying Vision, Text, and Layout for Universal Document Processing [105.36490575974028]
本稿では,テキスト,画像,レイアウトのモダリティを文書理解と生成を含むさまざまなタスク形式とともに統合するドキュメントAIモデルを提案する。
我々の手法は、財務報告、学術論文、ウェブサイトなど、さまざまなデータ領域にまたがって、文書理解やQAといった9つのドキュメントAIタスクの最先端を定めている。
論文 参考訳(メタデータ) (2022-12-05T22:14:49Z) - LayoutLMv3: Pre-training for Document AI with Unified Text and Image
Masking [83.09001231165985]
テキストと画像のマスキングを併用した文書AIのためのマルチモーダルトランスフォーマーを事前学習するためのLayoutLMv3を提案する。
単純な統一アーキテクチャとトレーニングの目的により、LayoutLMv3はテキスト中心および画像中心のDocument AIタスクの汎用的な事前トレーニングモデルになる。
論文 参考訳(メタデータ) (2022-04-18T16:19:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。