論文の概要: Thinking with Anchors: Grounded and Efficient Document Reasoning
- arxiv url: http://arxiv.org/abs/2608.04424v1
- Date: Wed, 05 Aug 2026 04:09:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.707229
- Title: Thinking with Anchors: Grounded and Efficient Document Reasoning
- Title(参考訳): アンカーで考える - 基礎的で効率的なドキュメント推論
- Authors: Sichen Zhu, Yuchen Zhu, Wenzhuo Xu, Jason Kuen, Wanrong Zhu, Jing Shi, Xuan Shen, Quanyi Wang, Yiwei Wang, Yujun Cai, Bing Shuai, Qin Zhang, Yongxin Chen, Shilong Liu, Molei Tao, Jiuxiang Gu,
- Abstract要約: ADOPD 2026はADOPDの拡張である。
ADOPD 2026はページ分解を空間的に基底化された文書理解に変換する。
ADOPD 2026は、ページ分解を検証可能なビジュアル・アンカー推論に接続することで、タスク・フレームワークを提供する。
- 参考スコア(独自算出の注目度): 87.07530736788898
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region semantics, spatial relations, and visual structure. We present ADOPD 2026, a reasoning-oriented extension of ADOPD that turns page decomposition into spatially grounded document understanding. ADOPD 2026 enriches page anchors inherited from ADOPD 2024 dataset with human-cleaned captions, semantic tags, and generated chain-of-thought (CoT) traces grounded to document regions. Instead of treating boxes, masks, and tags as independent supervision signals, we cast text blocks, visual entities, semantic labels, bounding boxes, and polygon masks as a shared vocabulary of visual anchors. This representation supports three connected capabilities. First, region-level semantic tagging asks models to identify document element types from both page context and local appearance, revealing long-tail semantic failures that standard layout benchmarks often hide. Second, unified vision-language grounding generates text regions and visual entities together with coordinates or polygonal outlines, transforming detection and segmentation outputs into structured anchors that can be reused by downstream reasoning systems. Third, current state-of-the-art models still struggle with dense counting tasks evaluated on DocCount, a benchmark derived from ADOPD 2026, highlighting the need for the Thinking-with-Anchors pipeline in document semantic understanding. By connecting page decomposition to verifiable visual-anchor reasoning, ADOPD 2026 provides a task framework that moves document understanding beyond localization toward anchor-grounded document intelligence.
- Abstract(参考訳): 既存の文書理解ベンチマークはページ要素の探索に重点を置いているが、現実の文書インテリジェンスには、領域の意味論、空間関係、視覚構造について共同で推論するモデルが必要である。
ADOPD 2026は、ページ分解を空間的に基底化された文書理解に変換するADOPDの拡張である。
ADOPD 2026は、ADOPD 2024データセットから継承されたページアンカーを、人間のクリーニングされたキャプション、セマンティックタグ、およびドキュメント領域に根ざした生成されたチェーン・オブ・シークレット(CoT)トレースで強化する。
ボックス、マスク、タグを独立した監視信号として扱う代わりに、テキストブロック、ビジュアルエンティティ、セマンティックラベル、バウンディングボックス、ポリゴンマスクを視覚アンカーの共有語彙としてキャストした。
この表現は3つの接続機能をサポートする。
まず、リージョンレベルのセマンティックタグは、ページコンテキストとローカル外観の両方からドキュメント要素タイプを識別するようモデルに要求する。
第二に、統合された視覚言語基底は、座標や多角形アウトラインと共にテキスト領域と視覚エンティティを生成し、検出とセグメンテーション出力を下流の推論システムで再利用可能な構造化アンカーに変換する。
第3に、現在の最先端モデルは、ドキュメントセマンティック理解におけるThinking-with-Anchorsパイプラインの必要性を強調した、DOPD 2026から派生したベンチマークであるDocCountで評価された密カウントタスクに依然として苦労している。
ADOPD 2026は、ページ分解を検証可能なビジュアルアンカー推論に接続することにより、文書理解をローカライゼーションを超えてアンカーグラウンドド文書インテリジェンスに移行するタスクフレームワークを提供する。
関連論文リスト
- MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing [74.84107522458798]
MPDocBench-Parseは、現実世界のアプリケーションにおけるマルチページ文書解析のためのベンチマークである。
433の注釈付き文書に3,246ページあり、英語と中国語の15種類の文書を網羅しており、レイアウトは様々である。
論文 参考訳(メタデータ) (2026-05-21T07:36:41Z) - Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation [19.889854990300595]
反復検索拡張生成(iRAG)は、複雑なマルチホップ問題に答える強力なパラダイムとして登場した。
Evidence (CoE) の textbfChain について述べる。
論文 参考訳(メタデータ) (2026-05-02T06:40:42Z) - ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment [28.897559367200376]
ビジュアルドキュメント検索は、視覚的にリッチなコレクションからクエリに関連するドキュメントページの集合を検索することを目的としている。
既存の手法では、クエリやビジュアルページを共有埋め込み空間にエンコードするために、VLM(Vision-Language Models)を用いることが多い。
そこで我々は,Reasoning-Guided Alignment (ReAlign)を提案する。
論文 参考訳(メタデータ) (2026-04-08T14:47:27Z) - MoDora: Tree-Based Semi-Structured Document Analysis System [62.01015188258797]
半構造化文書は、様々な不規則なレイアウトで配置された様々なインターリーブされたデータ要素を統合する。
MoDora は半構造化文書解析のための LLM を利用したシステムである。
実験では、MoDoraは5.97%-61.07%の精度でベースラインを上回っている。
論文 参考訳(メタデータ) (2026-02-26T14:48:49Z) - Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding [42.506971197471195]
ドキュメント解析のために約3.8Mの事前学習データペアで構成されるDocMark-Pileと、グラウンドド命令に従うための624kの微調整データアノテーションを備えたDocMark-Instructの2つのきめ細かい構造化データセットを紹介した。
提案手法は,様々なビジュアル文書理解ベンチマークにおいて,既存の最先端MLLMを著しく上回っている。
論文 参考訳(メタデータ) (2025-05-08T17:37:36Z) - Hypergraph based Understanding for Document Semantic Entity Recognition [65.84258776834524]
我々は,ハイパグラフアテンションを利用したハイパグラフアテンション文書セマンティックエンティティ認識フレームワークHGAを構築し,エンティティ境界とエンティティカテゴリを同時に重視する。
FUNSD, CORD, XFUNDIE で得られた結果は,本手法が意味的エンティティ認識タスクの性能を効果的に向上できることを示す。
論文 参考訳(メタデータ) (2024-07-09T14:35:49Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。