Fugu-MT 論文翻訳(概要): RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

論文の概要: RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

arxiv url: http://arxiv.org/abs/2606.23344v1
Date: Mon, 22 Jun 2026 13:48:39 GMT
ステータス: 翻訳完了
システム内更新日: 2026-06-24 22:29:34.90136
Title: RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild
Title（参考訳）: RT-DocLayout: リアルタイムのエンドツーエンドドキュメントレイアウト解析
Authors: Cheng Cui, Tingquan Gao, Xueqing Wang, Changda Zhou, Hongen Liu, Ting Sun, Yubo Zhang, Zelun Zhang, Jiaxuan Liu, Manhui Lin, Yue Zhang, Suyin Liang, Yiqing Xiang, Yi Liu,
Abstract要約: 文書レイアウト解析のための高効率なエンドツーエンドフレームワークRT-Docを提案する。提案モデルは,レイアウト要素の分類,画素レベルのセグメンテーション,幾何学的順序順予測を統一する。 RT-Docは、フルドキュメントの再構築品質を大幅に改善し、現実世界の文書インテリジェンスシステムのスケーラブルで実用的な基盤を提供する。
参考スコア（独自算出の注目度）: 14.715243408844058
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document layout elements, geometric distortions (\eg, paper warping and bending, perspective variations), and reading order within diverse layout structures. Existing approaches typically rely on fragmented multi-stage pipelines or computationally heavy generative Transformer architectures, leading to error propagation and limited efficiency. In this paper, we present RT-DocLayout, a highly efficient end-to-end framework for document layout analysis, designed as a front-end for document parsing tasks. The proposed model unifies classification, detection, pixel-level segmentation, and reading order prediction for layout elements within a single 33M-parameter architecture. Built upon the RT-DETR, our key contribution is a unified multi-task formulation within a single query-based decoder that simultaneously classifies, regresses bounding box, generates masks, and constructs relationship to reason reading order. By jointly learning geometric and structural representations, RT-DocLayout introduces multi-task optimization that substantially improves robustness under real-world document distortions. Extensive experiments on public benchmarks demonstrate state-of-the-art performance in document layout analysis while maintaining real-time inference speed(132.1 FPS). When coupled with downstream OCR engines, RT-DocLayout significantly improves full-document reconstruction quality, providing a scalable and practical foundation for real-world document intelligence systems.
Abstract（参考訳）: 不均一な文書レイアウト要素間の複雑な結合、幾何学的歪み(例えば、紙の反りと曲げ、視線の変化)、様々なレイアウト構造における読み順などにより、正確な文書レイアウト解析は文書解析システムにとって依然として重要なボトルネックとなっている。既存のアプローチは通常、断片化されたマルチステージパイプラインや計算的に重いトランスフォーマーアーキテクチャに依存しており、エラーの伝播と効率の制限につながっている。本稿では,文書レイアウト解析のための高効率なエンドツーエンドフレームワークであるRT-DocLayoutについて述べる。提案モデルでは,1つの33Mパラメータアーキテクチャ内のレイアウト要素の分類,検出,画素レベルのセグメンテーション,読み出し順序予測を統一する。 RT-DETRをベースとして構築されたキーコントリビューションは,単一クエリベースのデコーダ内でのマルチタスクの統一化である。 RT-DocLayoutは、幾何学的および構造的表現を共同学習することにより、実世界の文書歪みに対するロバスト性を大幅に改善するマルチタスク最適化を導入している。公開ベンチマークでの大規模な実験は、リアルタイム推論速度(132.1 FPS)を維持しながら、文書レイアウト解析における最先端の性能を示す。下流のOCRエンジンと組み合わせることで、RT-DocLayoutはフルドキュメントの再構築品質を大幅に改善し、現実世界の文書インテリジェンスシステムのスケーラブルで実用的な基盤を提供する。

論文の概要: RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

関連論文リスト