論文の概要: Knossos and Ariadne: Benchmarking and Learning Complete Diagram Topology Extraction with Vision-Language Models
- arxiv url: http://arxiv.org/abs/2610.04721v1
- Date: Sat, 03 Oct 2026 19:20:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-11 17:26:42.918418
- Title: Knossos and Ariadne: Benchmarking and Learning Complete Diagram Topology Extraction with Vision-Language Models
- Title(参考訳): KnossosとAriadne:視覚言語モデルによる全図トポロジ抽出のベンチマークと学習
- Abstract要約: Knossosは、6つの異なる領域にわたる19,200のダイアグラムのベンチマークである。
Ariadneは、タスクをノードインベントリ抽出とソース条件のエッジ予測に分解する構造化フレームワークである。
- 参考スコア(独自算出の注目度): 64.2756308051623
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Structural diagrams are widely used to represent complex systems and relational information across scientific, engineering, procedural, and spatial domains. Recent vision-language models (VLMs) have become increasingly capable of recognizing diagram elements and reasoning about their content, while complete diagram topology extraction remains comparatively underexplored. In this paper, we study diagram-to-graph topology extraction: extracting all diagram entities and the complete relations among them. To enable large-scale supervised training and systematic evaluation of this task, we introduce Knossos, a benchmark of 19,200 diagrams across six diverse domains, with 245,179 nodes and 439,740 edges. Its symbolic generation process provides exact alignment between rendered diagrams and annotations of complete topology, relation types, and connector geometry. To address the modeling challenge of complete topology extraction, we also present Ariadne, a structured framework that decomposes the task into node inventory extraction and source-conditioned edge prediction. Extensive experiments show that training on Knossos substantially improves complete topology extraction in smaller open-source VLMs. Ariadne further improves over one-step extraction under matched supervision, demonstrating the additional benefit of structured decomposition. It achieves the highest average Edge F1 among the evaluated methods on Knossos, while both backbone variants also improve over their unadapted counterparts on the real-world external benchmark. Code and benchmark are available at https://github.com/bangwayne/knossos_Ariadne_Public.
- Abstract(参考訳): 構造図は、科学的、工学的、手続き的、空間的な領域にまたがる複雑なシステムや関係情報を表現するために広く使われている。
近年の視覚言語モデル (VLM) では, 図要素の認識や内容の推論が可能になっているが, 完全な図形トポロジー抽出はいまだに未探索である。
本稿では,図形から図形までのトポロジ抽出について検討する。
Knossosは6つのドメインにまたがる19,200のダイアグラムのベンチマークであり、245,179のノードと439,740のエッジを持つ。
その記号生成プロセスは、描画図形と完全なトポロジー、関係型、コネクター幾何学のアノテーションの正確なアライメントを提供する。
完全トポロジ抽出のモデル化課題に対処するために,タスクをノードインベントリ抽出とソース条件付きエッジ予測に分解する構造化フレームワークであるAriadneを提案する。
大規模な実験により、Knossosのトレーニングは、より小さなオープンソースVLMにおける完全なトポロジー抽出を大幅に改善することが示された。
アリアドンは、一致した監督下での一段階の抽出をさらに改善し、構造化された分解のさらなる利点を示す。
Knossosの評価手法の中では、Edge F1の平均値が最も高く、両方のバックボーン変種は、実世界の外部ベンチマークにおいて、適応していないものよりも改善されている。
コードとベンチマークはhttps://github.com/bangwayne/knossos_Ariadne_Public.orgで公開されている。
関連論文リスト
- TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models [47.73707314147262]
TopoBench-180は、図-グラフトポロジー抽出のための人間検証ベンチマークである。
TopoAgentは、信頼できるトポロジ抽出のための構造認識と推論のためのフレームワークである。
論文 参考訳(メタデータ) (2026-08-27T15:40:27Z) - TopoFormer: Topology Meets Attention for Graph Learning [8.679678575739304]
Topoformerは、トポロジ的構造を注意に優しいシーケンスにエンコードするグラフ表現学習のフレームワークである。
提案手法のコアとなるTopo-Scanは,グラフを短い順序付きトポロジカルトークン列に分解する新しいモジュールである。
論文 参考訳(メタデータ) (2026-07-30T14:18:17Z) - Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning [32.78218766121055]
グラフ検索拡張生成(GraphRAG)は,複雑な推論において,大規模言語モデルを効果的に拡張した。
本稿では,フレームワーク全体を複雑な統合として結合する,垂直に統一されたエージェントパラダイムYoutu-GraphRAGを提案する。
論文 参考訳(メタデータ) (2025-08-27T13:13:20Z) - Human as Points: Explicit Point-based 3D Human Reconstruction from Single-view RGB Images [71.91424164693422]
我々はHaPと呼ばれる明示的なポイントベース人間再構築フレームワークを導入する。
提案手法は,3次元幾何学空間における完全明示的な点雲推定,操作,生成,洗練が特徴である。
我々の結果は、完全に明示的で幾何学中心のアルゴリズム設計へのパラダイムのロールバックを示すかもしれない。
論文 参考訳(メタデータ) (2023-11-06T05:52:29Z) - A Comprehensive Study on Large-Scale Graph Training: Benchmarking and
Rethinking [124.21408098724551]
グラフニューラルネットワーク(GNN)の大規模グラフトレーニングは、非常に難しい問題である
本稿では,既存の問題に対処するため,EnGCNという新たなアンサンブルトレーニング手法を提案する。
提案手法は,大規模データセット上でのSOTA(State-of-the-art)の性能向上を実現している。
論文 参考訳(メタデータ) (2022-10-14T03:43:05Z) - Spatio-Temporal Inception Graph Convolutional Networks for
Skeleton-Based Action Recognition [126.51241919472356]
我々はスケルトンに基づく行動認識のためのシンプルで高度にモジュール化されたグラフ畳み込みネットワークアーキテクチャを設計する。
ネットワークは,空間的および時間的経路から多粒度情報を集約するビルディングブロックを繰り返すことで構築される。
論文 参考訳(メタデータ) (2020-11-26T14:43:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。