論文の概要: CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings
- arxiv url: http://arxiv.org/abs/2608.00473v1
- Date: Sat, 01 Aug 2026 06:51:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:24.682777
- Title: CrossProjection: Geometric Grounding Beyond Viewpoint Change in Architectural Drawings
- Title(参考訳): CrossProjection: 建築図面の視点変化を超えた幾何学的グラウンド
- Authors: Kaho Li, Pengyu Zeng, Yuqin Dai, Jun Yin, Tianjing Feng, Shuai Lu,
- Abstract要約: CrossProjectionは、視覚言語モデルがコンポーネントのアイデンティティを保持し、不均一なビューにわたって幾何学を外部化するかどうかをアンカーグラウンドで診断する。
分類的判断、候補選択、自由点、線、領域のローカライゼーションを通じてマッチング、登録、幾何学的グラウンドを評価する。
- 参考スコア(独自算出の注目度): 13.010314444320384
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Architectural drawings violate the usual assumption behind multi-view reasoning: plans and sections are cuts, while elevations are facade projections, so corresponding components change appearance in ways camera motion cannot explain. We introduce CrossProjection, an anchor-grounded diagnostic of whether vision-language models preserve component identity and externalize geometry across heterogeneous architectural views. It evaluates Matching, Registration, and Geometric Grounding through categorical judgments, candidate selection, and free point, line, and region localization. Across 23 real drawing sets and 1,954 categorical conditions per model, GPT-5.5 scores 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%. A matched 200-target study crosses natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. Candidate-supported performance is often higher, but free localization remains fragile: on natural drawings, point/region PCK@.05 is 54-76% for GPT, 8-10% for Qwen, and 14-36% for GLM; line endpoint PCK@.05 is 22%, 4%, and 0%. A coordinate grid recovers some GPT point/region precision but not lines. Three architecture-trained participants reach 87.3-93.3% categorical accuracy and 76-92% GT-region hit, supporting task feasibility rather than a population-level human ceiling. Because the categorical families do not form a same-item Matching-Registration contrast and interface controls alter multiple burdens, we avoid mechanistic claims. The supported conclusion is narrower: closed-choice or marked-element success does not entail reliable explicit geometric grounding. For drawing-guided CAD/BIM systems, categorical correctness should not be treated as evidence of candidate-free spatial reliability. Reusable on-sheet anchors, fixed-denominator scoring, and hash-locked artifacts establish an audit trail for this gap.
- Abstract(参考訳): 建築図面は、計画とセクションはカットされ、標高はファサードプロジェクションであるため、対応するコンポーネントはカメラの動きが説明できない方法で外観を変える。
クロスプロジェクション(CrossProjection)は、視覚言語モデルがコンポーネントのアイデンティティを保持し、異種アーキテクチャビュー全体にわたって幾何学を外部化するかどうかを、アンカーグラウンドで診断する手法である。
分類的判断、候補選択、自由点、線、領域のローカライゼーションを通じてマッチング、登録、幾何学的グラウンドを評価する。
23の実際の描画セットと1,954のカテゴリ条件、GPT-5.5のスコアは82.4%、Qwen3-VL-32B-Instruct 62.2%、GLM-4.5V 57.2%である。
一致した200ターゲットの研究は、クローズド・クローズド・クローズド・アンド・フリー・ジオメトリー・アウトプットで自然およびベクトルテキスト圧縮された描画を横断する。
自然な描画では、ポイント/リージョンPCK@.05は54-76%、Qwenは8-10%、GLMは14-36%、ラインエンドポイントPCK@.05は22%、4%、0%である。
座標格子はGPT点/領域精度を回復するが、線は回復しない。
建築訓練を受けた3人の参加者は87.3-93.3%のカテゴリー精度と76-92%のGT領域のヒットに達し、人口レベルの人体天井よりもタスク実現可能性を支持している。
分類族は一致登録コントラストを形成せず、インタフェース制御は複数の負担を伴わないため、機械的クレームは避ける。
支持された結論はより狭く、クローズドチョイスやマーク要素の成功は、信頼できる明示的な幾何学的根拠を伴わない。
図面誘導CAD/BIMシステムでは、カテゴリー的正当性は、候補のない空間的信頼性の証拠として扱うべきではない。
再使用可能なオンシートアンカー、固定デノミネータスコア、ハッシュロックされたアーティファクトは、このギャップの監査パスを確立する。
関連論文リスト
- Evidence-Grounded Constraint Checking in Construction Documents [1.3787753513429983]
抽出された事実を正規化し、4状態ルールを決定的に実行するエビデンス基底パイプラインを提案する。
我々は,29の建設プロジェクトから160の参照ベースタスクに対するPDFエビデンスアロケータの評価を行った。
論文 参考訳(メタデータ) (2026-07-31T06:25:41Z) - Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction [55.11480729304395]
PointArena 2026は77.2%の精度でベンチマークで2位である。
ap proachは3つの障害モードをターゲットにしている。第一に、エージェント駆動のシンセシスは大きなセマンティクスとアンカー相対的な候補プールを構築する。
次に、determinis tic steerable-dataパイプラインは、認証された10,000サンプルのメインセットと、マスク、テンプレート、パス検証を使用するリザーブサンプルを生成する。
論文 参考訳(メタデータ) (2026-06-29T06:39:03Z) - Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search [50.16356451328644]
シャノン型エントロピーの不等式を証明することは情報理論の基本的な課題である。
我々は,原子実証のステップを微調整した小規模大規模言語モデルがこのプロセスを自動化することができるか検討する。
GPT-5.5は0ショットプロンプトで1.7%のサンプルを解き、Psitipは33.3%のサンプルを解いた。
論文 参考訳(メタデータ) (2026-06-04T05:43:12Z) - A Modelling and Evaluation Framework for EuroCrops-Driven Sentinel-2 Crop Segmentation [78.66324246922831]
本研究では,Sentinel-2イメージとEuroCropsパーセルレベルのアノテーションからセマンティックセグメンテーション対応農業データセットを生成するパイプラインを提案する。
このデータセットには、ヨーロッパ5カ国から67,337のパッチが含まれており、10種類の作物と背景の分類を減らしている。
The four-level U-Net with Group Normalization were training using 10 Sentinel-2 spectrum bands and a Composite loss with class-weighted cross-entropy and Dice loss。
論文 参考訳(メタデータ) (2026-05-30T11:20:29Z) - Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction [59.47026381165585]
視覚言語モデルは精度を上げて幾何学的問題を解くが、その中間状態は潜伏し、検証できないままである。
我々は,GeoGebra制約エンジンとのエージェント相互作用に潜時空間推論から幾何推論を再キャストするフレームワークDraw2Thinkを提案する。
論文 参考訳(メタデータ) (2026-05-20T05:46:14Z) - Global and Local Topology-Aware Attention with Persistent Homology and Euler Biases for Time-Series Forecasting [0.0]
時系列はしばしば、接続性、サイクル、シェルのような幾何学、方向変化、非線形近傍を含む予測幾何学構造を符号化する。
永続的ホモロジー(H0-H2)を用いた注意ログにそのような構造を加えるトポロジ対応アテンションフレームワークを提案する。
我々は,軽量アテンション/ライダー,PatchTSTForRegression,TimeSeriesTransformerForPredictionという3つのアーキテクチャファミリの保護されたトポロジ対応のバリエーションを評価した。
論文 参考訳(メタデータ) (2026-05-04T21:19:12Z) - ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data [0.0]
構成条件付きポストトレーニングは、モデルが学習した表現幾何学の構造化摂動として分析することができる。
グラフ, モデル, 基板間の構成による隠れ状態構造をトレースする, 幾何学第一のプログラムATLASを紹介する。
論文 参考訳(メタデータ) (2026-04-19T23:26:02Z) - Do Foundation Models Know Geometry? Probing Frozen Features for Continuous Physical Measurement [0.0]
視覚言語モデルは、それらのテキスト経路が表現できないような幾何学を符号化する。
ロラ微調整(r=16, 2,000枚)は、このギャップを6.5度に縮める。
これらの知見は、単一の凍結したバックボーンがマルチタスク幾何学的センサーとして機能することを可能にした。
論文 参考訳(メタデータ) (2026-03-06T16:48:27Z) - Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales [61.03549470159347]
視覚言語モデル (VLM) は急速に進歩しているが, オープンワールド環境における画像位置決め能力は, 網羅的に評価されていない。
我々は、視覚認識、ステップバイステップ推論、エビデンス利用を評価するVLM画像位置情報の総合ベンチマークであるEarthWhereを提示する。
論文 参考訳(メタデータ) (2025-10-13T01:12:21Z) - Generalized Focal Loss: Learning Qualified and Distributed Bounding
Boxes for Dense Object Detection [85.53263670166304]
一段検出器は基本的に、物体検出を密度の高い分類と位置化として定式化する。
1段検出器の最近の傾向は、局所化の質を推定するために個別の予測分岐を導入することである。
本稿では, 上記の3つの基本要素, 品質推定, 分類, ローカライゼーションについて述べる。
論文 参考訳(メタデータ) (2020-06-08T07:24:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。