論文の概要: HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition
- arxiv url: http://arxiv.org/abs/2607.08249v1
- Date: Thu, 09 Jul 2026 08:53:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-10 14:45:27.474196
- Title: HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition
- Title(参考訳): HSA:多粒度シーン分割のための階層的スロットアテンション
- Authors: Neelu Madan, Rongzhen Zhao, Andreas Mogelmose, Juho Kannala, Joni Pajarinen, Graham W. Taylor, Thomas B. Moeslund,
- Abstract要約: textbf$+41.5 ARI at holistic, textbf$+14.6 at semantic, textbf$+10.4 at panoptic level on COCO
textbf$+10.4 COCOでは、Pascal VOCでさらに大きな利益を上げ、3つではなく1つのモデルを必要とする。
- 参考スコア(独自算出の注目度): 44.65369749375938
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention. However, existing methods share two critical limitations: they decompose scenes into a flat set of slots at a single granularity, and this decomposition is based on appearance rather than semantics. Yet humans understand scenes through semantic hierarchies: separating foreground from background, recognizing object categories, and identifying individual instances. Crucially, such semantic hierarchies cannot emerge without supervision, because category names are human constructs, not visual patterns. We propose Hierarchical Slot Attention (HSA), which learns multi-granularity semantic scene decomposition from a single model. HSA decomposes scenes at three levels: holistic (foreground/background), semantic (object categories), and panoptic (individual instances). Using only 10\% labeled data, combined with hierarchical alignment loss, HSA learns all three levels jointly. We further introduce grouping purity and containment to measure whether the hierarchy is encoded in representation space, not just output masks. Experiments on COCO and PASCAL VOC demonstrate that HSA outperforms the strongest flat baseline by up to \textbf{$+$41.5} ARI at holistic, \textbf{$+$14.6} at semantic, and \textbf{$+$10.4} at panoptic level on COCO, with even larger gains on Pascal VOC, while requiring a single model instead of three. Code will be made available upon acceptance.
- Abstract(参考訳): スロットアテンションは、オブジェクト中心の学習のための強力なフレームワークであり、反復的な競争的注意を通して視覚シーンを潜在スロットに分解する。
しかし、既存の手法では、シーンを1つの粒度の平らなスロットに分解するなど、2つの重要な制限がある。
背景を背景から切り離し、対象のカテゴリを認識し、個々のインスタンスを識別する。
重要なのは、カテゴリー名は視覚的パターンではなく人間の構成物であるため、このような意味的階層は監督なしでは現れない。
本稿では,階層型スロット注意(HSA)を提案する。
HSAは、全体論(地上/背景)、意味論(対象カテゴリー)、パノプティクス(個人的事例)の3つのレベルでシーンを分解する。
階層的なアライメント損失と組み合わさったラベル付きデータのみを使用して、HSAは3つのレベル全てを共同で学習する。
さらに、出力マスクだけでなく、表現空間に階層がエンコードされているかどうかを測定するために、グループ純度と包含を導入する。
COCO と PASCAL VOC の実験では、HSA は、全体論では \textbf{$+$41.5} ARI 、意味論では \textbf{$+$14.6} 、COCO では panoptic レベルで \textbf{$+$10.4} に、Pascal VOC では 3 よりも大きく、さらに 1 つのモデルを必要とする。
コードは受理時に利用可能になる。
関連論文リスト
- CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation [54.53371540755023]
推論セグメンテーションは、複雑でしばしば暗黙的な指示によって参照されるターゲットに対して、ピクセル精度の高いマスクを求める。
我々は、限定ラベル付きデータとラベルなし画像の大きなコーパスから共同で学習する半教師付き推論セグメンテーションフレームワークCORAを提案する。
CORAは最先端の結果を達成し、都市景観理解のためのベンチマークデータセットであるCityscapesにラベル付きイメージを100個まで必要としています。
論文 参考訳(メタデータ) (2025-11-21T20:14:55Z) - Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses [31.85977999591524]
視覚言語モデルは、画像領域と大規模トレーニングデータの単語を暗黙的に関連付けることを学習する。
テキストモダリティ内のリッチな意味的構造と構文的構造は、監督の源として見過ごされている。
階層的構造化学習(HIST)は、追加の人間のアノテーションを使わずに、空間的視覚言語アライメントを強化する。
論文 参考訳(メタデータ) (2024-12-11T05:36:18Z) - Hierarchical Open-vocabulary Universal Image Segmentation [48.008887320870244]
Open-vocabulary Image segmentationは、任意のテキスト記述に従ってイメージをセマンティック領域に分割することを目的としている。
我々は,「モノ」と「スタッフ」の双方に対して,分離されたテキストイメージ融合機構と表現学習モジュールを提案する。
HIPIE tackles, HIerarchical, oPen-vocabulary, unIvErsal segmentation task in a unified framework。
論文 参考訳(メタデータ) (2023-07-03T06:02:15Z) - Deep Hierarchical Semantic Segmentation [76.40565872257709]
階層的セマンティックセマンティックセグメンテーション(HSS)は、クラス階層の観点で視覚的観察を構造化、ピクセル単位で記述することを目的としている。
HSSNは、HSSを画素単位のマルチラベル分類タスクとしてキャストし、現在のセグメンテーションモデルに最小限のアーキテクチャ変更をもたらすだけである。
階層構造によって引き起こされるマージンの制約により、HSSNはピクセル埋め込み空間を再評価し、よく構造化されたピクセル表現を生成する。
論文 参考訳(メタデータ) (2022-03-27T15:47:44Z) - Unsupervised Semantic Segmentation by Distilling Feature Correspondences [94.73675308961944]
教師なしセマンティックセグメンテーション(unsupervised semantic segmentation)は、アノテーションなしで画像コーパス内の意味論的意味のあるカテゴリを発見し、ローカライズすることを目的としている。
STEGOは、教師なし特徴を高品質な個別のセマンティックラベルに蒸留する新しいフレームワークである。
STEGOは、CocoStuffとCityscapesの両課題において、先行技術よりも大幅に改善されている。
論文 参考訳(メタデータ) (2022-03-16T06:08:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。