論文の概要: Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention
- arxiv url: http://arxiv.org/abs/2607.18112v1
- Date: Mon, 20 Jul 2026 16:08:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 18:48:37.701218
- Title: Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention
- Title(参考訳): 関節位置埋め込みと咬合レベル注意を併用したオクルージョン・アウェア・パノプティクス・セグメンテーション
- Authors: Wenbo Wei, Jun Wang, Shan Raza, Abhir Bhalerao,
- Abstract要約: textbfPosition textbfEmbedding textbfModulation with textbfOcclusion-textbfLevel textbfAttention (PEMOLA)を提案する。
PEMOLAはトランスフォーマーベースのパノプティカルセグメンテーションにシームレスに統合できる。
COCO-OLACとCityscapes-OLACの実験は、PEMOLAが一貫して汎視的セグメンテーションの品質を改善することを示した。
- 参考スコア(独自算出の注目度): 7.205486726857026
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Panoptic segmentation in complex scenes remains challenging because of occlusions, yet modern approaches often neglect occlusion modelling. In this paper, we propose \textbf{P}osition \textbf{E}mbedding \textbf{M}odulation with \textbf{O}cclusion-\textbf{L}evel \textbf{A}ttention (PEMOLA), a novel occlusion-aware module that can be seamlessly integrated into transformer-based panoptic segmentation. To obtain occlusion cues, we train an occlusion classifier on the COCO-OLAC dataset. The classifier derives the occlusion-level attention, which serves as spatial guidance, while the occlusion labels are encoded into a learnable embedding to produce channel-wise weights. Through joint modulation, PEMOLA elegantly introduces the occlusion priors into the position embedding, thereby improving the occlusion modelling. We further annotate the Cityscapes dataset with occlusion levels, termed Cityscapes Occlusion Labels for All Computer Vision Tasks (Cityscapes-OLAC), following the same labelling protocol as COCO-OLAC, to evaluate the cross-dataset generalisation ability of PEMOLA. Extensive experiments on COCO-OLAC and Cityscapes-OLAC demonstrate that PEMOLA consistently improves panoptic segmentation quality while introducing minimal computational overhead. These results highlight the importance of occlusion modelling, where incorporating occlusion-level attention helps deliver robust panoptic segmentation under occlusion. Code and dataset are available at https://github.com/wenbo-wei/PEMOLA.
- Abstract(参考訳): 複雑なシーンにおけるパノプティクスのセグメンテーションは、オクルージョンが原因で困難なままであるが、現代のアプローチはオクルージョンモデリングを無視することが多い。
本稿では,新しいオクルージョン対応モジュールである \textbf{P}osition \textbf{E}mbedding \textbf{M}odulation with \textbf{O}cclusion-\textbf{L}evel \textbf{A}ttention (PEMOLA)を提案する。
そこで我々は,COCO-OLACデータセット上でオクルージョン分類器を訓練する。
分類器は、空間誘導として機能するオクルージョンレベルの注意を導出し、オクルージョンラベルは学習可能な埋め込みに符号化され、チャネルワイドウェイトを生成する。
共同変調により、PEMOLAは位置埋め込みにオクルージョン先行をエレガントに導入し、オクルージョンモデリングを改善する。
さらに,Cityscapes Occlusion Labels for All Computer Vision Tasks (Cityscapes-OLAC) という,Cityscapes Occlusion Labels for All Computer Vision Tasks (Cityscapes-OLAC) をCOCO-OLACと同じラベルでアノテートし,PEMOLAのクロスデータセット一般化能力を評価する。
COCO-OLACとCityscapes-OLACの大規模な実験により、PEMOLAは最小の計算オーバーヘッドを導入しながら、一貫した汎視的セグメンテーション品質の向上を実証した。
これらの結果はオクルージョン・モデリングの重要性を強調し、オクルージョン・レベルの注意を取り入れることで、オクルージョンの下での堅牢なパノプティカルセグメンテーションを実現する。
コードとデータセットはhttps://github.com/wenbo-wei/PEMOLA.comで公開されている。
関連論文リスト
- LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination [66.06027569507403]
textbfViewpoint Imagination (VIM) は、観測された証拠と想像された証拠の両方について、隠蔽された一次観測と条件の行動予測から補完的な視点を生成する。
VIMは、追加のカメラをデプロイ時に必要とせずに、タスクスイート、オクルージョンタイプ、重大度レベルの堅牢性を改善する。
論文 参考訳(メタデータ) (2026-06-09T13:39:49Z) - OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation [49.49933361280724]
OcclusionFormerは、インスタンスを分離することで、Zオーダーの優先順位を明示的にモデル化する新しいフレームワークです。
また、個別のインスタンスを明示的に監視し、セマンティックな一貫性を高めるクエリアライメント損失を導入します。
提案手法は,重なり合う領域の曖昧さを効果的に低減し,正しい閉塞依存性を強制し,構造的整合性を保ち,様々な場面で精度の高い向上をもたらす。
論文 参考訳(メタデータ) (2026-05-20T16:10:44Z) - Harnessing Group-Oriented Consistency Constraints for Semi-Supervised Semantic Segmentation in CdZnTe Semiconductors [71.44213719783703]
ICAF(Intra-group Consistency Augmentation Framework)は、CdZnTe(Cadmium Zinc Telluride)半導体画像にラベルを付けるために開発された。
ICAF は View Augmentation Module (VAM) と View Correction Module (VCM) の2つの重要なモジュールで構成されている。
ICAFは、CdZnTeデータセット上の70.6% mIoUを2つのグループアノテートデータのみを用いて達成する。
論文 参考訳(メタデータ) (2025-08-18T09:40:36Z) - OccludeNet: A Causal Journey into Mixed-View Actor-Centric Video Action Recognition under Occlusions [37.79525665359017]
我々はOccludeNetを構築した。OccludeNetは大規模に隠蔽されたビデオデータセットで、実シーンと合成シーンの両方を含んでいる。
分析の結果,シーン関連度が低く,部分的な身体視認性が高いと精度が低下することが明らかとなった。
本稿では,背景調整と反事実推論を併用した因果認識(Causal Action Recognition, CAR)手法を提案する。
論文 参考訳(メタデータ) (2024-11-24T06:10:05Z) - COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding [8.261771972240778]
本稿では,COCO-OLAC (COCO Occlusion Labels for All Computer Vision Tasks) という大規模データセットを提案する。
COCO-OLACは、イメージを3つの知覚された閉塞レベルに手動でラベル付けすることで、COCOデータセットから派生した。
提案手法はベースラインモデルの性能を高め,COCO-OLACデータセット上でのSOTA性能を実現する。
論文 参考訳(メタデータ) (2024-09-19T13:26:28Z) - Object-level Scene Deocclusion [92.39886029550286]
オブジェクトレベルのシーン・デクルージョンのためのPArallel可視・コミュールト拡散フレームワークPACOを提案する。
PACOをトレーニングするために、500kサンプルの大規模なデータセットを作成し、自己教師付き学習を可能にします。
COCOAと様々な現実世界のシーンの実験では、PACOがシーンの排除に優れた能力を示し、芸術の状態をはるかに上回っている。
論文 参考訳(メタデータ) (2024-06-11T20:34:10Z) - CMU-Flownet: Exploring Point Cloud Scene Flow Estimation in Occluded Scenario [10.852258389804984]
閉塞はLiDARデータにおける点雲フレームのアライメントを妨げるが、シーンフローモデルでは不十分な課題である。
本稿では,CMU-Flownet(Relational Matrix Upsampling Flownet)を提案する。
CMU-Flownetは、隠されたFlyingthings3DとKITTYデータセットの領域内で、最先端のパフォーマンスを確立する。
論文 参考訳(メタデータ) (2024-04-16T13:47:21Z) - Peeking into occluded joints: A novel framework for crowd pose
estimation [88.56203133287865]
OPEC-NetはイメージガイドされたプログレッシブGCNモジュールで、推論の観点から見えない関節を推定する。
OCPoseは、隣接するインスタンス間の平均IoUに対して、最も複雑なOccluded Poseデータセットである。
論文 参考訳(メタデータ) (2020-03-23T19:32:40Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。