論文の概要: Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation
- arxiv url: http://arxiv.org/abs/2606.30244v1
- Date: Mon, 29 Jun 2026 12:54:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-30 18:07:16.235767
- Title: Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation
- Title(参考訳): リモートセンシング画像のセグメンテーションを基準とした効率的なクロスモーダルアライメントのための意味駆動尺度と空間選択
- Authors: Kun Li, Shengxi Gui, Francesco Nex, Michael Ying Yang,
- Abstract要約: Referring Remote Sensing Imageは、リモートセンシングイメージで自然言語表現によって指定された対象または領域をローカライズし、セグメンテーションする。
既存のRRSISモデルは大規模な基礎モデルの恩恵を受けているが、それらは主に完全な微調整に依存している。
本稿では,効率的なクロスモーダルアライメントのためのセマンティック・スケールと空間選択という新しい枠組みを提案する。
- 参考スコア(独自算出の注目度): 14.492779733437082
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression in a remote sensing image. While existing RRSIS models have benefited from large-scale foundation models, they predominantly rely on full fine-tuning. These approaches are computationally intensive and may weaken the generalization ability of pre-trained models, as extensive fine-tuning on significantly smaller downstream datasets can distort the well-structured feature representations learned during large-scale pre-training. Although Parameter-Efficient Tuning (PET) offers a potential alternative, existing PET frameworks primarily focus on single-modal optimization, failing to capture the complex cross-modal dependencies required for multimodal reasoning, while simultaneously struggling to bridge the substantial domain gap between natural scenes and aerial imagery. To address these limitations, we propose a novel framework, Semantic-driven Scale and Spatial Selection for Efficient Cross-modal Alignment (S4ECA), which enables effective and efficient cross-modal interaction through parameter-efficient adaptation. Specifically, we design a dual-encoder adapter architecture. The textual adapter employs learnable queries to distill highly semantic language proxies from word-level embeddings, facilitating early grounding. Simultaneously, the visual adapter refines hierarchical feature representations through a multi-scale dense extractor, followed by a language-guided scale and spatial selection mechanism that dynamically emphasizes relevant visual contexts, ensuring precise cross-modal alignment. By updating only 2.4% of the backbone parameters, our proposed model achieves state-of-the-art performance on the RRSIS-D and RefSegRS datasets, demonstrating superior efficiency and precision in complex aerial scenarios.
- Abstract(参考訳): Referring Remote Sensing Image Segmentation (RRSIS)は、リモートセンシング画像において自然言語表現によって指定された対象または領域をローカライズし、セグメンテーションする。
既存のRRSISモデルは大規模な基礎モデルの恩恵を受けているが、それらは主に完全な微調整に依存している。
これらのアプローチは計算集約的であり、大規模な事前学習中に学習したよく構造化された特徴表現を歪めることができるため、事前学習されたモデルの一般化能力を弱める可能性がある。
パラメータ効率チューニング(PET)は潜在的な代替手段を提供するが、既存のPETフレームワークは主に単一モーダル最適化に焦点を当てており、マルチモーダル推論に必要な複雑なクロスモーダル依存関係を捉えることができず、同時に自然のシーンと空中画像の領域ギャップを埋めることに苦労している。
これらの制約に対処するために,パラメータ効率の適応による効率的かつ効率的な相互モーダル相互作用を実現するための,S4ECA(Semantic-driven Scale and Spatial Selection for Efficient Cross-modal Alignment)を提案する。
具体的には、デュアルエンコーダアダプタアーキテクチャを設計する。
テキストアダプタは、学習可能なクエリを使用して、単語レベルの埋め込みから高度に意味のある言語プロキシを抽出し、早期の接地を容易にする。
同時に、マルチスケールの高密度抽出器を通じて階層的特徴表現を洗練させ、言語誘導のスケールと空間選択機構により、関連した視覚的コンテキストを動的に強調し、正確なクロスモーダルアライメントを確保する。
バックボーンパラメータの2.4%だけを更新することにより、RRSIS-DおよびRefSegRSデータセットの最先端性能を達成し、複雑な空域シナリオにおいて優れた効率と精度を示す。
関連論文リスト
- RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning [49.24254347822435]
リモートセンシング画像変化キャプション(RSICC)は、双方向のリモートセンシング画像間の変化を記述することを目的としている。
我々はRSICCLLMを提案する。RSICCLLMはRSSICCにおける大規模視覚言語モデルのための最初の後学習フレームワークである。
論文 参考訳(メタデータ) (2026-06-26T16:57:40Z) - Spatio-Semantic Expert Routing Architecture with Mixture-of-Experts for Referring Image Segmentation [0.3437656066916039]
画像セグメント化の参照は、自然言語表現によって記述された画像領域のためのピクセルレベルのマスクを作成することを目的としている。
画像セグメンテーションを参照するための空間分割型エキスパートルーティングアーキテクチャSERAを提案する。
SERAは、視覚言語フレームワーク内の2つの相補的な段階において、軽量で表現を意識した専門家の洗練を導入する。
論文 参考訳(メタデータ) (2026-03-13T00:37:20Z) - Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models [84.78794648147608]
永続的な幾何学的異常であるモダリティギャップが残っている。
このギャップを埋める以前のアプローチは、過度に単純化された等方的仮定によってほとんど制限されている。
固定フレームモダリティギャップ理論(英語版)を提案し、モダリティギャップを安定バイアスと異方性残差に分解する。
次に、トレーニング不要なモダリティアライメント戦略であるReAlignを紹介します。
論文 参考訳(メタデータ) (2026-02-02T13:59:39Z) - Scale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation [12.893224628061516]
リモートセンシング画像セグメンテーション(RRSIS)の目的は、自然言語表現を用いて、空中画像内の特定のピクセルレベル領域を抽出することである。
本稿では,これらの課題に対処するため,SBANet(Scale-wise Bidirectional Alignment Network)と呼ばれる革新的なフレームワークを提案する。
提案手法は,RRSIS-DとRefSegRSのデータセットにおける従来の最先端手法と比較して,優れた性能を実現する。
論文 参考訳(メタデータ) (2025-01-01T14:24:04Z) - Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation [50.433911327489554]
リモートセンシング画像セグメンテーション(RRSIS)の目標は、参照式によって識別された対象オブジェクトの画素レベルマスクを生成することである。
上記の課題に対処するため、クロスモーダル双方向相互作用モデル(CroBIM)と呼ばれる新しいRRSISフレームワークが提案されている。
RRSISの研究をさらに推し進めるために、52,472個の画像言語ラベル三重項からなる新しい大規模ベンチマークデータセットRISBenchを構築した。
論文 参考訳(メタデータ) (2024-10-11T08:28:04Z) - Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation [63.15257949821558]
Referring Remote Sensing Image (RRSIS)は、コンピュータビジョンと自然言語処理を組み合わせた新しい課題である。
従来の参照画像(RIS)アプローチは、空中画像に見られる複雑な空間スケールと向きによって妨げられている。
本稿ではRMSIN(Rotated Multi-Scale Interaction Network)を紹介する。
論文 参考訳(メタデータ) (2023-12-19T08:14:14Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。