論文の概要: HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space
- arxiv url: http://arxiv.org/abs/2609.24564v1
- Date: Mon, 21 Sep 2026 13:26:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-23 03:38:27.402314
- Title: HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space
- Title(参考訳): HyperCLIP++:ハイパーボリック空間におけるオープン語彙セマンティックセマンティックセグメンテーションのための微調整CLIP
- Abstract要約: 近年の研究では、CLIPのテキストと画像エンコーダの微調整により、セグメンテーション性能が向上することが示されている。
本稿では,新しいパラメータ効率適応戦略であるHyperCLIP++を提案する。
実験の結果,HyperCLIP++は3つのベンチマークで最先端のパフォーマンスを実現していることがわかった。
- 参考スコア(独自算出の注目度): 52.9767967184765
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its generalization capability, recent studies show that fine-tuning both CLIP's text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchy alignment, since during fine-tuning, the hierarchical level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encodes hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP's text embeddings decreases, facilitating better alignment with the pixel-level granularity of visual data. Building on this, we propose HyperCLIP++, a novel and parameter-efficient adaptation strategy. HyperCLIP++ directly adjusts the hyperbolic radius of CLIP's embeddings via scaling transformations to achieve a hierarchy alignment to the target task, i.e., segmentation. To ensure this hierarchy alignment is effected consistently across both modalities and preserves their cross-modal alignment during training, HyperCLIP++ integrates a Dual Cross-Relation Communication (DCRC) module that synchronizes these adjustments between the vision and text pathways. Our experiments show that HyperCLIP++ achieves state-of-the-art performance across three benchmarks while fine-tuning only approximately 5% of CLIP's total parameters. More importantly, we observe that after adjustment, CLIP's text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the hierarchical level required for this segmentation task might be quantified using the hyperbolic radius.
- Abstract(参考訳): 基本的な視覚言語モデルであるCLIPは、オープン語彙セマンティックセグメンテーションの強力なツールとして登場した。
最近の研究では、CLIPのテキストエンコーダと画像エンコーダの微調整が、特にオープンセットからのクラスにおいて、セグメンテーション性能を著しく向上させることが示されている。
本研究では,画像埋め込みの階層レベルが画像レベルからピクセルレベルに変化するため,階層的アライメントの観点からこの現象を説明する。
これを、自然に階層構造を符号化する双曲空間を活用することで実現する。
私たちのキーとなる観察は、微調整の間、CLIPのテキスト埋め込みの双曲半径が減少し、画像データのピクセルレベルの粒度との整合性が向上することです。
そこで我々は,新しいパラメータ効率適応戦略であるHyperCLIP++を提案する。
HyperCLIP++は、スケーリング変換を通じてCLIPの埋め込みの双曲半径を直接調整し、ターゲットタスク、すなわちセグメンテーションの階層的アライメントを達成する。
この階層的アライメントを両モード間で一貫して実施し、トレーニング中に相互アライメントを維持するために、HyperCLIP++は、ビジョンとテキストパス間の調整を同期するDual Cross-Relation Communication (DCRC)モジュールを統合する。
実験の結果,HyperCLIP++は3つのベンチマークで最先端のパフォーマンスを実現し,CLIPの総パラメータの約5%を微調整した。
さらに重要なことは、調整後、CLIPのテキスト埋め込みはデータセット全体にわたって比較的固定された双曲半径を示し、このセグメンテーションタスクに必要な階層レベルが双曲半径を用いて定量化されることを示唆している。
関連論文リスト
- SuperCLIP: CLIP with Simple Classification Supervision [88.86549733903314]
Contrastive Language-Image Pretrainingは、画像とテキストを共有埋め込み空間に整列させることにより、視覚言語タスクの強力な一般化を実現する。
近年,CLIP様モデルでは,テキスト中の微細なセマンティック信号が依然として使われていないことが報告されている。
分類に基づく教師付きコントラスト学習のフレームワークであるSuperCLIPを提案する。
論文 参考訳(メタデータ) (2025-12-16T15:11:53Z) - Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation [55.486872677160015]
本稿では,体としてのセグメンテーションバックボーンと,頭部としてのCLIPベースのセマンティックヘッドを統合したChimera-Segを提案する。
特に、Chimera-Segはトレーニング可能なセグメンテーションモデルとCLIPセマンティックヘッド(CLIP Semantic Head, CSH)を備えており、CLIP対応空間に高密度な特徴をマッピングする。
また,CLIP CLSトークンと高い類似性を示す濃厚な特徴から知識を抽出する選択的グローバル蒸留(SGD)を提案する。
論文 参考訳(メタデータ) (2025-06-27T09:26:50Z) - Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation [10.502680141980642]
オープンボキャブラリセマンティックセグメンテーションは、画像中の各ピクセルに任意のテキスト記述をラベル付けしようとする。
視覚言語基盤モデル、特にCLIPは、オープン語彙能力を取得するための強力なツールとして登場した。
H-CLIPは、CLIPの総パラメータの約4%を更新するだけで、新しいSOTAオープン語彙セマンティックセマンティックセマンティクス結果を達成する。
論文 参考訳(メタデータ) (2024-05-29T07:41:34Z) - Symmetrical Linguistic Feature Distillation with CLIP for Scene Text
Recognition [77.93678598476149]
CLIP-OCR(Symmetrical Linguistic Feature Distillation framework)を新たに構築する。
CLIP画像エンコーダを逆CLIPテキストエンコーダでカスケードすることにより、画像からテキストまでの特徴フローで対称構造を構築する。
大規模な実験では、CLIP-OCRが6つのSTRベンチマークで平均精度93.8%で有効であることが示されている。
論文 参考訳(メタデータ) (2023-10-08T04:00:20Z) - ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation [35.60888272729273]
近年、CLIPは2段階のスキームを用いて画素レベルのゼロショット学習タスクに適用されている。
このような方式は有効であるが、2つの画像エンコーダが必要であり、1つは提案生成用、もう1つはCLIP用であり、複雑なパイプラインと高い計算コストをもたらす。
本稿では,CLIPのゼロショット予測能力を画像からピクセルレベルまで直接拡張する,シンプルかつ効率的なワンステージソリューションを提案する。
論文 参考訳(メタデータ) (2022-12-07T12:05:00Z) - Context-self contrastive pretraining for crop type semantic segmentation [39.81074867563505]
提案したContext-Self Contrastive Loss (CSCL)は、セマンティックバウンダリをポップアップさせる埋め込み空間を学習する。
衛星画像時系列(SITS)からの作物型セマンティックセマンティックセグメンテーションでは,サテライト境界における性能が重要なボトルネックとなる。
より粒度の高い作物のクラスを得るための超解像における意味的セグメンテーションのプロセスを提案する。
論文 参考訳(メタデータ) (2021-04-09T11:29:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。