Fugu-MT 論文翻訳(概要): GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

論文の概要: GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

arxiv url: http://arxiv.org/abs/2510.22268v1
Date: Sat, 25 Oct 2025 12:16:10 GMT
ステータス: 翻訳完了
システム内更新日: 2025-10-28 15:28:15.012037
Title: GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification
Title（参考訳）: GSAlign:空中人物再同定のための幾何学的・意味的アライメントネットワーク
Authors: Qiao Li, Jie Li, Yukang Zhang, Lei Tan, Jing Chen, Jiayi Ji,
Abstract要約: Aerial-Ground person re-identification (AG-ReID) は、歩行者のイメージを根本的に異なる視点からマッチングすることを目的とした、新たな課題である。この課題は、極端に視点のずれ、ワープ、空中画像と地上画像の間の領域ギャップのために重大な課題を生じさせる。
参考スコア（独自算出の注目度）: 32.31970656501684
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to extreme viewpoint discrepancies, occlusions, and domain gaps between aerial and ground imagery. While prior works have made progress by learning cross-view representations, they remain limited in handling severe pose variations and spatial misalignment. To address these issues, we propose a Geometric and Semantic Alignment Network (GSAlign) tailored for AG-ReID. GSAlign introduces two key components to jointly tackle geometric distortion and semantic misalignment in aerial-ground matching: a Learnable Thin Plate Spline (LTPS) Module and a Dynamic Alignment Module (DAM). The LTPS module adaptively warps pedestrian features based on a set of learned keypoints, effectively compensating for geometric variations caused by extreme viewpoint changes. In parallel, the DAM estimates visibility-aware representation masks that highlight visible body regions at the semantic level, thereby alleviating the negative impact of occlusions and partial observations in cross-view correspondence. A comprehensive evaluation on CARGO with four matching protocols demonstrates the effectiveness of GSAlign, achieving significant improvements of +18.8\% in mAP and +16.8\% in Rank-1 accuracy over previous state-of-the-art methods on the aerial-ground setting. The code is available at: \textcolor{magenta}{https://github.com/stone96123/GSAlign}.
Abstract（参考訳）: Aerial-Ground person re-identification (AG-ReID) は、無人航空機(UAV)や地上監視カメラなど、非常に異なる視点から捉えた歩行者画像のマッチングを目的とした、新しくて困難なタスクである。この課題は、航空画像と地上画像の間の極端な視点の相違、閉塞、領域のギャップによって、重大な課題を生んでいる。それまでの作品は、クロスビュー表現を学習することで進歩してきたが、厳しいポーズのバリエーションや空間的ミスアライメントを扱うことには限界がある。これらの課題に対処するために,AG-ReIDに適した幾何学的・意味的アライメントネットワーク(GSAlign)を提案する。 GSAlignは、地上マッチングにおける幾何学的歪みと意味的ミスアライメント(semantic misalignment, 意味的ミスアライメント)に共同で取り組むための2つの重要なコンポーネント、LTPSモジュールと動的アライメントモジュール(Dynamic Alignment Module, DAM)を導入した。 LTPSモジュールは学習キーポイントのセットに基づいて歩行者の特徴を適応的にワープし、極端な視点変化による幾何学的変動を効果的に補償する。並行して、DAMは、目に見える身体領域をセマンティックレベルで強調する可視性対応の表現マスクを推定し、オークルージョンのネガティブな影響と、クロスビュー対応における部分的な観察を緩和する。 4つのマッチングプロトコルによるCARGOの総合評価では、GSAlignの有効性が示され、mAPでは+18.8\%、地上では+16.8\%の精度で+16.8\%の精度が向上した。コードは以下の通り。 \textcolor{magenta}{https://github.com/stone96123/GSAlign}。

論文の概要: GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

関連論文リスト