論文の概要: Harnessing Weak Pair Uncertainty for Text-based Person Search
- arxiv url: http://arxiv.org/abs/2604.08877v1
- Date: Fri, 10 Apr 2026 02:36:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-13 17:57:53.640914
- Title: Harnessing Weak Pair Uncertainty for Text-based Person Search
- Title(参考訳): テキストに基づく人物検索におけるハーネスング弱視の不確かさ
- Abstract要約: 本研究では,自然言語による興味ある人物を検索するテキストベースの人物検索について検討する。
弱陽性をフル活用するために,画像とテキストのペアの不確かさを明示的に推定する不確実性認識手法を提案する。
- 参考スコア(独自算出の注目度): 18.155384388834175
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: In this paper, we study the text-based person search, which is to retrieve the person of interest via natural language description. Prevailing methods usually focus on the strict one-to-one correspondence pair matching between the visual and textual modality, such as contrastive learning. However, such a paradigm unintentionally disregards the weak positive image-text pairs, which are of the same person but the text descriptions are annotated from different views (cameras). To take full use of weak positives, we introduce an uncertainty-aware method to explicitly estimate image-text pair uncertainty, and incorporate the uncertainty into the optimization procedure in a smooth manner. Specifically, our method contains two modules: uncertainty estimation and uncertainty regularization. (1) Uncertainty estimation is to obtain the relative confidence on the given positive pairs; (2) Based on the predicted uncertainty, we propose the uncertainty regularization to adaptively adjust loss weight. Additionally, we introduce a group-wise image-text matching loss to further facilitate the representation space among the weak pairs. Compared with existing methods, the proposed method explicitly prevents the model from pushing away potentially weak positive candidates. Extensive experiments on three widely-used datasets, .e.g, CUHK-PEDES, RSTPReid and ICFG-PEDES, verify the mAP improvement of our method against existing competitive methods +3.06%, +3.55% and +6.94%, respectively.
- Abstract(参考訳): 本稿では,自然言語記述による興味ある人物の検索を目的としたテキストベースの人物検索について検討する。
一般的な手法は、対照的な学習のような視覚とテキストのモダリティ間の厳密な1対1の対応ペアにフォーカスする。
しかし、このようなパラダイムは、同一人物である弱正のイメージテキストペアを意図せず無視するが、テキスト記述は異なる視点(カメラ)から注釈付けされる。
弱陽性をフル活用するために、画像とテキストのペアの不確実性を明示的に推定し、不確実性を最適化手順にスムーズな方法で組み込む不確実性認識手法を導入する。
具体的には,不確実性推定と不確実性正則化の2つのモジュールを含む。
1) 不確実性推定は,与えられた正の対に対する相対的な信頼を得るためのものであり,(2)予測された不確実性に基づいて,損失重量を適応的に調整する不確実性正則化を提案する。
さらに、弱いペア間の表現空間をさらに促進するために、グループワイドな画像テキストマッチング損失を導入する。
既存手法と比較して,提案手法はモデルが潜在的に弱い正の候補を退避させることを明示的に防止する。
広く使われている3つのデータセットに対する大規模な実験。
例えば、CUHK-PEDES、RSTPReid、ICFG-PEDESは、既存の競合手法+3.06%、+3.55%、+6.94%に対して、我々の手法のmAP改善を検証する。
関連論文リスト
- Rethinking Uncertainty in Segmentation: From Estimation to Decision [0.0]
医用画像のセグメンテーションでは、不確実性推定がしばしば報告されるが、意思決定を導くために使われることは稀である。
セグメンテーションを2段階のパイプラインとして定式化し、それに続く決定を行い、不確実性のみを最適化することは、達成可能な安全性向上のほとんどを達成できないことを示す。
以上の結果から,最も優れた手法とポリシーの組み合わせは,最大80%のセグメンテーション誤差をわずか25%のdeferralで除去できることが示唆された。
論文 参考訳(メタデータ) (2026-04-14T19:52:05Z) - Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic Learning [49.28548464288051]
Composed Image Retrieval (CIR)は、参照画像と修正テキストを組み合わせることで、画像検索を可能にする。
CIR三重項の内在ノイズは内在的不確実性を引き起こし、モデルの堅牢性を脅かす。
本稿では,これらの制約を克服するための不確実性誘導(HUG)パラダイムを提案する。
論文 参考訳(メタデータ) (2026-01-16T16:05:49Z) - Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images [16.071286104558393]
テキストに基づく歩行者探索 (TBPS) は, 対象歩行者の位置を自然言語で特定することを目的としている。
MUE(Multi-granularity Uncertainity Estimation)、PUD(Prototype-based Uncertainity Decoupling)、ReID(Cross-modal Re-identification)の3つのモジュールからなる新しいフレームワークであるUDD-TBPSを提案する。
論文 参考訳(メタデータ) (2025-05-06T14:25:30Z) - Token-Level Adversarial Prompt Detection Based on Perplexity Measures
and Contextual Information [67.78183175605761]
大規模言語モデルは、敵の迅速な攻撃に影響を受けやすい。
この脆弱性は、LLMの堅牢性と信頼性に関する重要な懸念を浮き彫りにしている。
トークンレベルで敵のプロンプトを検出するための新しい手法を提案する。
論文 参考訳(メタデータ) (2023-11-20T03:17:21Z) - Composed Image Retrieval with Text Feedback via Multi-grained
Uncertainty Regularization [73.04187954213471]
粗い検索ときめ細かい検索を同時にモデル化する統合学習手法を提案する。
提案手法は、強いベースラインに対して+4.03%、+3.38%、+2.40%のRecall@50精度を達成した。
論文 参考訳(メタデータ) (2022-11-14T14:25:40Z) - Reliability-Aware Prediction via Uncertainty Learning for Person Image
Retrieval [51.83967175585896]
UALは、データ不確実性とモデル不確実性を同時に考慮し、信頼性に配慮した予測を提供することを目的としている。
データ不確実性はサンプル固有のノイズを捕捉する」一方、モデル不確実性はサンプルの予測に対するモデルの信頼を表現している。
論文 参考訳(メタデータ) (2022-10-24T17:53:20Z) - ADDMU: Detection of Far-Boundary Adversarial Examples with Data and
Model Uncertainty Estimation [125.52743832477404]
AED(Adversarial Examples Detection)は、敵攻撃に対する重要な防御技術である。
本手法は, 正逆検出とFB逆検出の2種類の不確実性推定を組み合わせた新しい手法である textbfADDMU を提案する。
提案手法は,各シナリオにおいて,従来の手法よりも3.6と6.0のEmphAUC点が優れていた。
論文 参考訳(メタデータ) (2022-10-22T09:11:12Z) - Beyond Model Interpretability: On the Faithfulness and Adversarial
Robustness of Contrastive Textual Explanations [2.543865489517869]
本研究は、説明の忠実さに触発された新たな評価手法の基盤を築き、テキストの反事実を動機づけるものである。
感情分析データを用いた実験では, 両モデルとも, 対物関係の関連性は明らかでないことがわかった。
論文 参考訳(メタデータ) (2022-10-17T09:50:02Z) - Bayesian Triplet Loss: Uncertainty Quantification in Image Retrieval [10.743633102172236]
画像検索における不確かさの定量化は下流の決定に不可欠である。
本稿では,画像の埋め込みを決定論的特徴ではなく特徴とみなす新しい手法を提案する。
我々はベイズ三重項損失(Bayesian triplet loss)と呼ばれる後肢の変分近似を導出し、最先端の不確実性推定を導出する。
論文 参考訳(メタデータ) (2020-11-25T11:47:33Z) - Uncertainty-Aware Few-Shot Image Classification [118.72423376789062]
ラベル付き限られたデータから新しいカテゴリを認識できる画像分類はほとんどない。
画像分類のための不確実性を考慮したFew-Shotフレームワークを提案する。
論文 参考訳(メタデータ) (2020-10-09T12:26:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。