論文の概要: Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging
- arxiv url: http://arxiv.org/abs/2608.21970v1
- Date: Sat, 22 Aug 2026 14:17:48 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-25 18:24:36.939562
- Title: Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging
- Title(参考訳): 検索型ビジュアルプロンプティング:2光子イメージングの基礎モデル
- Authors: Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto, Daniela Giordano, Concetto Spampinato, Federica Proietto Salanitri,
- Abstract要約: 2光子カルシウムイメージングは基礎モデルにとって困難な設定である。
Retrieval-Augmented Visual Prompting (RAVP) は,各ターゲットタイルをアノテーション付き例で拡張するフレームワークである。
アレン・ブレイン天文台の実験では、強調された推論はゼロショットニューロンの検出とインスタンスのセグメンテーションを一貫して強化している。
- 参考スコア(独自算出の注目度): 12.045437765243419
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Two-photon calcium imaging presents a challenging setting for foundation models: image appearance varies substantially across recordings and experimental conditions, annotations are scarce, and rapid adaptation is often needed. Rather than adapting model weights through fine-tuning, we ask whether a foundation model can be guided at inference time by injecting external visual memory directly into its input. We implement this idea with SAM 3 and introduce Retrieval-Augmented Visual Prompting (RAVP), a framework in which each target tile is augmented with a retrieved annotated exemplar whose bounding box is used as a concept prompt. RAVP turns retrieval into a form of visual prompting and enables adaptation through input design alone. We study multiple exemplar selection strategies, including fluorescence-guided heuristics and a lightweight recall predictor trained to estimate which exemplar is most informative for a target tile. Experiments on the Allen Brain Observatory show that exemplar-augmented inference consistently strengthens zero-shot neuron detection and instance segmentation. Ablation studies further show that a single carefully selected exemplar is more effective than prompting with multiple retrieved examples. These results position inference-time visual memory injection as a simple and effective alternative to parameter adaptation for foundation models in specialized biomedical imaging.
- Abstract(参考訳): 2光子カルシウムイメージングは基礎モデルにとって難しい設定であり、画像の外観は記録や実験条件によって大きく異なり、アノテーションは乏しく、迅速な適応が必要とされることが多い。
微調整によりモデル重みを適応させるのではなく、外部視覚記憶を直接入力に注入することで、基礎モデルが推論時にガイドできるかどうかを問う。
このアイデアをSAM 3で実装し、Retrieval-Augmented Visual Prompting (RAVP)を導入します。
RAVPは、検索を視覚的なプロンプトの形式に変え、入力設計だけで適応できる。
蛍光誘導ヒューリスティックスや,対象タイルに最も有益であるものを推定するために訓練された軽量リコール予測器など,複数の代表的な選択戦略について検討した。
アレン・ブレイン天文台の実験では、強調された推論はゼロショットニューロンの検出とインスタンスのセグメンテーションを一貫して強化している。
アブレーション研究により、1つの慎重に選択された例は、複数のサンプルを抽出するよりも効果的であることが示されている。
これらの結果から, 生体イメージングの基礎モデルに対するパラメータ適応の簡便かつ効果的な代替手段として, 推測時ビジュアルメモリインジェクションが位置づけられた。
関連論文リスト
- Few-shot target-driven instance detection based on open-vocabulary object detection models [1.0749601922718608]
オープンボキャブラリオブジェクト検出モデルは、同じ潜在空間において、より近い視覚的およびテキスト的概念をもたらす。
テキスト記述を必要とせずに,後者をワンショットあるいは少数ショットのオブジェクト認識モデルに変換する軽量な手法を提案する。
論文 参考訳(メタデータ) (2024-10-21T14:03:15Z) - Dual-Image Enhanced CLIP for Zero-Shot Anomaly Detection [58.228940066769596]
本稿では,統合視覚言語スコアリングシステムを活用したデュアルイメージ強化CLIP手法を提案する。
提案手法は,画像のペアを処理し,それぞれを視覚的参照として利用することにより,視覚的コンテキストによる推論プロセスを強化する。
提案手法は視覚言語による関節異常検出の可能性を大幅に活用し,従来のSOTA法と同等の性能を示す。
論文 参考訳(メタデータ) (2024-05-08T03:13:20Z) - Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers [120.49126407479717]
本稿では,ゼロショットスケッチに基づく画像検索(ZS-SBIR)のためのテキスト・画像拡散モデルについて検討する。
スケッチと写真の間のギャップをシームレスに埋めるテキストと画像の拡散モデルの能力。
論文 参考訳(メタデータ) (2024-03-12T00:02:03Z) - Dual-View Data Hallucination with Semantic Relation Guidance for Few-Shot Image Recognition [49.26065739704278]
本稿では、意味的関係を利用して、画像認識のための二重視点データ幻覚を導出するフレームワークを提案する。
インスタンスビューデータ幻覚モジュールは、新規クラスの各サンプルを幻覚して新しいデータを生成する。
プロトタイプビューデータ幻覚モジュールは、意味認識尺度を利用して、新しいクラスのプロトタイプを推定する。
論文 参考訳(メタデータ) (2024-01-13T12:32:29Z) - Forgery-aware Adaptive Transformer for Generalizable Synthetic Image
Detection [106.39544368711427]
本研究では,様々な生成手法から偽画像を検出することを目的とした,一般化可能な合成画像検出の課題について検討する。
本稿では,FatFormerという新しいフォージェリー適応トランスフォーマー手法を提案する。
提案手法は, 平均98%の精度でGANを観測し, 95%の精度で拡散モデルを解析した。
論文 参考訳(メタデータ) (2023-12-27T17:36:32Z) - Anomaly Score: Evaluating Generative Models and Individual Generated Images based on Complexity and Vulnerability [21.355484227864466]
生成した画像の表現空間と入力空間の関係について検討する。
異常スコア(AS)と呼ばれる画像生成モデルを評価するための新しい指標を提案する。
論文 参考訳(メタデータ) (2023-12-17T07:33:06Z) - Prototype Learning for Explainable Brain Age Prediction [1.104960878651584]
回帰タスクに特化して設計された,説明可能なプロトタイプベースモデルであるExPeRTを提案する。
提案モデルでは,プロトタイプラベルの重み付き平均値を用いて,学習したプロトタイプのラテント空間における距離からサンプル予測を行う。
提案手法は,モデル推論プロセスに関する知見を提供しながら,最先端の予測性能を実現した。
論文 参考訳(メタデータ) (2023-06-16T14:13:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。