論文の概要: RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding
- arxiv url: http://arxiv.org/abs/2608.23928v1
- Date: Tue, 25 Aug 2026 00:20:29 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-26 14:09:34.668835
- Title: RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding
- Title(参考訳): RefineRank: 手術時空間的接地のためのジョイントボックスのリファインメントとランク付け
- Abstract要約: 既存のアプローチはトレードオフに直面している: 言語モデルは質問コンテキストを理解し、不正確な座標を生成する。
このギャップを候補ボックスレベルで埋めるRefineRankを紹介します。
コンパクトなトレーニング可能なモジュールであるRefineNetは、凍結された医療ビジョン言語モデルの言語と地域特性と、凍結されたオープンセット検出器の提案を組み合わせる。
- 参考スコア(独自算出の注目度): 1.1161075621843597
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Surgical spatio-temporal grounding (STG) requires locating, at each video time specified by a procedural question, the object that the question asks about. Existing approaches face a trade-off: vision language models understand the question context but produce imprecise coordinates, whereas open-set detectors provide localized candidate boxes whose confidence does not reflect which box answers the question. We introduce RefineRank, which closes this gap at the candidate-box level. A compact trainable module, RefineNet, combines the language and regional features of a frozen medical vision language model with the proposals of a frozen open-set detector: it predicts a bounded coordinate correction and a quality score for every candidate box, and a fixed decoding rule returns the original or refined box with the highest score. On the MedVidBench Official Rankings (Verified), RefineRank records 0.421 STG mIoU, the highest displayed STG score, while its global multi-metric rank is 11. In a controlled evaluation on separate training and evaluation videos, coordinate correction raises the candidate oracle upper bound from 0.6772 to 0.7302, and ranking the joint pool of original and refined candidates by their RefineNet scores improves STG mIoU from 0.2719 to 0.4534, whereas separately trained selectors over the same pool reach at most 0.4186. These results show that a small box-level module can reconcile question understanding with precise localization without retraining either backbone. Code is available at [https://github.com/linzhe001/RefineRank](https://github.com/linzhe001/RefineRank).
- Abstract(参考訳): 外科的時空間グラウンド(STG)は、手続き的質問によって指定された各ビデオ時間において、質問が問う対象の位置を特定する必要がある。
既存のアプローチはトレードオフに直面している: 視覚言語モデルは問題コンテキストを理解し、不正確な座標を生成する。
このギャップを候補ボックスレベルで埋めるRefineRankを紹介します。
コンパクトなトレーニング可能なモジュールであるRefineNetは、凍結された医療ビジョン言語モデルの言語と地域特性を、凍結されたオープンセット検出器の提案と組み合わせて、各候補ボックスに対する境界座標補正と品質スコアを予測し、固定復号規則は、元のまたは洗練されたボックスを最高スコアで返却する。
MedVidBench Official Rankings (Verified), RefineRank record 0.421 STG mIoU, the highest displayed STG score, the global multi-metric rank is 11。
個別のトレーニングおよび評価ビデオの制御された評価では、座標補正は候補オラクル上限を0.6772から0.7302に引き上げ、元の候補と洗練された候補のジョイントプールをRefineNetスコアでランク付けし、STG mIoUを0.2719から0.4534に改善する。
これらの結果から,小さいボックスレベルのモジュールは,いずれのバックボーンも再学習することなく,正確な局所化で質問理解を整合できることがわかった。
コードは[https://github.com/linzhe001/RefineRank](https://github.com/linzhe001/RefineRank]で入手できる。
関連論文リスト
- Decoupled Pipeline with Proposal Reranking and Score Fusion for Positive-Unlabeled Marine Species Detection [0.0]
DS@GT ARCのマルチステージシステムについて述べる。
このシステムは102チーム中12位にランクインした。
実験の結果、提案のリコールの保存、過剰な攻撃的なフィルタリングの回避、下流のランキングの改善は、検出器を微調整したり、ノイズの多い擬似ラベルを直接訓練するよりも効果的であった。
論文 参考訳(メタデータ) (2026-07-21T04:35:47Z) - Reference-Induced Consensus for Selective Posed-Reference Visual Localization [11.166886974945806]
RIC-Locは、シーントレーニングなしのポーズ参照ローカライザである。
主推定器では、SfM-point-map-freeである。
論文 参考訳(メタデータ) (2026-07-06T06:48:34Z) - SDR: Set-Distance Rewards for Radiology Report Generation [2.2462804525190427]
生成した埋め込みと参照埋め込みの間の集合間距離を連続的、置換不変な報酬として提案する。
2つのデータセットと3つのビジョン言語モデルにまたがって、GRPOによるセットからセットまでの距離に基づく報酬による後トレーニングは、教師付き微調整よりも一貫して優れています。
同じ設定距離はテストタイムのベスト・オブ・N$選択を可能にする。
論文 参考訳(メタデータ) (2026-05-30T00:10:51Z) - CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging [95.02210956333374]
本稿では,一対の選好評価にルーブリックに基づく判断を適応させる基準中心のフレームワークを提案する。
BigCodeRewardでは、CriterAlignはQwen2.5-VL-32Bモノリシック判事を60.4%から66.3%に改善した。
論文 参考訳(メタデータ) (2026-05-19T10:59:19Z) - OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation [53.88666485159289]
OpenDeepThinkは、集団ベースのテスト時間計算フレームワークで、ペアワイズBradley-Terryの比較によって選択する。
OpenDeepThinkはGemini 3.1 ProのCodeforces Eloを8回のLCMコールラウンドで+405ポイント引き上げる。
CF-73は、国際グランドマスターアノテーションによる73の専門家評価コードフォース問題と、公式判決に対する99%の地域評価合意のキュレートされたセットである。
論文 参考訳(メタデータ) (2026-05-14T17:57:40Z) - Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration [76.08899010904652]
CapCalは、ランキング決定から位置バイアスを機械的に分離する、トレーニング不要のフレームワークである。
シングルパス効率を保ちながら、トレーニング不要の手法で優れた性能を発揮する。
論文 参考訳(メタデータ) (2026-04-11T10:47:22Z) - HLTCOE Evaluation Team at TREC 2025: VQA Track [76.85337417923331]
HLT評価チームはTREC VQAのAnswer Generation (AG)タスクに参加した。
回答生成における意味的精度とランキングの整合性を改善することを目的としたリストワイズ学習フレームワークを開発した。
論文 参考訳(メタデータ) (2025-12-08T17:25:13Z) - Rank-DETR for High Quality Object Detection [52.82810762221516]
高性能なオブジェクト検出器は、バウンディングボックス予測の正確なランキングを必要とする。
本研究では, 簡易かつ高性能なDETR型物体検出器について, 一連のランク指向設計を提案して紹介する。
論文 参考訳(メタデータ) (2023-10-13T04:48:32Z) - Location-Aware Box Reasoning for Anchor-Based Single-Shot Object
Detection [19.669531374307805]
単発物体検出器は、ボックス提案の事前選択がないため、ボックスの品質を損なう。
境界ボックスに対する位置認識型アンカーベース推論(LAAR)を提案する。
LAARは、バウンディングボックスの品質評価を考慮して、位置と分類の信頼性の両方を考慮に入れている。
論文 参考訳(メタデータ) (2020-07-13T08:24:41Z) - Generalized Focal Loss: Learning Qualified and Distributed Bounding
Boxes for Dense Object Detection [85.53263670166304]
一段検出器は基本的に、物体検出を密度の高い分類と位置化として定式化する。
1段検出器の最近の傾向は、局所化の質を推定するために個別の予測分岐を導入することである。
本稿では, 上記の3つの基本要素, 品質推定, 分類, ローカライゼーションについて述べる。
論文 参考訳(メタデータ) (2020-06-08T07:24:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。