論文の概要: From Skeletons to Pixels: Few-Shot Precise Event Spotting via Representation and Prediction Distillation
- arxiv url: http://arxiv.org/abs/2604.22839v1
- Date: Tue, 21 Apr 2026 06:43:04 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-28 17:12:06.976859
- Title: From Skeletons to Pixels: Few-Shot Precise Event Spotting via Representation and Prediction Distillation
- Title(参考訳): 骨格からレンズへ:表現と予測蒸留による精密イベントスポッティング
- Authors: Zhong Han Ervin Yeoh, Jiang Kan,
- Abstract要約: 数発の精密イベントスポッティングのための2つの補完蒸留法について検討した。
AWD(Adaptive Weight Distillation)とAnnealed Multimodal Distillation for Few-Shot Event Detection(AMD-FED)を評価した。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Precise Event Spotting (PES) is essential in fast-paced sports such as tennis, where fine-grained events occur within very short temporal windows. Accurate frame-level localization is challenging because of motion blur, subtle action differences, and limited annotated data. We study two complementary distillation strategies for few-shot PES: Adaptive Weight Distillation (AWD), a prediction-level method that adaptively weights teacher supervision on unlabeled data, and Annealed Multimodal Distillation for Few-Shot Event Detection (AMD-FED), a representation-level framework that transfers robust skeleton knowledge into visual modalities through annealed pseudo-labeling. Both methods use multimodal distillation to improve generalization under limited supervision. We evaluate them on F3Set-Tennis(sub) under few-shot k-clip settings, where they consistently outperform single-modality baselines and prior PES approaches. After observing the stronger performance of representation-level distillation on tennis, we further validate AMD-FED on a second sports dataset, Figure Skating, where it also shows robust performance in the k-clip scenario. These results highlight the effectiveness of multimodal distillation, especially representation-level transfer, for few-shot precise event spotting.
- Abstract(参考訳): 精密イベントスポッティング(PES)は、非常に短い時間窓内できめ細かなイベントが発生するテニスのような速いペースのスポーツにおいて必須である。
フレームレベルの正確なローカライゼーションは、動きのぼかし、微妙な動作の違い、限られた注釈付きデータのために困難である。
アダプティブ・ウェイト蒸留(AWD)とアナールド・マルチモーダル蒸留(Annealed Multimodal Distillation for Few-Shot Event Detection,AMD-FED)の2つの相補的蒸留手法について検討した。
どちらの方法も、限定的な監督の下で一般化を改善するためにマルチモーダル蒸留を用いる。
我々は、F3Set-Tennis(sub)をk-clip設定で評価し、単一のモダリティベースラインと以前のPSSアプローチを一貫して上回っている。
テニスにおける表現レベル蒸留の強い性能を観察した後、第2のスポーツデータセットであるフィギュアスケートでAMD-FEDを検証する。
これらの結果は, マルチモーダル蒸留, 特に表現レベルの伝達が, 数発の正確なイベントスポッティングに有効であることを示す。
関連論文リスト
- Few-Shot Precise Event Spotting via Unified Multi-Entity Graph and Distillation [15.108898002423734]
イベントスポッティングはスポーツ分析の重要なコンポーネントである。
現在の手法は、大きなラベル付きデータセットによるドメイン固有のエンドツーエンドのトレーニングに依存している。
本稿では,MPSのためのUMEG-Net(Unified Multi-Entity Graph Network)を提案する。
論文 参考訳(メタデータ) (2025-11-18T06:45:42Z) - PSTTS: A Plug-and-Play Token Selector for Efficient Event-based Spatio-temporal Representation Learning [25.271901669843363]
イベントデータに対するPSTTS(Progressive Spatio-temporal Token Selection)を提案する。
PSTTSは、生のイベントデータに埋め込まれた時間的・時間的分布特性を利用して、冗長トークンを効果的に識別し、破棄する。
PSTTSはFLOPを29-43.6%削減し、FPSを21.6-41.3%増加させた。
論文 参考訳(メタデータ) (2025-09-26T15:30:00Z) - Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression [46.25518274714238]
アクションアセスメント(AQA)は、運動性能の自動的、公平な評価を目的としている。
現在の手法では、動画を固定フレームに分割することに集中しており、サブアクションの時間的連続性を損なう。
階層的なポーズ誘導型多段階コントラスト回帰による行動品質評価手法を提案する。
論文 参考訳(メタデータ) (2025-01-07T10:20:16Z) - Scale-Equivalent Distillation for Semi-Supervised Object Detection [57.59525453301374]
近年のSemi-Supervised Object Detection (SS-OD) 法は主に自己学習に基づいており、教師モデルにより、ラベルなしデータを監視信号としてハードな擬似ラベルを生成する。
実験結果から,これらの手法が直面する課題を分析した。
本稿では,大規模オブジェクトサイズの分散とクラス不均衡に頑健な簡易かつ効果的なエンド・ツー・エンド知識蒸留フレームワークであるSED(Scale-Equivalent Distillation)を提案する。
論文 参考訳(メタデータ) (2022-03-23T07:33:37Z) - Activation to Saliency: Forming High-Quality Labels for Unsupervised
Salient Object Detection [54.92703325989853]
本稿では,高品質なサリエンシキューを効果的に生成する2段階アクティベーション・ツー・サリエンシ(A2S)フレームワークを提案する。
トレーニングプロセス全体において、私たちのフレームワークにヒューマンアノテーションは関与していません。
本フレームワークは,既存のUSOD法と比較して高い性能を示した。
論文 参考訳(メタデータ) (2021-12-07T11:54:06Z) - EvDistill: Asynchronous Events to End-task Learning via Bidirectional
Reconstruction-guided Cross-modal Knowledge Distillation [61.33010904301476]
イベントカメラは画素ごとの強度変化を感知し、ダイナミックレンジが高く、動きのぼやけが少ない非同期イベントストリームを生成する。
本稿では,bfEvDistillと呼ばれる新しい手法を提案し,未ラベルのイベントデータから学生ネットワークを学習する。
EvDistillは、イベントとAPSフレームのみのKDよりもはるかに優れた結果が得られることを示す。
論文 参考訳(メタデータ) (2021-11-24T08:48:16Z) - Few-Shot Fine-Grained Action Recognition via Bidirectional Attention and
Contrastive Meta-Learning [51.03781020616402]
現実世界のアプリケーションで特定のアクション理解の需要が高まっているため、きめ細かいアクション認識が注目を集めている。
そこで本研究では,各クラスに付与されるサンプル数だけを用いて,新規なきめ細かい動作を認識することを目的とした,数発のきめ細かな動作認識問題を提案する。
粒度の粗い動作では進展があったが、既存の数発の認識手法では、粒度の細かい動作を扱う2つの問題に遭遇する。
論文 参考訳(メタデータ) (2021-08-15T02:21:01Z) - RMS-Net: Regression and Masking for Soccer Event Spotting [52.742046866220484]
イベントラベルとその時間的オフセットを同時に予測できる,軽量でモジュール化されたアクションスポッティングネットワークを開発した。
SoccerNetデータセットでテストし、標準機能を使用して、完全な提案は3平均mAPポイントで現在の状態を超えます。
論文 参考訳(メタデータ) (2021-02-15T16:04:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。