論文の概要: YOLOv14:Unified Cross-Domain Real-Time Object Detectionwith Adaptive Multi-View Representation
- arxiv url: http://arxiv.org/abs/2608.04720v1
- Date: Wed, 05 Aug 2026 11:39:06 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.855946
- Title: YOLOv14:Unified Cross-Domain Real-Time Object Detectionwith Adaptive Multi-View Representation
- Title(参考訳): YOLOv14:Adaptive Multi-View Representationを用いた統合クロスドメインリアルタイムオブジェクト検出
- Authors: Jinling Jia, Jian Lu, Jone Yawl, Chenbin Zhang,
- Abstract要約: リアルタイム物体検出器は、制御された条件下では顕著な精度を達成するが、非理想的な入力では急激に劣化する。
4つの相乗的革新を通じてこれらの課題に対処する統合検出フレームワークYOLOv14を提案する。
YOLOv14achieves 49.1 mAP on COCO val 2017 at 2.91 ms (T4 GPU) で、魚眼(+4.1 mAP)、パノラマ(+6.6 mAP)、ドローン(+6.4 mAP)、ゲームキャラクタ(+26.1 mAP)のベンチマークでかなりの利益を得ている。
- 参考スコア(独自算出の注目度): 8.040649850566211
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Real-time object detectors achieve remarkable accuracy under controlled conditions, yet degrade sharply on non-ideal inputs: fisheye distortion, game-rendered characters, aerial viewpoints, and 360° panoramas. We present YOLOv14, aunified detection framework addressing these challenges through four synergisticinnovations. (1) Deformable Area-Attention (D-AAttn) replaces rigid attentiongrids with learned 2D deformation fields, enabling adaptive sampling under geometric distortion. (2) Game2Real Domain Adaptation aligns rendered-game and photographic feature distributions via Adaptive Instance Normalization (AdaIN)and adversarial domain confusion, allowing game characters are detected as realhumans. (3) Multi-View Conditioning injects learned viewpoint embeddings intothe backbone with a cross-view contrastive loss that pulls same-class features fromdifferent perspectives closer. (4) An Adaptive Augmentation Policy automaticallyclassifies each input' scene type and routes to optimal augmentations, while a DynamicScaleRouter learns per-input feature pyramid weights. Together, YOLOv14achieves 49.1 mAP on COCO val2017 at 2.91 ms (T4 GPU), and delivers substantial gains on fisheye (+4.1 mAP), panorama (+6.6 mAP), drone (+6.4 mAP), andgame-character (+26.1 mAP) benchmarks
- Abstract(参考訳): リアルタイム物体検出器は、制御された条件下では顕著な精度を達成するが、魚眼歪み、ゲームレンダリング文字、空中視界、360度パノラマといった非理想的な入力によって著しく劣化する。
4つの相乗的革新を通じてこれらの課題に対処する統合検出フレームワークYOLOv14を提案する。
1) 変形性エリアアテンション (D-AAttn) は剛性アテンショングレードを学習された2次元変形場に置き換え, 幾何歪み下での適応サンプリングを可能にする。
2)ゲーム2Real Domain Adaptationは,AdaIN(Adaptive Instance Normalization, 適応インスタンス正規化)と対向ドメイン混乱を通じて,レンダリングゲームと写真の特徴分布を整列させ,ゲームキャラクタをリアル人間として検出する。
(3)マルチビューコンディショニングは、異なる視点から同一の特徴を引き出すクロスビューコントラスト損失で学習された視点埋め込みをバックボーンに注入する。
(4)Adaptive Augmentation Policyは各入力のシーンタイプとルートを自動的に分類し、DynamicScaleRouterは入力毎の特徴ピラミッド重みを学習する。
YOLOv14achieves 49.1 mAP on COCO val2017 at 2.91 ms (T4 GPU) and provide a significant gains on fisheye (+4.1 mAP), panorama (+6.6 mAP), drone (+6.4 mAP), andgame-character (+26.1 mAP) benchmarks
関連論文リスト
- Geometric-aware Pretraining for Vision-centric 3D Object Detection [77.7979088689944]
GAPretrainと呼ばれる新しい幾何学的事前学習フレームワークを提案する。
GAPretrainは、複数の最先端検出器に柔軟に適用可能なプラグアンドプレイソリューションとして機能する。
BEVFormer法を用いて, nuScenes val の 46.2 mAP と 55.5 NDS を実現し, それぞれ 2.7 と 2.1 点を得た。
論文 参考訳(メタデータ) (2023-04-06T14:33:05Z) - Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation [73.48323921632506]
パノラマ的セマンティックセマンティックセグメンテーションは2つの重要な課題により未探索である。
まず、変形性パッチ埋め込み(DPE)と変形性(DMLPv2)モジュールを備えたパノラマセマンティックトランス4PASS+を改良したトランスフォーマーを提案する。
第2に、教師なしドメイン適応パノラマセグメンテーションのための擬似ラベル修正により、Mutual Prototypeal Adaptation(MPA)戦略を強化する。
第3に、Pinhole-to-Panoramic(Pin2Pan)適応とは別に、9,080パノラマ画像を用いた新しいデータセット(SynPASS)を作成します。
論文 参考訳(メタデータ) (2022-07-25T00:42:38Z) - Decoupled Adaptation for Cross-Domain Object Detection [69.5852335091519]
クロスドメインオブジェクト検出は、オブジェクト分類よりも難しい。
D-adaptは4つのクロスドメインオブジェクト検出タスクで最先端の結果を達成する。
論文 参考訳(メタデータ) (2021-10-06T08:43:59Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。