Fugu-MT 論文翻訳(概要): Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning

論文の概要: Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning

arxiv url: http://arxiv.org/abs/2310.03670v1
Date: Mon, 25 Sep 2023 17:23:33 GMT
ステータス: 翻訳完了
システム内更新日: 2023-10-08 11:01:03.746558
Title: Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning
Title（参考訳）: 構築前の回帰:ポイントクラウドによる自己教師型学習のための回帰オートエンコーダ
Authors: Yang Liu, Chen Chen, Can Wang, Xulin King, Mengyuan Liu
Abstract要約: Masked Autoencoders (MAE) は、2Dおよび3Dコンピュータビジョンのための自己教師型学習において有望な性能を示した。我々は、ポイントクラウド自己教師型学習のための回帰オートエンコーダの新しいスキーム、Point Regress AutoEncoder (Point-RAE)を提案する。本手法は, 各種下流タスクの事前学習において効率よく, 一般化可能である。
参考スコア（独自算出の注目度）: 18.10704604275133
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Masked Autoencoders (MAE) have demonstrated promising performance in self-supervised learning for both 2D and 3D computer vision. Nevertheless, existing MAE-based methods still have certain drawbacks. Firstly, the functional decoupling between the encoder and decoder is incomplete, which limits the encoder's representation learning ability. Secondly, downstream tasks solely utilize the encoder, failing to fully leverage the knowledge acquired through the encoder-decoder architecture in the pre-text task. In this paper, we propose Point Regress AutoEncoder (Point-RAE), a new scheme for regressive autoencoders for point cloud self-supervised learning. The proposed method decouples functions between the decoder and the encoder by introducing a mask regressor, which predicts the masked patch representation from the visible patch representation encoded by the encoder and the decoder reconstructs the target from the predicted masked patch representation. By doing so, we minimize the impact of decoder updates on the representation space of the encoder. Moreover, we introduce an alignment constraint to ensure that the representations for masked patches, predicted from the encoded representations of visible patches, are aligned with the masked patch presentations computed from the encoder. To make full use of the knowledge learned in the pre-training stage, we design a new finetune mode for the proposed Point-RAE. Extensive experiments demonstrate that our approach is efficient during pre-training and generalizes well on various downstream tasks. Specifically, our pre-trained models achieve a high accuracy of \textbf{90.28\%} on the ScanObjectNN hardest split and \textbf{94.1\%} accuracy on ModelNet40, surpassing all the other self-supervised learning methods. Our code and pretrained model are public available at: \url{https://github.com/liuyyy111/Point-RAE}.
Abstract（参考訳）: マスク付きオートエンコーダ(mae)は、2dおよび3dコンピュータビジョンの自己教師あり学習において有望な性能を示している。それにもかかわらず、既存のmaeベースの手法には一定の欠点がある。まず、エンコーダとデコーダの間の関数的デカップリングは不完全であり、エンコーダの表現学習能力を制限する。次に、ダウンストリームタスクはエンコーダのみを使用し、プリテキストタスクでエンコーダ-デコーダアーキテクチャによって得られる知識を十分に活用できない。本稿では,ポイントクラウド自己教師型学習のための回帰オートエンコーダの新しい手法であるPoint Regress AutoEncoder (Point-RAE)を提案する。提案手法は,エンコーダが符号化した可視パッチ表現からマスクパッチ表現を予測し,デコーダが予測したマスクパッチ表現からターゲットを再構成するマスクレグレッサーを導入することで,デコーダとエンコーダとの間の機能を分離する。これにより、エンコーダの表現空間に対するデコーダ更新の影響を最小限に抑えることができる。さらに,可視パッチの符号化表現から予測されるマスクパッチの表現が,エンコーダから計算されたマスクパッチの表現と一致していることを保証するためにアライメント制約を導入する。事前学習段階で学習した知識をフル活用するために,提案したポイント-RAEのためのファインチューンモードを設計する。広範な実験により,我々のアプローチは事前学習時に効率的であり,様々な下流タスクをうまく一般化できることが証明された。具体的には、事前学習されたモデルは、scanobjectnn hardest split における \textbf{90.28\%} と modelnet40 における \textbf{94.1\%} の精度を高い精度で達成し、他の全ての自己教師付き学習方法を超える。私たちのコードと事前訓練されたモデルは、以下の通り公開されている。

論文の概要: Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning

関連論文リスト