論文の概要: On the Efficacy of Self-Supervised Point Cloud Encoders for Efficient 3D Large Language Models
- arxiv url: http://arxiv.org/abs/2607.29136v1
- Date: Fri, 31 Jul 2026 08:06:56 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-03 16:35:22.224816
- Title: On the Efficacy of Self-Supervised Point Cloud Encoders for Efficient 3D Large Language Models
- Title(参考訳): 高速3次元大言語モデルのためのセルフスーパービジョンポイントクラウドエンコーダの有効性について
- Authors: Yao Zheng, Tian Zhang,
- Abstract要約: 3Dポイントクラウド言語モデル(3D-LLM)は、大きな言語モデルとのペアリングポイントクラウドエンコーダによる3D理解を可能にする。
既存の方法は8倍のA100スケールの計算で画像テキストポイントのクラウドアライメントを必要とするコストのかかるマルチモーダルエンコーダに依存している。
特にPCP-MAE と Point-MAE は,低コストで自己管理型クラウドエンコーダが有効な代替手段となるか検討する。
- 参考スコア(独自算出の注目度): 7.625207441727006
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: 3D point cloud-language models (3D-LLMs) enable 3D understanding by pairing point cloud encoders with large language models, but existing methods rely on costly multi-modal encoders (e.g., ULIP-2) that require image-text-point cloud alignment on 8x A100-scale compute, creating high barriers for research and deployment. In this work, we systematically investigate whether low-cost self-supervised point cloud encoders, specifically PCP-MAE and Point-MAE, can serve as effective alternatives. Using MiniGPT-3D as our testbed, we evaluate 7 encoder initialization/pre-training setups (1 multi-modal baseline, 5 self-supervised, 1 random init) under frozen and unfrozen fine-tuning (12 total groups), across 2 architectures (MaskTransformer, PointTransformer), 3 objectives (PCP-MAE, Point-MAE, random init), and 2 datasets (Objaverse 660K, ShapeNet55-34 approximately 50K). Our experiments reveal three key findings: (1) The four-stage MiniGPT-3D pipeline can effectively train a 3D encoder from random initialization: an end-to-end trained random init encoder reaches 52.50% open-vocabulary accuracy and 44.45 captioning score, approaching top pre-trained variants; (2) Architecture and pre-training objective show strong crossover interaction: PCP-MAE + MaskTransformer achieves 59.00% accuracy (best self-supervised), while Point-MAE + MaskTransformer drops to 46.50%, with the pattern reversed for PointTransformer; (3) Closed-set ModelNet40 classification remains a core weakness of purely geometric encoders, reaching only ~13-18% accuracy vs. ~62% for the multi-modal baseline, even after end-to-end fine-tuning. Our results offer practical guidelines for cost-effective 3D-LLM design and reveal interaction patterns between self-supervised objectives and encoder architectures.
- Abstract(参考訳): 3Dポイントのクラウド言語モデル(3D-LLM)は、ポイントクラウドエンコーダと大きな言語モデルとのペアリングによる3D理解を可能にする。
本研究では,PCP-MAE や Point-MAE など,低コストの自己管理型クラウドエンコーダが有効な代替手段となるかどうかを系統的に検討する。
テストベッドとしてMiniGPT-3Dを用いて,2つのアーキテクチャ(PCP-MAE, Point-MAE, random init)と2つのデータセット(Objaverse 660K, ShapeNet55-34, 約50K)にまたがる7つのエンコーダの初期化/事前トレーニングセットアップ(マルチモーダルベースライン, 5つの自己教師, 1つのランダムinit)を評価した。
4段階のMiniGPT-3Dパイプラインは、ランダム初期化から効果的に3Dエンコーダを訓練できる: エンドツーエンドのトレーニングされたランダムなinitエンコーダは、52.50%のオープンボキャブラリ精度と44.45のキャプションスコアに達し、トップトレーニング済みの変種に近づき、アーキテクチャと事前学習対象が強いクロスオーバー相互作用を示す: PCP-MAE + MaskTransformerは59.00%の精度で、Point-MAE + MaskTransformerは46.50%の精度で、パターンはPoint Transformerに逆転する。
本研究は,費用対効果の高い3D-LLM設計のための実践的ガイドラインを提供し,自己監督対象とエンコーダアーキテクチャとの相互作用パターンを明らかにする。
関連論文リスト
- CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation [37.2660021156429]
CLIPoint3Dは、教師なしの3Dポイントクラウドドメイン適応のためのフレームワークである。
アプローチでは3Dサンプルを多重深度マップに投影し,凍結したCLIPバックボーンを活用する。
PointDA-10とGraspNetPC-10ベンチマークの実験では、CLIPoint3Dは3-16%の精度向上を達成した。
論文 参考訳(メタデータ) (2026-02-23T23:17:12Z) - TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP [52.79100775328595]
3Dビジュアルグラウンドティングは、人間の指示に基づいて現実世界の3D環境における視覚情報を理解するための具体的エージェントである。
既存の3Dビジュアルグラウンド法は、異なるモダリティの異なるエンコーダに依存している。
本稿では,3つのモードすべてを処理するために,統合された2次元事前学習型マルチモーダルネットワークを提案する。
論文 参考訳(メタデータ) (2025-07-20T10:28:06Z) - YOLOO: You Only Learn from Others Once [27.222676133154284]
我々は,新しいマルチモーダル3DMOTパラダイムである textbyoLOO を提案する。
YOLOOはポイントクラウドエンコーダに、ポイントクラウドや他のモダリティ(画像やテキストキューなど)から統一されたトリモーダル表現(UTR)を一度に学習する権限を与える。
特に、YOLOOは、2つのコアコンポーネント: 統一三モードエンコーダ(UTEnc)とフレキシブルな幾何学的制約(F-GC)モジュール。
論文 参考訳(メタデータ) (2024-09-01T05:09:32Z) - Joint Beam Search Integrating CTC, Attention, and Transducer Decoders [53.297697898510194]
4つのデコーダが同一のエンコーダを共有するような共同モデリング手法を提案する。
4Dモデルは共同で訓練され、モデルの正規化とモデルの堅牢性を最大化する。
さらに,3つのデコーダを組み合わせることで,新しい3つのビーム探索アルゴリズムを提案する。
論文 参考訳(メタデータ) (2024-06-05T05:18:20Z) - Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud
Pre-training [65.75399500494343]
Masked Autoencoders (MAE) は、2Dおよび3Dコンピュータビジョンのための自己教師型学習において有望な性能を示した。
自己監督型3次元点雲事前学習のための2D-3DジョイントMAEフレームワークであるJoint-MAEを提案する。
論文 参考訳(メタデータ) (2023-02-27T17:56:18Z) - Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud
Pre-training [56.81809311892475]
Masked Autoencoders (MAE) は、言語と2次元画像変換器の自己教師付き事前学習において大きな可能性を示している。
我々は3次元点雲の階層的自己教師型学習のための強力なマルチスケールMAE事前学習フレームワークであるPoint-M2AEを提案する。
論文 参考訳(メタデータ) (2022-05-28T11:22:53Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。