You Only Label Once: 3D Box Adaptation from Point Cloud to Image via
Semi-Supervised Learning
- URL: http://arxiv.org/abs/2211.09302v2
- Date: Tue, 12 Sep 2023 16:49:56 GMT
- Title: You Only Label Once: 3D Box Adaptation from Point Cloud to Image via
Semi-Supervised Learning
- Authors: Jieqi Shi, Peiliang Li, Xiaozhi Chen, Shaojie Shen
- Abstract summary: We propose a learning-based 3D box adaptation approach that automatically adjusts minimum parameters of the Lidar 3D bounding box to perfectly fit the image appearance of panoramic cameras.
We are the first to focus on image-level cuboid refinement, which balances the accuracy and efficiency well and dramatically reduces the labeling effort for accurate cuboid annotation.
- Score: 31.914887148307706
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: The image-based 3D object detection task expects that the predicted 3D
bounding box has a ``tightness'' projection (also referred to as cuboid), which
fits the object contour well on the image while still keeping the geometric
attribute on the 3D space, e.g., physical dimension, pairwise orthogonal, etc.
These requirements bring significant challenges to the annotation. Simply
projecting the Lidar-labeled 3D boxes to the image leads to non-trivial
misalignment, while directly drawing a cuboid on the image cannot access the
original 3D information. In this work, we propose a learning-based 3D box
adaptation approach that automatically adjusts minimum parameters of the
360$^{\circ}$ Lidar 3D bounding box to perfectly fit the image appearance of
panoramic cameras. With only a few 2D boxes annotation as guidance during the
training phase, our network can produce accurate image-level cuboid annotations
with 3D properties from Lidar boxes. We call our method ``you only label
once'', which means labeling on the point cloud once and automatically adapting
to all surrounding cameras. As far as we know, we are the first to focus on
image-level cuboid refinement, which balances the accuracy and efficiency well
and dramatically reduces the labeling effort for accurate cuboid annotation.
Extensive experiments on the public Waymo and NuScenes datasets show that our
method can produce human-level cuboid annotation on the image without needing
manual adjustment.
Related papers
- Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection [3.3062934610311436]
We propose an approach named MMAssist that improves the performance of 3D UDA with multi-modal assistance.<n>A method is designed to align 3D features between the source domain and the target domain by using image and text features as bridges.<n> Experimental results show that our approach achieves promising performance compared with state-of-the-art methods in three domain adaptation tasks.
arXiv Detail & Related papers (2025-11-11T08:27:22Z) - Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation [66.65719382619538]
Current methods for 3D semantic segmentation propose training models with limited annotations to address the difficulty of annotating large, irregular, and unordered 3D point cloud data.<n>We present a novel approach that maximizes the utility of sparsely available 3D annotations incorporating segmentation masks generated by 2D foundation models.
arXiv Detail & Related papers (2025-08-27T14:13:01Z) - General Geometry-aware Weakly Supervised 3D Object Detection [62.26729317523975]
A unified framework is developed for learning 3D object detectors from RGB images and associated 2D boxes.
Experiments on KITTI and SUN-RGBD datasets demonstrate that our method yields surprisingly high-quality 3D bounding boxes with only 2D annotation.
arXiv Detail & Related papers (2024-07-18T17:52:08Z) - Segment, Lift and Fit: Automatic 3D Shape Labeling from 2D Prompts [50.181870446016376]
This paper proposes an algorithm for automatically labeling 3D objects from 2D point or box prompts.
Unlike previous arts, our auto-labeler predicts 3D shapes instead of bounding boxes and does not require training on a specific dataset.
arXiv Detail & Related papers (2024-07-16T04:53:28Z) - OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding [54.981605111365056]
This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding.
Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing.
arXiv Detail & Related papers (2024-06-04T07:42:33Z) - View Selection for 3D Captioning via Diffusion Ranking [54.78058803763221]
Cap3D method renders 3D objects into 2D views for captioning using pre-trained models.
Some rendered views of 3D objects are atypical, deviating from the training data of standard image captioning models and causing hallucinations.
We present DiffuRank, a method that leverages a pre-trained text-to-3D model to assess the alignment between 3D objects and their 2D rendered views.
arXiv Detail & Related papers (2024-04-11T17:58:11Z) - Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance [72.6809373191638]
We propose a framework to study how to leverage constraints between 2D and 3D domains without requiring any 3D labels.
Specifically, we design a feature-level constraint to align LiDAR and image features based on object-aware regions.
Second, the output-level constraint is developed to enforce the overlap between 2D and projected 3D box estimations.
Third, the training-level constraint is utilized by producing accurate and consistent 3D pseudo-labels that align with the visual data.
arXiv Detail & Related papers (2023-12-12T18:57:25Z) - 3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with
2D Diffusion Models [102.75875255071246]
3D content creation via text-driven stylization has played a fundamental challenge to multimedia and graphics community.
We propose a new 3DStyle-Diffusion model that triggers fine-grained stylization of 3D meshes with additional controllable appearance and geometric guidance from 2D Diffusion models.
arXiv Detail & Related papers (2023-11-09T15:51:27Z) - Text2Control3D: Controllable 3D Avatar Generation in Neural Radiance
Fields using Geometry-Guided Text-to-Image Diffusion Model [39.64952340472541]
We propose a controllable text-to-3D avatar generation method whose facial expression is controllable.
Our main strategy is to construct the 3D avatar in Neural Radiance Fields (NeRF) optimized with a set of controlled viewpoint-aware images.
We demonstrate the empirical results and discuss the effectiveness of our method.
arXiv Detail & Related papers (2023-09-07T08:14:46Z) - Generating Images with 3D Annotations Using Diffusion Models [32.77912877963642]
We propose 3D Diffusion Style Transfer (3D-DST), which incorporates 3D geometry control into diffusion models.
Our method exploits ControlNet, which extends diffusion models by using visual prompts in addition to text prompts.
With explicit 3D geometry control, we can easily change the 3D structures of the objects in the generated images and obtain ground-truth 3D automatically.
arXiv Detail & Related papers (2023-06-13T19:48:56Z) - WeakM3D: Towards Weakly Supervised Monocular 3D Object Detection [29.616568669869206]
Existing monocular 3D detection methods rely on manually annotated 3D box labels on the LiDAR point clouds.
In this paper, we explore the weakly supervised monocular 3D detection. Specifically, we first detect 2D boxes on the image. Then, we adopt the generated 2D boxes to select corresponding RoI LiDAR points as the weak supervision.
This network is learned by minimizing our newly-proposed 3D alignment loss between the 3D box estimates and the corresponding RoI LiDAR points.
arXiv Detail & Related papers (2022-03-16T00:37:08Z) - Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning
of 3D Pose [10.028521796737314]
We study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data.
Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D pose annotation from the labelled to unlabelled images reliably.
arXiv Detail & Related papers (2021-10-27T06:53:53Z) - 3D Shape Segmentation with Geometric Deep Learning [2.512827436728378]
We propose a neural-network based approach that produces 3D augmented views of the 3D shape to solve the whole segmentation as sub-segmentation problems.
We validate our approach using 3D shapes of publicly available datasets and of real objects that are reconstructed using photogrammetry techniques.
arXiv Detail & Related papers (2020-02-02T14:11:16Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.