Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh
Reconstruction
- URL: http://arxiv.org/abs/2303.13796v3
- Date: Thu, 24 Aug 2023 16:18:35 GMT
- Title: Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh
Reconstruction
- Authors: Wenjia Wang, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Yanjun
Wang, Chunhua Shen, Lei Yang, Taku Komura
- Abstract summary: Zolly is the first 3DHMR method focusing on perspective-distorted images.
We propose a new camera model and a novel 2D representation, termed distortion image, which describes the 2D dense distortion scale of the human body.
We extend two real-world datasets tailored for this task, all containing perspective-distorted human images.
- Score: 66.10717041384625
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: As it is hard to calibrate single-view RGB images in the wild, existing 3D
human mesh reconstruction (3DHMR) methods either use a constant large focal
length or estimate one based on the background environment context, which can
not tackle the problem of the torso, limb, hand or face distortion caused by
perspective camera projection when the camera is close to the human body. The
naive focal length assumptions can harm this task with the incorrectly
formulated projection matrices. To solve this, we propose Zolly, the first
3DHMR method focusing on perspective-distorted images. Our approach begins with
analysing the reason for perspective distortion, which we find is mainly caused
by the relative location of the human body to the camera center. We propose a
new camera model and a novel 2D representation, termed distortion image, which
describes the 2D dense distortion scale of the human body. We then estimate the
distance from distortion scale features rather than environment context
features. Afterwards, we integrate the distortion feature with image features
to reconstruct the body mesh. To formulate the correct projection matrix and
locate the human body position, we simultaneously use perspective and
weak-perspective projection loss. Since existing datasets could not handle this
task, we propose the first synthetic dataset PDHuman and extend two real-world
datasets tailored for this task, all containing perspective-distorted human
images. Extensive experiments show that Zolly outperforms existing
state-of-the-art methods on both perspective-distorted datasets and the
standard benchmark (3DPW).
Related papers
- CameraHMR: Aligning People with Perspective [54.05758012879385]
We address the challenge of accurate 3D human pose and shape estimation from monocular images.
Existing training datasets containing real images with pseudo ground truth (pGT) use SMPLify to fit SMPL to sparse 2D joint locations.
We make two contributions that improve pGT accuracy.
arXiv Detail & Related papers (2024-11-12T19:12:12Z) - FAMOUS: High-Fidelity Monocular 3D Human Digitization Using View Synthesis [51.193297565630886]
The challenge of accurately inferring texture remains, particularly in obscured areas such as the back of a person in frontal-view images.
This limitation in texture prediction largely stems from the scarcity of large-scale and diverse 3D datasets.
We propose leveraging extensive 2D fashion datasets to enhance both texture and shape prediction in 3D human digitization.
arXiv Detail & Related papers (2024-10-13T01:25:05Z) - StackFLOW: Monocular Human-Object Reconstruction by Stacked Normalizing Flow with Offset [56.71580976007712]
We propose to use the Human-Object Offset between anchors which are densely sampled from the surface of human mesh and object mesh to represent human-object spatial relation.
Based on this representation, we propose Stacked Normalizing Flow (StackFLOW) to infer the posterior distribution of human-object spatial relations from the image.
During the optimization stage, we finetune the human body pose and object 6D pose by maximizing the likelihood of samples.
arXiv Detail & Related papers (2024-07-30T04:57:21Z) - Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs [15.017274891943162]
Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision.
Inertial sensor has been introduced to provide complementary source of information.
It remains challenging to integrate heterogeneous sensor data for producing physically rational 3D human poses.
arXiv Detail & Related papers (2024-04-27T09:02:42Z) - Personalized 3D Human Pose and Shape Refinement [19.082329060985455]
regression-based methods have dominated the field of 3D human pose and shape estimation.
We propose to construct dense correspondences between initial human model estimates and the corresponding images.
We show that our approach not only consistently leads to better image-model alignment, but also to improved 3D accuracy.
arXiv Detail & Related papers (2024-03-18T10:13:53Z) - SelfPose: 3D Egocentric Pose Estimation from a Headset Mounted Camera [97.0162841635425]
We present a solution to egocentric 3D body pose estimation from monocular images captured from downward looking fish-eye cameras installed on the rim of a head mounted VR device.
This unusual viewpoint leads to images with unique visual appearance, with severe self-occlusions and perspective distortions.
We propose an encoder-decoder architecture with a novel multi-branch decoder designed to account for the varying uncertainty in 2D predictions.
arXiv Detail & Related papers (2020-11-02T16:18:06Z) - Synthetic Training for Monocular Human Mesh Recovery [100.38109761268639]
This paper aims to estimate 3D mesh of multiple body parts with large-scale differences from a single RGB image.
The main challenge is lacking training data that have complete 3D annotations of all body parts in 2D images.
We propose a depth-to-scale (D2S) projection to incorporate the depth difference into the projection function to derive per-joint scale variants.
arXiv Detail & Related papers (2020-10-27T03:31:35Z) - Beyond Weak Perspective for Monocular 3D Human Pose Estimation [6.883305568568084]
We consider the task of 3D joints location and orientation prediction from a monocular video.
We first infer 2D joints locations with an off-the-shelf pose estimation algorithm.
We then adhere to the SMPLify algorithm which receives those initial parameters.
arXiv Detail & Related papers (2020-09-14T16:23:14Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.