World-Consistent Video-to-Video Synthesis
- URL: http://arxiv.org/abs/2007.08509v1
- Date: Thu, 16 Jul 2020 17:58:13 GMT
- Title: World-Consistent Video-to-Video Synthesis
- Authors: Arun Mallya, Ting-Chun Wang, Karan Sapra, Ming-Yu Liu
- Abstract summary: We introduce a novel vid2vid framework that efficiently utilizes all past generated frames during rendering.
This is achieved by condensing the 3D world rendered so far into a physically-grounded estimate of the current frame.
We propose a novel neural network architecture to take advantage of the information stored in the guidance images.
- Score: 35.617437747886484
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Video-to-video synthesis (vid2vid) aims for converting high-level semantic
inputs to photorealistic videos. While existing vid2vid methods can achieve
short-term temporal consistency, they fail to ensure the long-term one. This is
because they lack knowledge of the 3D world being rendered and generate each
frame only based on the past few frames. To address the limitation, we
introduce a novel vid2vid framework that efficiently and effectively utilizes
all past generated frames during rendering. This is achieved by condensing the
3D world rendered so far into a physically-grounded estimate of the current
frame, which we call the guidance image. We further propose a novel neural
network architecture to take advantage of the information stored in the
guidance images. Extensive experimental results on several challenging datasets
verify the effectiveness of our approach in achieving world consistency - the
output video is consistent within the entire rendered 3D world.
https://nvlabs.github.io/wc-vid2vid/
Related papers
- Efficient Camera-Controlled Video Generation of Static Scenes via Sparse Diffusion and 3D Rendering [15.79758281898629]
generative models can produce very realistic clips, but they are computationally inefficient, often requiring minutes of GPU time for just a few seconds of video.<n>This paper explores a new strategy for camera-conditioned video generation of static scenes.<n>Our approach amortizes the generation cost across hundreds of frames while enforcing geometric consistency.
arXiv Detail & Related papers (2026-01-14T18:50:06Z) - Endless World: Real-Time 3D-Aware Long Video Generation [57.411689597435334]
Endless World is a real-time framework for infinite, 3D-consistent video generation.<n>We introduce a conditional autoregressive training strategy that aligns newly generated content with existing video frames.<n>Our 3D injection mechanism enforces physical plausibility and geometric consistency throughout extended sequences.
arXiv Detail & Related papers (2025-12-13T19:06:12Z) - S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix [60.060882467801484]
We present a pose-free and training-free method that leverages an off-the-shelf monocular video generation model to produce immersive 3D videos.<n>Our approach first warps the generated monocular video into pre-defined camera viewpoints using estimated depth information, then applies a novel textitframe matrix inpainting framework.<n>We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, such as Sora, Lumiere, WALT, and Zeroscope.
arXiv Detail & Related papers (2025-08-11T14:50:03Z) - Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy [36.08715662927022]
We present Shape-for-Motion, a novel framework that incorporates a 3D proxy for precise and consistent video editing.<n>Our framework supports various precise and physically-consistent manipulations across the video frames, including pose editing, rotation, scaling, translation, texture modification, and object composition.
arXiv Detail & Related papers (2025-06-27T17:59:01Z) - Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation [66.95956271144982]
We present Voyager, a novel video diffusion framework that generates world-consistent 3D point-cloud sequences from a single image.<n>Unlike existing approaches, Voyager achieves end-to-end scene generation and reconstruction with inherent consistency across frames.
arXiv Detail & Related papers (2025-06-04T17:59:04Z) - Eye2Eye: A Simple Approach for Monocular-to-Stereo Video Synthesis [45.64047250474718]
Despite advances in video synthesis, creating 3D videos remains challenging due to the relative scarcity of 3D video data.
We propose a simple approach for transforming a text-to-video generator into a video-to-stereo generator.
Our framework automatically produces the video frames from a shifted viewpoint, enabling a compelling 3D effect.
arXiv Detail & Related papers (2025-04-30T19:06:09Z) - VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step [13.168559963356952]
VideoScene aims to distill the video diffusion model to generate 3D scenes in one step.
VideoScene achieves faster and superior 3D scene generation results than previous video diffusion models.
arXiv Detail & Related papers (2025-04-02T17:59:21Z) - Seeing World Dynamics in a Nutshell [132.79736435144403]
NutWorld is a framework that transforms monocular videos into dynamic 3D representations in a single forward pass.
We demonstrate that NutWorld achieves high-fidelity video reconstruction quality while enabling downstream applications in real-time.
arXiv Detail & Related papers (2025-02-05T18:59:52Z) - SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix [60.48666051245761]
We propose a pose-free and training-free approach for generating 3D stereoscopic videos.
Our method warps a generated monocular video into camera views on stereoscopic baseline using estimated video depth.
We develop a disocclusion boundary re-injection scheme that further improves the quality of video inpainting.
arXiv Detail & Related papers (2024-06-29T08:33:55Z) - Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting [94.84688557937123]
Video-3DGS is a 3D Gaussian Splatting (3DGS)-based video refiner designed to enhance temporal consistency in zero-shot video editors.
Our approach utilizes a two-stage 3D Gaussian optimizing process tailored for editing dynamic monocular videos.
It enhances video editing by ensuring temporal consistency across 58 dynamic monocular videos.
arXiv Detail & Related papers (2024-06-04T17:57:37Z) - Neural Video Fields Editing [56.558490998753456]
NVEdit is a text-driven video editing framework designed to mitigate memory overhead and improve consistency.
We construct a neural video field, powered by tri-plane and sparse grid, to enable encoding long videos with hundreds of frames.
Next, we update the video field through off-the-shelf Text-to-Image (T2I) models to text-driven editing effects.
arXiv Detail & Related papers (2023-12-12T14:48:48Z) - Flexible Techniques for Differentiable Rendering with 3D Gaussians [29.602516169951556]
Neural Radiance Fields demonstrated photorealistic novel view is within reach, but was gated by performance requirements for fast reconstruction of real scenes and objects.
We develop extensions to alternative shape representations, in particular, 3D watertight meshes and rendering per-ray normals.
These reconstructions are quick, robust, and easily performed on GPU or CPU.
arXiv Detail & Related papers (2023-08-28T17:38:31Z) - DiffSynth: Latent In-Iteration Deflickering for Realistic Video
Synthesis [15.857449277106827]
DiffSynth is a novel approach to convert image synthesis pipelines to video synthesis pipelines.
It consists of a latent in-it deflickering framework and a video deflickering algorithm.
One of the notable advantages of Diff Synth is its general applicability to various video synthesis tasks.
arXiv Detail & Related papers (2023-08-07T10:41:52Z) - ControlVideo: Training-free Controllable Text-to-Video Generation [117.06302461557044]
ControlVideo is a framework to enable natural and efficient text-to-video generation.
It generates both short and long videos within several minutes using one NVIDIA 2080Ti.
arXiv Detail & Related papers (2023-05-22T14:48:53Z) - Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video
Generators [70.17041424896507]
Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets.
We propose a new task of zero-shot text-to-video generation using existing text-to-image synthesis methods.
Our method performs comparably or sometimes better than recent approaches, despite not being trained on additional video data.
arXiv Detail & Related papers (2023-03-23T17:01:59Z) - 3D Video Loops from Asynchronous Input [22.52716577813998]
Looping videos are short video clips that can be looped endlessly without visible seams or artifacts.
In this paper, we propose a practical solution that enables an immersive experience on dynamic 3D looping scenes.
Experiments of our framework have shown promise in successfully generating and rendering 3D looping videos in real time even on mobile devices.
arXiv Detail & Related papers (2023-03-09T15:00:12Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.