GeoVerse.

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Kerui Ren1,2 Tao Lu2 Linning Xu3 Changjian Jiang4 Hunag Mu5 Chunhua Shen6,2 Mulin Yu2,† Bo Dai4,†
1Shanghai Jiao Tong University 2Shanghai Artificial Intelligence Laboratory 3The Chinese University of Hong Kong 4The University of Hong Kong 5Fudan University 6Zhejiang University

†Corresponding authors

TL;DR GeoVerse turns sparse images into world-consistent novel views by combining geometric latent diffusion, video generative priors, and persistent spatial memory.

What do we need for generation in geometric latent space?

  1. Stronger generative ability

    A geometric latent space provides strong structural priors, but limited training data can restrict its ability to complete unseen regions. We inject multilevel features from Wan2.2 VACE into geometric latent diffusion through a ControlNet-style adapter and expand training from 4 to 15 real and synthetic 3D datasets, enriching appearance priors and improving generalization across scenes and domains.

    Video priors + 15 training datasets
  2. Longer generation sequences

    Consistency within a single inference pass does not ensure consistency across successive generations. We maintain a global colored point-cloud memory that accumulates observed and generated content and reprojects it into requested cameras as target-aligned guidance, allowing new views to build on a shared scene representation while completing unseen regions.

    Persistent, target-aligned memory
  3. Faster generation

    High-step geometric latent diffusion can require over 100 seconds for a single inference pass. We apply reflow distillation to reduce sampling to four denoising updates: three at Level 1 and one in the Level 0 cascade. This reduces per-pass inference from over 100 seconds to under 10 seconds in our evaluation setting, while preserving visual detail.

    Four denoising updates
GeoVerse pipeline: spatial memory projects target-aligned guidance, Wan2.2 VACE provides video features, and geometric latent diffusion predicts target RGB and geometry to update memory.
Context observations initialize a global spatial memory that provides target-aligned RGB-D hints. Guided by the hints and Wan2.2 VACE features injected through a ControlNet-style adapter, geometric latent diffusion synthesizes target views in four denoising updates. The decoded RGB and geometry update the memory for subsequent view expansion.

Qualitative comparisons

Full resolution
Comparison with ViewCrafter, NeoVerse, GEN3C, MVGenMaster, Matrix3D, CAMEO, and GLD across indoor and outdoor scenes, followed by three rounds of long-sequence view generation.
Qualitative comparisons across multiple datasets show single-pass synthesis (top) and long-sequence generation over three successive rounds (bottom).

Few-step generation comparison

Full resolution
Figure 4: DL3DV comparisons across Level 1 and Level 0 denoising budgets, with enlarged crops showing fine texture and object boundaries.
Qualitative comparisons on DL3DV with different denoising budgets, showing how Level 1 generation and Level 0 cascade updates affect visual detail.

Citation

BibTeX
@misc{ren2026geoverse,
  title  = {GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space},
  author = {Ren, Kerui and Lu, Tao and Xu, Linning and Jiang, Changjian and
            Mu, Hunag and Shen, Chunhua and Yu, Mulin and Dai, Bo},
  year   = {2026},
  url    = {https://geoverse-nvs.github.io/}
}