Back to Academic/Research
Open WeightsSpecializedimagevideo3dUpdated June 15, 2026

World Tracing: Generative Multilayer 3D Geometry Representation

Model Overview

World Tracing is a generative 3D geometry representation framework developed by researchers from World Labs and University of Illinois Urbana-Champaign (UIUC) (Hao Zhang, Mohamed El Banani, et al.).

It bridges the gap between monocular depth estimation (high pixel alignment, incomplete geometry) and 3D generative models (complete geometry, poor pixel alignment). Given a single 2D image or short video clip, World Tracing predicts an ordered stack of camera-space 3D points (X, Y, Z) along camera rays for every pixel. The first layer represents visible surfaces, while subsequent layers generate occluded geometry behind them.


Key Features

  • Multilayer Pixel-Aligned Geometry: Predicts an ordered stack of 3D points per pixel, maintaining 1:1 alignment while capturing occluded surfaces.
  • WT-DiT Architecture: Utilizes a Flow-Matching Diffusion Transformer that treats geometry layers as coupled denoising tokens with factorized attention.
  • Unified Multi-Domain Support: Provides specialized variants for single objects (WT-O), scenes (WT-S), and dynamic video sequences.
  • Zero-Shot Downstream Integration: Enables text-guided 3D scene editing, novel-view video synthesis, and integration with mesh generators.
  • Faithfulness & Completeness Balance: Preserves high-precision visible surface depth accuracy while generating complete 3D structure.

Verified Project Links


Benchmarks & Evaluation

  • Objects & 3D-FRONT Benchmarks: Outperforms monocular depth and image-to-3D generative baselines in depth accuracy and occluded geometry reconstruction.

Key Features

Multilayer Pixel-Aligned Geometry: Predicts an ordered stack of 3D points per pixel, maintaining 1:1 alignment while modeling occluded surfaces

Feature 01

WT-DiT Architecture: Flow-Matching Diffusion Transformer treating geometry layers as coupled denoising tokens with global attention

Feature 02

Unified Multi-Domain Support: Features specialized variants for single objects (WT-O), indoor/outdoor scenes (WT-S), and dynamic videos

Feature 03

Zero-Shot Downstream Integration: Enables text-guided 3D editing, novel-view video synthesis, and mesh generation (TRELLIS)

Feature 04

Faithfulness & Completeness Balance: Preserves high-precision visible surface depth accuracy while generating complete 3D structures

Feature 05

You might also want to compare

Verified Sources

Tags

research-preview3d-reconstructionpoint-cloudworld-labsuiucoccluded-geometry

Model Specs

open-weights

Parameters

Undisclosed

Context Window

undisclosed

License

CC-BY-NC-ND-4.0

Deployment

self-hostable

Resources & Links

Curator Notes

Verified paper arXiv:2606.13652 and open-source project from World Labs and UIUC.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model