Back to NVIDIA
Open WeightsSpecializedimagevideo3dUpdated May 15, 2026

Déjà View (DVLT): Looping Transformers for Multi-View 3D Reconstruction

Model Overview

Déjà View (DVLT - Déjà View Looping Transformer) is a parameter-efficient model for multi-view 3D reconstruction developed by NVIDIA Dynamic Vision and Learning (DVL) Group, ETH Zürich, and University of Toronto (Vector Institute) (Alessandro Burzio, Tobias Fischer, Zan Gojcic, et al.).

Rather than scaling model depth via deep stacks of distinct transformer layers, DVLT employs a single, weight-tied transformer block applied recurrently across per-view features over $K$ refinement steps. From unposed RGB images or video, DVLT jointly predicts per-pixel depth, ray maps, and camera intrinsics/extrinsics.


Key Features

  • Weight-Tied Looping Transformer Block: Uses a single shared transformer block recurrently over $K$ refinement steps instead of deep stacks of unshared layers.
  • Inference-Time Compute Knob ($K$): Allows users to adjust refinement steps dynamically during inference to balance latency and accuracy.
  • Parameter Efficiency (~117M Parameters): Operates at ~117M parameters—achieving SOTA performance using 8–10× fewer parameters and 1.9–2.3× less compute than billion-parameter baselines.
  • Joint Geometry & Camera Estimation: Infers dense per-pixel depth, ray maps, and camera intrinsics/extrinsics simultaneously.
  • Cross-Domain Generalization: Robust performance across indoor (ScanNet), outdoor (Waymo), and object-centric (CO3D) environments.

Verified Project Links


Performance & Benchmarks

  • Outperforms feed-forward baselines (DUSt3R, MASt3R, VGGSfM) on ScanNet, ScanNet++, DTU, and RealEstate10K datasets.

Key Features

Weight-Tied Looping Transformer Block: Uses a single shared transformer block recurrently over K refinement steps

Feature 01

Inference-Time Compute Knob (K): Dynamically adjust the number of refinement steps during inference to balance latency and accuracy

Feature 02

Parameter Efficiency (~117M Parameters): Operates at ~117M parameters—achieving SOTA performance with 8-10x fewer parameters

Feature 03

Joint Geometry & Camera Estimation: Infers per-pixel depth, ray maps, and camera intrinsics/extrinsics from unposed RGB images

Feature 04

Cross-Domain Generalization: Robust performance across indoor, outdoor, object-centric, and driving environments

Feature 05

You might also want to compare

Verified Sources

Tags

3d-reconstructioncamera-pose-estimationlooping-transformernvidiadvl-lab

Model Specs

open-weights

Parameters

~117M

Context Window

undisclosed

License

Other/Custom

Deployment

self-hostable

Resources & Links

Curator Notes

Verified paper arXiv:2605.30215 and open-source project from NVIDIA DVL group.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model