MMeshFlow
MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer
MeshFlow is a state-of-the-art 3D generative model introduced as a CVPR 2026 Highlight by researchers from Meta AI and HKUST (Weiyu Li, Antoine Toisoul, Tom Monnier, Roman Shapovalov, Rakesh Ranjan, Ping Tan, and Andrea Vedaldi).
MeshFlow addresses the fundamental limitations of prior Auto-Regressive (AR) 3D mesh generators—such as slow token-by-token sequence generation and quantization artifacts from discretizing vertex coordinates—by establishing a continuous latent space representation and leveraging parallel flow matching diffusion transformers.
🔬 Architecture & Methodology
MeshFlow operates via a two-stage architecture:
- MeshVAE (Variational Autoencoder):
- Learns a compact continuous latent space for complex 3D triangular meshes.
- Encodes continuous 3D vertex positions, surface normals, and face connectivity without requiring coordinate discretization or quantization.
- Optimized with graph contrastive losses and topological regularization to ensure clean manifold mesh reconstruction.
- Flow-Based Diffusion Transformer:
- Built on Rectified Flow formulation over continuous MeshVAE latents.
- Replaces slow sequential AR token prediction with parallel generation of all mesh vertices and faces simultaneously.
- Conditioned on text prompts, reference images, or sparse point clouds via cross-attention mechanisms.
Input (Text / Image / Point Cloud) ──► Latent Condition Encoder ──► Flow-based DiT ──► Continuous Mesh Latent ──► MeshVAE Decoder ──► 3D Triangular Mesh (.obj / .ply)⚡ Key Features & Benchmarks
- 18× – 26× Faster Inference: Generates artist-grade 3D meshes in ~1 second per asset on modern GPUs, drastically outperforming autoregressive baseline models like MeshGPT.
- Zero Coordinate Quantization: Operates in continuous space, eliminating vertex snapping and staircase artifacts common in discrete coordinate models.
- Explicit Mesh Output: Outputs native 3D triangular meshes with clean connectivity and explicit vertices, ready for game engines (Unreal/Unity) and VFX suites (Blender/Maya).
- Flexible Conditioning: Supports text-to-3D, image-to-3D, and point-cloud-to-3D generation pipelines.
Benchmark Performance
| Metric / Benchmark | MeshFlow (Flow-DiT) | Autoregressive Baselines (e.g. MeshGPT) |
|---|---|---|
| Inference Time per Mesh | ~1.0 s | ~18 - 26 s |
| Generation Mechanism | Parallel Continuous Flow | Sequential Discrete Tokens |
| Quantization Artifacts | None (Continuous Space) | Present (Grid Discretization) |
| Chamfer Distance (Objaverse) | State-of-the-Art | Baseline |
🚀 Quickstart & Usage
# Clone official MeshFlow repository
git clone https://github.com/facebookresearch/meshflow.git
cd meshflow
# Create Conda environment & install dependencies
conda create -n meshflow python=3.10 -y
conda activate meshflow
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
# Inference: Text-to-3D Mesh Generation
python sample.py \
--config configs/meshflow_t23d.yaml \
--prompt "An artistic wooden chair with ornate carving" \
--output_dir ./outputs/🔗 Official Links & Resources
Key Features
MeshVAE Continuous Latent Space: Encodes 3D vertex positions, surface normals, and face topology into a compact continuous latent representation, avoiding coordinate quantization errors.
Flow-Based Diffusion Transformer: Employs Rectified Flow matching over continuous mesh latents to generate complete 3D meshes in parallel.
18x–26x Faster Generation: Drastically speeds up inference (~1s per asset) compared to traditional autoregressive mesh generation models like MeshGPT.
Explicit Artist-Grade Geometry: Outputs explicit 3D triangular meshes with clean connectivity suitable for downstream VFX, gaming, and 3D pipelines.
Multi-Modal Conditioning: Supports text-to-3D, image-to-3D, and point-cloud-conditioned mesh synthesis.
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
Undisclosed
Context Window
undisclosed
License
CC-BY-NC-4.0
Deployment
Resources & Links
Curator Notes
Verified CVPR 2026 Highlight paper by Meta AI and HKUST. Code available on GitHub (facebookresearch/meshflow) and paper on arXiv (arXiv:2606.04621).
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model