Back to Academic Research
Research PreviewMultimodaltextvideoimageUpdated June 1, 2026

GenCeption

Model Overview

GenCeption is a undisclosed-parameter model developed by Academic Research. Released on 2026-06-01.


📊 Quick Specs

Specification Table
SpecificationValue
Parametersundisclosed
Taskmultimodal-general
Modalitytext, video, image
LicenseOther/Custom
Typeresearch-preview

✨ Key Features

  • Unified Multi-Task Vision Model: Single feed-forward model handles depth, normals, pose, segmentation, keypoints, and 4D grounding from generative video pretraining
  • Text-Steered Inference: Tasks are specified via natural language text instructions without changing model weights
  • 4D Grounding: Supports spatiotemporal grounding in video (4D = 3D space + time)
  • No Task-Specific Training: Demonstrates that video generative pretraining serves as a general-purpose visual representation
  • SOTA Performance: Outperforms task-specific specialist models on multiple visual perception benchmarks

🔗 Resources


📜 License & Access

Other/Custom — See repository for specific license details.

Key Features

Unified Multi-Task Vision Model: Single feed-forward model handles depth, normals, pose, segmentation, keypoints, and 4D grounding from generative video pretraining

Feature 01

Text-Steered Inference: Tasks are specified via natural language text instructions without changing model weights

Feature 02

4D Grounding: Supports spatiotemporal grounding in video (4D = 3D space + time)

Feature 03

No Task-Specific Training: Demonstrates that video generative pretraining serves as a general-purpose visual representation

Feature 04

SOTA Performance: Outperforms task-specific specialist models on multiple visual perception benchmarks

Feature 05

You might also want to compare

Verified Sources

Tags

computer-visionvideo-diffusionmulti-taskdepth-estimationsegmentation

Model Specs

research-preview

Parameters

Undisclosed

Context Window

unknown

License

Other/Custom

Deployment

self-hostable

Resources & Links

Curator Notes

Partially enriched via migration on 2026-07-25. Manual review recommended.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model