Back to Academic/Research
Research PreviewVideo GentextvisionvideoUpdated May 19, 2026

CogOmniControl

Model Overview

CogOmniControl is a undisclosed-parameter model developed by Academic/Research. Released on 2026-05-19.


📊 Quick Specs

Specification Table
SpecificationValue
Parametersundisclosed
Taskvideo-generation
Modalitytext, vision, video
LicenseOther/Custom
Typeresearch-preview

✨ Key Features

  • Reasoning-Driven Controllable Video Generation: Decouples creative intent understanding from video synthesis
  • CogVLM Intent Cognition: Specialized VLM trained on authentic anime production data to turn sparse sketches into dense reasoning
  • CogOmniDiT Generation: Unified Diffusion Transformer supporting multi-condition in-context video synthesis
  • Reinforcement Learning Alignment: Aligns DiT generation paths with VLM reasoning outputs using RL
  • Closed-Loop Best-of-N Harness: Uses planned evaluators to select optimal generated video candidates
  • Professional Production Workflow Focus: Designed for storyboard sketches, clay renders, and complex multi-modal controls

🔗 Resources


📜 License & Access

Other/Custom — See repository for specific license details.

Key Features

Reasoning-Driven Controllable Video Generation: Decouples creative intent understanding from video synthesis

Feature 01

CogVLM Intent Cognition: Specialized VLM trained on authentic anime production data to turn sparse sketches into dense reasoning

Feature 02

CogOmniDiT Generation: Unified Diffusion Transformer supporting multi-condition in-context video synthesis

Feature 03

Reinforcement Learning Alignment: Aligns DiT generation paths with VLM reasoning outputs using RL

Feature 04

Closed-Loop Best-of-N Harness: Uses planned evaluators to select optimal generated video candidates

Feature 05

Professional Production Workflow Focus: Designed for storyboard sketches, clay renders, and complex multi-modal controls

Feature 06

You might also want to compare

Verified Sources

Tags

research-previewvideo-generationcontrollable-videovision-language-modeldiffusion-transformerreinforcement-learning

Model Specs

research-preview

Parameters

Undisclosed

Context Window

undisclosed

License

Other/Custom

Deployment

self-hostable

Resources & Links

Curator Notes

Partially enriched via migration on 2026-07-25. Manual review recommended.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model