Release Archive
A flat, chronological list of every model tracked in the Modelverse registry.
July 2026
Motion4Motion
Motion4Motion is a training-free framework for cross-subject motion transfer presented at SIGGRAPH 2026. Unlike prior skeleton-based approaches that only work for human-like characters with predefined skeleton topologies, Motion4Motion steps out of the skeleton framework entirely. It models the motion flow of a subject in a video rather than its skeleton, enabling motion transfer across entirely different species (e.g., human-to-animal) at inference time without any task-specific training or fine-tuning.
Kimi K3
Kimi K3 is Kimi's flagship 2.8T open-source model built on Kimi Delta Attention (KDA) and Attention Residuals. It features native visual understanding for images and video, a 1M-token context window with automatic caching, and an always-on thinking mode designed for frontier intelligence scenarios like long-horizon coding and knowledge work.
Claude Opus 5
For complex agentic coding and enterprise work. Claude Opus 5 is a step-change improvement over Claude Opus 4.8, with the largest gains in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.
KAT-Coder-V2.5-Dev
KAT-Coder-V2.5-Dev is a 34.7B-parameter Mixture-of-Experts (MoE) text generation model developed by Kwaipilot. Built on the Qwen3_5MoeForConditionalGeneration architecture using transformers. Supports en, zh language(s). Released on 2026-07-23 with 126 likes and 396 downloads on Hugging Face.
ARDY
ARDY is an autoregressive diffusion model for interactive human motion generation with online text prompting and flexible kinematic constraints.
MobileWan
MobileWan is a mobile video diffusion model developed by Qualcomm AI Research that aims to close the quality gap for on-device video generation.
Solar-Open2-250B
Solar-Open2-250B is a 250.3B-parameter Mixture-of-Experts (MoE) text generation model developed by upstage. Built on the SolarOpen2ForCausalLM architecture using transformers. Supports en, ko, ja language(s). Released on 2026-07-22 with 543 likes and 1,106 downloads on Hugging Face.
antares-1b
antares-1b is a 1.8B-parameter text generation model developed by fdtn-ai. Built on the GraniteMoeHybridForCausalLM architecture using transformers. Supports en language(s). Released on 2026-07-21 with 150 likes and 4,266 downloads on Hugging Face.
Gemini 3.5 Flash Cyber
Google DeepMind's specialized cybersecurity model, built to help defenders find, validate, and patch software vulnerabilities quickly and efficiently. Gemini 3.5 Flash Cyber is fine-tuned from the Gemini 3.5 Flash base model with deep expertise in vulnerability research, security analysis, exploit validation, and patch generation. It enables security teams to dramatically accelerate their defensive workflows compared to general-purpose LLMs.
Gemini 3.5 Flash-Lite
Google DeepMind's most cost-efficient model in the Gemini 3.5 family, designed for high-volume, low-latency applications where economy is paramount. Gemini 3.5 Flash-Lite delivers strong performance for everyday tasks at the lowest price point of the new July 2026 Gemini release batch, while retaining multimodal input capabilities and the 1M token context window.
Gemini 3.6 Flash
Google DeepMind's fastest frontier-class model, featuring built-in reasoning, native multimodal input (text, image, audio, video), and a 1M token context window. Ranked #1 in speed among all 186 models on Artificial Analysis at 275.5 tokens/second while scoring 50 on the Artificial Analysis Intelligence Index — placing it well above average among comparable models. Priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens with a 90% cache discount.
Mage-Flow
Mage-Flow is a 4.1B-parameter text to image model developed by microsoft. Released on 2026-07-21 with 242 likes and 891 downloads on Hugging Face.
Nanbeige4.2-3B
Nanbeige4.2-3B is a 4.2B-parameter text generation model developed by Nanbeige. Built on the NanbeigeForCausalLM architecture using transformers. Supports en, zh language(s). Released on 2026-07-21 with 374 likes and 8,169 downloads on Hugging Face.
Laguna-S-2.1-GGUF
Laguna-S-2.1-GGUF is a undisclosed-parameter text generation model developed by unsloth. Released on 2026-07-21 with 172 likes and 57,536 downloads on Hugging Face.
Motif-3-Beta
Motif-3-Beta is a 314.8B-parameter Mixture-of-Experts (MoE) text generation model developed by Motif-Technologies. Built on the MotifForCausalLM architecture using transformers. Supports en, ko language(s). Released on 2026-07-20 with 186 likes and 2,108 downloads on Hugging Face.
MiniCPM-RobotManip
MiniCPM-RobotManip is a 1.5B-parameter robotics model developed by openbmb. Built on the MiniCPMV_VLA architecture using transformers. Released on 2026-07-18 with 173 likes and 559 downloads on Hugging Face.
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF is a undisclosed-parameter image text to text model developed by DavidAU. Supports en, zh language(s). Released on 2026-07-17 with 496 likes and 407,421 downloads on Hugging Face.
Wan Streamer v0.3
Wan Streamer v0.3 learns video as a persistent world plus the events that unfold inside it, bringing free-form behavior to real-time audio-visual interaction.
Actionable World
WorldString (Actionable World) is a neural architecture developed by Tsinghua University, UCSD, CalTech, and NVIDIA for modeling physical object state manifolds (articulated, skinned, and soft objects) directly from point clouds or RGB-D video streams.
Agents A1
InternScience's open-weights 35B Mixture-of-Experts agentic model, designed for long-horizon planning, multi-teacher domain distillation, and complex tool execution.
Luna
Developed by researchers from HKUST, Tsinghua University, and Meta, Luna (Learning Universal 3D Human Animation Beyond Skinning) is an LBS-free neural animation model presented at ECCV 2026. It maps 2D controls directly into 3D Gaussian deformations for realistic human avatars.
MAMMA
MAMMA (Markerless Accurate Multi-person Motion Acquisition) is a state-of-the-art markerless multi-view human motion capture framework developed by MPI-IS and CMU (CVPR 2026 Oral). Powered by MammaNet (a ViT-Base transformer predicting 512 contact-aware and visibility-aware surface landmarks), it accurately fits SMPL-X body models to complex multi-person physical interactions.
MiniCPM5 1B
MiniCPM5 1B is a high-performance, dense 1.08B parameter language model developed by OpenBMB. Designed for on-device deployment and agentic workflows, it features dual-mode reasoning ('Think' CoT vs. 'NoThink' fast mode), a native 128K context window, native tool calling, and standard Llama architecture compatibility.
PaGeR
PaGeR (Panoramic Geometry Reconstruction) is a unified framework from ETH Zurich (Photogrammetry and Remote Sensing Lab) that adapts perspective 3D foundation models (Depth Anything 3) to single-pass panoramic geometry estimation, predicting scale-invariant depth, metric depth, surface normals, and sky masks.
RDM
Representation Distribution Matching (RDM / iRDM) is a paradigm for training state-of-the-art one-step visual generation models developed by EPFL (VITA Lab), Valeo.ai, and Sorbonne Université. It matches feature distributions between generated and reference images using Maximum Mean Discrepancy (MMD) with Nyström estimation across a battery of frozen pretrained encoders.
VideoMDM
VideoMDM is a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without requiring any 3D ground truth.
Qwen 3.7
Developed by Alibaba, Qwen 3.7 is a research preview model exploring new techniques in specialized research.
Magenta Realtime
An experimental agentic workflow research preview named Magenta Realtime, published by Google DeepMind to showcase novel methods.
WavFlow
Developed by Meta, WavFlow is a research preview model exploring new techniques in specialized research.
MAI Thinking
An experimental specialized research research preview named MAI Thinking, published by Microsoft to showcase novel methods.
GPT Dreaming
A research-preview model by OpenAI focusing on specialized research capabilities to invite community feedback.
GPT-Red
An automated AI red-teaming system developed by OpenAI, trained via self-play reinforcement learning to simulate prompt injection attacks and discover vulnerabilities in language models prior to deployment.
Step 3.7 Flash
An experimental specialized research research preview named Step 3.7 Flash, published by StepFun to showcase novel methods.
Stable Audio 3
Developed by Stability AI, Stable Audio 3 is a research preview model exploring new techniques in audio and speech.
Stable Layers
A research-preview model by Stability AI focusing on specialized research capabilities to invite community feedback.
HY-MT2
A research-preview model by Tencent focusing on specialized research capabilities to invite community feedback.
Inkling
Thinking Machines' flagship open-weights Mixture-of-Experts multimodal model, supporting Native MoE transformer architecture and up to 1M token context window.
GLM 5.2
An experimental specialized research research preview named GLM 5.2, published by Zhipu AI to showcase novel methods.
MoVerse
MoVerse is a undisclosed-parameter text to video model developed by Academic/Research. Released on 2026-07-14 with 0 likes and 0 downloads on Hugging Face.
Wan-Dancer
Wan-Dancer is a hierarchical framework for minute-scale coherent music-to-dance video generation developed by Tongyi Lab at Alibaba Group. Given a reference character image and music audio, Wan-Dancer generates long-duration (minute-scale), high-quality, rhythmically synchronized dance videos with global structural coherence and temporal continuity across multiple dance genres including Chinese Classical, K-pop, Street, Tap, and Latin styles.
codex-mini-latest
Fast reasoning model optimized for the Codex CLI
GPT-4.1
Smartest non-reasoning model
GPT-4.5 Preview
Deprecated large model.
GPT-4o Audio
GPT-4o models capable of audio inputs and outputs
GPT-4o mini Audio
Smaller model capable of audio inputs and outputs
GPT-4o mini Realtime
Smaller realtime model for text and audio inputs and outputs
GPT-4o mini Search Preview
Fast, affordable small model for web search
GPT-4o mini Transcribe
Speech-to-text model powered by GPT-4o mini
GPT-4o mini TTS
Text-to-speech model powered by GPT-4o mini
GPT-4o Realtime
Model capable of realtime text and audio inputs and outputs
GPT-4o Search Preview
GPT model for web search in Chat Completions
GPT-4o Transcribe
Speech-to-text model powered by GPT-4o
GPT-5.1 Chat
GPT-5.1 model used in ChatGPT
GPT-5.1-Codex
A version of GPT-5.1 optimized for agentic coding in Codex.
GPT-5.1
The best model for coding and agentic tasks with configurable reasoning effort
GPT-5.2 Chat
GPT-5.2 model used in ChatGPT
GPT-5.2-Codex
Our most intelligent coding model optimized for long-horizon, agentic coding tasks.
GPT-5.2
Previous frontier model for professional work with configurable reasoning effort
GPT-5.3 Chat
GPT-5.3 Instant model used in ChatGPT
GPT-5.3-Codex
The most capable agentic coding model to date.
GPT-5.4
A more affordable model for coding and professional work.
GPT-5.5
A new class of intelligence for coding and professional work.
GPT-5 Chat
GPT-5 model used in ChatGPT
GPT-5-Codex
A version of GPT-5 optimized for agentic coding in Codex
GPT-5
Previous intelligent reasoning model for coding and agentic tasks with configurable reasoning effort
gpt-audio-1.5
The best voice model for audio in, audio out with Chat Completions.
gpt-audio
For audio inputs and outputs with Chat Completions API
GPT Image 1.5
Our previous image generation model
GPT Image 1
Our previous image generation model
GPT Image 2
State-of-the-art image generation model
gpt-oss-120b
Most powerful open-weight model, fits into an H100 GPU
gpt-oss-20b
Medium-sized open-weight model for low latency
GPT-Realtime-1.5
The best voice model for audio in, audio out
GPT-Realtime-2.1
Reasoning model with tool use
GPT-Realtime-2
Reasoning model with tool use
GPT-Realtime-Translate
Streaming speech-to-speech translation model
GPT-Realtime-Whisper
Streaming speech-to-text model for realtime transcription
GPT-Realtime
Model capable of realtime text and audio inputs and outputs
o4-mini-deep-research
Faster, more affordable deep research model
o4-mini
Fast, cost-efficient reasoning model, succeeded by GPT-5 mini
omni-moderation
Identify potentially harmful content in text and images
Bonsai 27B
Bonsai 27B is a 27B-class multimodal model based on Qwen3.6 27B, optimized for on-device and local agentic workflows. It is available in two highly compressed variants: a 5.9 GB Ternary variant (1.71 bits/weight) and a 3.9 GB 1-bit variant (1.125 bits/weight), allowing it to fit within the memory budget of everyday laptops and phones (like the iPhone 17 Pro).
Sora 2
Flagship video generation with synced audio
text-moderation-stable
Previous generation text-only moderation model
text-moderation
Previous generation text-only moderation model
Inkling
Inkling is a 952.4B-parameter Mixture-of-Experts (MoE) image text to text model developed by thinkingmachines. Built on the InklingForConditionalGeneration architecture using transformers. Released on 2026-07-14 with 1,547 likes and 27,883 downloads on Hugging Face.
Laguna-S-2.1
Laguna-S-2.1 is a 117.6B-parameter Mixture-of-Experts (MoE) text generation model developed by poolside. Built on the LagunaForCausalLM architecture using transformers. Released on 2026-07-13 with 618 likes and 28,992 downloads on Hugging Face.
ABot-World
ABot-World is a real-time interactive world simulator developed by Alibaba AMAP CV Lab. Built around a 5-billion parameter causal video world model (ABot-World-0-5B-LF) fine-tuned from Wan2.2-TI2V-5B, it enables open-ended, action-conditioned video generation running in real time at 720p resolution @ 16 FPS on a single NVIDIA RTX 5090 GPU.
MuScriptor
MuScriptor is an open-weights multi-instrument audio-to-MIDI model developed by Mirelo AI and Kyutai.
LingBot-World 2.0
An interactive world model that supports continuous, hour-long real-time generation. It features a native dual-agent mechanism that allows for dynamically evolving environments responding to real-time user inputs without scene collapse.
Seedream 5 Pro
Seedream 5.0 Pro is ByteDance's flagship multimodal image generation and interactive precision editing model. Designed for professional creative workflows, it features native layer separation, interactive grounded editing, high-density multilingual typography, and multi-reference consistency up to 4K resolution.
krea2-identity-edit
krea2-identity-edit is a undisclosed-parameter text generation model developed by conradlocke. Released on 2026-07-07 with 534 likes and 0 downloads on Hugging Face.
Hy3
Tencent's flagship open-weights Mixture-of-Experts reasoning model, utilizing hybrid thinking and active chain-of-thought configuration.
Bonsai-27B-gguf
Bonsai-27B-gguf is a undisclosed-parameter text generation model developed by prism-ml. Released on 2026-07-04 with 632 likes and 2,028,115 downloads on Hugging Face.
Ternary-Bonsai-27B-gguf
Ternary-Bonsai-27B-gguf is a undisclosed-parameter text generation model developed by prism-ml. Released on 2026-07-04 with 1,009 likes and 595,415 downloads on Hugging Face.
Laguna-S-2.1-NVFP4
Laguna-S-2.1-NVFP4 is a 67.9B-parameter Mixture-of-Experts (MoE) text generation model developed by poolside. Built on the LagunaForCausalLM architecture using vllm. Released on 2026-07-02 with 130 likes and 89,186 downloads on Hugging Face.
Muse Image
Meta MUSE Image is a proprietary agentic image generation model developed by Meta Superintelligence Labs (MSL). It structures its generation process using inference-time reasoning and active background web searches to gather accurate real-world visual references prior to rendering, powering generative features across meta.ai, Instagram, and WhatsApp.
Muse Video
Meta's previewed video generation model showcasing complex physical comprehension and natively integrated audio.
Aspire
ASPIRE (Agentic Skill Programming through Iterative Robot Exploration) is a continual learning framework for robotics developed by NVIDIA GEAR Lab in collaboration with UMich, UIUC, UC Berkeley, and CMU. Using a code-as-policy approach, ASPIRE enables robot agents to autonomously generate, test, debug, and refine executable Python control programs.
CHORD
CHORD (Contact Wrench Guidance from Human Demonstration) is a framework by NVIDIA Isaac Team and GEAR Lab for learning long-horizon dexterous and whole-body robotic manipulation from human demonstrations using object-centric contact wrench space guidance in reinforcement learning.
GPT 5.6
OpenAI's latest flagship model engineered for prolonged, multi-step agentic coding tasks requiring minimal human oversight or handholding.
GPT Live
OpenAI's latest real-time voice interface designed to enable natural, flowing conversation. It supports dynamic interruptions, background task delegation to stronger frontier models, and can generate contextual visual responses in the UI.
Mira
A real-time multiplayer simulation model that acts entirely as a video generator responding to live user key presses, bypassing pre-designed game engines to render physics and collisions on the fly.
PixWorld
A 3D scene generator that reconstructs fully consistent 3D environments directly in pixel space from one or multiple reference images, avoiding the visual artifacts often caused by latent-space generation.
ProxyPose
A novel 6-DoF pose tracking system that reframes 3D position and rotation tracking as a video-to-video translation problem. It operates entirely at the pixel level without requiring 3D models or depth sensors.
Reve 2.1
A state-of-the-art image model known for extreme high-resolution outputs, precise text rendering, and robust targeted micro-editing via bounding box selections.
Wan Streamer 0.2
A real-time character simulation framework allowing users to converse with generated avatars—including humans, animals, or fictional characters—with highly responsive audio-visual sync.
High3
A large open-weights Mixture of Experts (MoE) model focused on reasoning, math, and agentic coding that punches above its weight class against much larger trillion-parameter models.
Grok 4.5
A highly efficient frontier model built for coding, engineering, and math. While its context window is smaller than immediate competitors, it prioritizes fast token generation and cost efficiency over topping raw intelligence leaderboards.
June 2026
Claude Sonnet 5
Built specifically as an execution layer for multi-step software engineering work. It handles sustained coding, browser tool use, and debugging autonomously at speeds matching previous Opus models.
VidiHand
VidiHand is a generative 4D hand motion reconstruction model developed by Nanyang Technological University (NTU) and Shanghai Jiao Tong University (SJTU). It recovers metric-scale 3D/4D two-hand pose trajectories from monocular egocentric video by fine-tuning internet-scale video diffusion models (Wan2.1-VACE) without needing external hand detectors or test-time optimization.
PhysiFormer
PhysiFormer is a diffusion transformer model developed by Visual Geometry Group (VGG), University of Oxford to simulate physically plausible 3D object dynamics directly in 3D world coordinates. Conditioned on initial 3D vertex positions, velocities, and material type descriptors, it treats full-horizon 3D trajectory prediction as a single denoising process.
Ornith 1.0
DeepReinforce's open-weights family of self-scaffolding models optimized for agentic coding, post-trained on Gemma 4 and Qwen 3.5.
OmniContact
OmniContact is a hierarchical framework for generalizable humanoid loco-manipulation developed by Noitom Robotics and HKUST. Centered on 'Contact Flow' (CF) representations, it combines high-level reference synthesis (CF-Gen) with 50Hz closed-loop tracking (CF-Track) to enable long-horizon meta-skill chaining, autonomous failure recovery, and real-time execution across tasks like carrying, pushing, sliding, and relocating.
Unlimited-OCR
Unlimited-OCR is a 3.3B-parameter Mixture-of-Experts (MoE) image text to text model developed by baidu. Built on the UnlimitedOCRForCausalLM architecture using transformers. Supports multilingual language(s). Released on 2026-06-19 with 3,029 likes and 2,500,391 downloads on Hugging Face.
Qwythos-9B-Claude-Mythos-5-1M-GGUF
Qwythos-9B-Claude-Mythos-5-1M-GGUF is a undisclosed-parameter image text to text model developed by empero-ai. Supports en language(s). Released on 2026-06-19 with 2,456 likes and 1,906,539 downloads on Hugging Face.
Flex4DHuman
Flex4DHuman is a generative framework from University of Washington, Zhejiang University, and Tencent for flexible multi-view video diffusion and 4D human reconstruction from monocular or sparse video streams without explicit geometry priors.
MuSViT
MuSViT (Music Score Vision Transformer) is the first foundation vision model specifically engineered for Optical Music Recognition (OMR) and sheet music representation. Developed by PRAIG at the University of Alicante (accepted at ECCV 2026), it uses a ViT-Base encoder pre-trained via Masked Autoencoders on 9.7 million IMSLP sheet music pages.
GLM-5.2
GLM-5.2 is a 753.3B-parameter Mixture-of-Experts (MoE) text generation model developed by zai-org. Built on the GlmMoeDsaForCausalLM architecture using transformers. Supports en, zh language(s). Released on 2026-06-16 with 4,422 likes and 667,403 downloads on Hugging Face.
Arbor
Generalist autonomous research agent framework developed by Renmin University of China (RUC-NLPIR) and Microsoft Research, utilizing Hypothesis-Tree Refinement (HTR) for long-horizon scientific research and system optimization.
Dots TTS
Dots TTS (dots.tts) is a 2-billion parameter, fully continuous, end-to-end autoregressive text-to-speech foundation model developed by RedNote HiLab (Xiaohongshu). It operates over continuous latent space via a 48 kHz AudioVAE, offering state-of-the-art zero-shot voice cloning across 24 languages.
i1
i1 is a 3-billion-parameter text-to-image diffusion model developed by ZLab at Princeton University. Introduced in 'i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models', it provides a fully open foundation—including model weights, PyTorch and JAX training/inference code, and curated caption datasets—built upon a LightningDiT cross-attention backbone with T5-Gemma-2B and FLUX.2 VAE.
LiveEdit
LiveEdit is a real-time, diffusion-based streaming video editing framework developed by Tsinghua University and HKUST (accepted at ECCV 2026). Built on top of Wan2.1, it achieves ~12.66 FPS causal streaming editing using a 3-stage distillation pipeline and an AR-oriented mask cache.
World Tracing
World Tracing is a generative 3D geometry representation framework developed by World Labs and UIUC (Hao Zhang, Mohamed El Banani, et al.). Given a single 2D image or short video clip, World Tracing predicts an ordered stack of 3D points along camera rays for every pixel, capturing both visible surfaces and occluded geometry behind them.
Kimi K2.7 Code
Moonshot AI's open-weights Mixture-of-Experts model optimized for software engineering, reducing thinking token usage and improving code generation efficiency.
Surflo
Surflo (Consistent 3D Surface Flow Model with Global State) is a feed-forward 3D surface reconstruction model developed by École Polytechnique, Kyoto University, Kyutai, and UC Berkeley. It encodes unposed RGB images into a 128-token global latent state and decodes 3D surface points using continuous flow matching with photometric rendering guidance.
DiffusionGemma
An experimental open-weights text diffusion model by Google DeepMind. Moving beyond traditional sequential autoregressive decoding, DiffusionGemma utilizes parallel block decoding to generate 256 tokens simultaneously, achieving up to 4x faster generation speeds on modern GPU hardware.
SCAIL 2
SCAIL 2 is an open-source end-to-end framework for controlled character animation and video-to-video motion transfer developed by Tsinghua University (KEG Group) and Z.ai. It performs motion transfer directly from driving video to reference character without relying on intermediate representations like 3D skeletons, OpenPose maps, or depth maps.
Claude Fable 5
Anthropic's most capable publicly available frontier model as of mid-2026. It is engineered for long-running, multi-agent autonomous workflows and sustained logical reasoning.
Claude Mythos 5
A highly restricted enterprise model available exclusively through Project Glasswing. It shares the core architecture of Fable 5 but contains advanced capabilities specifically optimized for life sciences, biology research, and high-security cyber contexts.
North Mini Code
North Mini Code is a specialized 30B total parameter, sparse Mixture-of-Experts (MoE) coding model with ~3B active parameters per token developed by Cohere and Cohere Labs. Designed for agentic software engineering and developer workflows, it achieves near-30B-scale reasoning with the compute overhead of a 3B parameter model.
AnchorWorld
AnchorWorld is an embodied egocentric world simulation framework that enables interactive, 3D motion-driven video generation with view-based evolution customization via spatial anchor views.
MeshFlow
MeshFlow is a flow-based diffusion transformer model with MeshVAE for efficient, high-quality artistic 3D mesh generation presented as a CVPR 2026 Highlight by Meta AI and HKUST.
Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is a flagship open-weight frontier LLM family engineered for agentic AI workloads, long-context understanding, complex reasoning, and tool calling. The family includes Nemotron 3 Ultra 550B (a hybrid Transformer-Mamba LatentMoE model with 1M context) and Llama-3.1-Nemotron-Ultra-253B-v1.
GenCeption
GenCeption repurposes a text-to-video generative diffusion model into a unified, feed-forward general-purpose vision learner steered by text instructions. Rather than training task-specific vision models, GenCeption demonstrates that a single model built on video generative pretraining can perform depth estimation, surface normal prediction, pose estimation, semantic segmentation, keypoint detection, and 4D grounding — all from the same weights — achieving SOTA performance across multiple tasks without task-specific fine-tuning.
StreamForce
StreamForce (Streaming Video Generation with Streaming Force Control) is a causal, unified framework for real-time streaming video generation developed by Northeastern University (NEU-VI Lab). It enables physically grounded interactive control over video rollouts via continuous, time-varying force inputs at up to 16.6 FPS.
WavTTS
WavTTS is an end-to-end zero-shot text-to-speech framework developed by Shanghai Jiao Tong University (SJTU) and ByteDance Seed. It generates speech directly in raw waveform space using flow matching and Diffusion Transformers (DiT), bypassing intermediate mel-spectrogram or neural codec token representations.
Qwen 3.7 Plus
Qwen 3.7 Plus is a multimodal interactive hybrid agent model developed by Alibaba Group. Designed to unify vision and language capabilities, it serves as a versatile foundation for agentic workflows, integrating Graphical User Interface (GUI) screen interaction with Command Line Interface (CLI) code execution and tool use.
MiniMax-M3
MiniMax's open-weights Mixture-of-Experts multimodal model, using MiniMax Sparse Attention (MSA) to enable low-latency long-context reasoning up to 1 million tokens.
OmniDreams
OmniDreams (NVIDIA Cosmos-Dreams) is an action-conditioned real-time generative world model developed by NVIDIA Spatial Intelligence Lab (SIL). It autoregressively synthesizes multi-camera photorealistic video observations conditioned on driving actions and simulator states for closed-loop autonomous vehicle simulation.
SeFi-Image
A text-to-image foundation model built on Semantic-First Diffusion, a paradigm that separates semantic layout streams from texture detail streams to improve generation quality while drastically reducing training compute requirements.
May 2026
Cosmos 3
NVIDIA Cosmos 3 is an open-source frontier foundation model platform engineered for Physical AI. Built on a Mixture-of-Transformers (MoT) architecture with a two-tower design, Cosmos 3 unifies an autoregressive Reasoner with a diffusion Generator to allow embodied agents to perceive, reason about, simulate, and act in the physical world.
Claude Opus 4.8
A flagship update to the Opus 4 lineage that served as the primary frontier model prior to the release of Fable 5. Highly effective at complex reasoning and deep technical analysis.
CubePart
CubePart is a undisclosed-parameter text to 3d model developed by Roblox Research. Released on 2026-05-27 with 19 likes and 0 downloads on Hugging Face.
PanoWorld
PanoWorld is a generative spatial world model developed by Ke Holdings Inc. (Beike) for consistent whole-house 360° panorama synthesis from 2D floorplans and style references using a floorplan 3D shell proxy and dynamic 3D Gaussian Splatting (3DGS) spatial memory.
Pantheon 360
Developed by USC, NYCU, Cornell, and Bosch Research (CVPR 2026), Pantheon 360 is a 3D-aware 360° video diffusion framework that utilizes an explicit 3D Cache to synthesize geometrically consistent panoramic videos and digital twins.
StreamChar
StreamChar is a real-time, long-horizon streaming framework for generating synchronized audio and video of talking characters from text transcripts developed by Alibaba Tongyi Lab (HumanAIGC Team). It decouples long-horizon orchestration from short-window audio-video denoising using a Joint Audio-Video Diffusion Transformer (DiT).
Bernini
Bernini is a unified framework for video generation and editing developed by ByteDance Research. It decouples high-level semantic planning (Qwen2.5-VL-7B MLLM) from pixel rendering (Wan2.2-A14B DiT), delivering state-of-the-art instruction-following video generation and editing with minimal spatial-temporal drift.
Lance
Lance is a lightweight, 3B-active-parameter native unified multimodal model developed by ByteDance. It handles image and video understanding, generation, and editing within a single dual-stream Mixture-of-Experts (MoE) architecture trained via multi-task synergy.
GenRecon
GenRecon (Bridging Generative Priors for Multi-View 3D Scene Reconstruction) is a 3D vision framework developed by researchers at Technical University of Munich (TUM). It reformulates multi-view 3D scene reconstruction as conditional 3D generation over overlapping spatial chunks, leveraging Trellis.2 generative priors to produce complete, editable PBR-ready meshes from sparse RGB images.
Scope
SCOPE (Simulating Cross-game Operations in Playable Environments) is an interactive real-time FPS world model developed by University of Chinese Academy of Sciences (UCAS). It incorporates a Spatial Action Decoupling module into video diffusion transformer blocks (Wan2.2) to achieve per-pixel temporal action conditioning, separating localized weapon/HUD effects from global environment rendering.
PiD
PiD (Pixel Diffusion Decoder) is a generative model and decoding paradigm developed by NVIDIA Spatial Intelligence Lab (SIL). It reformulates latent-to-pixel decoding as a conditional pixel-space diffusion process, replacing traditional VAE decoders to unify decoding and 4x/8x spatial super-resolution into a fast 4-step distilled diffusion process.
FashionChameleon
FashionChameleon is a real-time and interactive human-garment video customization framework developed by Xiamen University, Zhejiang University, and Alibaba Group. Built on the Wan2.2-TI2V-5B backbone, it achieves generation speeds of 23.8 FPS on a single GPU using streaming distillation and training-free KV-cache rescheduling.
Flash GRPO
Flash-GRPO is an efficient one-step policy optimization alignment framework for video diffusion models developed by Zhejiang University. It solves standard GRPO computational bottlenecks via Iso-Temporal Grouping and Temporal Gradient Rectification, achieving 6x training acceleration.
PhysX Omni
PhysX Omni is a unified generative framework developed by S-Lab, NTU in collaboration with ACE Robotics and NVIDIA. It creates simulation-ready 3D assets (rigid, deformable, articulated) with complete physical attributes (scale, mass, material stiffness, joint kinematics) directly from images or text for physics simulators like MuJoCo and Isaac Sim.
CogOmniControl
CogOmniControl is a reasoning-driven framework for controllable video generation developed by University of Macau. It factorizes generation into creative intent cognition (CogVLM) and in-context video diffusion (CogOmniDiT), optimized via RL and Best-of-N candidate selection.
MegaASR
MegaASR is a 1.7B foundation model for robust 'in-the-wild' speech recognition, developed by Tsinghua University, NTU, NUS, and Shanghai AI Lab. Built on Qwen3-ASR, it utilizes progressive acoustic-to-semantic SFT (A2S-SFT) and dual-granularity policy optimization (DG-WGPO) trained on 2.6M simulated samples.
Gemini 3.5 Flash
Google's most intelligent model for sustained frontier performance on agentic and coding tasks, optimized for fast agent loops and multi-step workflows.
PixlRelight
PIXLRelight is a feed-forward, transformer-based neural rendering framework for physically controllable single-image relighting. Developed at the University of Oxford, it bridges physically based rendering (PBR) and learned image synthesis via shared intrinsic conditioning.
ControlLight
ControlLight is a controllable, consistent, and generalizable low-light image enhancement framework built as a LoRA on FLUX.2-klein-9B. It enables continuous illumination adjustment while preserving scene structure and fine-grained visual details.
NAVA
NAVA (Native Audio-Visual Alignment for Generation) is a 6.3-billion parameter multimodal diffusion model developed by Baidu ERNIE Research. Utilizing an Align-then-Fuse MMDiT architecture, NAVA generates temporally synchronized and semantically coherent 720p video and stereo audio natively in a single unified pass.
Qwen Live Translate
Qwen Live Translate (Qwen3.5-LiveTranslate-Flash) is Alibaba's real-time multimodal simultaneous speech-to-speech and speech-to-text translation model family built on the Qwen-Omni architecture. It fuses audio, text, and visual input (lip movements and facial context) for low-latency interpretation and voice cloning.
Deja View
Déjà View (DVLT - Déjà View Looping Transformer) is a parameter-efficient model for multi-view 3D reconstruction developed by NVIDIA Dynamic Vision and Learning (DVL) group, ETH Zürich, and U of Toronto. Employing a single weight-tied transformer block recurrently over K refinement steps, DVLT jointly predicts per-pixel depth, ray maps, and camera parameters at ~117M parameters.
Gamma World
Gamma-World (γ-World) is a generative multi-agent world model developed by NVIDIA Spatial Intelligence Lab (SIL) and Toronto AI Lab. Unlike single-agent world models, Gamma-World models complex multi-agent environments where multiple independently controlled agents interact in real-time at 24 FPS with zero-shot generalization from 2 to 4+ agents.
LocateAnything
LocateAnything is a high-speed, high-accuracy vision-language model developed by NVIDIA Learning and Perception Research (LPR) Lab for universal visual grounding and spatial localization. It introduces Parallel Box Decoding (PBD) to predict bounding box coordinates atomically in a single forward pass, eliminating sequential coordinate token decoding bottlenecks to achieve up to 10x higher throughput.
ReactiveGWM
ReactiveGWM (Reactive Game World Models) is an interactive game world modeling framework developed by National University of Singapore (NUS) in collaboration with Tencent, HKPolyU, and HKUST-GZ. It decouples player action control from NPC autonomy in video generation, enabling steerable, prompt-aligned NPC behavior and zero-shot strategy transfer across different games.
L2P
L2P is a undisclosed-parameter text generation model developed by Academic/Research. Released on 2026-05-03 with 86 likes and 0 downloads on Hugging Face.
Lucida
Lucida is a BiRefNet-based image matting model fine-tuned specifically for the cases where general-purpose background removers fail: semi-transparent objects, camouflaged subjects, logos and typography with soft shadows, glow/VFX effects, illustrations, and print-style designs (stickers, tees). Trained across 9 specialized categories (camouflage, transparent, complex, thin, hair, text, fx, illustration, design) on ~52,882 image/alpha pairs, Lucida v7 achieves the best overall MAE (0.0257) across all models measured — including commercial references — while maintaining MIT licensing.
April 2026
Kimi K2.6
Moonshot AI's open-weights Mixture-of-Experts model with enhanced long-horizon coding capabilities and multi-agent orchestration for complex workflows.
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is a undisclosed-parameter Mixture-of-Experts (MoE) image text to text model developed by HauhauCS. Supports en, zh, multilingual language(s). Released on 2026-04-17 with 3,073 likes and 2,057,103 downloads on Hugging Face.
Claude Opus 4.7
A major reasoning upgrade specifically improving advanced software engineering and file-system memory usage. It accepts high-resolution images up to ~3.75 megapixels for pixel-perfect screenshot reading.
Muse Spark 1.1
A multimodal reasoning model built for agentic tasks. It can autonomously use browsers, write and deploy code, and visually verify changes by taking and analyzing application screenshots.
Claude Mythos Preview
The first highly restricted preview model of the Mythos class. It was utilized to test advanced cybersecurity and life sciences capabilities internally and with trusted enterprise partners under Project Glasswing.
Gemma 4 (31B Dense)
Google's most capable open-weights model, featuring a unified multimodal architecture that processes images and audio natively.
March 2026
GNM (Generative aNthropometric Model)
GNM (Generative aNthropometric Model) is Google's open ecosystem of state-of-the-art parametric statistical human models and associated perception stacks. Pronounced like 'genome' (/ˈdʒiː.noʊm/), GNM strives to be the most accurate and complete 3D parametric human model. The initial open-source release starts with GNM Head — a high-fidelity statistical 3D model of the human head with fine-grained, disentangled control over identity, expressions, and head pose, including controllable internal anatomy (eyeballs, teeth, tongue). Multi-framework support: NumPy, JAX, PyTorch, and TensorFlow.
Gemini 3.1 Flash-Lite
An extremely cost-effective model designed for high-frequency everyday tasks and bulk API queries, rivaling larger models.
February 2026
Gemini 3.1 Pro
Google's top-tier model focused on extreme reasoning, vibe coding, and complex agent-based workflows.
Claude Sonnet 4.6
A frontier-level model optimized for professional work at scale, specifically focusing on coding workflows and agent tooling.
Claude Opus 4.6
This release introduced native capabilities for coordinating agent teams directly and featured heavy architectural optimizations for deep causal reasoning.
January 2026
Kimi K2.5
Moonshot AI's open-weights Mixture-of-Experts multimodal model, introducing Agent Swarm orchestration and visual agentic intelligence.
MedGemma 1.5 4B
MedGemma 1.5 4B is an open-weight multimodal vision-language model developed by Google DeepMind and Google Research as part of the Health AI Developer Foundations (HAI-DEF) program. Built on Gemma 3 and integrating MedSigLIP, it natively processes 2D/3D radiology scans and histopathology slides.
LFM 2.5-1.2B
Liquid AI's device-optimized hybrid model utilizing Liquid Neural Networks, designed for local edge AI agent deployment.
December 2025
Manus 1.6 Max
Butterfly Effect's flagship autonomous agent and action engine, featuring wide parallel research, visual chain-of-thought, and end-to-end full-stack web and mobile development.
Relightable Chars
Relightable Holoported Characters (Relightable Chars / RHC) is a 3D vision framework developed by Max Planck Institute for Informatics (MPI-INF), VIA Research Center, and Google. It synthesizes photorealistic free-viewpoint renders and enables arbitrary relighting of full-body dynamic human performances using sparse-view RGB videos and 3D Gaussian splatting.
November 2025
Claude Opus 4.5
An update focusing heavily on coding, workplace tasks, and producing spreadsheets. It introduced the "Infinite Chats" feature to handle context window limits.
Gemini 3 Deep Think
Gemini 3 Deep Think is an inference-time compute-scaling reasoning mode integrated within Google DeepMind's Gemini 3 model family. It generates internal step-by-step reasoning chains, evaluates parallel hypothesis paths, and performs self-verification to solve complex multi-step problems in mathematics, physics, materials science, and software engineering.
October 2025
Claude Haiku 4.5
The fastest and most cost-efficient generation 4 model, designed for high-volume use, automation, and lightweight corporate tasks.
Veo 3.1
Veo 3.1 is Google DeepMind's flagship high-definition generative video model released in October 2025. It generates cinematic-quality, high-resolution videos (up to 4K) from text and image prompts with native synchronized multi-track audio (dialogue, ambient sound, and sound effects).
September 2025
Claude Sonnet 4.5
Designed for balanced performance and everyday professional workflows.
VaultGemma 1B
A 1B parameter open-weights model developed by Google DeepMind, trained from scratch using differential privacy (DP). VaultGemma is engineered specifically to prevent the memorization and leakage of sensitive training data, offering a mathematically proven approach to data privacy for secure local deployments.
August 2025
July 2025
May 2025
Claude Opus 4
Launched alongside Sonnet 4, providing the frontier reasoning and technical analysis backbone for early generation 4 deployments.
Claude Sonnet 4
The first of the generation 4 models targeting fast responses and general-purpose tasks with reliable, everyday reasoning.
AlphaEvolve
AlphaEvolve is an evolutionary AI coding agent developed by Google DeepMind designed for autonomous algorithm discovery and computational optimization. Operating as an iterative discovery engine, AlphaEvolve mutates candidate code, evaluates performance via client-side evaluators, and refines algorithms using Gemini Flash and Gemini Pro.
March 2025
LiTo
LiTo (Surface Light Field Tokenization) is a 3D latent representation developed by Apple Research that jointly captures object geometry and view-dependent appearance. Built upon this unified representation, a latent flow matching model enables high-quality image-to-3D generation from a single input image, bridging neural rendering and generative modeling.
Gemini 2.5 Pro
Gemini 2.5 Pro is Google DeepMind's flagship thinking AI model engineered for advanced reasoning, coding, and STEM tasks. Featuring internal deliberation and a 1M+ token context window, it natively processes text, image, audio, video, and code for software engineering and agentic automation.
February 2025
Claude 3.7 Sonnet
Anthropic's first hybrid reasoning model. It functions as both an ordinary LLM and an extended reasoning model, giving API users strict control over "thinking token" budgets for advanced math and science.
Brain2Qwerty
Brain2Qwerty is a non-invasive Brain-Computer Interface (BCI) deep learning system developed by Meta AI (FAIR) in collaboration with the Basque Center on Cognition, Brain and Language (BCBL). It decodes text directly from non-invasive neural recordings—specifically Magnetoencephalography (MEG) and Electroencephalography (EEG)—captured while participants type on a QWERTY keyboard.
Co-Scientist
Co-Scientist (AI Co-Scientist) is a Gemini 2.0-powered multi-agent AI system developed by Google DeepMind and Google Research. Designed to act as a virtual scientific collaborator, it generates, debates, evolves, and ranks novel research hypotheses with citation grounding across biomedical and physical sciences.
January 2025
December 2024
November 2024
Suno v4
Suno's flagship generative model for full-length music, vocal, and song production in high-fidelity stereo.
Qwen2.5-Coder 32B Instruct
Alibaba's state-of-the-art open-weights coding model, achieving coding capabilities comparable to GPT-4o on major benchmarks.
Claude 3.5 Haiku
A rapid, highly efficient model that matched or exceeded the performance of the original Claude 3 Opus on many tasks at a fraction of the cost.
October 2024
September 2024
Qwen 2.5 72B Instruct
Qwen2.5-72B-Instruct is the flagship 72.7-billion parameter instruction-tuned language model in Alibaba's Qwen2.5 series. Pre-trained on 18 trillion tokens, it delivers massive improvements in instruction following, coding, mathematical reasoning, multilingual understanding, and structured data generation.
o1
Previous full o-series reasoning model
August 2024
Imagen 3
Google DeepMind's highest-quality text-to-image model, featuring advanced detail rendering and robust prompt adherence.
Grok 2
xAI's frontier Mixture-of-Experts language model, featuring advanced reasoning, coding, and real-time information retrieval integrated on the X platform.
FLUX.1 schnell
A state-of-the-art 12B parameter flow transformer text-to-image generation model, optimized for speed and high prompt adherence.
July 2024
June 2024
Claude 3.5 Sonnet
The model that introduced the "Artifacts" interface in the Claude web app, allowing users to generate and interact with code, websites, and documents in a dedicated panel.
DeepSeek-Coder-V2 Instruct
An open Mixture-of-Experts coding model by DeepSeek that matches or exceeds GPT-4 Turbo in competitive programming and code reasoning.
Runway Gen-3 Alpha
Runway's next-generation video model, providing high fidelity, structural controls, and temporal consistency for professional workflows.
Stable Diffusion 3 Medium
Stability AI's latest open-weights image generation model using a novel Multimodal Diffusion Transformer (MMDiT) architecture with separate weights for image and text, delivering improved typography, prompt adherence, and image quality.
Kling 1.0
Cinematic text-to-video model by Kuaishou, supporting high-resolution generation, long durations, and realistic physical simulations.
May 2024
Codestral 22B
Mistral AI's specialized open-weights model optimized for code generation, supporting over 80 programming languages.
GPT-4o
OpenAI's multimodal flagship model accepting text, image, and audio inputs natively, with GPT-4-level intelligence at faster speed and lower cost.
Unitree G1
A mass-market humanoid robot designed for general-purpose applications, gymnastics, and high-precision task execution. It features advanced joint flexibility and has been successfully teleoperated via VR headsets to assist in surgical procedures.
April 2024
Mixtral 8x22B
A sparse mixture-of-experts model from Mistral AI, released with open weights under Apache 2.0, positioned as a strong open alternative to closed frontier models on reasoning and coding benchmarks.
Command R+
Cohere's most capable generation model optimized for enterprise RAG workflows and multilingual tasks, with strong performance on complex reasoning while maintaining efficiency for large-scale production deployments.
March 2024
Claude 3 Haiku
The fastest and most compact tier of the Claude 3 family, introducing multimodal vision capabilities to the lowest-cost tier.
Claude 3 Opus
The heavy-weight frontier model of the generation 3 family, taking the top spot on multiple reasoning benchmarks at the time of its release.
Claude 3 Sonnet
The balanced mid-tier model of the Claude 3 generation, offering strong capabilities at an efficient price point.
February 2024
Gemini 1.5 Pro
The flagship model of the Gemini 1.5 family, introducing a massive breakthrough with a 1 million token context window in production APIs, up to 10M tokens.
Sora Research Preview
OpenAI's text-to-video diffusion transformer model, generating high-fidelity video clips up to 60 seconds with physical simulation.
Gemini 1.0 Ultra
Google's most powerful offering in the original Gemini 1.0 family.
January 2024
December 2023
November 2023
Claude 2.1
An iterative improvement on Claude 2 that doubled the context window to an industry-leading 200K tokens and heavily reduced hallucinations.
GPT-4 Turbo
An older high-intelligence GPT model
Whisper large-v3
OpenAI's third-generation multilingual speech recognition and translation model, trained on extensive audio datasets.
TTS-1
Text-to-speech model optimized for speed
Embed v3 (English)
Cohere's state-of-the-art embedding model featuring noise-resistant retrieval, optimized for search engines and RAG workflows.
October 2023
September 2023
August 2023
July 2023
June 2023
March 2023
Claude 1
The inaugural foundational model from Anthropic, built on Constitutional AI principles and introduced alongside the initial Anthropic API.
GPT-4
An older high-intelligence GPT model
GPT-3.5 Turbo
Legacy GPT model for cheaper chat and non-chat tasks