Back to Timeline

Release Archive

A flat, chronological list of every model tracked in the Modelverse registry.

July 2026

Tsinghua University & StepFun

Motion4Motion

Motion4Motion is a training-free framework for cross-subject motion transfer presented at SIGGRAPH 2026. Unlike prior skeleton-based approaches that only work for human-like characters with predefined skeleton topologies, Motion4Motion steps out of the skeleton framework entirely. It models the motion flow of a subject in a video rather than its skeleton, enabling motion transfer across entirely different species (e.g., human-to-animal) at inference time without any task-specific training or fine-tuning.

Jul 25, 2026
research preview
Moonshot AI

Kimi K3

Kimi K3 is Kimi's flagship 2.8T open-source model built on Kimi Delta Attention (KDA) and Attention Residuals. It features native visual understanding for images and video, a 1M-token context window with automatic caching, and an always-on thinking mode designed for frontier intelligence scenarios like long-horizon coding and knowledge work.

Jul 25, 2026
open source
Anthropic

Claude Opus 5

For complex agentic coding and enterprise work. Claude Opus 5 is a step-change improvement over Claude Opus 4.8, with the largest gains in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.

Jul 24, 2026
api only
Kwaipilot

KAT-Coder-V2.5-Dev

KAT-Coder-V2.5-Dev is a 34.7B-parameter Mixture-of-Experts (MoE) text generation model developed by Kwaipilot. Built on the Qwen3_5MoeForConditionalGeneration architecture using transformers. Supports en, zh language(s). Released on 2026-07-23 with 126 likes and 396 downloads on Hugging Face.

Jul 23, 2026
open weights
NVIDIA

ARDY

ARDY is an autoregressive diffusion model for interactive human motion generation with online text prompting and flexible kinematic constraints.

Jul 22, 2026
open weights
Qualcomm

MobileWan

MobileWan is a mobile video diffusion model developed by Qualcomm AI Research that aims to close the quality gap for on-device video generation.

Jul 22, 2026
research preview
upstage

Solar-Open2-250B

Solar-Open2-250B is a 250.3B-parameter Mixture-of-Experts (MoE) text generation model developed by upstage. Built on the SolarOpen2ForCausalLM architecture using transformers. Supports en, ko, ja language(s). Released on 2026-07-22 with 543 likes and 1,106 downloads on Hugging Face.

Jul 22, 2026
open weights
fdtn-ai

antares-1b

antares-1b is a 1.8B-parameter text generation model developed by fdtn-ai. Built on the GraniteMoeHybridForCausalLM architecture using transformers. Supports en language(s). Released on 2026-07-21 with 150 likes and 4,266 downloads on Hugging Face.

Jul 21, 2026
open weights
Google DeepMind

Gemini 3.5 Flash Cyber

Google DeepMind's specialized cybersecurity model, built to help defenders find, validate, and patch software vulnerabilities quickly and efficiently. Gemini 3.5 Flash Cyber is fine-tuned from the Gemini 3.5 Flash base model with deep expertise in vulnerability research, security analysis, exploit validation, and patch generation. It enables security teams to dramatically accelerate their defensive workflows compared to general-purpose LLMs.

Jul 21, 2026
closed source
Google DeepMind

Gemini 3.5 Flash-Lite

Google DeepMind's most cost-efficient model in the Gemini 3.5 family, designed for high-volume, low-latency applications where economy is paramount. Gemini 3.5 Flash-Lite delivers strong performance for everyday tasks at the lowest price point of the new July 2026 Gemini release batch, while retaining multimodal input capabilities and the 1M token context window.

Jul 21, 2026
closed source
Google DeepMind

Gemini 3.6 Flash

Google DeepMind's fastest frontier-class model, featuring built-in reasoning, native multimodal input (text, image, audio, video), and a 1M token context window. Ranked #1 in speed among all 186 models on Artificial Analysis at 275.5 tokens/second while scoring 50 on the Artificial Analysis Intelligence Index — placing it well above average among comparable models. Priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens with a 90% cache discount.

Jul 21, 2026
closed source
microsoft

Mage-Flow

Mage-Flow is a 4.1B-parameter text to image model developed by microsoft. Released on 2026-07-21 with 242 likes and 891 downloads on Hugging Face.

Jul 21, 2026
open weights
Nanbeige

Nanbeige4.2-3B

Nanbeige4.2-3B is a 4.2B-parameter text generation model developed by Nanbeige. Built on the NanbeigeForCausalLM architecture using transformers. Supports en, zh language(s). Released on 2026-07-21 with 374 likes and 8,169 downloads on Hugging Face.

Jul 21, 2026
open weights
unsloth

Laguna-S-2.1-GGUF

Laguna-S-2.1-GGUF is a undisclosed-parameter text generation model developed by unsloth. Released on 2026-07-21 with 172 likes and 57,536 downloads on Hugging Face.

Jul 21, 2026
open weights
Motif-Technologies

Motif-3-Beta

Motif-3-Beta is a 314.8B-parameter Mixture-of-Experts (MoE) text generation model developed by Motif-Technologies. Built on the MotifForCausalLM architecture using transformers. Supports en, ko language(s). Released on 2026-07-20 with 186 likes and 2,108 downloads on Hugging Face.

Jul 20, 2026
open weights
openbmb

MiniCPM-RobotManip

MiniCPM-RobotManip is a 1.5B-parameter robotics model developed by openbmb. Built on the MiniCPMV_VLA architecture using transformers. Released on 2026-07-18 with 173 likes and 559 downloads on Hugging Face.

Jul 18, 2026
open weights
DavidAU

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF is a undisclosed-parameter image text to text model developed by DavidAU. Supports en, zh language(s). Released on 2026-07-17 with 496 likes and 407,421 downloads on Hugging Face.

Jul 17, 2026
open weights
Alibaba Group

Wan Streamer v0.3

Wan Streamer v0.3 learns video as a persistent world plus the events that unfold inside it, bringing free-form behavior to real-time audio-visual interaction.

Jul 16, 2026
research preview
Academic/Research

Actionable World

WorldString (Actionable World) is a neural architecture developed by Tsinghua University, UCSD, CalTech, and NVIDIA for modeling physical object state manifolds (articulated, skinned, and soft objects) directly from point clouds or RGB-D video streams.

Jul 15, 2026
research preview
Academic/Research

Agents A1

InternScience's open-weights 35B Mixture-of-Experts agentic model, designed for long-horizon planning, multi-teacher domain distillation, and complex tool execution.

Jul 15, 2026
research preview
Academic/Research

Luna

Developed by researchers from HKUST, Tsinghua University, and Meta, Luna (Learning Universal 3D Human Animation Beyond Skinning) is an LBS-free neural animation model presented at ECCV 2026. It maps 2D controls directly into 3D Gaussian deformations for realistic human avatars.

Jul 15, 2026
research preview
Academic/Research

MAMMA

MAMMA (Markerless Accurate Multi-person Motion Acquisition) is a state-of-the-art markerless multi-view human motion capture framework developed by MPI-IS and CMU (CVPR 2026 Oral). Powered by MammaNet (a ViT-Base transformer predicting 512 contact-aware and visibility-aware surface landmarks), it accurately fits SMPL-X body models to complex multi-person physical interactions.

Jul 15, 2026
research preview
Academic/Research

MiniCPM5 1B

MiniCPM5 1B is a high-performance, dense 1.08B parameter language model developed by OpenBMB. Designed for on-device deployment and agentic workflows, it features dual-mode reasoning ('Think' CoT vs. 'NoThink' fast mode), a native 128K context window, native tool calling, and standard Llama architecture compatibility.

Jul 15, 2026
open weights
Academic/Research

PaGeR

PaGeR (Panoramic Geometry Reconstruction) is a unified framework from ETH Zurich (Photogrammetry and Remote Sensing Lab) that adapts perspective 3D foundation models (Depth Anything 3) to single-pass panoramic geometry estimation, predicting scale-invariant depth, metric depth, surface normals, and sky masks.

Jul 15, 2026
research preview
Academic/Research

RDM

Representation Distribution Matching (RDM / iRDM) is a paradigm for training state-of-the-art one-step visual generation models developed by EPFL (VITA Lab), Valeo.ai, and Sorbonne Université. It matches feature distributions between generated and reference images using Maximum Mean Discrepancy (MMD) with Nyström estimation across a battery of frozen pretrained encoders.

Jul 15, 2026
research preview
Academic Research

VideoMDM

VideoMDM is a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without requiring any 3D ground truth.

Jul 15, 2026
open weights
Alibaba

Qwen 3.7

Developed by Alibaba, Qwen 3.7 is a research preview model exploring new techniques in specialized research.

Jul 15, 2026
research preview
Google DeepMind

Magenta Realtime

An experimental agentic workflow research preview named Magenta Realtime, published by Google DeepMind to showcase novel methods.

Jul 15, 2026
research preview
Meta

WavFlow

Developed by Meta, WavFlow is a research preview model exploring new techniques in specialized research.

Jul 15, 2026
research preview
Microsoft

MAI Thinking

An experimental specialized research research preview named MAI Thinking, published by Microsoft to showcase novel methods.

Jul 15, 2026
research preview
OpenAI

GPT Dreaming

A research-preview model by OpenAI focusing on specialized research capabilities to invite community feedback.

Jul 15, 2026
research preview
OpenAI

GPT-Red

An automated AI red-teaming system developed by OpenAI, trained via self-play reinforcement learning to simulate prompt injection attacks and discover vulnerabilities in language models prior to deployment.

Jul 15, 2026
research preview
Other

Step 3.7 Flash

An experimental specialized research research preview named Step 3.7 Flash, published by StepFun to showcase novel methods.

Jul 15, 2026
research preview
Stability AI

Stable Audio 3

Developed by Stability AI, Stable Audio 3 is a research preview model exploring new techniques in audio and speech.

Jul 15, 2026
research preview
Stability AI

Stable Layers

A research-preview model by Stability AI focusing on specialized research capabilities to invite community feedback.

Jul 15, 2026
research preview
Tencent

HY-MT2

A research-preview model by Tencent focusing on specialized research capabilities to invite community feedback.

Jul 15, 2026
research preview
Thinking Machines

Inkling

Thinking Machines' flagship open-weights Mixture-of-Experts multimodal model, supporting Native MoE transformer architecture and up to 1M token context window.

Jul 15, 2026
open weights
Zhipu AI

GLM 5.2

An experimental specialized research research preview named GLM 5.2, published by Zhipu AI to showcase novel methods.

Jul 15, 2026
research preview
Academic/Research

MoVerse

MoVerse is a undisclosed-parameter text to video model developed by Academic/Research. Released on 2026-07-14 with 0 likes and 0 downloads on Hugging Face.

Jul 14, 2026
research preview
Alibaba Group (Tongyi Lab)

Wan-Dancer

Wan-Dancer is a hierarchical framework for minute-scale coherent music-to-dance video generation developed by Tongyi Lab at Alibaba Group. Given a reference character image and music audio, Wan-Dancer generates long-duration (minute-scale), high-quality, rhythmically synchronized dance videos with global structural coherence and temporal continuity across multiple dance genres including Chinese Classical, K-pop, Street, Tap, and Latin styles.

Jul 14, 2026
open weights
OpenAI

codex-mini-latest

Fast reasoning model optimized for the Codex CLI

Jul 14, 2026
closed source
OpenAI

GPT-4.1

Smartest non-reasoning model

Jul 14, 2026
api only
OpenAI

GPT-4.5 Preview

Deprecated large model.

Jul 14, 2026
closed source
OpenAI

GPT-4o Audio

GPT-4o models capable of audio inputs and outputs

Jul 14, 2026
closed source
OpenAI

GPT-4o mini Audio

Smaller model capable of audio inputs and outputs

Jul 14, 2026
closed source
OpenAI

GPT-4o mini Realtime

Smaller realtime model for text and audio inputs and outputs

Jul 14, 2026
closed source
OpenAI

GPT-4o mini Search Preview

Fast, affordable small model for web search

Jul 14, 2026
closed source
OpenAI

GPT-4o mini Transcribe

Speech-to-text model powered by GPT-4o mini

Jul 14, 2026
api only
OpenAI

GPT-4o mini TTS

Text-to-speech model powered by GPT-4o mini

Jul 14, 2026
closed source
OpenAI

GPT-4o Realtime

Model capable of realtime text and audio inputs and outputs

Jul 14, 2026
closed source
OpenAI

GPT-4o Search Preview

GPT model for web search in Chat Completions

Jul 14, 2026
closed source
OpenAI

GPT-4o Transcribe

Speech-to-text model powered by GPT-4o

Jul 14, 2026
api only
OpenAI

GPT-5.1 Chat

GPT-5.1 model used in ChatGPT

Jul 14, 2026
closed source
OpenAI

GPT-5.1-Codex

A version of GPT-5.1 optimized for agentic coding in Codex.

Jul 14, 2026
closed source
OpenAI

GPT-5.1

The best model for coding and agentic tasks with configurable reasoning effort

Jul 14, 2026
api only
OpenAI

GPT-5.2 Chat

GPT-5.2 model used in ChatGPT

Jul 14, 2026
closed source
OpenAI

GPT-5.2-Codex

Our most intelligent coding model optimized for long-horizon, agentic coding tasks.

Jul 14, 2026
closed source
OpenAI

GPT-5.2

Previous frontier model for professional work with configurable reasoning effort

Jul 14, 2026
api only
OpenAI

GPT-5.3 Chat

GPT-5.3 Instant model used in ChatGPT

Jul 14, 2026
closed source
OpenAI

GPT-5.3-Codex

The most capable agentic coding model to date.

Jul 14, 2026
api only
OpenAI

GPT-5.4

A more affordable model for coding and professional work.

Jul 14, 2026
api only
OpenAI

GPT-5.5

A new class of intelligence for coding and professional work.

Jul 14, 2026
api only
OpenAI

GPT-5 Chat

GPT-5 model used in ChatGPT

Jul 14, 2026
closed source
OpenAI

GPT-5-Codex

A version of GPT-5 optimized for agentic coding in Codex

Jul 14, 2026
closed source
OpenAI

GPT-5

Previous intelligent reasoning model for coding and agentic tasks with configurable reasoning effort

Jul 14, 2026
api only
OpenAI

gpt-audio-1.5

The best voice model for audio in, audio out with Chat Completions.

Jul 14, 2026
api only
OpenAI

gpt-audio

For audio inputs and outputs with Chat Completions API

Jul 14, 2026
api only
OpenAI

GPT Image 1.5

Our previous image generation model

Jul 14, 2026
closed source
OpenAI

GPT Image 1

Our previous image generation model

Jul 14, 2026
closed source
OpenAI

GPT Image 2

State-of-the-art image generation model

Jul 14, 2026
api only
OpenAI

gpt-oss-120b

Most powerful open-weight model, fits into an H100 GPU

Jul 14, 2026
open weights
OpenAI

gpt-oss-20b

Medium-sized open-weight model for low latency

Jul 14, 2026
open weights
OpenAI

GPT-Realtime-1.5

The best voice model for audio in, audio out

Jul 14, 2026
api only
OpenAI

GPT-Realtime-2.1

Reasoning model with tool use

Jul 14, 2026
api only
OpenAI

GPT-Realtime-2

Reasoning model with tool use

Jul 14, 2026
api only
OpenAI

GPT-Realtime-Translate

Streaming speech-to-speech translation model

Jul 14, 2026
api only
OpenAI

GPT-Realtime-Whisper

Streaming speech-to-text model for realtime transcription

Jul 14, 2026
api only
OpenAI

GPT-Realtime

Model capable of realtime text and audio inputs and outputs

Jul 14, 2026
api only
OpenAI

o4-mini-deep-research

Faster, more affordable deep research model

Jul 14, 2026
closed source
OpenAI

o4-mini

Fast, cost-efficient reasoning model, succeeded by GPT-5 mini

Jul 14, 2026
closed source
OpenAI

omni-moderation

Identify potentially harmful content in text and images

Jul 14, 2026
api only
PrismML

Bonsai 27B

Bonsai 27B is a 27B-class multimodal model based on Qwen3.6 27B, optimized for on-device and local agentic workflows. It is available in two highly compressed variants: a 5.9 GB Ternary variant (1.71 bits/weight) and a 3.9 GB 1-bit variant (1.125 bits/weight), allowing it to fit within the memory budget of everyday laptops and phones (like the iPhone 17 Pro).

Jul 14, 2026
open weights
OpenAI

Sora 2

Flagship video generation with synced audio

Jul 14, 2026
closed source
OpenAI

text-moderation-stable

Previous generation text-only moderation model

Jul 14, 2026
closed source
OpenAI

text-moderation

Previous generation text-only moderation model

Jul 14, 2026
closed source
thinkingmachines

Inkling

Inkling is a 952.4B-parameter Mixture-of-Experts (MoE) image text to text model developed by thinkingmachines. Built on the InklingForConditionalGeneration architecture using transformers. Released on 2026-07-14 with 1,547 likes and 27,883 downloads on Hugging Face.

Jul 14, 2026
open weights
poolside

Laguna-S-2.1

Laguna-S-2.1 is a 117.6B-parameter Mixture-of-Experts (MoE) text generation model developed by poolside. Built on the LagunaForCausalLM architecture using transformers. Released on 2026-07-13 with 618 likes and 28,992 downloads on Hugging Face.

Jul 13, 2026
open weights
Alibaba

ABot-World

ABot-World is a real-time interactive world simulator developed by Alibaba AMAP CV Lab. Built around a 5-billion parameter causal video world model (ABot-World-0-5B-LF) fine-tuned from Wan2.2-TI2V-5B, it enables open-ended, action-conditioned video generation running in real time at 720p resolution @ 16 FPS on a single NVIDIA RTX 5090 GPU.

Jul 9, 2026
open weights
Mirelo AI

MuScriptor

MuScriptor is an open-weights multi-instrument audio-to-MIDI model developed by Mirelo AI and Kyutai.

Jul 9, 2026
open weights
Other

LingBot-World 2.0

An interactive world model that supports continuous, hour-long real-time generation. It features a native dual-agent mechanism that allows for dynamically evolving environments responding to real-time user inputs without scene collapse.

Jul 9, 2026
open weights
ByteDance

Seedream 5 Pro

Seedream 5.0 Pro is ByteDance's flagship multimodal image generation and interactive precision editing model. Designed for professional creative workflows, it features native layer separation, interactive grounded editing, high-density multilingual typography, and multi-reference consistency up to 4K resolution.

Jul 8, 2026
closed source
conradlocke

krea2-identity-edit

krea2-identity-edit is a undisclosed-parameter text generation model developed by conradlocke. Released on 2026-07-07 with 534 likes and 0 downloads on Hugging Face.

Jul 7, 2026
open weights
Tencent

Hy3

Tencent's flagship open-weights Mixture-of-Experts reasoning model, utilizing hybrid thinking and active chain-of-thought configuration.

Jul 6, 2026
open weights
prism-ml

Bonsai-27B-gguf

Bonsai-27B-gguf is a undisclosed-parameter text generation model developed by prism-ml. Released on 2026-07-04 with 632 likes and 2,028,115 downloads on Hugging Face.

Jul 4, 2026
open weights
prism-ml

Ternary-Bonsai-27B-gguf

Ternary-Bonsai-27B-gguf is a undisclosed-parameter text generation model developed by prism-ml. Released on 2026-07-04 with 1,009 likes and 595,415 downloads on Hugging Face.

Jul 4, 2026
open weights
poolside

Laguna-S-2.1-NVFP4

Laguna-S-2.1-NVFP4 is a 67.9B-parameter Mixture-of-Experts (MoE) text generation model developed by poolside. Built on the LagunaForCausalLM architecture using vllm. Released on 2026-07-02 with 130 likes and 89,186 downloads on Hugging Face.

Jul 2, 2026
open weights
Meta

Muse Image

Meta MUSE Image is a proprietary agentic image generation model developed by Meta Superintelligence Labs (MSL). It structures its generation process using inference-time reasoning and active background web searches to gather accurate real-world visual references prior to rendering, powering generative features across meta.ai, Instagram, and WhatsApp.

Jul 1, 2026
closed source
Meta

Muse Video

Meta's previewed video generation model showcasing complex physical comprehension and natively integrated audio.

Jul 1, 2026
research preview
NVIDIA

Aspire

ASPIRE (Agentic Skill Programming through Iterative Robot Exploration) is a continual learning framework for robotics developed by NVIDIA GEAR Lab in collaboration with UMich, UIUC, UC Berkeley, and CMU. Using a code-as-policy approach, ASPIRE enables robot agents to autonomously generate, test, debug, and refine executable Python control programs.

Jul 1, 2026
open weights
NVIDIA

CHORD

CHORD (Contact Wrench Guidance from Human Demonstration) is a framework by NVIDIA Isaac Team and GEAR Lab for learning long-horizon dexterous and whole-body robotic manipulation from human demonstrations using object-centric contact wrench space guidance in reinforcement learning.

Jul 1, 2026
open weights
OpenAI

GPT 5.6

OpenAI's latest flagship model engineered for prolonged, multi-step agentic coding tasks requiring minimal human oversight or handholding.

Jul 1, 2026
closed source
OpenAI

GPT Live

OpenAI's latest real-time voice interface designed to enable natural, flowing conversation. It supports dynamic interruptions, background task delegation to stronger frontier models, and can generate contextual visual responses in the UI.

Jul 1, 2026
closed source
Other

Mira

A real-time multiplayer simulation model that acts entirely as a video generator responding to live user key presses, bypassing pre-designed game engines to render physics and collisions on the fly.

Jul 1, 2026
research preview
Other

PixWorld

A 3D scene generator that reconstructs fully consistent 3D environments directly in pixel space from one or multiple reference images, avoiding the visual artifacts often caused by latent-space generation.

Jul 1, 2026
open source
Other

ProxyPose

A novel 6-DoF pose tracking system that reframes 3D position and rotation tracking as a video-to-video translation problem. It operates entirely at the pixel level without requiring 3D models or depth sensors.

Jul 1, 2026
open weights
Other

Reve 2.1

A state-of-the-art image model known for extreme high-resolution outputs, precise text rendering, and robust targeted micro-editing via bounding box selections.

Jul 1, 2026
closed source
Other

Wan Streamer 0.2

A real-time character simulation framework allowing users to converse with generated avatars—including humans, animals, or fictional characters—with highly responsive audio-visual sync.

Jul 1, 2026
research preview
Tencent

High3

A large open-weights Mixture of Experts (MoE) model focused on reasoning, math, and agentic coding that punches above its weight class against much larger trillion-parameter models.

Jul 1, 2026
open weights
xAI

Grok 4.5

A highly efficient frontier model built for coding, engineering, and math. While its context window is smaller than immediate competitors, it prioritizes fast token generation and cost efficiency over topping raw intelligence leaderboards.

Jul 1, 2026
closed source

June 2026

Anthropic

Claude Sonnet 5

Built specifically as an execution layer for multi-step software engineering work. It handles sustained coding, browser tool use, and debugging autonomously at speeds matching previous Opus models.

Jun 30, 2026
api only
Academic/Research

VidiHand

VidiHand is a generative 4D hand motion reconstruction model developed by Nanyang Technological University (NTU) and Shanghai Jiao Tong University (SJTU). It recovers metric-scale 3D/4D two-hand pose trajectories from monocular egocentric video by fine-tuning internet-scale video diffusion models (Wan2.1-VACE) without needing external hand detectors or test-time optimization.

Jun 29, 2026
open weights
Academic/Research

PhysiFormer

PhysiFormer is a diffusion transformer model developed by Visual Geometry Group (VGG), University of Oxford to simulate physically plausible 3D object dynamics directly in 3D world coordinates. Conditioned on initial 3D vertex positions, velocities, and material type descriptors, it treats full-horizon 3D trajectory prediction as a single denoising process.

Jun 25, 2026
open weights
Other

Ornith 1.0

DeepReinforce's open-weights family of self-scaffolding models optimized for agentic coding, post-trained on Gemma 4 and Qwen 3.5.

Jun 25, 2026
open weights
Academic/Research

OmniContact

OmniContact is a hierarchical framework for generalizable humanoid loco-manipulation developed by Noitom Robotics and HKUST. Centered on 'Contact Flow' (CF) representations, it combines high-level reference synthesis (CF-Gen) with 50Hz closed-loop tracking (CF-Track) to enable long-horizon meta-skill chaining, autonomous failure recovery, and real-time execution across tasks like carrying, pushing, sliding, and relocating.

Jun 24, 2026
research preview
baidu

Unlimited-OCR

Unlimited-OCR is a 3.3B-parameter Mixture-of-Experts (MoE) image text to text model developed by baidu. Built on the UnlimitedOCRForCausalLM architecture using transformers. Supports multilingual language(s). Released on 2026-06-19 with 3,029 likes and 2,500,391 downloads on Hugging Face.

Jun 19, 2026
open weights
empero-ai

Qwythos-9B-Claude-Mythos-5-1M-GGUF

Qwythos-9B-Claude-Mythos-5-1M-GGUF is a undisclosed-parameter image text to text model developed by empero-ai. Supports en language(s). Released on 2026-06-19 with 2,456 likes and 1,906,539 downloads on Hugging Face.

Jun 19, 2026
open weights
Academic/Research

Flex4DHuman

Flex4DHuman is a generative framework from University of Washington, Zhejiang University, and Tencent for flexible multi-view video diffusion and 4D human reconstruction from monocular or sparse video streams without explicit geometry priors.

Jun 18, 2026
research preview
Academic/Research

MuSViT

MuSViT (Music Score Vision Transformer) is the first foundation vision model specifically engineered for Optical Music Recognition (OMR) and sheet music representation. Developed by PRAIG at the University of Alicante (accepted at ECCV 2026), it uses a ViT-Base encoder pre-trained via Masked Autoencoders on 9.7 million IMSLP sheet music pages.

Jun 16, 2026
open weights
zai-org

GLM-5.2

GLM-5.2 is a 753.3B-parameter Mixture-of-Experts (MoE) text generation model developed by zai-org. Built on the GlmMoeDsaForCausalLM architecture using transformers. Supports en, zh language(s). Released on 2026-06-16 with 4,422 likes and 667,403 downloads on Hugging Face.

Jun 16, 2026
open weights
Academic/Research

Arbor

Generalist autonomous research agent framework developed by Renmin University of China (RUC-NLPIR) and Microsoft Research, utilizing Hypothesis-Tree Refinement (HTR) for long-horizon scientific research and system optimization.

Jun 15, 2026
research preview
Academic/Research

Dots TTS

Dots TTS (dots.tts) is a 2-billion parameter, fully continuous, end-to-end autoregressive text-to-speech foundation model developed by RedNote HiLab (Xiaohongshu). It operates over continuous latent space via a 48 kHz AudioVAE, offering state-of-the-art zero-shot voice cloning across 24 languages.

Jun 15, 2026
research preview
Academic/Research

i1

i1 is a 3-billion-parameter text-to-image diffusion model developed by ZLab at Princeton University. Introduced in 'i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models', it provides a fully open foundation—including model weights, PyTorch and JAX training/inference code, and curated caption datasets—built upon a LightningDiT cross-attention backbone with T5-Gemma-2B and FLUX.2 VAE.

Jun 15, 2026
research preview
Academic/Research

LiveEdit

LiveEdit is a real-time, diffusion-based streaming video editing framework developed by Tsinghua University and HKUST (accepted at ECCV 2026). Built on top of Wan2.1, it achieves ~12.66 FPS causal streaming editing using a 3-stage distillation pipeline and an AR-oriented mask cache.

Jun 15, 2026
research preview
Academic/Research

World Tracing

World Tracing is a generative 3D geometry representation framework developed by World Labs and UIUC (Hao Zhang, Mohamed El Banani, et al.). Given a single 2D image or short video clip, World Tracing predicts an ordered stack of 3D points along camera rays for every pixel, capturing both visible surfaces and occluded geometry behind them.

Jun 15, 2026
open weights
Moonshot AI

Kimi K2.7 Code

Moonshot AI's open-weights Mixture-of-Experts model optimized for software engineering, reducing thinking token usage and improving code generation efficiency.

Jun 12, 2026
open weights
Academic/Research

Surflo

Surflo (Consistent 3D Surface Flow Model with Global State) is a feed-forward 3D surface reconstruction model developed by École Polytechnique, Kyoto University, Kyutai, and UC Berkeley. It encodes unposed RGB images into a 128-token global latent state and decodes 3D surface points using continuous flow matching with photometric rendering guidance.

Jun 11, 2026
open weights
Google DeepMind

DiffusionGemma

An experimental open-weights text diffusion model by Google DeepMind. Moving beyond traditional sequential autoregressive decoding, DiffusionGemma utilizes parallel block decoding to generate 256 tokens simultaneously, achieving up to 4x faster generation speeds on modern GPU hardware.

Jun 10, 2026
research preview
Academic/Research

SCAIL 2

SCAIL 2 is an open-source end-to-end framework for controlled character animation and video-to-video motion transfer developed by Tsinghua University (KEG Group) and Z.ai. It performs motion transfer directly from driving video to reference character without relying on intermediate representations like 3D skeletons, OpenPose maps, or depth maps.

Jun 9, 2026
open weights
Anthropic

Claude Fable 5

Anthropic's most capable publicly available frontier model as of mid-2026. It is engineered for long-running, multi-agent autonomous workflows and sustained logical reasoning.

Jun 9, 2026
api only
Anthropic

Claude Mythos 5

A highly restricted enterprise model available exclusively through Project Glasswing. It shares the core architecture of Fable 5 but contains advanced capabilities specifically optimized for life sciences, biology research, and high-security cyber contexts.

Jun 9, 2026
closed source
Cohere

North Mini Code

North Mini Code is a specialized 30B total parameter, sparse Mixture-of-Experts (MoE) coding model with ~3B active parameters per token developed by Cohere and Cohere Labs. Designed for agentic software engineering and developer workflows, it achieves near-30B-scale reasoning with the compute overhead of a 3B parameter model.

Jun 9, 2026
open weights
Academic/Research

AnchorWorld

AnchorWorld is an embodied egocentric world simulation framework that enables interactive, 3D motion-driven video generation with view-based evolution customization via spatial anchor views.

Jun 5, 2026
research preview
Meta AI / HKUST

MeshFlow

MeshFlow is a flow-based diffusion transformer model with MeshVAE for efficient, high-quality artistic 3D mesh generation presented as a CVPR 2026 Highlight by Meta AI and HKUST.

Jun 5, 2026
open weights
NVIDIA

Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is a flagship open-weight frontier LLM family engineered for agentic AI workloads, long-context understanding, complex reasoning, and tool calling. The family includes Nemotron 3 Ultra 550B (a hybrid Transformer-Mamba LatentMoE model with 1M context) and Llama-3.1-Nemotron-Ultra-253B-v1.

Jun 4, 2026
open weights
Academic Research

GenCeption

GenCeption repurposes a text-to-video generative diffusion model into a unified, feed-forward general-purpose vision learner steered by text instructions. Rather than training task-specific vision models, GenCeption demonstrates that a single model built on video generative pretraining can perform depth estimation, surface normal prediction, pose estimation, semantic segmentation, keypoint detection, and 4D grounding — all from the same weights — achieving SOTA performance across multiple tasks without task-specific fine-tuning.

Jun 1, 2026
research preview
Academic/Research

StreamForce

StreamForce (Streaming Video Generation with Streaming Force Control) is a causal, unified framework for real-time streaming video generation developed by Northeastern University (NEU-VI Lab). It enables physically grounded interactive control over video rollouts via continuous, time-varying force inputs at up to 16.6 FPS.

Jun 1, 2026
research preview
Academic/Research

WavTTS

WavTTS is an end-to-end zero-shot text-to-speech framework developed by Shanghai Jiao Tong University (SJTU) and ByteDance Seed. It generates speech directly in raw waveform space using flow matching and Diffusion Transformers (DiT), bypassing intermediate mel-spectrogram or neural codec token representations.

Jun 1, 2026
open weights
Alibaba

Qwen 3.7 Plus

Qwen 3.7 Plus is a multimodal interactive hybrid agent model developed by Alibaba Group. Designed to unify vision and language capabilities, it serves as a versatile foundation for agentic workflows, integrating Graphical User Interface (GUI) screen interaction with Command Line Interface (CLI) code execution and tool use.

Jun 1, 2026
closed source
MiniMax

MiniMax-M3

MiniMax's open-weights Mixture-of-Experts multimodal model, using MiniMax Sparse Attention (MSA) to enable low-latency long-context reasoning up to 1 million tokens.

Jun 1, 2026
open weights
NVIDIA

OmniDreams

OmniDreams (NVIDIA Cosmos-Dreams) is an action-conditioned real-time generative world model developed by NVIDIA Spatial Intelligence Lab (SIL). It autoregressively synthesizes multi-camera photorealistic video observations conditioned on driving actions and simulator states for closed-loop autonomous vehicle simulation.

Jun 1, 2026
open weights
Other

SeFi-Image

A text-to-image foundation model built on Semantic-First Diffusion, a paradigm that separates semantic layout streams from texture detail streams to improve generation quality while drastically reducing training compute requirements.

Jun 1, 2026
open weights

May 2026

NVIDIA

Cosmos 3

NVIDIA Cosmos 3 is an open-source frontier foundation model platform engineered for Physical AI. Built on a Mixture-of-Transformers (MoT) architecture with a two-tower design, Cosmos 3 unifies an autoregressive Reasoner with a diffusion Generator to allow embodied agents to perceive, reason about, simulate, and act in the physical world.

May 31, 2026
open weights
Anthropic

Claude Opus 4.8

A flagship update to the Opus 4 lineage that served as the primary frontier model prior to the release of Fable 5. Highly effective at complex reasoning and deep technical analysis.

May 28, 2026
api only
Roblox Research

CubePart

CubePart is a undisclosed-parameter text to 3d model developed by Roblox Research. Released on 2026-05-27 with 19 likes and 0 downloads on Hugging Face.

May 27, 2026
open weights
Academic/Research

PanoWorld

PanoWorld is a generative spatial world model developed by Ke Holdings Inc. (Beike) for consistent whole-house 360° panorama synthesis from 2D floorplans and style references using a floorplan 3D shell proxy and dynamic 3D Gaussian Splatting (3DGS) spatial memory.

May 26, 2026
research preview
Academic/Research

Pantheon 360

Developed by USC, NYCU, Cornell, and Bosch Research (CVPR 2026), Pantheon 360 is a 3D-aware 360° video diffusion framework that utilizes an explicit 3D Cache to synthesize geometrically consistent panoramic videos and digital twins.

May 26, 2026
research preview
Alibaba

StreamChar

StreamChar is a real-time, long-horizon streaming framework for generating synchronized audio and video of talking characters from text transcripts developed by Alibaba Tongyi Lab (HumanAIGC Team). It decouples long-horizon orchestration from short-window audio-video denoising using a Joint Audio-Video Diffusion Transformer (DiT).

May 25, 2026
research preview
ByteDance Research

Bernini

Bernini is a unified framework for video generation and editing developed by ByteDance Research. It decouples high-level semantic planning (Qwen2.5-VL-7B MLLM) from pixel rendering (Wan2.2-A14B DiT), delivering state-of-the-art instruction-following video generation and editing with minimal spatial-temporal drift.

May 24, 2026
open weights
Academic/Research

Lance

Lance is a lightweight, 3B-active-parameter native unified multimodal model developed by ByteDance. It handles image and video understanding, generation, and editing within a single dual-stream Mixture-of-Experts (MoE) architecture trained via multi-task synergy.

May 24, 2026
research preview
Academic/Research

GenRecon

GenRecon (Bridging Generative Priors for Multi-View 3D Scene Reconstruction) is a 3D vision framework developed by researchers at Technical University of Munich (TUM). It reformulates multi-view 3D scene reconstruction as conditional 3D generation over overlapping spatial chunks, leveraging Trellis.2 generative priors to produce complete, editable PBR-ready meshes from sparse RGB images.

May 22, 2026
research preview
Academic/Research

Scope

SCOPE (Simulating Cross-game Operations in Playable Environments) is an interactive real-time FPS world model developed by University of Chinese Academy of Sciences (UCAS). It incorporates a Spatial Action Decoupling module into video diffusion transformer blocks (Wan2.2) to achieve per-pixel temporal action conditioning, separating localized weapon/HUD effects from global environment rendering.

May 22, 2026
open weights
NVIDIA

PiD

PiD (Pixel Diffusion Decoder) is a generative model and decoding paradigm developed by NVIDIA Spatial Intelligence Lab (SIL). It reformulates latent-to-pixel decoding as a conditional pixel-space diffusion process, replacing traditional VAE decoders to unify decoding and 4x/8x spatial super-resolution into a fast 4-step distilled diffusion process.

May 22, 2026
open weights
Academic/Research

FashionChameleon

FashionChameleon is a real-time and interactive human-garment video customization framework developed by Xiamen University, Zhejiang University, and Alibaba Group. Built on the Wan2.2-TI2V-5B backbone, it achieves generation speeds of 23.8 FPS on a single GPU using streaming distillation and training-free KV-cache rescheduling.

May 20, 2026
research preview
Academic/Research

Flash GRPO

Flash-GRPO is an efficient one-step policy optimization alignment framework for video diffusion models developed by Zhejiang University. It solves standard GRPO computational bottlenecks via Iso-Temporal Grouping and Temporal Gradient Rectification, achieving 6x training acceleration.

May 20, 2026
research preview
NVIDIA

PhysX Omni

PhysX Omni is a unified generative framework developed by S-Lab, NTU in collaboration with ACE Robotics and NVIDIA. It creates simulation-ready 3D assets (rigid, deformable, articulated) with complete physical attributes (scale, mass, material stiffness, joint kinematics) directly from images or text for physics simulators like MuJoCo and Isaac Sim.

May 20, 2026
open weights
Academic/Research

CogOmniControl

CogOmniControl is a reasoning-driven framework for controllable video generation developed by University of Macau. It factorizes generation into creative intent cognition (CogVLM) and in-context video diffusion (CogOmniDiT), optimized via RL and Best-of-N candidate selection.

May 19, 2026
research preview
Academic/Research

MegaASR

MegaASR is a 1.7B foundation model for robust 'in-the-wild' speech recognition, developed by Tsinghua University, NTU, NUS, and Shanghai AI Lab. Built on Qwen3-ASR, it utilizes progressive acoustic-to-semantic SFT (A2S-SFT) and dual-granularity policy optimization (DG-WGPO) trained on 2.6M simulated samples.

May 19, 2026
research preview
Google DeepMind

Gemini 3.5 Flash

Google's most intelligent model for sustained frontier performance on agentic and coding tasks, optimized for fast agent loops and multi-step workflows.

May 19, 2026
api only
Academic/Research

PixlRelight

PIXLRelight is a feed-forward, transformer-based neural rendering framework for physically controllable single-image relighting. Developed at the University of Oxford, it bridges physically based rendering (PBR) and learned image synthesis via shared intrinsic conditioning.

May 18, 2026
research preview
Academic/Research

ControlLight

ControlLight is a controllable, consistent, and generalizable low-light image enhancement framework built as a LoRA on FLUX.2-klein-9B. It enables continuous illumination adjustment while preserving scene structure and fine-grained visual details.

May 15, 2026
research preview
Academic/Research

NAVA

NAVA (Native Audio-Visual Alignment for Generation) is a 6.3-billion parameter multimodal diffusion model developed by Baidu ERNIE Research. Utilizing an Align-then-Fuse MMDiT architecture, NAVA generates temporally synchronized and semantically coherent 720p video and stereo audio natively in a single unified pass.

May 15, 2026
research preview
Alibaba

Qwen Live Translate

Qwen Live Translate (Qwen3.5-LiveTranslate-Flash) is Alibaba's real-time multimodal simultaneous speech-to-speech and speech-to-text translation model family built on the Qwen-Omni architecture. It fuses audio, text, and visual input (lip movements and facial context) for low-latency interpretation and voice cloning.

May 15, 2026
open weights
NVIDIA

Deja View

Déjà View (DVLT - Déjà View Looping Transformer) is a parameter-efficient model for multi-view 3D reconstruction developed by NVIDIA Dynamic Vision and Learning (DVL) group, ETH Zürich, and U of Toronto. Employing a single weight-tied transformer block recurrently over K refinement steps, DVLT jointly predicts per-pixel depth, ray maps, and camera parameters at ~117M parameters.

May 15, 2026
open weights
NVIDIA

Gamma World

Gamma-World (γ-World) is a generative multi-agent world model developed by NVIDIA Spatial Intelligence Lab (SIL) and Toronto AI Lab. Unlike single-agent world models, Gamma-World models complex multi-agent environments where multiple independently controlled agents interact in real-time at 24 FPS with zero-shot generalization from 2 to 4+ agents.

May 15, 2026
open weights
NVIDIA

LocateAnything

LocateAnything is a high-speed, high-accuracy vision-language model developed by NVIDIA Learning and Perception Research (LPR) Lab for universal visual grounding and spatial localization. It introduces Parallel Box Decoding (PBD) to predict bounding box coordinates atomically in a single forward pass, eliminating sequential coordinate token decoding bottlenecks to achieve up to 10x higher throughput.

May 15, 2026
open weights
Academic/Research

ReactiveGWM

ReactiveGWM (Reactive Game World Models) is an interactive game world modeling framework developed by National University of Singapore (NUS) in collaboration with Tencent, HKPolyU, and HKUST-GZ. It decouples player action control from NPC autonomy in video generation, enabling steerable, prompt-aligned NPC behavior and zero-shot strategy transfer across different games.

May 14, 2026
open weights
Academic/Research

L2P

L2P is a undisclosed-parameter text generation model developed by Academic/Research. Released on 2026-05-03 with 86 likes and 0 downloads on Hugging Face.

May 3, 2026
research preview
Ege Orcun

Lucida

Lucida is a BiRefNet-based image matting model fine-tuned specifically for the cases where general-purpose background removers fail: semi-transparent objects, camouflaged subjects, logos and typography with soft shadows, glow/VFX effects, illustrations, and print-style designs (stickers, tees). Trained across 9 specialized categories (camouflage, transparent, complex, thin, hair, text, fx, illustration, design) on ~52,882 image/alpha pairs, Lucida v7 achieves the best overall MAE (0.0257) across all models measured — including commercial references — while maintaining MIT licensing.

May 1, 2026
open weights

April 2026

March 2026

February 2026

January 2026

December 2025

November 2025

October 2025

September 2025

August 2025

July 2025

May 2025

March 2025

February 2025

January 2025

December 2024

November 2024

October 2024

September 2024

August 2024

July 2024

June 2024

May 2024

April 2024

March 2024

February 2024

January 2024

December 2023

November 2023

October 2023

September 2023

August 2023

July 2023

June 2023

March 2023

December 2022

September 2022

April 2022

April 2020