Back to Newsroom

Weekly Model Recap: 27 Frontier & Open-Weight Releases (2026-08-02)

By Modelverse Editorial·August 2, 2026·8 min read
Weekly Model Recap: 27 Frontier & Open-Weight Releases (2026-08-02)

Welcome to the Modelverse Weekly Intelligence Digest for 2026-08-02.

Over the past week, the global AI ecosystem experienced rapid architectural iteration, with 27 new foundation, open-weight, and domain-specialized models indexed in the Modelverse repository.

Executive Summary & Macro Trends:

  • Test-Time Compute & Long Context Standardized: Closed frontier labs (Anthropic's Claude Opus 5, Moonshot AI's Kimi K3) are prioritizing adaptive reasoning budgets alongside 1M+ token context windows.
  • Video World Models & Spatial Motion: Open-weights research saw major breakthroughs in skeleton-free motion transfer (Motion4Motion) and mobile real-time video diffusion (MobileWan).
  • Domain Specialization Expansion: Dedicated cybersecurity (Gemini 3.5 Flash Cyber) and healthcare vision-language models (MedGemma 1.5 4B) are outperforming general-purpose LLMs on specialized benchmarks.

🏆 Top Model Spotlights of the Week

🌟 Motion4Motion

Developer: Tsinghua University & StepFun • Task: video-generationParameters: undisclosedAccess: research-preview

Motion4Motion is a training-free framework for cross-subject motion transfer presented at SIGGRAPH 2026. Unlike prior skeleton-based approaches that only work for human-like characters with predefined skeleton topologies, Motion4Motion steps out of the skeleton framework entirely. It models the motion flow of a subject in a video rather than its skeleton, enabling motion transfer across entirely different species (e.g., human-to-animal) at inference time without any task-specific training or fine-tuning.

Key Technical Innovations:

  • Training-Free: No fine-tuning or task-specific training required — runs at inference time on any video pair
  • Skeleton-Free Motion Transfer: Models subject motion flow instead of skeleton topology, enabling transfer across different species and body types
  • Cross-Species Transfer: Works across diverse morphologies — human, animal, and non-human character motion transfer
  • TransPE (Transferring Positional Encoding): Novel technique for mapping motion patterns between subjects at different spatial scales

👉 Explore Full Specs & Benchmarks for Motion4Motion


🌟 Kimi K3

Developer: Moonshot AI • Task: chat-reasoningParameters: 2.8TAccess: open-source

Kimi K3 is Kimi's flagship 2.8T open-source model built on Kimi Delta Attention (KDA) and Attention Residuals. It features native visual understanding for images and video, a 1M-token context window with automatic caching, and an always-on thinking mode designed for frontier intelligence scenarios like long-horizon coding and knowledge work.

Key Technical Innovations:

  • Max Output: 131K tokens
  • Built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes)
  • Stable LatentMoE framework efficiently activates 16 out of 896 experts
  • 1M-token context window with automatic context caching

👉 Explore Full Specs & Benchmarks for Kimi K3


🌟 Claude Opus 5

Developer: Anthropic • Task: chat-reasoningParameters: undisclosedAccess: api-only

For complex agentic coding and enterprise work. Claude Opus 5 is a step-change improvement over Claude Opus 4.8, with the largest gains in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.

Key Technical Innovations:

  • Max Output: 128K tokens
  • 1M token context window (default and maximum)
  • Thinking on by default with adjustable effort levels (low to max)
  • Mid-conversation tool changes while preserving prompt cache

👉 Explore Full Specs & Benchmarks for Claude Opus 5


🌟 Gemini 3.5 Flash Cyber

Developer: Google DeepMind • Task: code-generationParameters: undisclosedAccess: closed-source

Google DeepMind's specialized cybersecurity model, built to help defenders find, validate, and patch software vulnerabilities quickly and efficiently. Gemini 3.5 Flash Cyber is fine-tuned from the Gemini 3.5 Flash base model with deep expertise in vulnerability research, security analysis, exploit validation, and patch generation. It enables security teams to dramatically accelerate their defensive workflows compared to general-purpose LLMs.

Key Technical Innovations:

  • Specialized for cybersecurity: vulnerability detection, validation, and patching
  • Designed to help defenders — not exploit — with a safety-first approach
  • Fine-tuned from Gemini 3.5 Flash with deep security domain expertise
  • Accelerates security workflows: from vulnerability identification to patch generation

👉 Explore Full Specs & Benchmarks for Gemini 3.5 Flash Cyber


🌟 Gemini 3.6 Flash

Developer: Google DeepMind • Task: chat-reasoningParameters: undisclosedAccess: closed-source

Google DeepMind's fastest frontier-class model, featuring built-in reasoning, native multimodal input (text, image, audio, video), and a 1M token context window. Ranked #1 in speed among all 186 models on Artificial Analysis at 275.5 tokens/second while scoring 50 on the Artificial Analysis Intelligence Index — placing it well above average among comparable models. Priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens with a 90% cache discount.

Key Technical Innovations:

  • Max Output: 66K tokens
  • Ranked #1 fastest model among 186 models at 275.5 output tokens/second (Artificial Analysis)
  • Intelligence Index score of 50, above the 31-point average for comparable models
  • Built-in hybrid reasoning mode with extended thinking capabilities

👉 Explore Full Specs & Benchmarks for Gemini 3.6 Flash


🌟 Qwen3.8 Max Preview

Developer: Alibaba • Task: chat-reasoningParameters: undisclosedAccess: api-only

Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows

Key Technical Innovations:

  • Max Output: 131K tokens
  • Native reasoning capability
  • Tool / function calling support

👉 Explore Full Specs & Benchmarks for Qwen3.8 Max Preview


🌟 1. Frontier Foundation & Reasoning Breakthroughs

Frontier laboratories continue to push the boundaries of reasoning performance, extended context caching, and agentic code generation.

ModelDeveloperPrimary TaskModalityParametersStatus
Claude Opus 5Anthropicchat-reasoningtext, image, codeundisclosedapi-only
Gemini 3.5 Flash CyberGoogle DeepMindcode-generationtext, codeundisclosedclosed-source
🚀 Claude Opus 5
  • Developer: Anthropic
  • Summary: For complex agentic coding and enterprise work. Claude Opus 5 is a step-change improvement over Claude Opus 4.8, with the largest gains in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.
  • Highlights: Max Output: 128K tokens. 1M token context window (default and maximum)
🚀 Gemini 3.5 Flash Cyber
  • Developer: Google DeepMind
  • Summary: Google DeepMind's specialized cybersecurity model, built to help defenders find, validate, and patch software vulnerabilities quickly and efficiently. Gemini 3.5 Flash Cyber is fine-tuned from the Gemini 3.5 Flash base model with deep expertise in vulnerability research, security analysis, exploit validation, and patch generation. It enables security teams to dramatically accelerate their defensive workflows compared to general-purpose LLMs.
  • Highlights: Specialized for cybersecurity: vulnerability detection, validation, and patching. Designed to help defenders — not exploit — with a safety-first approach

🎥 2. Generative Vision, Video & Spatial World Models

Generative video models are rapidly advancing toward real-time temporal consistency, low-latency mobile inference, and skeleton-free motion control.

ModelDeveloperPrimary TaskModalityParametersStatus
Motion4MotionTsinghua University & StepFunvideo-generationvideoundisclosedresearch-preview
Kimi K3Moonshot AIchat-reasoningtext, image, video2.8Topen-source
ARDYNVIDIAchat-reasoningtext, motionundisclosedopen-weights
MobileWanQualcommvideo-generationtext, videoundisclosedresearch-preview
Gemini 3.5 Flash-LiteGoogle DeepMindchat-reasoningtext, image, audio, videoundisclosedclosed-source
Gemini 3.6 FlashGoogle DeepMindchat-reasoningtext, image, audio, videoundisclosedclosed-source
Qwen3.8 Max PreviewAlibabachat-reasoningtext, image, videoundisclosedapi-only
🎥 Motion4Motion
  • Developer: Tsinghua University & StepFun
  • Summary: Motion4Motion is a training-free framework for cross-subject motion transfer presented at SIGGRAPH 2026. Unlike prior skeleton-based approaches that only work for human-like characters with predefined skeleton topologies, Motion4Motion steps out of the skeleton framework entirely. It models the motion flow of a subject in a video rather than its skeleton, enabling motion transfer across entirely different species (e.g., human-to-animal) at inference time without any task-specific training or fine-tuning.
  • Highlights: Training-Free: No fine-tuning or task-specific training required — runs at inference time on any video pair. Skeleton-Free Motion Transfer: Models subject motion flow instead of skeleton topology, enabling transfer across different species and body types
🎥 Kimi K3
  • Developer: Moonshot AI
  • Summary: Kimi K3 is Kimi's flagship 2.8T open-source model built on Kimi Delta Attention (KDA) and Attention Residuals. It features native visual understanding for images and video, a 1M-token context window with automatic caching, and an always-on thinking mode designed for frontier intelligence scenarios like long-horizon coding and knowledge work.
  • Highlights: Max Output: 131K tokens. Built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes)
🎥 ARDY
  • Developer: NVIDIA
  • Summary: ARDY is an autoregressive diffusion model for interactive human motion generation with online text prompting and flexible kinematic constraints.
  • Highlights: Interactive human motion generation. Autoregressive diffusion
🎥 MobileWan
  • Developer: Qualcomm
  • Summary: MobileWan is a mobile video diffusion model developed by Qualcomm AI Research that aims to close the quality gap for on-device video generation.
  • Highlights: On-device video generation. Mobile-optimized diffusion model
🎥 Gemini 3.5 Flash-Lite
  • Developer: Google DeepMind
  • Summary: Google DeepMind's most cost-efficient model in the Gemini 3.5 family, designed for high-volume, low-latency applications where economy is paramount. Gemini 3.5 Flash-Lite delivers strong performance for everyday tasks at the lowest price point of the new July 2026 Gemini release batch, while retaining multimodal input capabilities and the 1M token context window.
  • Highlights: Max Output: 66K tokens. Most cost-efficient model in the new Gemini batch — lowest price per token
🎥 Gemini 3.6 Flash
  • Developer: Google DeepMind
  • Summary: Google DeepMind's fastest frontier-class model, featuring built-in reasoning, native multimodal input (text, image, audio, video), and a 1M token context window. Ranked #1 in speed among all 186 models on Artificial Analysis at 275.5 tokens/second while scoring 50 on the Artificial Analysis Intelligence Index — placing it well above average among comparable models. Priced at $1.50 per 1M input tokens and $7.50 per 1M output tokens with a 90% cache discount.
  • Highlights: Max Output: 66K tokens. Ranked #1 fastest model among 186 models at 275.5 output tokens/second (Artificial Analysis)
🎥 Qwen3.8 Max Preview
  • Developer: Alibaba
  • Summary: Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows
  • Highlights: Max Output: 131K tokens. Native reasoning capability

⚡ 3. Open-Weight & High-Efficiency Foundation Models

Open-weight foundation models offer strong reasoning capabilities, privacy guarantees, and local deployment efficiency across edge and cloud infrastructure.

ModelDeveloperPrimary TaskModalityParametersStatus
DeepSeek-V4-Flash-0731deepseek-aichat-reasoningtext304.2Bopen-weights
Inkling-SmallThinking Machinesmultimodal-generaltext, image266.0Bopen-weights
Mage-VLMicrosoftmultimodal-generaltext, image4.7Bopen-weights
KAT-Coder-V2.5-DevKwaipilotchat-reasoningtext34.7Bopen-weights
Solar-Open2-250Bupstagechat-reasoningtext250.3Bopen-weights
XYZ-Aquila-miniXYZAILabchat-reasoningtext35.1Bopen-weights
antares-1bfdtn-aichat-reasoningtext1.8Bopen-weights
Mage-Flow-Edit-TurboMicrosoftmultimodal-generaltext4.1Bopen-weights
Mage-FlowMicrosoftimage-generationtext, image4.1Bopen-weights
Nanbeige4.2-3BNanbeigechat-reasoningtext4.2Bopen-weights
Motif-3-BetaMotif-Technologieschat-reasoningtext314.8Bopen-weights
DeepSeek-V4-Flash-0731
  • Developer: deepseek-ai
  • Summary: DeepSeek-V4-Flash-0731 is a 304.2B-parameter Mixture-of-Experts (MoE) text generation model developed by deepseek-ai. Built on the DeepseekV4ForCausalLM architecture using transformers. Released on 2026-07-31 with 301 likes and 0 downloads on Hugging Face.
  • Highlights: 304.2B parameters with sparse MoE architecture for efficient inference. Built on DeepseekV4ForCausalLM architecture (transformers)
Inkling-Small
  • Developer: Thinking Machines
  • Summary: Inkling-Small is a 266.0B-parameter Mixture-of-Experts (MoE) image text to text model developed by thinkingmachines. Built on the InklingForConditionalGeneration architecture using transformers. Released on 2026-07-27 with 160 likes and 840 downloads on Hugging Face.
  • Highlights: 266.0B parameters with sparse MoE architecture for efficient inference. Built on InklingForConditionalGeneration architecture (transformers)
Mage-VL
  • Developer: Microsoft
  • Summary: Mage-VL is a 4.7B-parameter image text to text model developed by microsoft. Built on the MageVLForConditionalGeneration architecture using transformers. Released on 2026-07-25 with 121 likes and 2,951 downloads on Hugging Face.
  • Highlights: 4.7B parameters. Built on MageVLForConditionalGeneration architecture (transformers)
KAT-Coder-V2.5-Dev
  • Developer: Kwaipilot
  • Summary: KAT-Coder-V2.5-Dev is a 34.7B-parameter Mixture-of-Experts (MoE) text generation model developed by Kwaipilot. Built on the Qwen3_5MoeForConditionalGeneration architecture using transformers. Supports en, zh language(s). Released on 2026-07-23 with 126 likes and 396 downloads on Hugging Face.
  • Highlights: 34.7B parameters with sparse MoE architecture for efficient inference. Built on Qwen3_5MoeForConditionalGeneration architecture (transformers)
Solar-Open2-250B
  • Developer: upstage
  • Summary: Solar-Open2-250B is a 250.3B-parameter Mixture-of-Experts (MoE) text generation model developed by upstage. Built on the SolarOpen2ForCausalLM architecture using transformers. Supports en, ko, ja language(s). Released on 2026-07-22 with 543 likes and 1,106 downloads on Hugging Face.
  • Highlights: 250.3B parameters with sparse MoE architecture for efficient inference. Built on SolarOpen2ForCausalLM architecture (transformers)
XYZ-Aquila-mini
  • Developer: XYZAILab
  • Summary: XYZ-Aquila-mini is a 35.1B-parameter text generation model developed by XYZAILab. Built on the Qwen3_5MoeForConditionalGeneration architecture using transformers. Released on 2026-07-22 with 344 likes and 447 downloads on Hugging Face.
  • Highlights: 35.1B parameters. Built on Qwen3_5MoeForConditionalGeneration architecture (transformers)

🛠️ 4. Specialized Domain Models (Cyber, Medical, Code)

Domain-specific fine-tuning delivers state-of-the-art accuracy in security vulnerability detection, clinical medical comprehension, and interactive robotics.

ModelDeveloperPrimary TaskModalityParametersStatus
Audio8-TTS-Preview-0.6bAudio8audio-speechtext, audio601Mopen-weights
VibeVoice-ASR-BitNetMicrosoftaudio-speechaudio, text323Mopen-weights
🛠️ Audio8-TTS-Preview-0.6b
  • Developer: Audio8
  • Summary: Audio8-TTS-Preview-0.6b is a 601M-parameter text to speech model developed by Audio8. Built on the ArkttsModel architecture using transformers. Supports yue, zh, nl, en, fr, de, it, ja, ko, pl, es language(s). Released on 2026-07-28 with 126 likes and 225 downloads on Hugging Face.
  • Highlights: 601M parameters. Built on ArkttsModel architecture (transformers)
🛠️ VibeVoice-ASR-BitNet
  • Developer: Microsoft
  • Summary: VibeVoice-ASR-BitNet is a 323M-parameter automatic speech recognition model developed by microsoft. Built on the VibeVoiceForASRTraining architecture using ggml. Supports en, zh, fr, it, ko, pt, vi language(s). Released on 2026-07-24 with 115 likes and 3,864 downloads on Hugging Face.
  • Highlights: 323M parameters. Built on VibeVoiceForASRTraining architecture (ggml)

📦 5. Community Open-Source Fine-tunes & Quantizations

The Hugging Face open-source community continues to publish optimized GGUF quantizations, NVFP4 low-bit representations, and specialized instruction-tuned fusions.

ModelDeveloperPrimary TaskModalityParametersStatus
DeepSeek-V4-Flash-0731-GGUFunslothchat-reasoningtext284B parametersopen-weights
Kimi-K3-GGUFunslothmultimodal-generaltext, imageundisclosedopen-weights
Solar-Open2-250B-Nota-NVFP4nota-aichat-reasoningtext144.6Bopen-weights
Laguna-S-2.1-GGUFunslothchat-reasoningtextundisclosedopen-weights
GLM-5.2-Vision-NVFP4basetenmultimodal-generaltext, image381.0Bopen-weights

📊 Strategic Outlook & Key Takeaways

  1. Enterprise Deployment: Production teams are increasingly pairing fast Flash/Lite tier models for high-throughput routing with heavy reasoning models for complex multi-step execution.
  2. Open-Weights Parity: The gap between closed API benchmarks and open-weight models continues to narrow, with open models matching previous-generation frontier baselines.
  3. Multimodal Native Architecture: Audio, vision, and text are no longer separate adapters; newly designed models incorporate unified multimodal tokenizers natively.

Stay up to date with real-time AI benchmarks, model comparisons, and release tracking at Modelverse.

weekly-digestmodel-releasesai-roundup

Footnotes & Primary References

Related content

Disrupting a Criminal Scam Operation

OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.

Read article

End-to-End Forecasting with TimesFM 2.5: Backtesting, Covariates, Anomaly Detection, and Scalable Colab Deployment

In this tutorial, we build an advanced end-to-end time-series forecasting workflow with TimesFM 2.5. We begin by configuring the runtime, installing the required dependencies, dete...

Read article

NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework

Agentic RL research is constant algorithm modification, and in mainstream frameworks every change threads through trainer, distributed backend, and rollout glue. NVIDIA's Molt targ...

Read article