Back to Newsroom

Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM

By MarkTechPost / Modelverse Editorial·July 26, 2026·3 min read
Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM

NewsHub NewsHub Premium Content Read our exclusive articles FacebookInstagramX Home Open Source/Weights AI Agents Tutorials Voice AI Robotics Newsletter → Partner with Us NewsHub Search Home Open Source/Weights AI Agents Tutorials Voice AI Robotics Newsletter → Partner with Us Home Editors Pick Agentic AI Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and... Editors PickAgentic AITechnologyAI ShortsArtificial IntelligenceApplicationsLanguage ModelLarge Language ModelMachine LearningNew ReleasesSecuritySoftware EngineeringStaffTech News Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM By Asif Razzaq - July 25, 2026 Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family. It is not just a new frontier model. It is a third endpoint on the Fugu orchestrator, tuned for security reasoning. Sakana launched that orchestrator a month earlier.

Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM. It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Preview.

The two evaluations sit at opposite ends of a security workflow:

Together the pair spans 'find and prove the bug' and 'turn intel into a detection.' That framing is the most defensible part of Sakana's announcement.

When the CyberGym researchers published their first results, the best agent-model pairing reached roughly 20%. Anthropic reported 83.1% for Claude Mythos Preview under Project Glasswing in April 2026. OpenAI reported 85.6% for its updated GPT-5.5-Cyber, against 81.8% for GPT-5.5. Sakana's 86.9% is therefore a small step past the reported frontier, not a jump.

CTI-REALM is a different story. Microsoft's own evaluation put the top three configurations, all Claude, in a band from 0.624 to 0.685. Fugu-Cyber's 72.1% would sit above that band. One caveat matters. CTI-REALM is scored as a trajectory reward between 0 and 1. It is not a pass/fail rate. Sakana calls it a success rate anyway.

Fugu is itself a language model. It is trained to read a query and build an agentic scaffold on the fly. It then delegates sub-tasks to specialist models in a pool.

The approach is documented in the Fugu technical report and two ICLR 2026 papers, TRINITY and the Conductor. TRINITY assigns Thinker, Worker, and Verifier roles across multiple LLMs. The Conductor learns natural-language coordination strategies through reinforcement learning.

Official Announcement

Read the full update directly from the official source at MarkTechPost News.

Stay tuned to Modelverse for real-time model analysis and benchmark coverage.

ai-newsbreakingmarktechpost

Related content

Are brain waves the next unlock for physical AI?

Forget YouTube videos—frontier physical AI models need multiple camera angles, dense annotation, and soon, brain wave readings.

Read article

Ben Bernanke

Official announcement from Anthropic: Ben Bernanke

Read article

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government and internet servic...

Read article
© 2026 Modelverse®. All rights reserved.Modelverse Newsroom