From too Dangerous for to Available to the Public in a few Months

On April 7, 2026, Anthropic announced project Glasswing: “we’ve observed [capabilities] in a new frontier model trained by Anthropic that we believe could reshape cybersecurity.” The aim of the project was to give select companies access to Claude Mythos, originally a general-purpose LLM that was later developed to specialize in cybersecurity. At the time, many sources, such as MindStudio, Futurism, and myself, claimed Mythos was too dangerous to release to the public. Mythos set a new standard for top-end AI models: be on par with or better than the best humans in the world.

Two months later, on June 9th, Anthropic announced the release of Claude Fable 5 and Mythos 5. The former is the publicly available model with safeguards to redirect to Claude Opus 4.8 (Anthropic’s most capable public model after Fable 5) if queries on some topics (chiefly cyberattacks) are mentioned. The latter is the same underlying model as Fable 5, but has fewer safeguards as the public model and is only accessible to companies within Project Glasswing, Anthropic’s project to help companies like Microsoft, AWS, and Apple, among others, harden their cyber systems. By releasing Mythos, and more importantly Fable, Anthropic implicitly said they can control a tool capable of causing extreme harm in the form of cyberattacks. Many people have doubted their abilities given that in less than a week, Amazon researchers had already found a way past Fable’s safeguards.

Illustration of two suited executives, pinned with Amazon and Apple logos, each signing a shared rulebook across a conference table, flanked by a glowing molten orb sealed in a glass case labeled "Too Dangerous," beneath a U.S. government seal.
A model too dangerous for the public to touch! now the companies closest to it are the ones writing the rules for what ‘dangerous’ means.

Three days after Fable’s public launch, on June 12th, Anthropic announced they had received a directive from the US government to suspend access to Fable 5 and Mythos 5 to all foreign nationals (including Anthropic employees), citing national security concerns. According to the Wall Street Journal, the ban came following a conversation between Amazon CEO Andy Jassy and Treasury Secretary Scott Bessent. Jassy relayed that researchers at Amazon, a large investor in Anthropic, were able to use a series of prompts to get Fable to provide information which could potentially be useful to cyber attackers and was supposed to be off limits for the model. Since Anthropic can’t easily verify which users are US nationals or foreign, they decided to suspend access to Fable 5 and Mythos 5 to all users.Most recently, on June 30th, Anthropic announced they would be redeploying Fable 5 the next day. They stated they trained a new classifier for detecting behaviors identified in the Amazon report which they say can block over 99% of attempts to circumvent Fable’s safeguards. In the rare event Fable does provide information to a cyberattacker, it will not be detailed enough to help them. Anthropic states that the increased security comes at the cost of increased false positives from the new classifier, meaning if someone is requesting something non-malicious from Fable, there is a chance the request still gets delegated to Opus 4.8. Anthropic plans to further refine the classifier.

I question how long until Fable 5 (or now similarly skilled models like GPT-5.6 or Kimi K3) has its safeguards overcome. Anthropic themselves acknowledge that no defensive tool is going to be 100% accurate: “Any method of scoring jailbreaks will be imperfect.” Eventually, one of these models will be utilized in a cyberattack. This will respark intense conversations around AI safeguards and the release of Fable 5 will likely be a moment pointed to as an ‘I told you so’ moment for proponents of the toughest regulations.

Anthropic is making an effort to not have their safeguards overcome with their HackerOne program. HackerOne is a cybersecurity company dedicated to using AI and human researchers to find vulnerabilities in cyber systems. Anthropic has a page on HackerOne for hackers to report any vulnerabilities they find with Fable 5’s safeguards. Though the page is listed as having a 100% response rate, the page also has 20 reports within the last 90 days and 0 resolved (as of this writing).

Beyond Fable 5

On June 26, 2026, OpenAI announced it would be releasing ChatGPT 5.6 preview to a small group of trusted partners. The GPT-5.6 line includes three models: Sol, Terra, and Luna—the flagship model, the balanced everyday model, and the efficient model. Notably, they claim GPT-5.6 Sol in “ultra” mode outperforms Claude Mythos 5 on every demonstrated benchmark test.

As part of their cooperation with Trump’s AI executive order, OpenAI only released GPT-5.6 to trusted partners before launching the model to the public. The executive order asks companies, on a voluntary basis, to provide early access of frontier models to the government and allow the government to select trusted partners to receive early access. The executive order does not create any regulations, agencies, or obligations for companies developing frontier models. Though they are not required to, it is in the AI companies’ best interests to cooperate with the government in order to stay on Trump’s good side.

Despite cooperating, OpenAI opposes the executive order as the long-term norm. In the announcement of the GPT-5.6 preview, they state, “It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them.” The executive order drastically slows the release of frontier models. If the US still had a monopoly on AI, this may be less of a problem, but this executive order could actually be a national security risk itself: China now has several AI labs competing with the leading US companies. Recently, Chinese AI company Moonshot AI has made headlines with Kimi K3 Swarm, a model they claim competes with and even beats Fable 5 and GPT-5.6 Sol in certain benchmarks at a far lower cost. If the US were to slow down its native AI models, they would give China the upper hand in AI development, significantly harming the US’ industry and economy.

On July 9th, OpenAI announced it would be releasing the GPT-5.6 line to the public. The company states that although GPT-5.6 is very cybersecurity, it is more effective at improving defenses than engineering an end-to-end threat. OpenAI takes a similar approach to talking about safeguards for 5.6 as Anthropic does Fable. They state that they have multiple layers of security to ensure minimal misuse, speak about the dual-use conundrum (see my past piece on Mythos), and allow prompts to be retried on lower tier models.

The Rapidly Changing Industry

I often like to say “If anybody tells you they’re an expert on AI for any longer than three months, they’re lying.” The events covered in this blog are exactly why I say that. Within two months a model many said was too dangerous to be given to the public was publicly released, unreleased, and then, rereleased. Within a few weeks of the back and forth, not one, but two new frontier models came out promising similar or greater capabilities. All that’s to say this industry moves so fast that it’s impossible to always be on top of the most important skills regarding AI and the methods to get the most useful outcomes.

I had a conversation with Dan about the constantly changing landscape and he made a very good point: a lot of people are either waiting till tomorrow to learn about AI or trying to learn about tomorrow’s AI because they know it’ll be different then. The problem with the former is that the field keeps changing and if you keep waiting, you’ll never end up learning. The problem with the latter is that nobody completely knows what tomorrow’s AI is going to look like. He said that learning AI today is the best time to do it since you have all the skills, at least for a time, you build an intuition both about how to use AI and how to learn new skills, and some of your skills from today will transfer to tomorrow. For example, as Claude users were introduced to Fable 5, many coding skills had to be relearned, but nearly every other skill like those in document procurement or tone adjustment remained mostly the same. As long as you can stay adaptable, there’s no better time to learn about AI than now, even though we know the landscape will completely change in just a few months. After all, it only took 2 months after Mythos was called too dangerous for the public, and then for those who stood to profit from it most decided what ‘dangerous’ meant going forward.

Read more of Eddie’s investigations into recent corporate intrigue here.

[In this post, we used AI for polish, not purpose.]


Leave a Reply

Your email address will not be published. Required fields are marked *