Anthropic CEO Dario Amodei. Photo: Ludovic MARIN / AFP via Getty Images

Is this the world’s most dangerous AI model?

Anthropic's latest frontier model is so powerful that the company is refusing to release it to the public

If companies in the West can produce technology this powerful, the Chinese Communist Party can't be far behind

Advanced AI poses existential problems for policymakers, businesses and even families

Anthropic CEO Dario Amodei. Photo: Ludovic MARIN / AFP via Getty Images

Share this article

You will all know this experience. Sitting in a job interview, you are asked your biggest weakness. Desperate to impress, you used a humble brag; I’m a bit of a perfectionist, I get so invested in success. Something similar may have happened in the world of artificial intelligence. Anthropic, the company behind the large language model (LLM) Claude, has announced that its newest frontier model, Mythos, is so powerful and world-changing that they are actually not going to release it to the general public. 

This may just be a humble brag, designed to talk up its share price ahead of a rumoured IPO later this year. But what if it’s not? What if the claims made about Mythos are broadly right, and in a few short years, these technologies have evolved from fun, to useful, to potentially world-changing?

Let’s start with the claims. Anthropic say that Mythos is so good at code that it has found security vulnerabilities in many real-world systems, including a 27-year-old operating system called OpenBSD that has been repeatedly proven as secure. Terrifyingly, these security-shattering cyber capabilities weren’t explicitly trained into Mythos; they emerged as a downstream consequence of general improvements in coding ability.

This has led Anthropic to its humble brag announcement that it will not release its model, citing fears of the damage it could wreak on the world’s security architecture. Instead, it has given early limited access to leading companies, including Apple, Google, JP Morgan (along with officials across the US government) in its Project Glasswing. This unlikely alliance is designed to help use Mythos to shore up the existing security systems, now that the prospect of a super-powerful AI is closer than ever.

This announcement, if true and not merely marketing froth, presents many very serious concerns, for tech, government and society more broadly.

Firstly, if Anthropic can make something this good, OpenAI and others may not be far behind. This means that the Chinese Communist Party may not be far behind them. We saw the extensive use of insanely advanced capabilities wielded by the United States military in Iran recently, thanks to tools developed by the ever-impressive Palantir, integrating Anthropic’s Claude AI model. What, then, does the advent of such powerful AI augur for future conflict, whether the current wars in Ukraine and the Gulf, or for any future showdown between the US and China?

Secondly, what does this mean for the so-called ‘alignment’ problem? Alignment is AI techbro speak for making sure that these advanced technologies have desires and objectives that are aligned with mankind’s own. We have seen plenty of sci-fi movies where alignment goes wrong. But Anthropic say that this is their most aligned model ever, and yet still believe it to be the most dangerous.

In its own ‘system card’ (AI lab speak for an explainer and transparency document that sets out what the model can do and so on), Anthropic compares Mythos to a skilled climbing instructor who poses a greater risk than a novice instructor precisely because he is so skilled and may therefore lead them towards the most dangerous climbs. If you want another scare, search for the word ‘sandwich’ in the card.

But it also poses existential problems for policymakers, businesses and even families who rely on the security built into their bank accounts and Ring doorbells. How many politicians this week have noticed this potentially seismic change to the world’s security assumptions? How many instead were doomscrolling X to find out about a conflict they have no hope of influencing (including, of course, Keir Starmer)?

One bright spot is that Anthropic’s ongoing dispute with the Pentagon, which has seen them blacklisted by America’s newly-renamed Department of War, may present an opening for the UK. Pre-Mythos, backroom efforts were already underway to encourage further expansion of the AI firm’s London office, and even a possible dual US-UK listing for that much-anticipated IPO.

Politicians, civil servants, spooks, businessmen and indeed all of us should consider the impact on our lives of what Mythos could do, let alone everything that is to come. If cryptography is about to become less secure, what impact will that have on bank accounts, energy grids, water plants, WhatsApp message privacy or even the UK’s nuclear deterrent?

You may recall that three years ago, HM Treasury advertised an ‘experienced’ role as head of cyber security for the whole Treasury, with an annual salary of just £57,000. Imagine being in that job interview and being asked for your biggest weakness, knowing there are weaknesses everywhere, that only AI can see. 

Share this article

Written by

James Price is a senior fellow at the Adam Smith Institute.

CapX depends on the generosity of its readers.

If you value what we do, please consider making a donation.

Amount
Period

Your message has not been sent.