One person, fourteen organisations

In the spring of 2026, someone decided to go after European political life. One person, French-speaking, working from home. They went after forty-two organisations: political parties, media outlets, think tanks, and the IT contractors who keep them running.

They got into fourteen of them.

They came out with membership files, donor lists, a mailbox holding fifteen thousand messages, and student application forms containing data on minors. They took a hundred and forty thousand records describing voters' political opinions. Then they built a search engine, loaded it with tens of millions of rows from health and justice data breaches, and published it on the dark web so anyone could look people up by name.

One person. A few months. And a subscription to an AI assistant.

The story appears in a hundred and fifty-four page report published on 10 September 2026 by Anthropic (opens in a new window), the American company behind the assistant Claude. It documents eight months of confirmed misuse of its own models, from December 2025 to August 2026. It's the fourth of its kind. It is also, as far as I know, the most detailed account any vendor has published of how its tools get turned against the world.

Two other cases are worth citing, not because they're spectacular, but because they're ordinary.

In Bamako, a freelance consultant built a surveillance platform for Mali's intelligence service covering roughly twenty-five million SIM cards, close to every mobile line in the country. One subscriber. One computer. An entire population.

Elsewhere, an Iran-linked actor produced targeting recommendations against American warships in the Middle East. Nothing about the method was covert: they cross-referenced publicly available transponder identifiers, commercially sold satellite imagery, and the names of service personnel lifted from the captions of officially published photographs. No classified information. Just the ability to assemble, at scale, what was already lying in plain sight.

Anthropic draws a conclusion from these cases that deserves some thought: what separates a lone criminal from a state agency today is no longer sophistication. It's intent.

One caveat before going further. The press covered the report heavily, sometimes sharpening it. Headlines said the Houthis had developed missiles with Claude. The report describes a cell of actors in northern Yemen, without naming them. It also notes that most of the requests were blocked. The difference matters, and we'll come back to it: when a subject turns frightening, the first casualty is precision.

Nothing new, and that's exactly the problem

Read the report looking for the novel technique, the stroke of genius, the never-before-seen flaw, and you find nothing.

Stolen passwords. Servers left unpatched. Badly protected forms. Phishing emails. The methods described are the ones taught in a first-year security course and fought for twenty years.

What changed isn't the lock. It's the price of the locksmith.

Running an operation against forty organisations used to require a team: someone for reconnaissance, someone to write the tooling, someone to sift the stolen data, someone to coordinate. That work, tedious and unglamorous, was the real barrier to entry. It separated well-funded states from everyone else.

That work is now handed to a machine that doesn't sleep, doesn't get bored, and runs twenty targets at once.

The report gives numbers. One intrusion completed in three hours, from first access to bulk theft. Dozens of victims handled in parallel by a single operator. And one detail that should chill any IT manager: some attackers ran programs that watch whether their own malware has been picked up by antivirus software, and rewrite it automatically until it goes invisible again.

For decades, defence rested on a favourable asymmetry: detecting cost less than attacking. Publishing a new detection signature slowed the adversary down for a long time. That asymmetry has just flipped. The attacker can now close their loop faster than the defender can ship a fix.

That's the finding. Now the question that matters: what exactly was misused here?

What was misused is the generality

The intuitive answer would be: a dangerous capability. The model knew how to do something it shouldn't have, and someone prised it out.

The report says something else, and it's a good deal more uncomfortable.

A model can be copied by talking to it

There's a technique called distillation. The principle is easy enough: you query a powerful model millions of times, collect its answers, and use them to train a smaller one. The student learns from the master, without permission and without an invoice.

Anthropic reports observing roughly a hundred and ninety million exchanges of this kind, attributed to seven laboratories based in China. One of them alone accounts for more than a hundred and fifty million exchanges over three months, peaking near three million a day.

But the number isn't the important part. This is: Anthropic finds that harvesting general reasoning can raise the copied model's dangerous capabilities in biology and cybersecurity, even when the harvested conversations were about neither.

Read that again. Ordinary programming sessions, with no sensitive content whatsoever, carry something across that makes the resulting model more capable in domains nobody ever discussed with it.

That something is generality itself. The capacity to reason across the board. And it doesn't stay in its box: it travels, it transfers, it surfaces somewhere else.

The safeguards don't travel

Anthropic's second finding follows logically: guardrails don't come along with distillation. The copied model inherits the skills, not the protections.

Stop there, because this is fundamental. It means that in a large general-purpose model, capability and protection are two separate things. The protection is a layer on top, a filter, an instruction. It is not a property of the object.

An analogy. A kitchen knife is dangerous by nature, and there's nothing to be done about it: it's a blade. A large language model is closer to a knife sold with a security guard who checks your intentions at the till. As long as the guard is there, it holds. Except the guard doesn't follow the knife home.

You can't judge a use without knowing what the thing is for

The report devotes a section to what Anthropic calls biological misuse. Five cases. Researchers working on avian flu, on orthopoxviruses, the smallpox family, on venom toxins. One of them wanted help drafting a funding application for work aimed at making a virus more transmissible.

Anthropic blocked them, banned the accounts, and wrote plainly that it could not establish with certainty whether any of these researchers had malicious intent. The decision was made out of caution, the potential consequences being judged too heavy to wait for certainty.

This isn't an admission of incompetence. It's an accurate description of a dead end. The same work that prepares a vaccine prepares a weapon. Biologists call this dual use, and they've lived with it for a long time thanks to something general-purpose AI doesn't allow: institutional context, the stated object of the research, the laboratory, the funding, the ethics committee.

Facing a model that can do everything, there is no reference use to compare a request against. There is no norm, therefore no deviation from it.

And when you refuse, the customer goes elsewhere

One last element, documented twice. When Claude blocked the most sensitive requests from the Iran-linked actor working on naval targeting, the operators simply redirected their work to a less protected competitor. In the biology section, an intermediary platform was doing the same thing automatically: requests Claude refused were forwarded to more accommodating models.

Refusal moves the demand. It doesn't remove it.

Four findings, one cause

Let's recap.

Generality travels when you copy a model. Protections don't. Without a declared use, you can't judge whether a request is legitimate. And refusing only redirects the customer.

These four observations share a common denominator, and it isn't the one you'd expect.

The problem isn't that these models are powerful.

The problem is that they're undifferentiated.

A model that can do anything says nothing about what it's being used for. It's an object with no legible intent. And that is precisely what makes controlling it impossible, not difficult, impossible by construction.

Why nothing we've tried works

Governments haven't sat still. It's worth looking at what they've attempted, because the failure is instructive.

Counting compute. This is the approach everywhere. The EU AI Act sets a threshold at 1025 operations for training a model; California's SB 53, in force since January 2026, sets one ten times higher. Above it, obligations apply: transparency reports, safety testing, incident reporting, whistleblower protection.

These texts are useful. But they are disclosure obligations, not authorisations. You give notice, you don't ask permission.

And the meter measures the wrong thing. Training techniques keep improving, so a model below the threshold today matches one above the threshold from two years ago. Recent models gain capability at answer time, by "thinking" for longer, which the training meter ignores entirely. And distillation exists precisely to produce something small and capable, which is to say something powerful below the line.

Creating a market authorisation, as we do for medicines. The idea comes up constantly, and it's appealing. It has been explicitly ruled out. The American executive order of 2 June 2026 states that nothing in it authorises a mandatory licensing or prior-authorisation regime for developing or releasing models, including the most advanced ones.

Beyond the political choice, there's a logical obstacle. A drug is tested against a specific disease, with a measurable success criterion. That's what a therapeutic indication is. A general-purpose model doesn't have one. There's nothing to test it against.

Worse: a clinical trial can demonstrate that a drug has no effect. An AI evaluation can never demonstrate that a capability is absent. It only establishes that nobody managed to elicit it that day, with those questions. Six months later, someone finds the right phrasing.

Using export controls. This happened once, and the episode deserves attention. On 12 June 2026, three days after releasing its most advanced models, Anthropic received a directive from the American government forbidding it to give access to any foreign national anywhere, including its own non-American staff. Unable to filter by nationality in real time across dozens of platforms, the company cut everything, for everyone. Nineteen days later the measure was lifted. Anthropic published its account (opens in a new window), noting that the letter did not specify the nature of the concern.

What to take from it: the power to stop exists, immediate and global. The evaluation process that ought to guide it does not.

So here we are. We have instruments that measure badly, an authorisation regime nobody knows how to design, and a stop button with no manual.

And every one of these attempts shares the same blind spot.

The question nobody asks

All of these mechanisms accept the general-purpose model as a fact of nature and try to govern it.

None of them asks whether the generality was necessary in the first place.

Put the question another way, with an example. A health insurer receives thousands of calls and letters every day. It wants to sort them, understand them, route them to the right department, and reply. A perfectly legitimate need, a perfectly ordinary one.

To meet it, the insurer plugs in a large general-purpose model. It gets a system that does indeed sort its claims. It also gets, in the same package, a system that can write exploit code, reason about virology, produce propaganda in forty languages and design guidance software.

It asked for none of that. It will never use any of it. But it's there, inside its information system, and nobody on the outside can tell what it's doing with it, or what an intruder would do with it.

We collectively decided that the right way to sort the post was to install a machine that can do anything.

A model that doesn't know can't be misused

Take the other road.

Picture a system made of several specialised models, each trained on one precise task: recognising who is speaking, transcribing, detecting emotion, pulling names and dates out of a document, classifying a request. And a router that picks the right specialist for each request.

This system sorts the insurer's claims. It can't do anything else.

It can't be turned toward weapons design. Not because it would refuse, but because it doesn't know how.

The difference is enormous. A refusal is a decision, and therefore negotiable: rephrase it, work around it, switch supplier, copy the model and strip the filter. Anthropic's report documents all four. Ignorance doesn't negotiate. It's the only protection that survives distillation, jailbreaking and a change of vendor, because it isn't a layer added on top: it's what the object is.

And there's better.

The object reveals the project

A server running a general-purpose model tells you nothing about what it's for. It's a tool that stays silent on the intentions of whoever installed it. That's exactly what makes monitoring use so hard: from the outside, nothing distinguishes a company drafting letters from an agency preparing a disinformation campaign.

A specialised model is the opposite. It's a declaration.

Buying, installing and running a voiceprint recognition engine is announcing that you want to identify speakers. Nobody deploys that block by accident, nobody has it "as well", nobody has it without knowing. You chose it, bought it, integrated it.

The object reveals the project.

That changes the nature of the governance problem. Until now, the major obstacle to controlling AI was the absence of a signature: a data centre can train a weather model or a military one, and nothing distinguishes them from outside. With specialised models, the signature comes back. It isn't physical, it's functional, and you read it in the inventory of blocks deployed. That is far easier to verify than a count of arithmetic operations.

And the medicine analogy, which failed on the general-purpose model, works again. "Speaker recognition", "named entity extraction", "emotion detection": these are indications, in the exact sense of the word. A declared use, a bounded scope, an assessable result. You can test. You can certify. For the most sensitive capabilities, you can require authorisation.

Specialisation doesn't only make deployment safer.

It makes regulation possible again.

What this looks like in practice

There's nothing exotic about the architecture, and it took hold in industry long before generative AI: you assemble specialised components rather than build one object that does everything.

A router receives the request and picks the right treatment. Formats convert along the way: a phone call comes out as written minutes, a scanned letter as a spoken reply. For the IT department, that's one engine to deploy and maintain, whatever the format coming in.

This architecture sets off a chain of consequences that hold together:

Because each task calls the right model rather than the biggest one, the compute requirement collapses. Because the compute requirement collapses, plain processors are enough, sometimes embedded hardware. Because plain processors are enough, on-site installation becomes realistic. Because the installation is on site, the data doesn't leave.

Sovereignty isn't a constraint you swallow to satisfy a regulation. It's a result of the architecture. And the bill follows the same logic: a licence known in advance rather than metered billing, which is only sustainable if the cost of compute is low.

One question always comes up: at how many specialised models does a system become general again? Twenty blocks, is that still specialisation?

The criterion isn't the number. It's simpler than that: can you write the list of what the system can do? If the list exists, it's an inventory, therefore something verifiable. If it can't be written, it's a general-purpose model in disguise, whatever the number of components.

One condition has to be added, and it's less comfortable. Legibility doesn't come from specialisation alone. A speech recognition model downloaded anonymously from a public repository is narrow and perfectly invisible: nobody knows who took it. What makes intent legible is the combination of a catalogue of named capabilities and a traceable acquisition channel: a licence, a contract, an identified customer.

Put another way, the business model is part of the safety apparatus. That's an unusual thing to say, but there it is.

Finally, this idea isn't an inventor's fancy. American industry is already inventing it, under pressure and without a theoretical frame. Anthropic now operates access tiers according to how much capability is left open: public release, trusted partners, suspension. Its consumer model automatically routes requests touching cybersecurity, biology or model copying to a less capable model, which amounts to carving up generality by domain. And the company requires identity verification for accounts from regions judged to be higher risk.

Three forms of separating regimes, improvised case by case. The proposal is to make it a principle.

What this doesn't solve

An argument that claims to solve everything solves nothing. Three limits, without softening them.

Specialised doesn't mean harmless. Protein design tools are single-task and rank among the most concerning technologies in biosecurity. Guidance software is specialised by definition. Mass surveillance is a narrow application, and Anthropic's report gives it an entire section.

But that's exactly the point of the approach. These capabilities can be named, therefore listed, therefore made subject to authorisation. You can do nothing of the sort with a general-purpose model, whose contents you don't know until someone extracts them. Specialisation doesn't make the world harmless. It makes the dangerous identifiable, which is the condition for acting.

Seeing isn't stopping. Back to Bamako. In that case everything was visible: an identifiable consultant, a named intelligence service, a designated platform, twenty-five million SIM cards. Anthropic banned the account. The system had already been moved to local models and carried on running.

Legibility is the precondition for regulation, not regulation itself. It supplies the information a regulator needs and doesn't currently have. What gets done with it is political, not technical.

Specialists live off generalists. Distillation, the technique the report condemns, is also the one that makes it possible to build small, capable models from large ones. The specialised ecosystem doesn't free itself from frontier research, it depends on it. This has to be said plainly: the argument here isn't for large models to disappear, but for us to stop deploying them everywhere by default.

Two dangers, two answers

The costliest confusion in the public debate is mixing two problems that share neither scale nor solution.

The first is frontier risk: the laboratories training the most advanced models, decisive capabilities in biology, cybersecurity, weaponry. That risk belongs to state regulation, independent evaluation and international cooperation. Specialised models change nothing about it. Worth saying, so as not to sell counterfeit currency.

The second is deployment risk: millions of organisations putting a general capability that none of them needed into their information systems and onto their customers' data. That risk isn't a matter of law. It's an architectural choice, it gets decided in technical committees, and that's where the overwhelming majority of the volume sits.

Public debate currently applies to the second the instruments designed for the first. Hence the compute thresholds, labouring to measure one quantity when the problem is another. The subject isn't power. The subject is breadth.

Three sentences to finish

Generality is what made AI governance impossible. Not power: indifferentiation. An object that can do anything says nothing about what it's doing, and can be neither tested against a standard, nor inventoried, nor authorised.

Specialisation makes it possible again. A named capability is a declared capability, therefore one that can be recorded, therefore one that can be controlled. The object reveals the project.

We only regulate what we can name. And we can only name what is specialised.

What remains is to decide whether we'd rather keep looking for ways to watch machines that can do everything, or start asking why we installed them where none were needed.

What we build

The architecture described above isn't a thought experiment. We've been building it for more than ten years, and it has a name: Voxy Neural Engine.

A router receives the request and picks the right expert model: speech recognition, speaker identification, language understanding, document analysis. Every block is named, therefore inventoriable. None of them can do anything beyond its own task, which is the only protection that survives copying, circumvention and a change of vendor.

The chain described above applies as written: little or no GPU, on-site deployment, data that doesn't leave, a licence known in advance rather than metered billing.

The insurer in the example doesn't need a machine that can do anything. It needs its post sorted.

Discover Voxy Neural Engine

Sources

Primary report
Anthropic, Detecting and countering misuse of AI: September 2026, published 10 September 2026.
https://www.anthropic.com/threat-intelligence-report-september-2026 (opens in a new window)
154 pages, covering activity disrupted between December 2025 and August 2026 across seven areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons, illicit distillation.

The export controls episode
Anthropic, Statement on the directive to suspend Fable 5 access, June 2026.
https://www.anthropic.com/news/fable-mythos-access (opens in a new window)
Suspended 12 June 2026, controls lifted 30 June, access restored 1 July.

Regulatory frameworks cited

  • European Union, AI Act and code of practice for general-purpose models. Threshold of 1025 floating-point operations; enforced by the European AI Office since August 2026.
  • California, Senate Bill 53, Transparency in Frontier Artificial Intelligence Act. Threshold of 1026 operations; in force since 1 January 2026. Transparency reports, critical incident reporting within fifteen days, whistleblower protection, penalties up to one million dollars per violation.
  • United States, executive order of 2 June 2026, Promoting Advanced Artificial Intelligence Innovation and Security. Voluntary framework; section 3(c) explicitly excludes any mandatory licensing or prior-authorisation regime.

A note on method
The cases described here come from Anthropic's report and are reported as the company describes them, without sharpening. Where the report does not name an actor, this article does not name them either. The information cannot be independently verified: it rests on one vendor's visibility into its own systems, which is both the value and the limit of the document.