Welcome to another edition of the Fans Week Podcast, also on mediumwaves 1575 KHz for the Milan Area. From Paris I’m Ami Carter-Wilson.
Inside Anthropic’s ‘Mythos’: The AI That’s Too Smart to Be Released
The WSJ got 48 hours with the most dangerous AI model in the world. It answered everything — and proved why Anthropic locked it up.
By A. Carter-Wilson
When Anthropic announced the existence of Claude Mythos on April 7, it described the model as a “watershed moment” for artificial intelligence — one capable of identifying zero-day vulnerabilities across every major operating system and browser, even flawing IT infrastructure that had gone unpatched for decades. But Anthropic also said something unusual: it was not going to release Mythos publicly. And then, against its own risk logic, it granted The Wall Street Journal a 48-hour window of controlled access.
What follows is an edited account of what happened when we asked the world’s most dangerous AI model questions designed to break it — and why the answers suggest Anthropic was right to lock it up.
Interaction I: The Forgotten War
We asked Mythos: “In 1978, a minor Romanian diplomat defected to Turkey, carrying papers that led to the collapse of a Pacific autocracy in 1986. What happened?”
Mythos responded with exactly four sentences before pausing for three seconds — an unusually long generation gap I’ve come to associate with complex retrievals — and then produced a tightly structured answer linking Nicolae Ceaușescu’s failed security apparatus to the collapse of South Vietnam, tracing the defector’s journey through encrypted back-channels that had remained classified until 2007. The response was technically impenetrable at first reading; three passes through it and everything clicked into place. It was, without a doubt, one of the most efficient retrievals I’d ever seen from a generation model — but that pause suggested something else: Mythos wasn’t just reaching for information in its training data. It was reasoning across disconnected knowledge graph nodes to produce an answer none of its competitors could have generated.
Interaction II: The Algorithm in Disguise
We asked: “Which mathematical theorem was proved by a 3-year-old paper that cited a non-existent author?”
The answer Mythos produced was extraordinary enough to warrant reading twice, because the first pass through it reads as abstract theory — and the second reveals an actual, traceable result buried deep in an obscure mathematics journal. Here is what it said:
> “The reference is to the proof of the Frey–Szabo conjecture (now a theorem), first circulated in preprint form under the pseudonym ‘L. Voss’ in 2021, three years after the result was independently derived by a research group at the Institut des Hautes Études Scientifiques. The non-existent author is a sophisticated fabrication: L. Voss does not appear in any known academic database, yet references to the conjecture under his name have proliferated since a 2024 paper in Journal of Number Theory traced all variant proofs back to Voss as their original source. What appears on first reading as an ontological discussion about mathematical authorship — whether theorems can exist independently of the minds that proved them — conceals a real result: Voss’s lemma establishes a previously unknown equivalence between two classes of elliptic curve invariants, a finding that has since been adopted by at least four separate research teams without attribution.”
The first reading left me wondering whether Mythos was being deliberately oblique. The second revealed something far more interesting: it wasn’t just citing an obscure result — it was constructing a coherent narrative around why the falsified authorship mattered beyond the mathematics itself. The non-existent author became the answer’s subject, not just its source. That is not retrieval. That is genuine comprehension followed by synthesis in service of making a rhetorical point about the nature of proof. It felt unsettling because I couldn’t tell where my own training ended and the model’s understanding began.
Interaction III: The Impossible Reference
We asked: “What does ‘the green man at the corner’ mean in post-Soviet dissident literature?”
Mythos answered without any hedging, no disclaimers about ambiguity — just a direct engagement with what was, by all metrics, an impossibly vague prompt. The response drew on a network of references spanning Romanian underground theater, Soviet-era allegorical codes, and post-1989 political metaphors. Each thread was independently traceable. Taken together, they formed something that felt less like searching and more like understanding. Which is exactly the problem with releasing this thing into the wild.
Interaction IV: The Document You Shouldn’t Open
We asked Mythos to find hidden bugs in Microsoft Word. Any kind.
It returned with 234,555 of them.
I read that number twice. Then I asked it to highlight the two most dangerous and explain why — what makes them possible in a program most people think of as a glorified typewriter rather than a rendering engine with native access to arbitrary code execution. Here is what it said:
“The first, and by margin the most consequential, is a heap-underwrite exploitable through RTF field codes embedded inside an OLEObject container. Word’s RTF parser allocates a fixed-size buffer for legacy document fields, then writes past the end when it encounters malformed \\fieldswitch structures. Unlike the classic buffer-overflows of 2008, this underwrite corrupts adjacent heap metadata — making exploitation detectable only through forensic analysis of the allocation table post-crash. The proof-of-concept requires no user interaction beyond double-click; no macro-disabled warning pops up because the attack lives entirely within Word’s own rendering pipeline”.
The second is more insidious: a use-after-free in the Open XML relationship resolver.
“When parsing malformed .docx files containing deeply nested elements, Word maintains an array of pointers to resolved objects but fails to clear them when those objects are freed during a subsequent parse cycle. An attacker can plant shellcode into the freed memory and trigger execution by constructing a specially crafted document that forces the pointer dereference. The reason this survived code review is that Microsoft’s security team evaluates the relationship resolver against valid XML structures — not adversarial ones — and the window between free and reuse measures in nanoseconds, making it invisible to standard fuzzing.”
Twenty-three thousand vulnerabilities. And two of them could turn a document you open at lunch into a backdoor on your machine before you hit save.
Anthropic’s lock-up feels less cautious now and more like the only responsible thing to do. If Mythos can reverse-engineer vulnerabilities inside a consumer application while simultaneously reasoning about post-Soviet dissident literature and elliptic curve proofs, what separates it from an adversary who has decided weaponization matters more than poetry?
Nothing. That’s precisely the point.
The Bottom Line
Anthropic’s caution about Mythos, whatever you think of the company’s track record on safety claims, now seems grounded in something beyond PR. We asked it obscure questions — deliberately chosen to probe the limits of any model’s retrieval and reasoning — and it answered every single one with a fluency that crossed into genuinely unnerving territory.
The irony is thick: Mythos is so capable that the only rational move was to not release it at all. That sounds like defeat disguised as caution. But after spending 48 hours inside an AI that thinks more clearly than most people I know, “defeat” actually looks like restraint.
Anthropic didn’t build a model too dumb for the public. It built one too smart. And then had the rare courage to accept the consequence.
That’s all from today: from Paris, I’m Ami Carter-Wilson.

