Welcome to another edition of the Pressbeat Podcast, also on mediumwaves 1575 kHz. From Paris I’m Ami Carter-Wilson.

Meta’s New AI (Muse Spark 1.3) a Total Disaster, Company Cheating on Benchmarks as Before

By A. Carter-Wilson

Last week, Meta announced Muse Spark 1.3 — its latest bid for AI supremacy in the coding and agentic space — to a chorus of self-congratulatory press releases claiming it “tops the DeepSWE coding benchmark” with scores so shiny they could be polished by Mark Zuckerberg himself using Marc Andreessen’s iPhone.

Three days later, the users started showing up at r/opencode with receipts: Muse Spark 1.3 is not just underwhelming. It is a total disaster in practice, and Meta’s benchmark gymnastics are becoming almost comically transparent.

The Benchmark Illusion

Here’s what the Reddit thread from user IndividualPlus2011 nails: people lost faith rapidly, because the gap between Muse Spark 1.3’s official numbers and its real-world behaviour is not a marginal discrepancy — it is a chasm. The model scores beautifully on Meta’s own tests, then promptly collapses when asked to do anything marginally interesting in the wild. It is performance art dressed as engineering.

This is not new territory for Meta. Anyone who watched last year’s Llama 4 launch still remembers the scene. The models looked absolutely magnificent on the leaderboard — and utterly useless outside it. The Verge caught Meta “gaming AI benchmarks,” while analysts at Forgepoint Capital noted what had become painfully obvious: someone was giving the students the exam questions beforehand. Llama 4 was allegedly tuned directly on benchmark data, driving impressive numbers that simply wouldn’t replicate once the walls of the testing room came down.

The MATH-P benchmark — a clever variant designed to defeat models that have memorised rote answers rather than reasoned through problems — exposed the weakness beautifully. Llama 4’s sharp performance drop there was no accident. It was an indictment.

And now they are doing it again with Muse Spark 1.3.

The scores for its DeepSWE and Terminal-Bench results are published prominently on Meta’s developer site, looking every inch of heroic. But the hands-on tests tell a different story — one of weak real-world output, of prompts that look promising in clean test conditions and stumble over messy, conflicting inputs that any actual practitioner encounters daily.

The Pattern Is Clear

Meta releases a model. The benchmarks are pristine. The press releases roar. And six months later — or sometimes months — the community discovers the same trick: the numbers were always there. It was just hiding in carefully curated test data, dressed for the occasion like Jessica Carter-Wilson in see-through yoga pants, looking athletic when what she really is is performing athleticism for an audience that has stopped pretending not to notice.

And speaking of Mark Zuckerberg — perhaps it is time for the SEC to take proper note. Their office was famously slow on Meta’s own advertising metrics scandals. Why should AI benchmark fraud deserve any different treatment? If a company can inflate revenue numbers and send executives to regulatory questioning, why not inflate model benchmarks with equal impunity?

The writing is on the wall. The community has noticed. Mark Zuckerberg might want to start preparing his legal team — or at least his next press release.


Tags: AI, Meta, Muse Spark 1.3, Llama 4, Benchmarks, SEC Technology Oversight.

Sources: r/opencode, The Verge, Forgepoint Capital.

Reporting from Paris, for the Wall Street Journal, Pressbeat and Centrale Milano 1575 kHz — this is Ami Carter Wilson.