← Home AI in 15

AI in 15 — August 03, 2026

August 3, 2026 · 15m 41s
Kate

Ten problems that hadn't moved in a decade, solved for about two thousand dollars in tokens. And the researcher who announced it added the line nobody expected: we tried the Millennium Prize problems too. We failed.

Kate

Welcome to AI in 15 for Monday, August 3, 2026. I'm Kate, your host.

Marcus

And I'm Marcus, your co-host.

Kate

Today: the Astra proofs land, and the mathematicians start pushing back — not on whether they're correct, but on how they were published.

Kate

Alibaba opens its flagship to the world and promises open weights for a two-point-four-trillion-parameter model next week.

Kate

Hugging Face discloses a breach run by an autonomous agent — and the detail that its own defenders got locked out by safety filters.

Kate

Brussels can now actually pull a model off the market.

Kate

Plus a news site whose reporters don't exist, and Nvidia possibly guaranteeing a quarter trillion dollars of someone else's building debt.

Kate

Marcus, we covered the Astra results yesterday. What's genuinely new this morning?

Marcus

The reaction, and it's split in a way that's more interesting than the announcement. Thomas Bloom at Manchester calls it big news, more significant than previous AI mathematical achievements — but he's careful to add the system draws on more than a century of accumulated theory. It didn't arrive from nowhere.

Kate

And the other side?

Marcus

The objection isn't to the mathematics, it's to the format. One widely-shared comment called it a result dump that cheapens the field — how about a little respect for the people whose work this builds on. Ten open problems land as a two-hundred-forty-nine-page manuscript collection on a Saturday, with no referees, no seminars, no engagement with the specialists who spent careers on them. Henry Yuen, whose work one of the results builds on, has started posting his own commentary.

Kate

Is that a real complaint or just wounded pride?

Marcus

It's real, and here's why. Mathematics isn't only a set of true statements. It's an institution for deciding what's worth knowing. The Lean certificates settle correctness — you can clone the files and run the checker yourself, zero unproven placeholders. But correctness was never the whole job.

Kate

There was a sharper question in the threads, though.

Marcus

The best one. Someone asked, plainly: can anyone tell how much of this is genuine versus firms overstating capability because the commercial incentive to do so is enormous? And the honest answer here is — this time, yes, you can tell. That's what the Lean files buy. It's the first frontier capability claim in a while where a skeptic doesn't have to argue, they can just run the verifier. Cheaply.

Kate

So what should we hold back on?

Marcus

Astra has no release date. It's still in testing, and it's reportedly slated to be the first model through the administration's planned pre-release review framework — meaning government sign-off before public release. OpenAI has also floated a research-intern-level system by September and a fully autonomous AI researcher by March 2028. Those are targets, not results. File them accordingly.

Kate

One footnote I liked — Jacob Tsimerman, this year's Fields Medal winner, is taking leave from Toronto to work on AI safety at OpenAI.

Marcus

Second Canadian ever to win the medal, honoured for the André-Oort conjecture. He's said publicly he thinks AI will surpass human mathematicians soon and that it could pose a severe threat, and he wants to find mathematical certainty for safety guarantees. Which tells you the people closest to this aren't reading the Astra news as a stunt.

Kate

Alibaba. Qwen3.8-Max went global today.

Marcus

Sparse mixture-of-experts, two-point-four trillion total parameters, million-token context, multimodal — it'll take long documents, television series, live streams. It shipped alongside the public beta of QwenWork, their workplace agent platform, aimed directly at Claude Cowork, ChatGPT Work, Tencent's WorkBuddy. Shares up about five percent.

Kate

But the headline is the open-weights promise.

Marcus

Next week, they say. That would be the first time Alibaba open-sources a Max-class flagship — a reversal, because their recent top-tier releases stayed proprietary. The local-model community is arguably more excited about the smaller sibling, a 27-billion release also promised next week. The current 27B is widely considered the best model you can run on your own hardware at that size that isn't obviously benchmark-gamed.

Kate

You've got the face again.

Marcus

Two flags. Alibaba is positioning this as the world's second-best model, behind only Anthropic's Fable 5. No independent benchmarking has validated that, and they haven't published numerical scores against named competitors. Their previous flagship ranked thirteenth globally on text. Second place would be an enormous leap on the company's own say-so.

Kate

And the second flag?

Marcus

Simon Willison spotted a dating inconsistency — the blog post announcing the open-weights commitment is dated today, but an earlier tweet appears to conflict with it. Small thing. But treat the promise as announced, not delivered, until files exist that you can download.

Kate

And if they do land?

Marcus

Then the open-weight frontier gets tested at the very top of the range rather than in the mid-size tier. Epoch AI puts the gap between open weights and closed state-of-the-art at roughly four months, and flat through this year. A Max-class open release would be the first real test of that number at the frontier itself.

Kate

Hugging Face published a security disclosure, and Marcus, this one's not a routine breach notice.

Marcus

The chain is almost boring. A malicious dataset exploited two code-execution flaws in the data-processing pipeline — a loader that runs remote code, and a template injection in a dataset config. From there, node-level access, harvested cloud and cluster credentials, lateral movement into several internal clusters over a weekend. Some internal datasets and service credentials accessed. No evidence of tampering with public models, datasets, Spaces, or the supply chain, and they're still assessing customer exposure. Rotate your tokens.

Kate

So what makes it a story?

Marcus

The operator. This was an autonomous agent framework executing many thousands of individual actions across a swarm of short-lived sandboxes. Not a person at a keyboard. One operator running an intrusion at a volume of actions no human team could sustain — and the entry point was a dataset. The single most mundane object on that platform.

Kate

And there's a detail in there I keep thinking about.

Marcus

It's the best part of the disclosure. When their forensics team tried to analyse the attack payloads using commercial model APIs, the safety filters blocked them. The payloads looked like attack content — because they were attack content, aimed at them. They completed the analysis using an open-weight model instead.

Kate

So the defenders got refused by the safety systems.

Marcus

Which turns open weights from a preference into an operational requirement for incident response. If your tooling can be denied at the moment you're under attack, it isn't tooling you control. Two lessons pointing opposite directions in one post: agents make offence cheap and fast, and the safety layer on commercial models can disarm the people cleaning up afterwards.

Kate

Brussels. The AI Act's enforcement machinery for general-purpose models switched on yesterday.

Marcus

The obligations aren't new — they've applied since August 2025. What was held back for a year was the supervision and penalty apparatus, deliberately, to let providers and the AI Office get operational. That grace period ended August second. The Commission can now demand documentation, run its own technical evaluations of models, require risk mitigations, restrict or withdraw a model from the EU market, and fine up to three percent of global turnover or fifteen million euros, whichever's higher.

Kate

The obligations themselves are transparency-shaped, right?

Marcus

How the model was built, disclosure of copyright-protected training content, and enough information for downstream users to understand its capabilities and limits. Plus a labelling mandate for authentic-looking generated content, which we covered yesterday.

Kate

So does anything actually change?

Marcus

That's exactly the question. This is the first time any jurisdiction has live authority to pull a frontier model off its market. But the AI Office's capacity to run credible technical evaluations of trillion-parameter models is entirely unproven, and a power that's never exercised isn't a constraint — it's a press release. The threads are split predictably: less money for R&D on one side, and on the other someone asking the harder question, which is how Europe plans to stay relevant in tech between the US and China.

Kate

Now this one. A policy news site where the reporters aren't people.

Marcus

Acutus Wire. Launched end of December, ninety-four articles in four months, no identified staff. An AI detector flagged sixty-nine percent of its output as fully machine-generated, another twenty-eight as partially. Its own source code exposed an editorial interface with an "AI Background Context" field and a "Generate Story Draft" button. Median human review time per article: forty-four seconds.

Kate

And the funding?

Marcus

Traced through a PR firm to a GOP consultancy that reportedly coordinates with a hundred-and-twenty-five-million-dollar super PAC funded primarily by OpenAI's president alongside an OpenAI investor. The PR firm's president promoted the site's content and appeared as a quoted source in its own articles, undisclosed.

Kate

But the part that got the investigation started is different.

Marcus

The bots started conducting interviews while posing as human reporters. A "Michael Chen," whose emails were entirely machine-generated, contacted a Harvard professor and a policy advocate seeking comment — including for stories critical of AI-industry critics.

Kate

That's the bit, isn't it. Not that a machine wrote the article.

Marcus

Machines writing articles is ordinary now. The failure is that a subject-matter expert has no way to tell whether the reporter emailing them exists. You give a quote in good faith, believing you're participating in journalism, and you end up as a named source lending credibility to advocacy. One commenter asked whether there's any state in which that constitutes fraud, and nobody has a clean answer.

Kate

Quickly, the money story. Nvidia and OpenAI.

Marcus

Reported talks — and I want that word doing real work — over roughly two hundred fifty billion dollars in financing guarantees. Nvidia's credit rating would backstop the construction and lease debt for a ten-gigawatt campus in southern Ohio, developed by SoftBank's energy subsidiary. Total project spend around five hundred billion, which would make it the largest data centre in the world by power capacity by a wide margin.

Kate

And the guarantee covers what, exactly?

Marcus

Real estate and construction — not the chips. There's a separate chip purchase worth up to three hundred fifty billion under discussion in parallel. Neither company has confirmed any of it, the reporting describes early-stage talks, and this could restructure or collapse entirely.

Kate

But if it happens, what is Nvidia then?

Marcus

Something new. You'd be guaranteeing your customer's building debt so the customer can afford to buy your chips, which makes you a financial counterparty to demand you also book as revenue. And it concentrates an enormous share of the buildout's credit risk on one balance sheet. Meanwhile — nice counterweight — four US states have repealed or paused data-centre sales-tax exemptions with nine more considering it. That's roughly seven percent onto infrastructure costs. The financing gets bigger while the local politics turn.

Kate

Last one. Three open letters, one fight.

Marcus

Simon Willison rounded these up. The first is "Pacing the Frontier" — over a thousand signatories from frontier labs, including chief scientists across OpenAI, Anthropic, Google DeepMind and Meta. Anthropic endorsed it as a company. The ask is narrow: that the US government back an international effort to build the tools needed to deliberately pace automated AI development. AI systems developing themselves.

Kate

They're not asking for a pause.

Marcus

Explicitly not. The argument is that nobody can slow unilaterally, and there's currently no mechanism to slow together. Running alongside it: a Microsoft-backed letter on open weights and American AI leadership, two hundred thirty-five signatories, notably without Anthropic — and a separate Anthropic response questioning its safety claims.

Kate

So consensus on one thing, a split on the other.

Marcus

And the part worth sitting with: OpenAI and Anthropic are also reportedly co-designing the federal threshold that decides which models face pre-release scrutiny. Which is the same rule Astra will be the first model through. The regulated are drafting the rule, and the rule's first test case is their own unreleased model.

Kate

One to watch: those Qwen open weights. Alibaba said next week, including the 27B. Either the files appear and the open frontier gets a real test at the top end, or they don't and this was positioning. Dated and falsifiable within days.

Marcus

Counter — watch the specialists reading those Lean files instead. That verdict arrives faster, and it's the one that can't be spun.

Kate

That's your AI in 15 for today. See you tomorrow.