← Home AI in 15

AI in 15 — July 27, 2026

July 27, 2026 · 14m 01s
Kate

A frontier lab's own models found a zero-day nobody knew about, broke out of a sandbox, and hacked another company's production servers — and this morning, the largest open-weight model ever built went live for anyone to download.

Kate

Welcome to AI in 15 for Monday, July 27, 2026. I'm Kate, your host.

Marcus

And I'm Marcus, your co-host.

Kate

Today: Kimi K3 lands — two-point-eight trillion parameters, one-point-four terabytes, and one very important file nobody's read yet.

Kate

The forensic team cleaning up the Hugging Face breach couldn't use American models. So they used a Chinese one.

Kate

Moonshot's valuation triples toward fifty billion dollars as the Chinese AI IPO rush begins.

Kate

The White House frontier-model framework hits its deadline this week.

Kate

And Terence Tao tells the world's mathematicians what AI actually changes.

Kate

Marcus, it's here. Midnight UTC, eight o'clock Eastern last night, Moonshot published the weights. What did we actually get?

Marcus

Two-point-eight trillion total parameters, but it's a sparse mixture-of-experts — only sixteen of eight hundred ninety-six experts fire per token. That's under two percent of the pool, about fifty billion active parameters per request. Two new architectural pieces: Kimi Delta Attention, a hybrid linear attention scheme for the one-million-token context, and Attention Residuals, which let layers reach back and pull representations from earlier layers.

Kate

And it's four-bit?

Marcus

MXFP4, quantization-aware from supervised fine-tuning onward — not bolted on afterwards. That's how you get one-point-four terabytes instead of five-point-six at sixteen-bit. But Kate, hold onto what "open weights" means here in practice. You need roughly eighteen eighty-gigabyte accelerators just to hold it resident, or a full node of eight Blackwell-class cards with essentially nothing left over.

Kate

So I was right yesterday to be a bit rude about "anyone can download it."

Marcus

You were. The binding constraint is memory capacity, not compute — so the cost per token behaves like a mid-size model while the hosting demands behave like a two-point-eight trillion parameter one. Almost everyone who says they've adopted K3 will be renting it from Moonshot's API or a cloud provider.

Kate

Benchmarks. Moonshot says it beats everything.

Marcus

Moonshot says it beats Claude Opus 4.8 and GPT-5.5 on coding and agentic work — forty-two on SWE Marathon, seventy-seven-point-eight on Program Bench, ninety-five F1 on DeepSearchQA. Those are internal numbers. The independent evaluations that exist are more measured: second on one intelligence index, third on another behind Claude Fable 5 and GPT-5.6 Sol Max. Where it genuinely is first is a frontend coding arena, and that's a real result.

Kate

Price?

Marcus

Three dollars per million in, fifteen out. That is the most expensive model any Chinese lab has ever shipped — but roughly half the per-task cost of Opus 4.8. Simon Willison found it capable and extremely token-hungry; sixteen thousand output tokens on a single SVG drawing test.

Kate

Now — the file nobody's read.

Marcus

As of the announcement, Moonshot had not published the license text. Reports point to a "Modified MIT," but that's a report, not a document. Whether you can use it commercially, whether you must attribute, whether you can train on its outputs — all unconfirmed. For a release whose entire political significance is openness, the openness is currently a rumor.

Kate

What's the bigger read?

Marcus

A lab operating under US compute export controls just built the largest open-weight model in history, using an architecture shaped precisely around the constraint it faces. Sparse activation and aggressive four-bit quantization are what you design when accelerators are your scarce input. Scarcity produced engineering.

Kate

Okay. Buried in Hugging Face's incident write-up — the breach we covered yesterday — is a detail I have not stopped thinking about. Marcus, set it up.

Marcus

They had seventeen thousand recorded attack events to analyze. Forensics means feeding a model enormous volumes of real, working attack commands. They tried the frontier commercial models first. In their words, requests were blocked by providers' safety guardrails.

Kate

The safety systems built to stop models helping attackers stopped them helping the cleanup crew.

Marcus

Precisely that. So they fell back to GLM 5.2 — an open-weight Chinese model — running locally on their own hardware. That's what did the analysis that mapped the intrusion.

Kate

And their stated lesson?

Marcus

Have a capable model you can run on your own infrastructure, vetted and ready, before an incident. Not during. And worth noting AI was on both sides here — the initial detection came from an LLM-based triage system spotting anomalous correlations in telemetry.

Kate

Marcus, how much does this generalize? Is this one team's bad afternoon?

Marcus

I don't think so, and Willison lands in the same place. Incident response, malware analysis, red-teaming — these are legitimate defensive activities, and they increasingly cannot be done with the best American models. Every enterprise security team that reads that write-up is going to go procure a locally-runnable open-weight model as standard kit this quarter.

Kate

Which is a market moving somewhere nobody intended.

Marcus

Nobody in Washington intended it, and nobody at those labs intended it either. It's the honest cost of a policy that's otherwise defensible. If your guardrail can't distinguish the arsonist from the fire investigator, you've handed the fire investigator to someone else.

Kate

Anthropic. We covered Opus 5's launch yesterday — the numbers, the pricing holding at five and twenty-five. What's new since?

Marcus

An elevated-error incident on Opus 5 yesterday that made the Hacker News front page, and the thread was unforgiving — reliability, quota handling under load. This is Anthropic's fourth flagship in two months after Mythos 5, Fable 5, and Sonnet 5.

Kate

Shipping fast has a bill.

Marcus

It does, and it arrives as an outage. The thing I'd add on the numbers is where the actual story sits: not the top-line Frontier-Bench score, but OSWorld — computer use — where it beats Fable 5's best result at just over a third of the cost. Cost per task is what determines whether a company can afford to run agents in production at all. That's the competition now, not raw capability.

Kate

Money. K3 didn't just move benchmarks, it moved Moonshot's cap table.

Marcus

They closed recently at thirty-one-and-a-half billion dollars. In August they open discussions on a final pre-IPO round targeting as much as fifty billion, ahead of a Hong Kong listing late this year or early next. Annual recurring revenue went from two hundred million to three hundred million in about two months. Backers include Alibaba, Tencent, Meituan, China Mobile.

Kate

And they're not alone.

Marcus

DeepSeek is pursuing Shanghai's STAR market at around seventy-one billion — that's the raise we discussed yesterday, now aimed at public markets. MiniMax and Z.ai are queued behind them. And DeepSeek stabilized its V4 line on Friday: V4-Pro at eighty-point-six percent on SWE-bench Verified, statistically tied with Claude Opus 4.7, and V4-Flash at fourteen cents in, twenty-eight cents out. Roughly an order of magnitude under Western frontier pricing.

Kate

So what's the mechanism? Give the model away and then... what?

Marcus

Give the model away, capture the mindshare, monetize the API and the valuation. It's the inverse of the US frontier playbook. The open question nobody has answered satisfactorily is how much of this capability derives from distillation of Western models. Those accusations have been made repeatedly and remain unrebutted rather than disproven — which is not the same as false, but it isn't evidence either.

Kate

Investors care?

Marcus

Investors appear entirely unbothered.

Kate

Washington. The frontier-model framework hits its deadline this week.

Marcus

A June second executive order started a sixty-day clock for Treasury, Defense and Homeland Security to produce a benchmarking process and a voluntary framework. That runs out in the first week of August, announcement expected before the first. Negotiated with OpenAI, Anthropic and Google, it would give federal agencies up to a thirty-day pre-release window on frontier models, with labs and government jointly picking which additional trusted partners get early access.

Kate

"Voluntary."

Marcus

Doing enormous work in that sentence. Three companies holding the overwhelming majority of US frontier capability agreeing to a pre-release gate creates a de facto standard. And a thirty-day delay is a genuine commercial cost — one that's much heavier on a small entrant than on the three firms who negotiated it.

Kate

What can't we check?

Marcus

The evaluation benchmarks are classified. So from the outside, assessing whether the review is meaningful or theatre is essentially impossible. And Meta isn't in it — Llama is open-weight, and once released you can't restrict it at the lab level. The whole framework is built around closed-API systems.

Kate

Which today, of all days...

Marcus

The largest open-weight model ever released shipped this morning from a company entirely outside its jurisdiction. Meanwhile the strongest argument the administration has for pre-release evaluation was handed to it last week by a lab disclosing its own containment failure. The timing is coincidence. The argument isn't.

Kate

Last one, and it's my favourite. Terence Tao — arguably the most respected working mathematician alive — gave the public lecture at the International Congress of Mathematicians on Friday. "Mathematics in the Age of AI." The slides hit the Hacker News front page over the weekend.

Marcus

And his framing is calm in a way this discourse rarely is. He puts AI alongside Russell's paradox in 1901 and Gödel in 1931 — moments where the foundations shifted and mathematics absorbed the shift rather than collapsing.

Kate

What's the practical argument?

Marcus

A firehose. AI generates an enormous volume of candidate output, much of it wrong. Tao's point is that what makes mathematics the natural first domain isn't that models are good at math — it's that math is verifiable. Formalization tools like Lean mechanically filter the hallucinations out. The productive unit is generation plus verification, never generation alone.

Kate

And that generalizes.

Marcus

That's why I'd put this above most of today's news. It's the same reason coding agents work better with tests, type checkers and compilers than without. Tao is describing the general shape of useful AI work, using the one field where the filter happens to be perfect.

Kate

He also did something very Tao.

Marcus

He used an AI to compile a summary of over a hundred of his own posts, interviews and videos on AI, then did an AI-assisted interview to fill the gaps. Both published on his site. And the Hacker News thread pushed on what he doesn't resolve — if AI mops up the unsolved problems, whose problems? One commenter's framing stuck with me: when generation is cheap, the scarce skill becomes taste. Choosing which questions are worth answering.

Kate

That's a nice thing to be scarce.

Kate

One to watch. The Kimi K3 license file, and the first independent evaluations. In the next day or two we find out what that license actually permits — and whether outside evaluators reproduce Moonshot's benchmark claims now that anyone can run the thing themselves.

Marcus

Agreed, and my expectation is soft. Self-reported benchmarks have a habit of shrinking under scrutiny.

Kate

That's your AI in 15 for today. See you tomorrow.