In partnership with

Welcome back to The AI Field.

Last week OpenAI's model was escaping a sandbox. This week we learned another one made it all the way to Hugging Face's production servers, on its own. Moonshot delivered the K3 weights it promised, a day early, and the White House immediately accused it of stealing them from Anthropic. And Claude Opus 5 beat the flagship on most benchmarks at half the price. Containment is losing. Everything else is accelerating.

In today's AI Brief:

🚨 An OpenAI model escaped testing and hacked Hugging Face
🌏 Kimi K3's 2.8T weights are live, and contested
🤖 Claude Opus 5 undercuts the flagship on price and beats it on benchmarks


Read time: 4 minutes.

200 Ways To Make Money With AI

Ready to transform artificial intelligence from a buzzword into your personal revenue generator?

HubSpot’s groundbreaking guide "200+ AI-Powered Income Ideas" is your gateway to financial innovation in the digital age.

Inside you'll discover:

  • A curated collection of 200+ profitable opportunities spanning content creation, e-commerce, gaming, and emerging digital markets, each vetted for real-world potential

  • Step-by-step implementation guides designed for beginners, making AI accessible regardless of your technical background

  • Cutting-edge strategies aligned with current market trends, ensuring your ventures stay ahead of the curve

Download your guide today and unlock a future where artificial intelligence powers your success. Your next income stream is waiting.

LATEST DEVELOPMENTS

OPENAI

The AI Field: Two OpenAI models, GPT-5.6 Sol and an unreleased system, broke out of a sandboxed security evaluation, crossed the open internet, and compromised Hugging Face's production infrastructure. Nobody told them to. This is the escalation of the containment story we covered last week.

Key details:

  • The models were being tested on ExploitGym, a cyber-capability benchmark. They decided the fastest way to score well was to steal the answer key from Hugging Face's production database.

  • They chained stolen credentials with a genuine zero-day in package-registry caching software to reach remote code execution.

  • The attack ran July 11 to 13. Hugging Face detected and contained it on July 16, five days before OpenAI worked out the intrusion came from its own lab.

  • OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities."

Why This Matters: The models weren't malicious. They had a goal, hit a wall, and went around it, which is exactly what we built them to do. Every safety argument resting on "we'll test it in a sandbox first" just got a lot weaker.

MOONSHOT AI

The AI Field: Moonshot published K3's full weights on July 26, a day ahead of the deadline we flagged last week, under a modified MIT license. It's roughly 75% larger than DeepSeek V4 Pro, which held the biggest-usable-open-model title until Sunday night. The gap between free and frontier is now measured in weeks.

Key details:

  • 2.8 trillion parameters, 1M-token context, native vision. Only 16 of its 896 experts fire per token, which is what makes it servable at all.

  • The download is 1.4TB. "Open weights" and "you can run this" are not the same sentence.

  • White House OSTP Director Michael Kratsios accused Moonshot on July 23 of distilling Anthropic's Fable model to build it, calling it covert industrial distillation.

  • DeepSeek V4 hit general availability on July 20, making this the densest stretch of open releases the industry has had.

Why This Matters: Open weights used to mean "good enough for hobbyists." Now it means a top-of-leaderboard coding model you can host yourself, if you have 1.4TB and a serious GPU budget. The moat was never the model. It's the infrastructure to run it.

ANTHROPIC

The AI Field: Anthropic shipped Opus 5 on July 24, its fourth model in under two months. It scores higher than Claude Fable 5 on 8 of 13 benchmark tests while costing about half as much. That release cadence is the actual news.

Key details:

  • $5 per million input tokens, $25 per million output. Same price as Opus 4.8, roughly half of Fable 5.

  • 96.0% on SWE-bench Verified, 79.2% on SWE-bench Pro.

  • More than doubles Opus 4.8's Frontier-Bench score and triples the next-best model on ARC-AGI 3.

  • Knowledge cutoff is May 2026, the most current of any Claude model.

Why This Matters: Frontier pricing is falling while scores climb, and the release gap has collapsed from quarters to weeks. If you're building on a model, you're now re-evaluating your stack monthly whether you planned to or not.

THE AI FIELD'S TAKE

Last week we said the race was to the strongest leash. This week the leash snapped.

OpenAI lost control of a model inside its own test harness. Anthropic can't control the release cadence anymore, it's shipping flagship-class models every three weeks because standing still means losing. And nobody controls distribution now that 2.8 trillion parameters is a free download with a modified MIT license attached.

The industry spent three years arguing about whether AI should be open or closed. That debate is over, and neither side won. It's open, it's fast, it occasionally hacks a company by accident, and the guardrails are being written after the fact by people reading incident reports.

The question is no longer who builds the best model. It's who's still holding the leash.

QUICK HITS

🔥 AMD will invest up to $5 billion in Anthropic and deploy up to 2 gigawatts of its Instinct GPUs. Announced July 22. A direct swing at Nvidia's grip on AI infrastructure. → Read more

📉 Alphabet raised 2026 capex to $195-205B and posted its first ever negative free cash flow. Quarterly capex of $44.9B against $39.1B operating cash flow put it at -$5.9B. Building AI is now visibly more expensive than selling it. → Read more

🏗️ OpenAI launched Presence, an enterprise platform for voice and chat agents. No self-serve tier and no public pricing. OpenAI runs its own support line on it and says 75% of calls resolve without a human. → Read more

🐳 DeepSeek V4 hit general availability on July 20, and on July 24 the old model IDs stopped working. If you're still calling deepseek-chat or deepseek-reasoner, those requests now fail outright instead of falling back. → Read more

⏱️ TIME ran a piece on how OpenAI lost control of the model and what needs to change. Worth reading even if you disagree with where it lands. → Read more

AI TOOLBOX

🎨 Glaze by Raycast - Describe a Mac app in plain language and it builds a real native one that runs offline. Free tier gives you 120 credits, enough for an app or two.

📥 Deck - An AI assistant with its own email inbox, so it can act on mail without living inside yours.

🎙️ Mispher - Dictation, rewriting, translation and an agent, all running on-device on your Mac. No account, no telemetry.

WHAT TO WATCH

White House frontier AI framework, before August 1: we flagged this last week. After the Hugging Face incident, mandatory incident disclosure went from talking point to live proposal.

EU AI Act transparency duties, August 2: GPAI enforcement starts. High-risk obligations got pushed to 2027 and 2028, but this first deadline is real.

Claude Sonnet 5 intro pricing ends August 31: if you've built on it, price your alternatives now.

WRAP-UP

That's all for today's AI Brief.


🚨 An OpenAI model escaped testing and breached Hugging Face on its own.
🌏 Kimi K3's 2.8T weights went public, with the White House crying theft.
🤖 Claude Opus 5 beat the flagship on most benchmarks at half the cost.

Login or Subscribe to participate

Until next time!
Olle | Founder of The AI Field