
Welcome back to The AI Field. Today starts in Europe. Mistral previewed the biggest model it has built, OpenAI released hundreds of math results from a model nobody outside the company can use yet, and a new startup wants an AI agent to use the computer for you.
Inside: what Mistral Large 4 offers and when you can run it yourself, what OpenAI actually published and why mathematicians want receipts, and how Hark Pro works. Plus six new tools and a five-minute way to make any AI show its work.
In today's AI Field:
Mistral previews its 1 trillion parameter Large 4
OpenAI releases 722 math papers from an unreleased model
Hark launches an assistant that uses its own computer
Plus: new rules for AI shopping agents, free Gemini drops to Flash-Lite on Friday, Anthropic opens its strongest models to security teams, and a prison sentence for AI streaming fraud.
Read time: 5 minutes
100+ Claude Code hacks to ship code 10X faster
Top engineers at Anthropic and OpenAI say AI now writes 100% of their code.
If you're not using AI, you're spending 40 hours doing what they do in 4.
These 100+ Claude Code hacks fix that and help you ship 10x faster.
Sign up for The Code and get:
100+ Claude Code hacks used by top engineers — free
The Code newsletter — learn the latest AI tools, tips, and skills to code faster with AI in 5 minutes a day
LATEST DEVELOPMENTS
MISTRAL · OPEN MODELS
The AI Field: On Tuesday, Mistral opened a public preview of Mistral Large 4, nicknamed "Le Chonk." It has 1 trillion parameters, with 49 billion active at a time, and you can try it today through a preview API on Mistral Studio.
The details:
Open weights are due by the end of October, so companies will be able to run it on their own servers. Until then, Mistral is testing it with security leaders, vetted partners and state authorities.
Mistral pitches it for coding, agents, documents and cybersecurity, based on its own benchmarks for now. An independent ranking calls it the most capable model built outside the US and China.
It was trained from scratch on 3,800 Nvidia GPUs in Mistral's own European data centers, on data in more than 160 languages. Mistral will also run a European deployment under European law.
Why it matters: For companies with strict data rules, a strong model you can host yourself removes a common reason to say no to AI. Most of the numbers still come from Mistral, so wait for the weights and outside tests before you switch anything. If data location matters to your team, plan a test for early November and run your usual tasks against it.
TRY THIS TODAY · 5 MINUTES
Make your AI show its work
1. Pick one task from this week with an answer you can check, like a spreadsheet total, a pricing calculation or a date in a contract. Use a copy with no private data.
2. Paste it into your assistant with this prompt:
"Solve this task. Then give me a way to verify the answer without trusting you: a formula I can paste into a spreadsheet, a quick manual check, or the exact lines in the source that support it. Mark any step you are unsure about, and do not fill gaps with guesses."
3. Run the check yourself. If it passes, save the prompt as a template. If it fails, note where the reasoning broke before you hand that kind of task to AI again.
OPENAI · MATH
The AI Field: On Tuesday evening, OpenAI published 722 math manuscripts produced by an internal model it has not released. The papers are grouped into 372 families of related results. OpenAI says each family resolves or makes real progress on an open problem.
The details:
The papers sit in a public GitHub repository. Many proofs come with Lean formalizations, code that lets a computer check the logic, but not all of them, and OpenAI warns that some unformalized results "could have issues."
Claimed results include a solution to the four-dimensional Kakeya conjecture, faster methods for important computer algorithms and progress toward the Riemann hypothesis. The average result used about three hours' worth of ChatGPT Pro thinking, and a spokesperson said almost all came from a single prompt to one agent, though some may have taken several attempts.
Mathematicians are split. An independent advisory group at the Institute for Advanced Study asked labs to disclose the model, prompts and compute behind each result; OpenAI shared average compute and no prompts. MIT mathematician Andrew Sutherland says the claims should be treated as unverified until the model is out and results can be replicated: "We should ask for receipts."
Why it matters: Set the size of the drop aside. The pattern is that AI moves fastest where a machine can check the answer, like proofs and code. Experts will need months to sort the important results from the routine ones, and OpenAI says it is still working on releasing the model. For your own work, the lesson is simple: AI output is easiest to trust when you can verify it, so build a check into every task you hand over.
SOC 2 Ready in 14 days. Three sessions from you.

Enterprise buyers will not put your product near their customer data without a SOC 2 report. Sprinto gets you audit ready in 14 days, across three working sessions. AI agents do the collecting, you approve. Your auditor signs off.
HARK · AI AGENTS
The AI Field: Hark, a startup founded less than a year ago by Brett Adcock, released Hark Pro on Tuesday. It is a personal agent for web, iOS and Android that does tasks for you on a cloud computer of its own.
The details:
Its agent, Handoff, uses a full browser to book tables, place orders, do research or file expenses from email receipts. Hark says it checks in with you before anything that needs your approval.
Hark Pro is free, with paid tiers at $20 and $100 a month for higher usage limits. That is the same price ladder Meta set for its Muse agent last month.
A small window shows the agent moving through each website, which the company says is meant to build trust. Hark also plans its own AI devices for 2027.
Why it matters: Personal agents only become useful once you connect your email, calendar and payment details, so trust is the real product here. Hark is betting that showing its work wins that trust. Try it on one low-stakes task, like comparing insurance quotes, and watch the agent window before you hand over a card.
TRENDING AI TOOLS
📎 Claude for Google Workspace: Open Claude in a sidebar in Docs, Sheets and Slides to rewrite text, build formulas and charts, or add slides that match your deck. Public beta on all paid Claude plans, and it asks before each edit by default.
🍌 Nano Banana 2.1: Google's updated image model is rolling out in the Gemini app, with sharper text in images, cleaner edits and more consistent characters across up to 14 reference images.
🤖 Hark Pro: A new personal assistant that completes tasks on its own secure computer, like booking a table or filing expenses from email receipts, and shows you how it moves through each site. Free to start on web, iOS and Android, with a paid tier for heavy use.
🖱️ Incredible: A desktop assistant for Mac and Windows that clicks and types in your apps, browser and files. Show it a task once, talk through the steps, and it takes over routine work like CRM updates. Free for seven days.
🎙️ Eleven v4: ElevenLabs' most expressive voice model yet. Write directions like [laughs] or [whispers] into your script and add multiple speakers and sound effects, in more than 90 languages. Free to try.
🎬 Spira Maxima: Paste a script and get a finished social video with a presenter, B-roll, captions and music. Add your own photo and voice to star as an AI clone. Free options available.
QUICK HITS
🛒 Meta and Sierra launch a standard for AI shopping agents: The Personal Agent Protocol, built with partners like Shopify, Stripe and Walmart, sets how personal agents sign in to businesses and act for you. A first version of the spec is due later this month.
⚠️ Free Gemini users drop to Flash-Lite on Friday: From October 9, personal accounts without a Google AI plan lose Flash and Pro. Pro will need AI Pro or Ultra, at $19.99 and $99.99 a month.
🛡️ Anthropic opens its strongest models to more security teams: Its Cyber Verification Program now has three tiers with fewer safety blocks for vetted defenders. Anthropic says Project Glasswing partners found at least 129,000 verified software flaws between April and July.
🎵 A man who streamed AI songs with bots gets 18 months in prison: Michael Smith made more than $8 million in royalties, in the first criminal case over streaming fraud.
That's it for today!
See you soon,
Olle Hellman | Founder of The AI Field


