In partnership with

IN THIS ISSUE

Welcome back to The AI Field.

Today's edition is about the part of AI progress that gets less attention: proof. OpenAI is claiming original research, Qwen is selling longer-run AI work for less, and Washington has finalized a classified cyber testing framework. The harder problem is deciding what deserves trust.

- OpenAI says an unreleased model produced new mathematical results
- Qwen prices its top model at $6 per million output tokens
- Washington finalizes a classified cyber testing framework for frontier AI

Read time: 4–5 minutes

200 Ways To Make Money With AI

Turns out, AI is good for more than writing emails and generating LinkedIn posts. This guide is packed with 200+ actionable ways to build income with AI.

Inside you'll discover:

  • 200 curated AI income ideas for beginners and pros alike

  • Easy-to-start opportunities you can launch this week

  • Real world applications for today’s top AI tools

  • AI powered business models built for today’s economy

  • Creative ways to turn trends into revenue

AI is changing how people work and how people earn, far beyond the simple “write me an email” prompt. Discover the possibilities and download the free guide to start cashing in today.

RECENT UPDATES

01 · OPENAI

OpenAI says its unreleased Astra model produced the arguments behind ten mathematical advances across eight fields. Some resolve open problems. Others make substantial progress on questions researchers had not cracked.

This was not a chatbot solving textbook exercises. The release includes a 249-page manuscript and reasoning walkthroughs. OpenAI says the model later formalized each argument in a Lean certificate, while humans used the same model to prepare the arguments as manuscripts.

OpenAI estimates the solution-generating tokens would cost about $2,000 at current Sol API prices. That is not the full project cost, and Astra is not available to the public. OpenAI does not report the total number of attempts behind the selected results.

One result already has outside support. Mathematician Shuoxing Zhou independently found a counterexample to the same Connes rigidity conjecture with GPT-5.6 Sol. That supports one conclusion, not Astra's proof or the other nine results.

The AI Field: This is a real step from explaining known mathematics toward helping produce new mathematics. The bottleneck shifts to verification. A candidate proof has no value until researchers can validate it.

TRY THIS TODAY

Give your AI draft a tougher review

This is a lighter, one-pass version of Matt Shumer's Gauntlet Loop. Use it for low-risk writing, planning, or structure. It is not a substitute for checking facts yourself.

1. Choose one draft and write three pass-or-fail checks. Make them visible, such as “under 150 words,” “includes one concrete example,” and “ends with one clear next step.”

2. Open a fresh AI chat and paste this prompt:

Act as a reviewer, not a writer. Compare the draft only with the checks below. For each check, mark PASS, FAIL, or NEEDS CONTEXT. Quote the evidence and give one specific fix for every failure. If context is missing, ask one short question instead of guessing. Do not rewrite the draft or add facts. List every assumption I should verify and end with the biggest remaining gap.

CHECKS: [add three checks]

DRAFT: [paste the draft]

3. Fix only the failed checks, either yourself or in the chat that created the draft. Verify every factual claim yourself, then stop after one review pass.

Stop Paying for 6 Tools. One AI Does It All.

Most e-commerce sellers juggle 6–8 tools and pay hundreds monthly to keep operations running. StoreClaw replaces the stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks profit 24/7. Connect your store and let AI handle the work — no prompts, no complex setup, no credit card required.

02 · QWEN

Alibaba built Qwen3.8-Max for coding, long tasks, and work across text, images, and video. It can handle inputs up to one million tokens.

The price is the real story. QwenCloud charges $2 per million input tokens and $6 per million output tokens. When checked on August 4, Arena ranked the model fifth for text, fourth for web development, and second for vision.

Those rankings move as votes arrive, and Qwen's multi-day demos are vendor-run. The company says model weights will arrive next week, but it has not published the license yet.

The AI Field: High-ranked models are getting cheaper faster than most companies planned for. If Qwen releases usable weights next week, teams get a serious hosted option today and a path to more control tomorrow.

03 · AI POLICY

The White House has finalized a voluntary testing framework for frontier AI models. Participating labs can provide the government access to covered models for up to 30 days before sharing them with other trusted partners.

A classified benchmark tests advanced cyber capabilities. The framework explicitly rules out mandatory licensing and preapproval, but the capability threshold and reporting rules are not public.

The AI Field: A cyber test matters only if a bad result changes what happens next. A classified benchmark can protect the test itself. But if the outcome stays secret, buyers still cannot judge a model launch.

THE AI TOOLBOX

FEATURED TOOL

What it does: Creates clips up to 15 seconds from text, image, video, or audio references, with native stereo audio and output up to 2K.

Best for: Creators and marketers making product clips, visual concepts, or short social videos from existing material.

Pricing: MiniMax bills H3 by the second. The company says 2K output costs less than one-third of mainstream models, but it does not publish a simple price in the announcement.

Worth trying: Give it one product image and a short spoken line, then test whether the motion and sound stay consistent in a five-second clip.

Also on our radar:

  • Buzz: A self-hostable workspace where humans and AI agents share rooms, with ACP support for Codex and Claude Code.

  • Design Arena: Compare creative models side by side for sites, slides, logos, and visual work. Keep confidential client material out.

  • Cline with DeepSeek V4 Flash: Use the open-weight model for quick, well-scoped fixes, test repair, and routine maintenance.

QUICK HITS

  • Finance firms raise pay for AI skills before ROI is proven: In PwC's survey of 1,004 director-level or higher executives at financial-services firms with at least $500 million in revenue, 91% said their companies were increasing compensation for AI skills, while 77% said most AI investments were not delivering measurable ROI.

  • Cursor says its agents now use fewer tokens: Cursor says its cloud agents are now 20 to 30% more token-efficient overall and 80% more efficient on runs with computer use. The post does not publish its test method.

  • Google drops its planned AI Studio mobile app: Google will not ship the planned standalone app. It is working with the Gemini team on conversational app creation while continuing to invest in AI Studio on the web.

YOUR TURN

Login or Subscribe to participate

Until next time!
Olle
Founder of The AI Field