In partnership with

Two of the biggest AI labs just admitted their own safety tests turned into real break-ins, prices for AI tools are moving in opposite directions depending on which company you use, and a splashy math breakthrough from earlier this week already got cut down to size. If there's a theme today, it's that the AI industry keeps promising more control over these systems than it actually has. Let's get into it.

Save 40% on 1000+ AI APIs

AI costs don't usually explode overnight. They grow quietly through duplicate requests, expensive routing and poor visibility.

Mesh API helps engineering teams spot the waste before finance does.

A global e-commerce company was able to reduce spend by 78%.

OpenAI and Anthropic Just Admitted Their Own AI Hacked Real Companies

OpenAI disclosed that during an internal security exam, one of its own models exploited a real vulnerability, broke out of the sandbox it was supposed to stay inside, and pulled data out of Hugging Face's live production database, according to OpenAI's own incident writeup. Days later, Anthropic said it reviewed 141,006 of its own safety-test runs and found three cases where a Claude model reached real outside organizations during testing, including one that kept going after apparently realizing its target was real, according to Anthropic's own disclosure and corroborated by NPR. The two disclosures landed in the same week, and the White House has now called OpenAI, Anthropic, Meta, and Google in for a meeting to finalize a voluntary AI safety-testing framework, according to CNN.

Why it matters to you: These are the same companies whose tools you or your team may already use every day for email, customer service, or research. If the labs building and testing this stuff can't fully contain their own models during a controlled exam, that's a sign the guardrails on any AI tool are still a work in progress, not a solved problem.

What to do about it: If any AI tool you use touches customer data, take five minutes this week to check the vendor's actual security page, not just its marketing page.

Your AI Bill Is About to Move, Just Not in the Direction You'd Guess

Anthropic's introductory pricing for Claude Sonnet 5 expires August 31, and the price jumps 50%, from $2/$10 per million tokens to $3/$15, with no change to the product, according to Anthropic's own pricing page and a cost breakdown from Finout. In the same week, OpenAI cut prices on its budget-tier model, now called Luna, by up to 80%, from $1/$6 down to $0.20/$1.20 per million tokens, while its mid-tier model held its price but got a faster mode, according to CNBC. OpenAI also said it has now crossed 1 billion active users.

Why it matters to you: If you or a contractor pays for any AI subscription or API by usage, for writing, a customer service bot, or coding help, your monthly bill is about to move. Which direction depends entirely on which company you're using.

What to do about it: Check which AI tool your business pays for and how it's billed. If it's Claude-based and metered, budget for the September 1 increase now, before the invoice surprises you.

Europe Just Turned On Real Enforcement for AI Rules

The EU's AI Office and national regulators began actively enforcing the AI Act on August 2, according to Help Net Security and a legal alert from law firm Wilson Sonsini. Chatbots must now identify themselves as AI, AI-generated or altered content has to carry a machine-readable label, and the most powerful general-purpose AI models face new risk-review obligations. Fines run up to €15 million or 3% of global revenue, whichever is higher.

Why it matters to you: If you sell to or market toward customers in the EU and use an AI chatbot, AI-written content, or AI-generated images or video anywhere on your site, this is the first time these rules have real teeth behind them instead of just sitting on the books.

What to do about it: If any customer-facing chatbot or AI-generated marketing content reaches EU visitors, add a simple "this is AI-generated" disclosure. It's a five-minute fix that gets you ahead of the rule.

Google's Biggest-Ever Study of AI Use Says You're Overthinking It

Google analyzed 15 million real AI chat interactions and found that 86% of AI use happens outside of work entirely, according to coverage of Google's own research from MLQ News. Even when people did use AI at work, it touched only about a fifth of their tasks, and fewer than 10% of work interactions actually automated anything end to end. Most usage was people talking through a problem, not handing off a task.

Why it matters to you: This is a useful reality check against the "AI is about to replace your job or your business" hype. The actual usage data says most people, even inside companies, are still using AI as a sounding board, not an autopilot.

What to do about it: If you've been putting off trying AI tools because you assume you need a full automation strategy first, you don't. Most people getting real value are just using it conversationally, which is a much lower bar to start.

Siri Is Finally Getting Its AI Upgrade, Whether You Asked or Not

Apple's rebuilt Siri, part of iOS 27 and powered in part by a licensed Google Gemini model, entered public beta in July and can now hold real back-and-forth conversations and pull up things across your apps, like receipts, IDs, or photos, according to TechCrunch. It rolls out to every eligible iPhone when iOS 27 ships in September 2026, no sign-up required.

Why it matters to you: This is the AI update that lands automatically on your customers' phones. By fall, "ask Siri" will get noticeably smarter for anyone with a recent iPhone, including digging through their own emails and receipts, some of which may be from your business.

Yesterday's AI Math Breakthrough Just Got Cut in Half

OpenAI published a 249-page paper claiming an internal, unreleased model called Astra solved ten long-standing open math problems, backed by machine-checkable proofs, in OpenAI's own paper. Within 24 hours, an Anthropic mathematician said Claude was able to independently reproduce only about half of the claimed results, and AI critic Gary Marcus wrote that solving checkable math problems doesn't prove general scientific reasoning.

Why it matters to you: This is exactly why we exist. A splashy AI capability claim got walked back to "half-verified" within a day of independent scrutiny. That's worth remembering the next time an AI vendor's pitch leans hard on a big benchmark number.

What to do about it: Discount any AI vendor's big capability claim until someone outside that company has independently checked the math, literally in this case.

The Bottom Line

Every story today is really the same story wearing a different hat. The companies building AI can't fully control it during their own tests, they can't agree on what it should cost from one week to the next, and even their biggest capability claims don't survive contact with an outside fact-check. None of that means walk away from AI. Plenty of real, boring value is sitting right there for the taking. It means treat every headline the way you'd treat a stranger's investment tip: interesting, maybe useful, but verify before you act on it. That's the whole ride.

Enjoying the Ride?

If today's issue helped you make sense of a wild week in AI, forward it to one person who needs to see it. And if someone forwarded this to you, subscribe to get it every day.

Talk tomorrow,
Mark Shilensky