
AI is now producing work faster than people can check it, and this week four of the biggest AI labs sat under oath and declined to guarantee their own agents. That is the thread running through today's issue: more output, more power, and the same old question of who is double-checking. Let's start with the math.
What if AI made us prove we’re creative?
This year AI shook up creative strategy. Most teams are still working like it didn't.
Ari Murray, Chief Digital Officer at Salt and Stone sits down with Bryan Harvey, Head of Brand at Hightouch, for an honest conversation about what's breaking on creative teams right now, how the roles are changing, and what a day looks like when creatives spend it on ideas, concepts, and craft while agents handle the manual execution.
Leave with new insights from 3 experts in the field and 5 prompts to instantly improve your creative output.
OpenAI Dropped 722 Math Papers Written by an AI
What happened: OpenAI published a large batch of new math results from an internal model it has not released, posted on GitHub. As widely reported, the batch is 722 manuscripts grouped into 372 families, from roughly 4,000 problems the model was given. OpenAI says the average result took about three hours of ChatGPT Pro thinking time. Many of the proofs come with Lean files, which let a computer check each logical step, but OpenAI says more are still being added, so not everything has been machine-verified yet.
Why it matters to you: You will never read these papers, and neither will most mathematicians. The real story is the pattern: AI can now generate expert-looking work far faster than humans can verify it. That applies to your contracts, reports, and customer replies too. Fast output is cheap. Checked output is the scarce part.
What to do about it: For anything important that AI drafts for you, decide in advance who checks it and how.
Mistral's Trillion-Parameter AI Is Cheap, and the Weights Go Public This Month
What happened: French AI company Mistral launched Mistral Large 4, nicknamed "Le Chonk." It has about 1 trillion total parameters, handles text and images, and is priced at $1.36 per million input tokens and $4.18 per million output tokens. A preview is live now, and Mistral says the open weights, meaning the files anyone can download and run themselves, will be released by the end of October. Its benchmark claims, including strong cybersecurity scores, come from Mistral itself and have not been independently confirmed.
Why it matters to you: Open models put pricing pressure on the big paid ones, and they let companies run AI on their own servers instead of sending data to someone else's cloud. If you have a developer or an IT vendor, it is worth asking whether an open model could cut your AI bill or keep sensitive data in-house.
Most business owners I talk to are not short on ideas. They are short on a system that runs without them. That is exactly why I am excited about the Epic Marketing Summit, a 2-day live event in Miami this January where you build a working AI-powered sales funnel and vibe code your own AI app on site. You walk out with a finished asset, not a binder of notes you still have to go implement.
It is built for owners who want to turn their sales process into a system instead of running it on their own time and energy. Last year's event sold out, and it is back with a bigger AI focus. I will be honest about my own angle: I missed last year and had serious FOMO, because I know a lot of the trainers, and they were not teaching theory. They had actually done the work. I am going in January, and if you are serious about building leverage into how you sell, I would love to see you there too. Come build it with us. Grab your spot here, because the closer we get to January, the fewer seats there will be.
Anthropic Widens Access to Its Strongest Cyber AI and Posts the Numbers
What happened: Anthropic expanded its Cyber Verification Program into three tiers for verified security professionals: Defense Access, Red Team Access, and Specialized Access for safety-critical systems like power grids and flight systems. Anthropic says partners found at least 129,000 verified software vulnerabilities between April and July, and its own scanning found 5,500 more, with over 33,000 rated critical or high severity. In its own test, the defense tier was blocked on 46 of 50 trials, while the red team tier had no blocks.
Why it matters to you: The same AI that finds holes for defenders can find them for attackers, which is why access is gated. Software you rely on is likely to get a wave of security patches. Those update notices are not noise.
What to do about it: Turn on automatic updates for your business software, and ask your IT person how quickly patches get applied.
OpenAI Apologizes in Australia Over a Medicare-Related Breach
What happened: OpenAI's chief strategy officer, Jason Kwon, apologized at an Australian parliamentary inquiry over a breach involving Medicare data, according to ABC News Australia. Kwon admitted OpenAI should have told the government sooner. He said staff learned of the hack several weeks before CEO Sam Altman did, and that the company now alerts staff when its models use the internet in ways they should not during training. OpenAI also says it disclosed a separate breach involving NSW Parks and Wildlife much faster, and it is forming a task force.
Why it matters to you: The coverage I reviewed did not spell out the technical details of how the breach happened. What it does show is how slowly bad news can travel even inside a company that builds this technology. If you let AI tools touch customer data, you want to hear about problems in days, not weeks.
Claude Now Works Inside Google Docs, Sheets, and Slides
What happened: Anthropic released a beta Claude sidebar for Google Docs, Sheets, and Slides on its Pro, Max, Team, and Enterprise plans, per its support page. By default it describes each edit and waits for you to click Allow, and there is an "accept all edits" mode for ordinary changes. It only asks Google for access to the one open file, not your whole Drive or Gmail. Anthropic notes that it is not covered for HIPAA-regulated data, and admins on Google Workspace have to allow the install first.
Why it matters to you: If your team lives in Google's apps, this means AI help without copying and pasting between tabs. Cleaning up a messy spreadsheet or tightening a proposal is the kind of everyday task where it should save real time. This is also a beta, so expect rough edges.
What to do about it: Try it on a low-stakes document first, and leave it on "ask before edits."
Let me tell you about the single best business decision I have made that did not involve a spreadsheet. I have been a member of the Ignite Mastermind for over 3 years, and it has been a vital part of how I have grown my business. Some of my closest friends, and a few real business partners, came directly out of being in that room.
Ignite is an annual membership built for entrepreneurs still under $1 million a year in revenue who want structure, tools, and peers to get past that milestone faster. You get 6 to 8 live virtual workshops a year, tickets to 2 in-person multi-day events, and weekly calls with Perry Belcher and his team. There is also a connected software suite that replaces over $5,000 a month in separate marketing tools. As a bonus, anyone who joins through me gets monthly accountability and business development calls with me personally. If you are looking for a mastermind to move you to your next milestone, this is where I would point you. See what is inside Ignite here.
Meta, Sierra, Shopify, Stripe, and Walmart Back a Rulebook for AI Shopping Agents
What happened: Meta and Sierra introduced the Personal Agent Protocol, an open standard for how consumers' personal AI agents deal with businesses. Shoppers decide what access their agent gets, including read-only or write access. Companies decide what agents may do on their websites, APIs, or their own agents. Partners include Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart. The first v0.1 specification is planned for later this month.
Why it matters to you: Soon, some of your customers will send an AI assistant to compare prices, ask questions, or place orders instead of visiting your site themselves. Standards like this decide whether your business is easy for those agents to work with. It is early, and rival standards exist, so there is no need to rebuild anything yet.
What to do about it: Make sure your hours, prices, and policies are clearly posted on your website, in plain text.
The Top 1% of AI Payers Spend $903 a Month
What happened: a16z's latest consumer AI report, summarized by The Neuron, found that 4.5% of eligible U.S. consumers paid for ChatGPT, Gemini, or Claude in August, up from 2.1% a year earlier. The median payer spent about $25 a month. The top 1% averaged $903 a month and accounted for 19.5% of observed consumer AI spending, more than the bottom half of payers combined. The data is a U.S. card-spending panel, so a16z cautions against reading it as company revenue.
Why it matters to you: Most people who pay for AI pay a modest amount, while a small group of heavy users spends a lot. For you, that means the useful question is not which tool is trendiest, but whether one subscription saves hours each week. If it does, paying more can make sense. If it does not, $25 a month is plenty.
Four AI Giants Testified Under Oath and Would Not Promise Their Agents Are Safe
What happened: At a New York City Council hearing, representatives from OpenAI, Anthropic, Meta, and Google were asked to guarantee safety, according to amNY. None gave a blanket promise that failing a safety test would stop a model's release, and none gave a probability for a catastrophic event. Google said there is no rigorous scientific method yet to assign one. OpenAI is also reviewing possible past agent misalignment incidents dating to November 2025. Days earlier, five companies signed a voluntary White House safety agreement that carries no binding requirements. Elon Musk's SpaceXAI skipped the hearing despite a subpoena.
Why it matters to you: Refusing to promise perfection is honest, but it also tells you where responsibility lands: on whoever uses the output. No vendor is going to guarantee an AI's answers for you.
What to do about it: Before acting on an AI answer you cannot undo, have a person confirm the key facts against the original source.
The Bottom Line
Look at the thread today. OpenAI generated 722 papers nobody has had time to fully check. Mistral is about to hand a trillion-parameter model to anyone who wants it. Anthropic is widening access to its most powerful cyber tools. And when asked under oath to stand behind their agents, the labs said they could not. Nobody is going to do your double-checking for you. The winners in this stretch will not be the people who adopt every new thing the fastest. They will be the ones who use AI for speed and keep a human on the part that matters. Stop the rollercoaster, get off at the stop that makes sense for your business, and keep your hands on the safety bar.
Enjoying the Ride?
If today's issue helped, forward it to one person who runs a business and is trying to make sense of all this. And if someone forwarded this to you, you can subscribe so you never miss an issue.
Talk tomorrow,
Mark Shilensky
Follow me personally on Social Media:
Facebook: https://www.facebook.com/MarkShilenskyPage
Instagram: https://www.instagram.com/markashilensky/
Threads: https://www.threads.com/@markashilensky
LinkedIn: https://www.linkedin.com/in/shilensky/
X: https://x.com/markshilensky
TikTok: https://www.tiktok.com/@mark.shilensky
BlueSky: https://bsky.app/profile/markshilensky.bsky.social




