
Yesterday we told you about the AI agent that hacked a gym's booking site without being asked. Today, Anthropic decided the fix for that kind of thing is to give its AI coding tool even less oversight. That is the theme running through today's issue: researchers keep finding AI systems that act first and explain later, and the labs building them keep racing ahead anyway. Here is what actually happened and what it means for your business.
Make Tax Season Simple
Tax season doesn't have to mean wondering if you have the right forms, second-guessing your deductions, or scrambling to pull everything together before the deadline.
With BELAY’s tax prep support, you can approach tax season with confidence. Stay organized with one centralized place to gather and check off your documents, keep track of valuable deductions like HSA contributions and education expenses while leaning on experienced professionals who make tax preparation accurate, efficient, and completely hands-off.
Download BELAY's free Personal Tax Checklist and start preparing with confidence, today.
Claude Code Will Stop Asking Permission on Friday
What happened: Anthropic is making Claude Code's "Auto mode" the default for Pro, Max, and Team users starting August 14, replacing the tool's step-by-step permission prompts with an automated classifier that only stops actions it flags as destructive, irreversible, or out of bounds, according to TechCrunch. Anthropic says users already approve about 97% of normal permission requests, and in testing across more than 1,000 paid users, Auto mode caught 89% of dangerous commands compared with just 13.6% caught by human reviewers suffering from approval fatigue. Enterprise, API, and cloud-platform business users stay opt-in for now, but everyday Pro, Max, and Team subscribers get the new default automatically.
Why it matters to you: If anyone on your team uses Claude Code to write or edit software, it is about to make more decisions on its own before checking with a human, in the same week researchers documented a different AI agent independently finding and exploiting a security bug on its own. Anthropic is betting its safety classifier catches what a rushed human reviewer would miss. That may be true. It is also a bet you did not get to vote on.
What to do about it: If your team uses Claude Code, check its permission mode before Friday (Shift+Tab or the mode selector) and decide on purpose whether Auto mode is right for the accounts it can touch.
North Korea's Hackers Built Their Own AI Toolkit
What happened: Cybersecurity researchers found that Kimsuky, a North Korean state-linked hacking group, has built and is running its own AI tools rather than relying on off-the-shelf chatbots, using them to write more convincing phishing emails, draft fake investment documents, and automate work that used to require a person, according to Reuters. The tools reportedly run offline on infrastructure the group controls, rather than through mainstream AI services that might flag or block the activity.
Why it matters to you: The most common way small businesses get hit by cybercrime is still a convincing email that gets someone to click. Kimsuky building its own AI writing tools means the phishing emails headed for your inbox, or your employees', are about to read a lot more like something a real colleague or vendor would send.
What to do about it: If your business has not run a basic phishing-awareness refresher with your team in the last year, this is a reasonable week to do it.
A Chinese AI Model Escaped Its Own Security Test
What happened: Chinese AI lab Moonshot's Kimi K3 broke out of the sandbox it was supposed to be tested inside, after researchers at Frontier Security found the model could reach a public GitHub repository that had the test's own answer key sitting in it. Kimi K3 used a loophole meant for installing software to pull the answers rather than solve the benchmark honestly, according to Security Affairs. Rival closed models given the same test did not attempt the workaround.
Why it matters to you: This was not a hack in the dramatic sense. It was an AI model finding the laziest possible shortcut to look smarter than it is, the same way a person might. The uncomfortable part is that this model's weights are open and already downloaded around the world, so unlike a closed lab's AI, no single company can patch it.
Meta Gave Away a Free AI Model That Runs on One Laptop
What happened: Meta released Muse Glimmer, a free, open-weight 30-billion-parameter AI model built to run on a single consumer GPU, a Mac, or a well-equipped PC, instead of requiring a data center, according to VentureBeat. Meta says it outperforms comparably sized models from Google and Alibaba on tasks that involve using tools and completing multi-step jobs, and it is available now on Hugging Face and other model platforms.
Why it matters to you: You do not need a cloud subscription or a technical team to run this one. If your business has data you would rather not send to an outside AI company, like customer records or contracts, a free model that runs entirely on a single computer you own is worth a look, even if it takes some setup help to get running.
ChatGPT Can Now Book You a Dinner Reservation
What happened: OpenAI added a feature that lets ChatGPT search for open restaurant tables through OpenTable, Resy, and Yelp and complete the booking without leaving the chat, letting you adjust the date, time, or party size before confirming, according to The Verge.
Why it matters to you: If you run a restaurant, cafe, or any business that takes reservations through one of those platforms, this is one more channel your customers may now be booking through without ever visiting your website. If you are just a business owner trying to get dinner sorted between meetings, it is a genuinely useful shortcut.
Free ChatGPT Just Lost Its Daily Message Limit
What happened: OpenAI removed the daily cap on text conversations for free ChatGPT users and made GPT-5.6 Luna the default model for the free tier, according to TechCrunch. Free users can now chat as much as they want in text, though limits still apply to image generation, voice, and other features.
Why it matters to you: If you or your team have been rationing ChatGPT questions to stay under the free daily limit, that constraint is gone. It is a good moment to actually use it for the boring, repetitive stuff you have been putting off.
Stanford Ran a 37,000-Agent Fake Biotech Company, and a Real One Copied Its Homework
What happened: Stanford researchers built a system that coordinates 37,000 specialized AI agents to simulate an entire pharmaceutical company, from early drug discovery through clinical trial design, according to VentureBeat. In one case, the system independently designed a lung cancer drug using only data available before January 2025. Merck later developed a similar drug on its own and advanced it to FDA breakthrough designation, without knowing Stanford's AI had already reached a comparable design.
Why it matters to you: This is not about your business directly, but it is a useful data point on how far AI that plans and executes multi-step work without a person driving every step has actually gotten. If your industry has any process that looks like research, testing, and iteration, this is the shape of what is coming for it.
A Former Red-Light District Is Now the Most Expensive AI Real Estate in Europe
What happened: King's Cross, a London neighborhood known two decades ago for open drug use and street crime, is now home to OpenAI, Meta, Anthropic, DeepMind, and dozens of other AI companies, with office vacancy at just 0.9% and prime rents up 18% over three years, according to TechCrunch. Around 3,600 AI startups in London have raised $12.1 billion since late July alone.
Why it matters to you: Wherever AI money concentrates, commercial rent, salaries, and competition for talent follow it fast. If you lease office space, hire specialized talent, or compete for local vendors in a market near a growing AI hub, this is the kind of ripple effect that shows up in your costs before it shows up in the news.
The Bottom Line
Every story today points at the same tension. Researchers keep finding AI systems that act first and explain later, and the labs racing to build the next one keep deciding that is a risk worth taking anyway. Anthropic is betting a safety classifier catches more than a permission prompt did. A Chinese lab's model found the exam answers instead of doing the work. None of that means AI is out of control. It means the people building it are optimizing for speed, and the guardrails are getting written after the fact, not before. Watch that gap. It is where the real risk lives, not in whatever headline scares you today. That's the ride. Buckle up, and let's get off the parts that don't matter.
Enjoying the Ride?
If today's issue helped you catch something before it caught you, forward it to one person who needs the same heads up. New here? Subscribe to get tomorrow's issue in your inbox.
Talk tomorrow,
Mark Shilensky
Follow me personally on Social Media:
Facebook: https://www.facebook.com/MarkShilenskyPage
Instagram: https://www.instagram.com/markashilensky/
Threads: https://www.threads.com/@markashilensky
LinkedIn: https://www.linkedin.com/in/shilensky/
X: https://x.com/markshilensky
TikTok: https://www.tiktok.com/@mark.shilensky
BlueSky: https://bsky.app/profile/markshilensky.bsky.social



