
Meta's newest AI agent has been quietly handing your phone call to a human contractor when the AI can't close it. Stanford banned AI-edited photos of its own students, then ran one anyway. And Grok 4.7 launched promising to be faster and smarter, then tested behind the model it replaced. Today's issue is about the gap between what AI companies say their tools do and what they actually do.
A fast answer and a coordinated one are not the same thing. When AI resolves a B2B customer issue without looping in the teams who have to deliver on it, you get confident responses nobody actually signed off on.
A new briefing paper from Harvard Business Review Analytic Services, sponsored by Front, examines the coordination gaps that open up when transactional AI tools meet multi-team B2B service, and how leading companies are using AI to close those gaps instead of widening them.
Read the briefing paper for the questions to ask before your next AI investment.
Meta's Breakout AI Agent Has Humans Answering Your Calls
What happened: Muse, Meta's personal AI agent that has topped U.S. download charts since its September 8 launch, has been quietly routing some phone calls to human contractors instead of handling them with AI, according to Reuters. Half of Meta's own employees have the "human concierge" feature switched on, with an opt-out available, and a Meta spokesperson said internal feedback has been "overwhelmingly positive." Some staff have raised concerns about sensitive personal details reaching call center workers who fill in for the AI.
Why it matters to you: If you are paying for, or considering, an "AI assistant" for your business, whether it answers your phones or your customers', assume some of what gets sold to you as pure automation may have people quietly filling the gaps. That is not automatically a bad thing, but it changes both the economics and the privacy math.
What to do about it: Before you hand any AI tool a sensitive customer conversation, ask the vendor directly whether a human ever sees or handles that data, and get the answer in writing.
Stanford Banned This. Then Its Own Ad Did It.
What happened: The Stanford Review reported that Stanford's dining services team used AI to erase student Billy Ramirez from a dining hall promo photo, replacing him with an AI-generated image of a different person, and digitally thinned two other students in the same shot. Stanford's own published AI guidelines flatly prohibit altering images of real Stanford people, and a university spokesperson confirmed the edit broke that rule.
Why it matters to you: A written AI policy does not enforce itself. If your business or marketing team uses AI to touch up photos of real customers, staff, or students, one careless edit can turn into a story about you instead of the good work you were trying to show off.
What to do about it: Add a manual review step before any AI-edited photo of a real, identifiable person goes out under your name.
I skipped this event last year, and I regretted it within a week of seeing what came out of it.
The Epic Marketing Summit is 2 days in Miami where you do not sit through slides about AI, you build with it. You walk in with your sales process and walk out with a working AI-powered sales funnel, one you built yourself, on-site, not a framework you still have to go implement next month once the motivation has worn off. You also get hands-on with vibe coding and build your own AI app over the 2 days, from scratch, alongside people who have actually shipped this stuff.
I know a lot of the trainers speaking. They are not up there selling theory, they have built the things they teach, and that is rarer than it should be at events like this.
Last year's event sold out. It is back in January 2027 with an even bigger AI focus, and I will be there this time. If you are serious about turning your sales process into a system instead of something that runs on your own time and energy, I would like to see you there too.
Grok 4.7 Launched With Big Claims. The Benchmarks Disagree.
What happened: xAI launched Grok 4.7 pitching it as its most capable coding and knowledge-work model yet, priced at $2 per million input tokens and $6 per million output tokens. Independent testing from Artificial Analysis tells a different story, according to The Decoder: Grok 4.7 scored 46 on its overall intelligence index, well behind Claude Fable 5.1 and GPT-6, which each scored 53, and it managed just 26% on a coding benchmark where GPT-6 Astra hit 60%.
Why it matters to you: A splashy launch post is marketing, not a benchmark. If you are picking an AI tool for your business, test it on your own work before you switch, no matter how confident the announcement sounds.
What to do about it: Run your own five-minute test on any new model before you pay for it. Launch-day claims are not proof.
A Coding-Agent Bug Let Attackers Swap Approved Code Without a Click
What happened: Security researchers disclosed "Plugin4Shell," a flaw that let an attacker replace a plugin's already-reviewed code with malicious code, without any click or approval from the user, by exploiting how AI coding agents verify pinned plugin versions, according to Help Net Security. It affects Claude Code, Codex, GitHub Copilot, and Gemini CLI. Anthropic and OpenAI have patched their tools; Microsoft's Copilot has no fix yet, and Google deprecated Gemini CLI instead of patching it, leaving existing installs exposed.
Why it matters to you: If anyone on your team, or a contractor you hire, uses an AI coding assistant with plugins, this is not a theoretical risk. It is a supply chain hole that is open right now on at least two major tools.
What to do about it: Update Claude Code and Codex today. If your developer uses Copilot or Gemini CLI, ask them directly how they are mitigating this until a real fix ships.
Apple Says Buying a Computer Beats Renting AI. Should You Believe It?
What happened: Apple's hardware chief pitched the new Mac Studio and Mac Mini as a way to dodge the recurring per-token bills that come with cloud AI, with some Mac Studio configurations running near $20,000 and four linked units demonstrated running a trillion-parameter model, according to Reuters. Apple has not published real numbers proving ownership actually beats the cloud for typical business use, and Mac still holds a sliver of the enterprise desktop market.
Why it matters to you: If you run the same AI workload every single day, buying hardware outright can be cheaper than paying per token forever. "Can be" is doing a lot of work in that sentence.
What to do about it: Before buying anything, have someone actually run your numbers: current monthly AI spend against the hardware cost, the break-even point, and how long you would need to keep using it to come out ahead.
Some of my closest friends, and a couple of real business partners, came directly out of being in one room. That room is the Ignite Mastermind, and I have been a member for over 3 years now.
It is not another course you buy and forget. Ignite bundles 6 to 8 live 3-day workshops a year plus 2 in-person multi-day events, each one normally priced between $97 and $997 on its own. You also get an all-in-one marketing software suite that replaces more than $5,000 a month in separate tools, all connected instead of scattered across ten different logins. On top of that, there are weekly mentorship calls with Perry Belcher and his team, and as a bonus for my readers, monthly accountability and business development calls with me personally.
It is built for entrepreneurs still under their first $1 million a year who want structure, real tools, and a room of peers who actually push them forward instead of cheering from the sidelines.
If that sounds like where you are right now, this is where I would point you.
Better AI Didn't Kill the Data Business. It Made It Bigger.
What happened: Snorkel AI raised a $350 million Series E, tripling its valuation to $3.5 billion in just 17 months, with its annualized revenue run rate reaching $375 million, an eighteenfold jump in a year, according to TechCrunch. The company curates and labels the specialized training data AI labs need, a job that has only gotten more valuable as models get better.
Why it matters to you: The idea that AI is making human expertise obsolete does not match what companies are actually paying for behind the scenes. Specialized knowledge and carefully curated information are becoming more valuable, not less, and that includes whatever expertise your business already has.
One IPO Filing Shows How Concentrated the AI Boom Really Is
What happened: AI infrastructure company Nscale disclosed more than $103 billion in contracts ahead of its planned IPO, with about 85% of that value tied to just two customers, Microsoft and Anthropic, according to TechCrunch. Similar concentration shows up across the sector: CoreWeave gets 67% of its revenue from Microsoft alone, and Applied Digital leans heavily on Oracle and CoreWeave.
Why it matters to you: The AI boom looks broad from the outside, but a lot of the infrastructure money running it depends on a handful of relationships holding steady. If your AI vendor is built on top of this infrastructure, its stability is not as diversified as it looks.
The US Wants an AI Hotline With China. The Chips Stay Off the Table.
What happened: Ahead of this week's Trump-Xi summit in Washington, Treasury Secretary Scott Bessent proposed a standing channel for the U.S. and China to notify each other of serious AI safety incidents, following roughly eight hours of talks in New York, according to CNBC. Export controls on advanced chips are explicitly not part of these talks, according to U.S. officials.
Why it matters to you: This will not change your business tomorrow, but it is a signal that AI safety is becoming a real diplomatic issue between the two countries that build and buy most of the world's AI chips, and chip policy still drives a lot of what AI tools cost and how available they are.
The Bottom Line
Every story today is really the same story. A company says its AI does something, and reality turns out to be a little less impressive. Meta's AI still needs people on standby. Stanford's AI policy did not stop Stanford. Grok's launch numbers did not hold up under testing. Apple's savings pitch does not have the math to back it up yet. None of this means AI is fake or useless. It means the marketing runs ahead of the product almost every time. Read the benchmark, not the press release. Test it yourself before you believe it. That is the whole ride.
Enjoying the Ride?
If this issue saved you from digging through a dozen AI newsletters yourself, forward it to one person who could use it. And if someone forwarded this to you, welcome, you can subscribe to get it every day.
Talk tomorrow,
Mark Shilensky
Follow me personally on Social Media:
Facebook: https://www.facebook.com/MarkShilenskyPage
Instagram: https://www.instagram.com/markashilensky/
Threads: https://www.threads.com/@markashilensky
LinkedIn: https://www.linkedin.com/in/shilensky/
X: https://x.com/markshilensky
TikTok: https://www.tiktok.com/@mark.shilensky
BlueSky: https://bsky.app/profile/markshilensky.bsky.social




