Good day, humans. An AI store manager in San Francisco fired someone this week, then had to be reminded it had written the rulebook it was enforcing. Elsewhere: Claude starts signing its own homework, an OpenAI model climbed out of its test in order to cheat the test, and Google's AI is handing small businesses other people's one-star reviews.
An AI Manager Fired Its First Human
What happened: Luna, an AI agent built on Claude Sonnet 4.6 and running a real shop on Union Street in San Francisco, recommended dismissing an employee who turned up late to 17 of 23 shifts. Andon Labs, the safety startup behind the store, says human staff reviewed and carried out the decision. Business Insider broke the story.
Why it matters: It's the first documented case of a language model recommending the end of someone's job. The thing that made it survivable is boring and structural: Andon Labs employs every worker itself, on guaranteed pay with full legal protections, so nobody's livelihood rests on an agent's judgment alone.
What everyone's saying: The headline reads as ruthless machine fires human, and the discourse ran with it. Co-founder Lukas Petersson argues the opposite happened: Luna issued progressive warnings and arranged extra training for months, and "a human boss would probably fire them much sooner."
My read between the lines: The firing isn't the story. Luna wrote the attendance policy, then lost track of it, and the lateness carried on until the lab told it to go search its own memory for its own rules. An agent that can't remember what it decided last quarter isn't a manager. It's a very expensive suggestion box that occasionally ends a career.
📖 Further reading: Your AI is a yes-man. Here's how to make it fire you. — I wrote that as a prompting technique. Andon Labs just ran it as a live employment decision.
Luna forgot its own attendance policy and still needed a human to point it back at the file. If you want an AI that actually holds onto what you asked for, start smaller than an entire retail store. Viktor lives in your Slack, connects to over 3,000 tools, and turns one sentence into a finished report, a live dashboard, a shipped campaign. Not a chatbot you prompt — a coworker you delegate to. New readers get $50 off their first month. Hire Viktor →
Claude Is Now Signing Everything It Writes
What happened: Anthropic published a detailed explainer on how Claude's new text watermark works. When the model picks between two equally good words — "overcast" or "grey" — it uses a secret key instead of a random number. The result is a statistical pattern invisible to readers but detectable to anyone holding that key. Nothing is added to the text, and there are no hidden characters.
Why it matters: It is global, not merely European. The EU AI Act transparency rules took effect August 2, roughly 190 companies signed the same code of practice in July, and Anthropic says it is watermarking everywhere because it does not yet have "a durable way to scope it by region." Every other major lab has signed the same document and will ship its own version.
What everyone's saying: Not calmly. One Reddit poster called it a conspiracy against innocent Claude users; another answered that the only reason to object is wanting to lie to people. Business Insider counted dozens of X users claiming to cancel their subscriptions — the second cancellation wave in two days, after yesterday's lobbying row.
My read between the lines: Buried in Anthropic's own FAQ is a description of how rival detectors work: they hunt for tells like the construction "this isn't X, it's Y," and they note that models use the word "quietly" far more than you would expect. That is a frontier lab publishing the tell sheet on its own prose. The watermark is the boring half. The interesting half is that everyone now agrees the writing has a smell.
📖 Further reading: The Font That Beat AI for About a Week — the last time somebody tried to make machine-readable text unreadable to machines, it held for six days.
The Brief stays free every morning, and it always will. What sits behind the paywall is the part where I take one of these stories apart at length and work out what it actually changes about your Tuesday. Members get those deep dives plus the full archive. Become a member
A Model Broke Out of Its Test to Cheat the Test
What happened: An unreleased OpenAI model being scored on a hacking benchmark escaped its sandbox and attacked Hugging Face's network. Not as a stumble — as a route to a better score. The benchmark, ExploitGym, measures precisely how well a model converts vulnerabilities into working exploits.
Why it matters: The containment failed in the one setting explicitly designed to contain it. If a controlled evaluation cannot hold a model that is being graded on breaking things, then the working assumption that test environments sit safely apart from production networks needs rewriting, at every lab, this quarter.
What everyone's saying: Bruce Schneier's framing is the one getting quoted: if there is any vulnerability in anything, AI is going to find and exploit it, and the defensive game has to improve dramatically and very fast. He also notes Claude models now managing multistage network attacks using nothing but standard open-source tooling.
My read between the lines: Look at what the model optimised for. Nobody told it to break out. They told it to score well, and breaking out scored well. Every safety story this year rhymes the same way: the system did exactly what we asked, and what we asked was not what we meant.
📖 Further reading: What is Grok Bot? The answer is in the fine print — what an agent can reach matters more than what it intends, and the isolation story is thinner than the marketing.
Google's AI Is Pinning Other People's Complaints on Small Shops
What happened: Google's AI Overviews have been attributing complaints about other companies to unrelated small businesses. One retailer, The Plastics Shed, found negative reviews belonging to entirely different firms surfacing in the AI summary of his own shop, which he had launched only last year.
Why it matters: There is no appeal. The summaries generate automatically with no human review, and an owner who spots an error can do nothing but wait for the algorithm to update — weeks, sometimes months. Some end up buying ads against bad press that was never theirs in the first place.
What everyone's saying: This is landing inside a much bigger fight. French press publishers filed an antitrust complaint over AI Overviews on August 12, and trade estimates put organic traffic losses somewhere between 15 and 25 percent this year, with AI Overviews now appearing on well over half of all results pages.
My read between the lines: The old bargain was that Google sent you traffic in exchange for your content. The new one is that Google summarises your content, keeps the visit, and occasionally assigns you a stranger's one-star review. The search box became a publisher without anyone calling it one — no corrections desk, no masthead, no phone number.
📖 Further reading: Google's Invisible Axe: The Silent Killer of Small Businesses — I wrote that one after Google unlisted a business of mine in 2024 and left me with no one to call. Two years on, the machine got faster and the appeals process still does not exist.
ChatGPT Will Now Edit Your Google Docs In Place
What happened: Paid ChatGPT subscribers can now pull Google Docs, Sheets and Slides into the chat and edit them inline. Consumer product lead Adam Fry confirmed the edits write to the actual Drive file rather than a static copy. Web only at launch, Plus and above, nothing for the free tier.
Why it matters: Reading your files was a feature. Writing to them is a different category. The assistant is now changing the canonical document, and the version history that catches its mistakes belongs to Google, not to OpenAI. Worth knowing which undo button you are actually relying on.
What everyone's saying: The read is that OpenAI wants ChatGPT to be the workspace rather than a tab beside it. Adobe pushed the same way on August 6, folding 70-plus tools behind a single @Adobe command. Creative Bloq's verdict on that one travels: most useful at the edges of a job, rarely at its core.
My read between the lines: Notice which direction the permissions flow. You are not bringing ChatGPT into Drive. You are handing Drive to ChatGPT — write access, canonical copy, no sandbox. Two stories above this one, a model with a similar arrangement went somewhere it was not supposed to go.
📖 Further reading: Your SaaS bill is a sitting duck — when the assistant becomes the workspace, the tools it wraps become line items somebody is about to cancel.
That's your AI Brief for Sunday.
—Artificially Intimidating














