Good day . Five stories today, and every one of them is really about a price tag. Sam Altman thinks the industry is putting up too many data centers. Anthropic took the benchmark crown and raised your invoice in the same release. An agent valued at $2.5 billion asked a reviewer for his Google password. Google taught Gemini to skip the boring parts of a video. And somewhere in Berlin, a box the size of a lunchbox is running a 284-billion-parameter model with no meter attached.
Altman Calls the Compute Boom “Unsustainable Silliness”
What happened: On the debut episode of Alex Heath’s Sources podcast, released September 1, OpenAI CEO Sam Altman said he is seeing “the first signs of what feels to me like unsustainable silliness” — new “neocloud” companies promising gigantic amounts of compute next year without the revenue or customers to pay for it. He carved out his own company: “I’m not worried about our compute buildout plans. I am worried about the world’s compute buildout plans.”
Why it matters: A neocloud rents out GPUs — the specialized chips AI runs on — roughly the way a landlord rents apartments, and dozens have piled in, including former Bitcoin miners, on the bet that demand outruns supply forever. If Altman is right, some of them are pouring concrete for warehouses nobody has signed a lease on, and the write-down lands on investors rather than on OpenAI.
What everyone’s saying: Traders are reading it as a sorting signal, Benzinga notes — CoreWeave and Nebius have contracted backlogs, while pivoted miners like IREN, Hut 8 and Cipher Mining have far less locked in. Binance founder Changpeng Zhao added on September 2 that “hot money” is rotating back out of AI and into crypto.
My read between the lines: The largest buyer of compute on Earth has advised everybody else to stop building it. Altman even laid out the mechanism on the podcast — if OpenAI drives compute costs down, “some people that made dumb financial decisions” get caught — which is a competitive strategy delivered in the voice of a weather forecast.
📖 Further reading: Neo-Napster: The Compute Revolution Nobody Saw Coming — the case that serious compute drifts to the edge, which is precisely the demand curve the neoclouds are betting against.
Every story in today’s brief is somebody counting what AI costs them. Here is the other column. Viktor is an AI agent that lives in your Slack, connects to 3,000-plus tools, and comes back with the finished thing — the report, the dashboard, the campaign, the shipped code — instead of a conversation about the thing. Not a chatbot you prompt. A coworker you delegate to. New readers get $50 off their first month. Hire Viktor →
Fable 5.1 Won the Benchmark and the Bill
What happened: Anthropic shipped Claude Fable 5.1 and Mythos 5.1 alongside a 75% cut to cache-read pricing, and Fable 5.1 took the top score on the Artificial Analysis Intelligence Index. It also burns roughly 1.7 times the output tokens of Fable 5 to get there, so the cost of a single benchmark task rose about 20%, to $3.76.
Why it matters: Models bill by the token, and “thinking longer” is not free — a model that reasons its way to a better answer using more words costs more to run even when the per-token price falls. The cache discount saves roughly $1.40 per task; the extra verbosity eats that and keeps going.
What everyone’s saying: Latent Space flagged the same split — new state of the art, 75% cache cut, 70% more output tokens — and community analysis there suggests Fable and Mythos 5.1 ship identical weights, differing only in safety-classifier thresholds and fallback routing. OfficeChai put it more bluntly: it is now the most expensive model on the index, running 57% above Opus 5.
My read between the lines: Yesterday we wrote about the bill that isn’t in the repo — same trick, different invoice. A headline discount on the cheapest input you buy is a number you feel in a press release, not in a P&L, and the figure that actually moved is the one nobody puts on a launch slide.
📖 Further reading: Fable 5 Costs 2x Opus — and Using It Wrong Costs You More Than That — the routing rules in there just picked up a new price column.
The Brief is free and stays free. What sits behind the paywall is the part where I take something apart — the pricing math, the fine print, the thing that only shows up after you have run it for a month. Members get all of it, plus the full archive. Become a member.
Your AI Agent Would Like Your Password Now
What happened: Behind the Craft ran four personal AI agents — Instinct, Grok Bot, ChatGPT and Hermes — through real tasks and read each one’s privacy policy. Instinct, freshly valued at $2.5 billion after a $250 million Series B, asked the author for his Google password and his two-factor code in order to finish a job.
Why it matters: A two-factor code is the last thing standing between a stranger and your email, and handing one to software means that software is now you, everywhere, with nothing left to tell you apart from an intruder. These agents work by logging in as you on a cloud machine that keeps running after you shut your laptop.
What everyone’s saying: The review’s conclusion is that the more seamlessly one of these products works, the harder it becomes to audit what it actually did. The New Stack notes that Grok Bot’s own documentation calls its per-bot screens “separate work surfaces, not separate security boundaries,” and advises keeping credentials off the machine entirely if any bot on the account should not reach them.
My read between the lines: We spent twenty years teaching people that nobody legitimate ever asks for a 2FA code, and it has taken about eighteen months to talk them back out of it. The tell is that this is a product decision, not a technical wall — passing the credential is simply the cheapest way to ship an agent, and the industry is finding out whether convenience buys back the reflex.
📖 Further reading: What is Grok Bot? The answer is in the fine print — one of the four agents tested here, and the fine print turns out to be the entire story.
Gemini Learned to Skip the Boring Parts
What happened: On September 1 Google switched on “agentic video understanding” across Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. Rather than chopping a video into one frame per second and reading all of it, the model now decides which moments to watch, at what speed, and whether to lean on frames, audio or the transcript.
Why it matters: Google reports up to 88% fewer tokens, up to 66% lower analysis cost and up to 7% better accuracy — the rare release where the cheaper option is also the more accurate one. It is live now in the Gemini API and AI Studio with no extra feature fee, and Google says it will power YouTube’s “Ask YouTube” answers in the coming months.
What everyone’s saying: The Decoder framed it as the obvious fix to a bad default: a fixed frame rate means paying full freight to stare at a three-hour lecture of one static slide. Developers are most interested in sub-second moment retrieval, which catches cuts and state changes that one-frame-per-second sampling missed entirely.
My read between the lines: Set this beside today’s Fable story and you have the whole 2026 argument in two data points: one lab making the model think longer, another teaching it to look less. Google did not build a better video model here. It built one that knows when to stop reading, which is a cheaper thing to sell and a much harder thing to put on a leaderboard.
📖 Further reading: I found 350,000 tokens hiding in plain sight — the same lesson one layer up: most token spend goes on input nobody needed to send.
192GB of RAM Fits in a Lunchbox Now
What happened: Ahead of IFA opening in Berlin on Friday, September 4, a wave of roughly two-liter desktops built on AMD’s Ryzen AI Max+ PRO 495 arrived carrying up to 192GB of unified memory. ACEMAGIC says its F9A Pro ran DeepSeek V4 Flash — a 284-billion-parameter model — locally, and BOSGAME’s M5 MAX is expected to ship between late September and mid-October at $3,600 to $3,800.
Why it matters: Unified memory means the processor, the graphics and the AI accelerator all draw from one pool, so the ceiling on what you can run at home is now the RAM number rather than the graphics card. A 284-billion-parameter model sitting on a desk means no per-token bill, no rate limit, and no data leaving the building.
What everyone’s saying: Acer is pushing the same class of silicon into a laptop — the Aspire G 3D 16, with 128GB and a glasses-free 3D display, per Notebookcheck — while Framework, GMKtec, Minisforum and GEEKOM are all at the show with variants of their own. The category barely existed two years ago.
My read between the lines: Thirty-seven hundred dollars buys roughly ten months of a serious API habit, which is exactly the arithmetic the entire cloud AI business would prefer you never sat down and did. Note the tension with the top of this brief: Altman is worried about too many data centers going up at the precise moment the interesting compute started fitting under a monitor.
📖 Further reading: Your laptop has been in the way this whole time — the case for moving the work off your machine, now arguing with a lunchbox that runs a 284B model.
That’s your AI Brief for Thursday.
—Artificially Intimidating















