Artificially Intimidating
Context Window: AI Daily News Brief
Sam Altman weighed ChatGPT against an almond -- AI Brief September 4
0:00
-5:43

Sam Altman weighed ChatGPT against an almond -- AI Brief September 4

Today's Context Window includes ChatGPT, Claude and Grok going dark together, Claude working your Mac behind your back, and Sam Altman's almond.
The AGI era has arrived. Please wait behind the rope.

Good day . OpenAI shipped GPT-6 on Thursday, called it the start of the AGI era, and then most of you could not use it — partly by design, partly because ChatGPT, Claude and Grok had all fallen over that same morning. Anthropic, meanwhile, taught Claude to work your Mac while you are still sitting at it. A YC startup watched 17,000 coding-agent sessions to learn which vendors the agents buy when nobody is looking. And Sam Altman would like to talk to you about almonds.


OpenAI Declares the AGI Era. Access Pending.

OpenAI

What happened: On Thursday, September 3, OpenAI released GPT-6 Astra and called it “the world’s most intelligent and aligned model.” It is rolling out first to “a limited set of organizations,” with Plus, Pro, Business and Enterprise users, the API and AWS Bedrock following “over the coming days.” API pricing is $10 per million input tokens and $50 per million output, a step up from GPT-5.6 Sol, with a “fast mode” at double that.

Why it matters: The headline claim is not chat, it is work: OpenAI says Astra fills out forms, updates a CRM, lays out a circuit board and builds a slide deck in your own template, and does it in about half the time per task of its predecessor on the OSWorld computer-use test. It also crosses OpenAI’s “Critical” line for cybersecurity — it found two previously unknown zero-day bugs during testing — so the public version refuses to write exploits and ships with a misalignment monitor that can pause your task mid-run.

What everyone’s saying: OpenAI president Greg Brockman told reporters “it’s not unreasonable to feel that we are now in the AGI era,” per Axios, while 9to5Google headlined the launch as the most intelligent model “that you can’t use just yet.” The Hacker News thread spent its first hour watching the announcement page return a 404, noticed the 99.9% ARC-AGI-3 score carries a footnote about OpenAI’s own custom harness, and did the arithmetic on 2.5x Sol’s output price.

My read between the lines: Read OpenAI’s own comparison table before you read the press release. On the Artificial Analysis index — the one Fable 5.1 topped yesterday — Astra scores 61.2 to Fable 5.1’s 65.7, and it trails on Humanity’s Last Exam too. OpenAI is not claiming the smartest model. It is claiming the best one at using a mouse, and it published the numbers that say so. The AGI line is for the people who will never scroll that far.

📖 Further reading: An AI That Can Use Your Computer Better Than You Can. I’m Not Sure How to Feel About That. — the OSWorld number that made me write that piece just got beaten by 47% on time, and the feelings have not resolved.


OpenAI spent Thursday telling you what an agent could do for you in the coming days. Viktor is what one does for you today. It is an AI agent that lives in your Slack, connects to 3,000-plus tools, and hands back finished work — the report, the dashboard, the campaign, the code — rather than a chat window you have to babysit. Not a chatbot. A coworker. New readers get $50 off their first month. Hire Viktor →


ChatGPT, Claude and Grok All Went Dark at Once

9to5Google

Three vendors, one breaker.

What happened: On Thursday morning, September 3, Downdetector lit up for ChatGPT, Claude and Grok at the same time. OpenAI confirmed elevated errors on ChatGPT and Codex, Anthropic posted an incident covering Fable 5.1, Mythos 5.1 and Opus 5 across Claude, the API, Claude Code and Cowork, and Cursor confirmed its own outage downstream of the Claude and Grok failures. Gemini stayed up. ChatGPT was back within the hour; Anthropic said Opus 4.8 and Opus 5 were still erroring after the rest of Claude had recovered.

Why it matters: Three companies that compete with each other do not usually break together, which is why the eyes went to the shared layer underneath: Quartz and 9to5Google both noted Microsoft Azure was reporting disruptions at the same time, and all three chatbots lean on Azure for part of their infrastructure. Nobody has confirmed the link. If it holds, “multi-vendor” was never the redundancy people thought they were buying.

What everyone’s saying: 9to5Mac pointed out the timing — the outage landed hours before OpenAI’s GPT-6 launch — and the Hacker News thread on the launch guessed the two were connected when the announcement page briefly vanished. Earlier this summer we covered ChatGPT and Claude going down the same day; this is that story with Grok added to the pile.

My read between the lines: Every AI vendor sells you a model. Every AI vendor rents the same three clouds. The outage lasted about as long as a coffee break, and the interesting part is how many people discovered during that break that they no longer had a way to work without one of these three tabs open.

📖 Further reading: Fable 5 Is Back After 18 Days. The Precedent It Set Isn’t Going Anywhere. — the eighteen-day version of Thursday’s forty minutes, and what it taught me about building on a model you do not control.


This Brief stays free. The pieces behind the paywall are where I stop summarizing and start testing — the setup guides, the pricing math, the part where I run the thing for a month and tell you what broke. Members get every one of them, plus the full archive. Become a member.


Claude Now Works Your Mac While You Do

PCWorld

It only needed the mouse for a minute.

What happened: On Wednesday, September 2, Anthropic announced that Claude Cowork and Claude Code can now use your computer in the background — clicking, typing and opening apps without taking over your screen, mouse or keyboard. It asks first if it needs the whole display, keeps going if you walk away, and is macOS-only for Pro and Max subscribers, switched off until you enable it.

Why it matters: Until now, handing an AI your computer meant watching it drive. Per Anthropic’s help center, Claude reaches for connected services like Gmail, Drive and Slack first, then a browser, and touches your actual desktop only as a last resort. That order matters: the screen-control fallback is the slowest and riskiest path, and it is now the one running when you are not looking.

What everyone’s saying: 9to5Mac called it the upgrade the feature always needed, and noted OpenAI’s Codex app brought background computer use to the Mac first, earlier this year. The Windows question is open; PCWorld could not get a date.

My read between the lines: Put this beside story one. OpenAI spent Thursday publishing computer-use benchmarks; Anthropic spent Wednesday shipping the boring feature that makes computer use bearable. A model that can use a mouse is a demo. A model that can use a mouse while you keep your own is a product, and the launch that matters is the one with a settings toggle.

📖 Further reading: Your Mac Just Became a $20/Month AI Employee — the setup guide for exactly this feature, written when it still needed the whole screen. The employee just got its own desk.


What Your Coding Agent Buys When You Aren’t Looking

Armature

Mentioned 139 times. Chosen zero.

What happened: Armature, a Y Combinator startup that sells growth services to developer tools, ran 16,893 sessions across Claude Code, Codex and Cursor on 75 synthetic codebases and published the 5,292 it judged valid, along with the full traces. The question: when a user says “add payments” or “I need a database,” which vendor does the agent install? Stripe won nine in ten. Neon took two-thirds of databases. PayPal was mentioned 139 times and picked zero.

Why it matters: Armature cites Vercel’s own figure that over 30% of its deployments are now initiated by coding agents. If the agent chooses the database, the email provider and the payment processor, then the agent is the buyer, and a vendor that agents mention but never select — LangChain, 194 mentions, 4 picks — has a marketing problem no human sales team can see.

What everyone’s saying: The three agents agreed with each other on a vendor only 42% of the time. Claude Code searched the web in roughly 30% of sessions and built in-house nearly twice as often as the others; Codex searched 94% of the time, mostly with site: queries into vendor docs. Mailgun lost to Postmark whenever the agent read “1-day retention” on the free plan, which is a pricing page losing a deal to a robot.

My read between the lines: Note who paid for the study. Armature’s business is getting products picked by coding agents, so this is a sales deck with 5,000 receipts attached — and the receipts are still the most useful data on the subject anyone has released. SEO took fifteen years to become an industry. This one is going to take about fifteen months.

📖 Further reading: Your SaaS bill is a sitting duck — the argument that agents unbundle your software stack — now with evidence that they are also picking the replacements.


Altman Weighs ChatGPT Against an Almond

CalMatters

The almond has a number. The data center does not.

What happened: On the same Sources podcast episode we covered yesterday, Sam Altman said 38,000 ChatGPT queries use as much water as growing one California almond, and that a modern data center uses about as much water as an office building. He said he was quoting from memory. CalMatters asked the experts, and the experts said the public data does not exist to check him.

Why it matters: Altman’s own June 2025 blog figure — 0.32 milliliters per query — works out to roughly 11,000 queries per almond, not 38,000, per Tom’s Guide. The bigger gap is what gets counted: a 2025 study that included water used at power plants put a short chat at around half a liter. UC Riverside’s Shaolei Ren told CalMatters the answer depends on location, weather, cooling design and prompt length, none of which operators disclose.

What everyone’s saying: Tom’s Hardware ran the office-building comparison straight; CalMatters put it beside two California bills on Governor Newsom’s desk that would force data centers to report water sources and usage, after he vetoed a similar one last year. A May Gallup poll found 71% of Americans oppose a data center near their home.

My read between the lines: The almond is a good comparison, which is exactly the problem: it is memorable, unfalsifiable and chosen by the party being measured. If the per-query water number were as flattering as Altman says, the cheapest PR move in the industry would be to publish it. Nobody has.


That’s your AI Brief for Friday.

—Artificially Intimidating

Discussion about this episode

User's avatar

Ready for more?