Artificially Intimidating
Context Window: AI Daily News Brief
Big AI Just Started Grading Its Own Homework -- AI Brief September 27
0:00
-5:42

Big AI Just Started Grading Its Own Homework -- AI Brief September 27

Today's Context Window: ads die when agents replace browsing, Grok Bot gets your bank login, AI aces homework at CMU, and DeepSeek's 3M-sandbox stack.
A hand-drawn illustration of three AI-lab towers -- OpenAI, Anthropic and Google -- with a robotic arm reaching out of the OpenAI tower to stamp its own report card A+, while an unplugged, chained-shut U.S. Capitol sits dark below.
The report card is theirs. So is the stamp.

Good day . OpenAI, Google, and Anthropic are building their own AI safety regulator -- no government required. Grok Bot just got the keys to your bank account. And a Carnegie Mellon professor explains, in exhausting detail, what happens to a class once AI can do all the homework. Let's get into it.


Big AI Is Writing Its Own Safety Rules TheStreet

  • What happened: Google, OpenAI, and Anthropic are building a self-regulatory body called the Standards Authority for Frontier AI (SAFA), targeting a late-2026 or early-2027 launch. The three labs originally wanted a public-private partnership with federal oversight, but that stalled after a draft White House executive order was put on hold and the administration told them to reach industry consensus first.

  • Why it matters: If SAFA becomes the industry's real rulebook, the companies being tested help write the test: third-party model checks, incident reporting, and what counts as “responsible AI” all get defined by the labs building the thing being judged, before any regulator weighs in.

  • What everyone's saying: Backers pitch it as FINRA-for-AI, complete with a bipartisan CEO shortlist (a former Biden science adviser and a current White House AI adviser). Meta, xAI, and Nvidia pushed back publicly at Dreamforce, arguing SAFA is a deal among three labs, not a consensus across the field -- and that the biggest labs auditing themselves against a bar their smaller rivals never helped set is the whole problem, not the solution.

  • My read between the lines: SAFA can set standards, but without government registration it almost certainly can't punish anyone for missing them -- which makes it less a regulator than an expensive permission slip the industry writes for itself, then grades with a ruler it also drew. Today's other story about a professor giving up on take-home grading because it can't be trusted anymore is, unintentionally, the perfect footnote.

📖 Further reading: Washington Found the Off Switch for Anthropic -- worth a re-read now that the industry wants to write its own switch instead of waiting for Washington to flip it.


Three AI labs just spent months arguing over who gets to grade their own homework. Your team doesn't need a committee -- it needs the homework done. Viktor is an AI agent that lives in Slack, plugs into 3,000+ of your tools, and hands back the finished thing: the report, the dashboard, the campaign, the code fix. Not a chatbot waiting on your next prompt. A coworker you assign work to. New readers get $50 off their first month. Hire Viktor →


The Ex-Twitter CEO Betting Ads Die With Agents 20VC

A hand-drawn illustration of a roadside ADS billboard flickering off above a human tollbooth, while a line of small robots walks through a hole cut in the fence beside it, straight into a glowing globe-shaped tunnel.
The humans still pay the toll. The robots found the hole in the fence.
  • What happened: Former Twitter CEO Parag Agrawal told 20VC that online advertising stops working once AI agents, not humans, do most of the web browsing -- there's no eyeball left to sell to. His startup Parallel, already valued at $2 billion after a Sequoia-led raise, is building what he calls “AdSense for agents”: a system that pays content owners per agent visit instead of per human click.

  • Why it matters: Nearly every website that survives on ad revenue today is betting that human traffic keeps showing up. If agents really do browse the web “1,000 times more than humans,” as Agrawal claims, the economic floor under the free internet shifts from eyeballs to API calls, and it shifts fast.

  • What everyone's saying: The skeptics point at Agrawal's own math: he admits current web-search pricing needs to drop roughly 10x before agent-scale browsing pencils out at all -- so the “AdSense for agents” pitch is as much a bet on future compute costs collapsing as it is on agents actually showing up in force.

  • My read between the lines: The guy who ran a company built entirely on human attention is now the guy telling you human attention is about to become worthless online -- which is either the most honest pivot in tech, or the best-timed I-told-you-so a founder has ever pre-loaded for himself.

📖 Further reading: A Publisher's Perspective on The Bleak Future of Google's AI-Powered Search -- the view from the other side of the traffic cliff Agrawal is now betting his company on.


Grok Bot Can Now Touch Your Bank Account wccftech

A hand-drawn illustration of a giant bank vault door embossed with an X mark, wide open, as a small faceless robot wheels away a wheelbarrow of credit cards and cash after being handed a giant brass key.
The vault door has an X on it. So does the plan for who's watching.
  • What happened: In an interview with China Media Group, Elon Musk admitted Grok trails Anthropic's models “by years” -- even as he announced that Grok Bot, xAI's autonomous agent platform, can now link directly to users' bank accounts, credit cards, and investment accounts through a new Finance integration.

  • Why it matters: This is the first time a mainstream consumer AI agent has been handed standing, direct access to someone's actual money rather than just advice about it -- a meaningfully different risk category than a chatbot that suggests you make a budget.

  • What everyone's saying: Coverage keeps circling back to Musk's own concession that Grok “trails Anthropic by years” on raw model quality, with Musk framing xAI's bet as physical-world engineering, using Tesla and SpaceX data, rather than trying to out-code a rival that's already ahead on software tasks.

  • My read between the lines: Musk just told the world his AI isn't the smartest one in the room, then handed that same AI standing access to everyone's checking account -- which either means Grok Bot's finance tooling has been genuinely bulletproofed, or “smart enough” became a much lower bar than “connected enough.”

📖 Further reading: What is Grok Bot? The answer is in the fine print -- we read the fine print on the cloud computer and the logins two days in -- here's what the new Finance integration adds to that picture.


The Brief is free and always will be. Today's self-appointed AI safety regulator is exactly the kind of story I take apart properly behind the paywall -- who's really accountable, and to whom. Members get those deep dives plus the full archive. Become a member →


A CMU Professor Rebuilt His Class Around AI Christian Kästner

A hand-drawn illustration of a conveyor-belt machine that turns handwritten homework into stacks of perfect A+ papers, as a bow-tied professor stamps them with a SEE A TA rubber stamp instead of grading them.
The machine can produce the A+. The professor just wants to talk to the kid who turned it in.
  • What happened: Carnegie Mellon professor Christian Kästner writes that AI agents can now complete his entire “Machine Learning in Production” course -- reflections, coding assignments, even the write-ups -- so he's redesigned most of his assessments around 15-minute in-person TA check-ins and oral debriefs instead of graded take-home work.

  • Why it matters: This isn't a professor banning AI or pretending it doesn't exist -- he explicitly lets students use it everywhere except written exams. It's a working blueprint for what “grading” looks like once take-home work can no longer prove that anyone actually learned anything.

  • What everyone's saying: The Hacker News thread, past 190 points, splits into two camps: instructors trading their own workarounds (quizzes tied to homework, penalties on resubmission, oral defenses) and a smaller, louder group arguing the whole system should stop pretending and just test understanding exclusively in person.

  • My read between the lines: The most telling line in the whole piece is a teaching assistant who felt “silly” grading a solution whose own commit message read “Authored by Claude Code.” Once the paper trail admits the human didn't do the work, pretending to grade the human's understanding of it stops being a teaching exercise and starts being theater.


DeepSeek's Machine Runs 3 Million Sandboxes a Day arXiv

  • What happened: DeepSeek published a 31-page paper, with more than 130 listed authors including founder Liang Wenfeng, detailing DSec (DeepSeek Elastic Compute) -- the sandbox infrastructure it uses to train AI agents at scale. A single production unit spans roughly 160 nodes, handles about 3 million disposable practice environments a day, and can spin up more than 5,000 new ones per second.

  • Why it matters: Training an AI agent well means giving it thousands of realistic, throwaway environments to practice in -- file systems, terminals, fake company software -- and DSec is DeepSeek showing its work on the unglamorous plumbing that makes large-scale agent training possible at all.

  • What everyone's saying: Researchers are treating the paper less as a model announcement and more as an infrastructure playbook -- the kind of granular operational detail Western labs rarely publish about their own training stacks, which is part of why it's drawing attention on its own.

  • My read between the lines: Everyone keeps score by counting parameters and benchmark points, but the company that keeps out-shipping labs with ten times its funding just told you, in exhausting technical detail, that its actual edge is 5,000 sandboxes a second -- which is a very unglamorous thing to be winning at, and a very hard one to copy.


That's your AI Brief for Sunday.

—Artificially Intimidating

Discussion about this episode

User's avatar

Ready for more?