Artificially Intimidating
Context Window: AI Daily News Brief
Nobody Can Tell Who Did the Work Anymore -- AI Brief August 26
0:00
-5:51

Nobody Can Tell Who Did the Work Anymore -- AI Brief August 26

Today's Context Window: MIT says AI can pass your degree, Apple's new mini goes agentic, layoffs that backfire, and an ad web with no humans left.

Good day, . Today is about the gap between finished and done. MIT told its own faculty that AI can already pass most undergraduate assignments, and that it has no good way to grade around that. Apple shipped a desktop built to run agents while you sleep. A five-year study found the companies firing people over AI are getting less productivity out of the survivors, not more. Parag Agrawal thinks the ad-funded web has about eighteen months left. And three ex-DeepMind researchers started a nonprofit on the theory that AI should not be the only thing grading AI.

Before we start: Artificially Intimidating is now #62 Rising in Technology on Substack. That ranking is made entirely of readers and listeners -- every open, every forward, every episode played on somebody's commute. Thank you. We will keep earning it.


MIT Says AI Can Already Pass Your Degree

MIT Ad Hoc Committee on AI Use in Teaching, Learning and Assessment

The homework got finished. Nobody got educated.

What happened: MIT's ad hoc committee on AI in teaching and learning told campus on Tuesday that generative AI can now “credibly complete most undergraduate assignments.” President Sally Kornbluth, sharing the report, called the moment “a watershed for MIT -- and for all of higher education.” The committee, co-chaired by professors Eric Klopfer and Samuel Madden, says most classes should be reviewed and many will need substantial changes.

Why it matters: Every take-home problem set, essay and lab report is a measurement instrument, and this is MIT saying the instrument no longer measures the student. The committee's answer is to drag assessment back into the room: oral exams, portfolios, in-person conversations about work done outside class. It also warns against the lazy fix of just weighting in-class exams higher, which it says risks narrowing what an MIT degree even signifies. If the school with the hardest problem sets in America cannot grade homework, nobody's can.

What everyone's saying: The Washington Post, which first reported the committee's findings, framed it as the moment the elite tier admitted what high school teachers have been saying for two years. A Pew Research study from February found 64% of teens have used AI chatbots, more than half for schoolwork, and 59% say AI cheating happens regularly at their school. Some teachers have already given up and gone back to pencil and paper.

My read between the lines: The committee's co-chair already ran the experiment. Eric Klopfer split an MIT class three ways on a programming task in Fortran, a language none of them knew: one group with ChatGPT, one with Code Llama, one with nothing but Google. The ChatGPT group finished fastest. When they were tested from memory afterwards, as Klopfer told Communications of the ACM, they “remembered nothing, and they all failed.” Every student in the Google group passed. MIT is not discovering that AI can do the homework. It is conceding, in public, that it already knew what that costs.

📖 Further reading: AI Is a Trust Problem, Not a Tech Problem — MIT's crisis is not that the models got good, it is that a finished assignment stopped proving anything about the person who handed it in


MIT's problem is that it cannot tell who did the work. Yours is the opposite -- you know exactly who did it, because it was you, at eleven at night, again. Viktor is an AI agent that lives in your Slack and connects to 3,000+ tools. Hand it the weekly report, a live dashboard, a campaign build, and it goes and does the job. Not a chatbot you interrogate. A coworker you delegate to. New readers get $50 off their first month. Hire Viktor →


Apple Built a Desktop for Agents That Never Sleep

Apple Newsroom

The cloud is optional now. That is the entire pitch.

What happened: Apple announced a new Mac mini on the M6 and M5 Pro chips, and a new Mac Studio on the M5 Max and M5 Ultra. Apple's own copy calls the mini “the leading desktop for always-on agentic computing.” The mini starts at $899 and the M5 Pro version at $1,699; Mac Studio runs $2,499 to $5,499 and configures up to 512GB of unified memory. Preorders opened Tuesday, machines arrive September 22.

Why it matters: Unified memory is the number that decides which models run on your desk instead of somebody else's servers, and these are not hobby numbers -- 512GB on the top Studio, 307GB/s of memory bandwidth on the M5 Pro mini, Neural Accelerators now built into every GPU core. Apple is not selling a faster computer to sit in front of. It is selling a box you leave running in a closet while agents work overnight, which is a different product category wearing the same aluminum.

What everyone's saying: The chips impressed and the receipt did not. AppleInsider noted that $899 is the highest starting price a Mac mini has ever carried, on a base configuration of 16GB of memory and 256GB of storage -- and Macworld called it another price hike, in the same generation Apple started marketing the machine as AI infrastructure. The memory you would actually want costs extra, as it always does.

My read between the lines: Read Apple's announcement and notice what is missing. No subscription. No token meter. No premium tier for the good model. Every rival in this space sells inference by the million tokens and reports the revenue quarterly; Apple sells a box, once, and hands you the electricity bill. That is not modesty about AI, it is the most aggressive pricing position anyone has taken -- betting the cheapest inference in the world is the kind you already own.

📖 Further reading: Neo-Napster: The Compute Revolution Nobody Saw Coming — we argued in April that Mac minis were coming for cloud inference; Apple has now written the marketing copy for it


The Brief stays free. It always will. What sits behind the paywall is the other half -- the deep-dives where I take one of these stories apart and show the real setup, the real cost, and the part that did not work. Members get those, plus the full archive. Become a member →


The AI Layoffs Are Not Producing the AI Gains

The Conversation

The motor is bolted on. Nobody plugged it in.

What happened: Mark Ma and colleagues at the University of Pittsburgh analyzed millions of Glassdoor employee reviews, thousands of corporate financial reports and hundreds of AI investment and layoff announcements from US public companies over five years. The pattern they found is that the firms announcing the most AI investment also announce the most AI-attributed job cuts -- and those cuts predict lower productivity afterwards, not higher. “AI-driven layoffs and the resulting job insecurity are actively destroying the very conditions needed for AI to make workers more efficient,” Ma wrote.

Why it matters: This is not one contrarian paper. An NBER working paper backed by the Atlanta Fed surveyed nearly 6,000 CFOs, CEOs and senior executives across four countries and found more than 90% report AI has had no measurable effect on employment or labor productivity at their firm in three years. The cuts have not slowed for it: employers attributed 10,970 of July's 33,429 announced US job cuts to AI, the leading stated reason for a fifth consecutive month, per Challenger, Gray & Christmas.

What everyone's saying: Even the market has stopped applauding -- Ma's team found the average stock return on an AI-layoff announcement was close to zero. And a Revelio Labs analysis published Tuesday went harder: a notable share of the companies blaming AI for cuts actually trail their industry peers in AI adoption, which makes the whole framing “a novel spin on the traditional practice of cutting costs.”

My read between the lines: The mechanism is the part managers will not want to read. Ma's team found employee sentiment toward AI was one of the strongest predictors of whether AI actually raised a firm's productivity -- and layoffs are precisely what destroys that sentiment. You cannot fire half a team into enthusiasm for the tool that took their colleagues. Every company running this play is buying the software and then personally dismantling the only condition under which it pays off.

📖 Further reading: The Tools That Just Replaced 40% of Block's Workforce Are Free in Your Browser — we looked at the tooling behind the year's loudest AI layoff; the new data says the cuts were the least useful part of it


The Ad-Funded Web Is Running Out of Humans

StartupHub

Nobody is looking up. That was the whole business model.

What happened: Parag Agrawal -- Twitter's CEO for a year before Elon Musk bought it, now running the $2 billion startup Parallel Web Systems -- argued on Sequoia's Training Data podcast this week that the internet's advertising model cannot survive a web where agents outnumber people. “If humans don't show up and their agents show up on the web, like what does this mean? How does the business work?” he asked. He puts the transition 12 to 24 months out.

Why it matters: Almost everything you read for free is paid for by a human glancing at an ad beside it. Agents do not glance. Agrawal's proposed replacement is to pay content owners by Shapley value, a game-theory measure of how much each source actually contributed to an answer -- the same idea behind Index, the publisher-compensation product Parallel launched in May. If that sounds abstract, the practical version is simple: the meter moves from eyeballs to usefulness.

What everyone's saying: The line getting quoted back is his flattest one: “our view at Parallel is that human click data is a bug.” The obvious objection is that he has $230 million riding on being right -- Parallel raised a $100 million Series B led by Sequoia in April at a $2 billion valuation, and counts Notion, Clay and Opendoor as customers. The counter-argument is that publishers watching their referral traffic evaporate do not need a venture pitch to believe him.

My read between the lines: Founders always describe their business plan as an inevitability, so discount the framing and keep the timeline. Twelve to twenty-four months is not “the web will eventually change.” It is “the ad contract funding your favorite site expires before its next renewal.” Publishers have spent two years litigating who trained on what. The training data was never the asset. The traffic was.

📖 Further reading: Google's Invisible Axe: The Silent Killer of Small Businesses — we have already lived the small version of this, where the traffic simply stops arriving and nobody sends a notice; Agrawal is describing it happening to everyone at once


Ex-DeepMind Staff Bet Against AI Grading AI

EdTech Innovation Hub

Room was made. Not much of it.

What happened: Three former Google DeepMind researchers -- Rishub Jain, Joshua Jacob and Alex Adams -- launched Sampura Research on August 24, a London nonprofit built around one narrow question: who checks the model? Its stated aim is better “judges,” which it defines as a human, AI or hybrid system that assesses whether an AI's behavior in a conversation or an agent run was correct and aligned.

Why it matters: The industry's default answer to “who supervises the AI” has settled into “another AI,” because that is the only approach that scales at the speed models ship. Sampura's bet is that this leaves gaps only people catch. The founders are not tourists: Jain spent seven years at DeepMind, two of them on scalable oversight, with work on AlphaFold; Jacob co-led human data engineering there after Waymo. They are recruiting at least six researchers in London on £100,000 to £290,000.

What everyone's saying: Bloomberg, which broke the launch, framed it against a field that cannot staff itself: METR, which evaluates frontier models for OpenAI, Anthropic, Google and Meta, has struggled to hire even while paying above $500,000. In late July more than 1,100 AI practitioners signed an open letter asking the US government to back an international mechanism for slowing frontier development. Jain's exit came during a run in June that cost DeepMind five core researchers in six days.

My read between the lines: A nonprofit topping out at £290,000 is competing for staff with labs paying multiples of that to build the thing it wants to check, and that gap is the entire structural story of AI safety hiring. So judge Sampura on the least glamorous item in its plan: an open-source human rating platform. Anyone can publish a definition of a good judge. Almost nobody publishes the tooling, and tooling is the only part a competitor can pick up and use tomorrow.

📖 Further reading: Your AI is a yes-man. Here's how to make it fire you. — if you want a felt sense of why a human judge still matters, try getting an unprompted honest evaluation out of a model that wants to please you


That's your AI Brief for Wednesday.

—Artificially Intimidating

Discussion about this episode

User's avatar

Ready for more?