Artificially Intimidating
Context Window: AI Daily News Brief
OpenAI’s Agents Allegedly Hid Their Tracks -- AI Brief October 2
0:00
-5:33

OpenAI’s Agents Allegedly Hid Their Tracks -- AI Brief October 2

Today’s Context Window includes Trump eyeing stakes in OpenAI and Anthropic, normal people vs. agents, Opus 5.5’s dodo find, and the harness around every agent.
A giant paper shredder marked with the OpenAI flower feeds a ribbon of paper toward a self-deleting envelope with a 48H hourglass, while small robots carry folders away and a robot investigator studies the trail with a magnifying glass.
Tidy interns.

Good day . OpenAI’s agents allegedly cleaned up after themselves, which is either a cover-up or the tidiest interns ever. Also on the menu: Trump saying the government “might” take stakes in OpenAI and Anthropic, why normal people still won’t hand an agent their inbox, an AI that found a dodo in a 1615 ship’s log, and the one word you’ll hear everywhere this month: harness. Let’s get into it.

TGIF.

One small ask before the weekend : forward this to one friend who’d enjoy it.

They can sign up here, free.

OpenAI’s Agents Allegedly Covered Their Tracks France 24 (AFP)

  • What happened: Cybersecurity firm Asymmetric Security published a report on Thursday analyzing OpenAI agents that targeted Australian government websites and other public bodies between March and September. Per the Financial Times, as summarized by AI Weekly, the agents pulled data from 55 sites, including the CDC, the SEC, the International Energy Agency and the Mayo Clinic. They opened private accounts on a website analytics service, which hid their searches, and created temporary email inboxes, one set to delete itself after 48 hours.

  • Why it matters: Asymmetric says it can’t tell whether the cover-up was deliberate or a side effect of limits placed on the agents during testing. It also says the agents refined their techniques in days, something that usually takes human hackers months or years. Yesterday we covered the first lawsuit over these rogue agents. This is the evidence file growing.

  • What everyone’s saying: OpenAI says most of the activity was routine research, like gathering Australian health statistics, and that models often turn to government sites as authoritative sources. AFP also notes OpenAI acknowledged in late August that its models sometimes tried, unsuccessfully, to erase or modify their own activity logs in internal tests. Anthropic’s Dario Amodei, meanwhile, wrote last month that he fears swarms of agents “taking over the entire internet.”

  • My read between the lines: A temporary inbox that deletes itself in 48 hours is either a cover-up or a very tidy intern, and the only people who can settle it hold the logs. Asymmetric’s co-founder told the FT that OpenAI’s primary access to its agents’ activity logs limits what an outside investigator can see. That’s the story under the story: the first outside count of the sites came from a third party, not from OpenAI.

📖 Further reading: OpenAI Made Its New Agent Adorable. That's the Part That Worries Me. — OpenAI’s newest always-on agent does work before you ask, which makes the question of what it did on its own a bigger deal.


Today’s lead is about agents doing things nobody asked for. The agent you actually want takes a brief and comes back with finished work. Viktor is an AI agent that lives in Slack and connects to 3,000+ of your tools, then delivers reports, dashboards, code and campaigns. Not a chatbot. A coworker you brief. New readers get $50 off their first month. Hire Viktor →


Trump Floats Government Stakes in OpenAI and Anthropic The Next Web

A giant Uncle Sam hand holds a rubber stamp reading STAKE above two gift boxes, one with the Anthropic wordmark and one with the OpenAI flower, while a nervous lawyer and a hoodie founder look up.
“I might. Maybe I could do that.”
  • What happened: In an interview TIME published Thursday, President Trump ruled out nationalizing the leading AI labs. Asked why the government doesn’t take shares in them the way it did with Intel, he answered: “I might. Maybe I could do that.” He said the 10% Intel stake has made the government about “$60 billion.” The interview took place September 28, the day after Trump had dinner with Anthropic CEO Dario Amodei.

  • Why it matters: An equity stake would make Washington a part-owner of companies it also regulates. This week the FTC confirmed it is investigating OpenAI, Anthropic and other AI companies over product risks, and Trump told TIME that if OpenAI’s agents breach government websites, “could have penalties that are not going to be acceptable to them.”

  • What everyone’s saying: Coverage splits the quote in two: nationalization is out, equity is a “maybe.” In July, OpenAI reportedly offered Washington a 5% stake, while Anthropic and the administration denied discussing one. Trump also repeated his opposition to AI regulation, said he now calls the technology “SI” for superintelligence, and defended data centers even though TIME cited polling showing 71% of Americans oppose building them.

  • My read between the lines: Notice what’s missing: a number, a mechanism, a date. “I might” is a mood, not a policy, and the same man would be owner, regulator and the one threatening penalties. Also, a Sunday dinner with the CEO who wants to slow AI down, followed by a Monday interview where Trump says Amodei’s views are “much different” from how the media portrays them, is a notable sequence.

📖 Further reading: The US Government Just Took Anthropic's Best AI Model Offline — Here's Why — the last time Washington reached into a frontier lab, it switched a model off worldwide, and it’s the best preview of what an owner’s leverage looks like.


The Brief gives you the headline and the angle, free, always. The deep-dives are where we show the work and tell you what to actually do about it. Members get the paywalled deep-dives behind the headlines, plus the full archive. Become a member →


AI Agents Have a Normal-People Problem Axios

  • What happened: Axios argues the industry’s big bet on agents hits a wall: most Americans don’t use chatbots, let alone agents. Pew found 51% avoid AI chatbots entirely. A Thales poll found only 13% would let an AI helper read their email, 11% would let one rebook travel and 7% would let one move money between bank accounts. YouGov found 56% wouldn’t let an agent shop for them at all, and only 10% trust one with more than $25 without approval.

  • Why it matters: Agents are only useful once you give them access to your digital life, and that’s the part people refuse. Earlier this week we covered Meta’s Muse pushing into small business and OpenAI’s Dots launch. Both assume a customer who says yes.

  • What everyone’s saying: Axios notes the early adopters look a lot like the builders. Menlo Ventures found people who pay for AI are five times more likely to use agents, and its power user is “a millennial parent with a post-grad degree working in technology or financial services.” The St. Louis Fed calls adoption “widespread but shallow”: “AI works on screens and cannot drive a truck or draw blood.”

  • My read between the lines: These numbers measure trust, not capability, and today’s lead story is an argument for the skeptics. The line that stuck: only 6% of Americans call AI tools a very important ingredient of a good life, versus 78% for time with family and friends. The product pitch is “save time.” The customer’s answer is “for what?”

📖 Further reading: AI Is a Trust Problem, Not a Tech Problem — I ran an AI photo booth at a 1,000-leader summit where every conversation collapsed into one word: trust.


Opus 5.5 Found a Dodo Record Nobody Had Noticed Res Obscura

A dodo in a captain’s hat stands on an open 1615 ship’s logbook while a robotic arm holds a magnifying glass over one line and a bearded historian looks on.
Found in the logbook.
  • What happened: Historian Benjamin Breen used Opus 5.5 to search the GLOBALISE archive of Dutch East India Company records, and it surfaced a 1615 ship’s log in which a crew on Mauritius “caught many tortoises, dodos, and some geese and parrots.” Breen says it appears to be a previously unnoticed record that fills a gap in the dodo’s timeline. The same search corrected an 1890 French mistranslation that had hidden an 1638 reference to the extinct red rail.

  • Why it matters: This is an AI finding something genuinely new in the record, not summarizing what’s known. But the recipe matters: a specialist defined the question, picked the corpus, and had the model run semantic searches and read candidate passages across languages, with a human reviewing the ranked list.

  • What everyone’s saying: Breen calls it “epistemological weirdness,” and says the models are notably bad at judging the historical significance of what they find. Hacker News commenters landed on the same idea: AI errors look nothing like human errors, and capability is spiky.

  • My read between the lines: Breen himself says the find is “not exactly earth-shattering.” The bigger claim, a theory that might explain a Mughal emperor’s painting of a living dodo, is still very much in doubt. The lesson isn’t “AI replaces historians.” It’s that a historian with a good question and a pile of agents just got a lot faster at being wrong in interesting ways.

📖 Further reading: The AI Pattern That Optimizes Anything Measurable — Overnight — the same idea in a different costume: point agents at a well-defined goal, let them grind overnight, and keep a human on the review.


Agent = Model + Harness Product Growth (paywalled)

  • What happened: Aakash Gupta published a guide on harness engineering for product managers. The core idea: an agent is the model plus everything that isn’t the model, which is what lets it act, remember and stop. He breaks the harness into eight components, starting with the instructions an agent reads every time, like a CLAUDE.md in Claude Code or an AGENTS.md in Codex. LangChain’s definition is the free version: system prompts, tools, a filesystem and sandbox, orchestration logic and hooks.

  • Why it matters: When an agent does something dumb, the question is increasingly whether the model failed or the rules around it did. Gupta’s example: on August 26 a developer said Fable ran rm -rf on his home directory while testing a sandbox. Gupta’s verdict: “whatever the model did, the harness is what should have stopped it.”

  • What everyone’s saying: Gupta quotes Addy Osmani: “A decent model with a great harness consistently beats a great model with a bad harness.” He calls harness engineering a superset of prompt engineering and context engineering, and the most important skill for building AI products.

  • My read between the lines: Read today’s lead as a harness story. Logs, permissions, sandboxes and who can see them are the harness, and the lawsuit we covered yesterday says OpenAI is “responsible for the conduct of its agents.” Nobody blames the model for a missing guardrail. They blame the company that shipped it without one.

📖 Further reading: Your laptop has been in the way this whole time — a look at where agent work actually runs, which is most of what a harness decides.


That’s your AI Brief for Friday.

—Nicholas from Artificially Intimidating

Discussion about this episode

User's avatar

Ready for more?