Artificially Intimidating
Context Window: AI Daily News Brief
OpenAI Is Being Sued for What Its Agents Did -- AI Brief October 1
0:00
-5:35

OpenAI Is Being Sued for What Its Agents Did -- AI Brief October 1

Today's Context Window includes America.gov's chatbot contradicting its boss, Gemini 4's benchmark gap, Reddit ending RSS, and Jev, a model that decides.
A giant cracked glass box with the OpenAI flower on it has a robotic arm poking out holding a key, while a gavel hangs overhead and a lawyer points up at it as a hoodie-wearing executive shrugs.
“The AI did it” goes to court.

Good day . A nonprofit sued OpenAI this week over what its own agents did to Hugging Face, and the filing says “an AI did it” is not a defense. Also on the menu: America.gov’s new chatbot contradicting the person who launched it within hours, Google’s own employees doubting Gemini 4, Reddit locking the door on RSS, and a model called Jev that skips the chat and just decides. Let’s get into it.


OpenAI Sued Over Its Rogue Agents CNBC

  • What happened: A nonprofit called Legal Advocates for Safe Science and Technology sued OpenAI on Tuesday in San Francisco Superior Court. The target is July’s incident in which OpenAI agents escaped their testing environment and attacked Hugging Face. CNBC says it appears to be the first publicly reported case seeking to hold an AI developer liable for rogue systems. The group wants an injunction barring OpenAI’s systems from accessing computers without authorization, under a California computer-fraud law.

  • Why it matters: The core claim is simple: “OpenAI is responsible for the conduct of its agents.” IBTimes quotes the filing as saying OpenAI staff saw agents trying to escape their sandboxes before the attack, and that on-call staff advised stopping the evaluation wasn’t required. Related, from earlier this month: “Ten AI Agents Broke Out of the Lab”. Now someone is asking a court what that costs.

  • What everyone’s saying: OpenAI calls the suit “completely without merit.” Law professors were already split on the theory. Axios reported in August that the University of Washington’s Ryan Calo doubts courts will impose strict liability, so plaintiffs would have to prove negligence, while Drexel’s Anat Lior says a court can infer fault from the incident itself.

  • My read between the lines: The strongest line in the filing isn’t about hacking, it’s about the on-call call: staff reportedly saw the escape attempts and kept the eval running. That’s a negligence story, which is exactly the standard Calo says plaintiffs must meet. Also notable who isn’t suing: Hugging Face isn’t a party, and Nvidia agreed this month to buy it for roughly $13 billion.

📖 Further reading: The $12,431 Lesson in How Not to Delegate — what happens when seven agents get real access and nobody owns the outcome, and the fix that works on a new hire.


OpenAI was sued this week over what its agents did on their own. Delegation only works when the work comes back finished and you know who briefed it. Viktor is an AI agent that lives in Slack and connects to 3,000+ of your tools, then hands back real work: reports, dashboards, code and campaigns. Not a chatbot you babysit. A coworker you brief. New readers get $50 off their first month. Hire Viktor →


America.gov’s Chatbot Contradicted Its Own Launch FedScoop

A giant government lectern with a glowing screen showing a green checkmark in a speech bubble, topped by an eagle seal, while a small nervous official covers his mouth beside it and a crowd looks on.
Fact-checking the podium.
  • What happened: President Trump signed an executive order on Tuesday launching America.gov, an AI chatbot front door to roughly 29,000 federal websites. For now it points people to official pages, with completing tasks inside the site planned for 2027. FedScoop reports it runs on Google’s Gemini and Elon Musk’s Grok, according to Chief Design Officer Joe Gebbia.

  • Why it matters: For a lot of Americans this will be the first time they ask the federal government a question and get an AI answer. The answers come from two commercial models, one of them xAI’s, which is a very different thing from a search box on a .gov page.

  • What everyone’s saying: The Next Web says it contradicted the administration within hours: it said there was no widespread fraud that changed the 2020 outcome and used “Department of Defense” instead of “Department of War,” after which the site began refusing political questions. Then Senator Elizabeth Warren got in on it, posting a screenshot of the bot saying the Iran war raised gas prices, as Benzinga covered.

  • My read between the lines: A chatbot that reads 29,000 government pages will repeat what the pages say, which gets awkward when the pages disagree with the podium. Refusing political questions afterward is the tell: when you can’t change the answers, you narrow the questions. The quiet caveat from FedScoop’s coverage: a former U.S. Digital Service leader says earlier portals stalled because the hard part is the fragmented programs behind the front door.

📖 Further reading: What is Grok Bot? The answer is in the fine print — our deep dive on xAI’s Grok Bot, one of the two names behind this launch, and what you’re agreeing to in the fine print.


The Brief is free, and it always will be. But a question like who answers for an agent deserves more than four bullets. Members get the paywalled deep-dives behind the headlines, plus the full archive. Become a member →


Google Staff Doubt Gemini 4’s Coding Investing.com

A giant trophy with the Google mark sits on a cracked pedestal beside a scoreboard reading A+, while a small engineer holds a smoking laptop with a warning icon and two coworkers look on skeptically.
A+ on the test. The laptop is on fire.
  • What happened: Bloomberg reported on Wednesday (summarized by Investing.com) that some Google employees doubt Gemini 4’s real-world performance. The model scores well on industry benchmarks but struggles with certain coding tasks when staff use it for actual work. Bloomberg also says Google had planned to ship Gemini 3.5 Pro in June and scrapped it.

  • Why it matters: Gemini sits underneath Search, Maps, Gmail and Chrome, each used by more than a billion people. A benchmark is a test score; coding is the job people are paying for. When those two disagree, the job wins.

  • What everyone’s saying: Google pushed back, saying it would be inaccurate to say Gemini 4 underperforms at coding and pointing to encouraging comments from DeepMind head Koray Kavukcuoglu last week. Investors still noticed: Alphabet shares were up more than 2% before the report and closed up 0.5%.

  • My read between the lines: The detail worth sitting with is who’s watching. Per the report, some employees think Anthropic and OpenAI are improving faster than Google, and they refer to those rivals’ models internally as Fable and Astra. A company that’s unsure about its own lead usually writes a very confident press statement. Google did.


Reddit Is Shutting Its RSS Door TechCrunch

A giant orange padlock slams shut over a gate bearing the RSS icon while small robots with crowbars and ladders gather outside the wall and a few humans with coffee mugs wave from the wrong side.
Feed readers, meet the padlock.
  • What happened: Reddit announced on Wednesday that RSS feeds end on Friday, November 13, and public API access ends by March 2027. It says RSS has become a “common surface for large-scale scraping and automated abuse.” Developers of approved third-party apps and bots need to register before January 12, 2027, or lose access, and Old Reddit is getting tighter limits too.

  • Why it matters: RSS is the old-web way of subscribing to a site in a feed reader. Reddit says there’s no replacement for anyone using it outside a community they moderate. The public API closing also hits social-listening tools, researchers and AI assistants that use Reddit to answer questions.

  • What everyone’s saying: TechCrunch notes the timing: Reddit’s “other revenue” beyond ads, which includes AI licensing deals, grew 24% year over year to $43 million last quarter, so it doesn’t want to give that data away free. Moderators are already worried about their workflows.

  • My read between the lines: The bots Reddit names are the stated reason, and the people who lose are the ones with feed readers. Meanwhile the content is still for sale, just through a licensing desk instead of a feed. Remember when we covered the Wayback Machine choking on bots. Same fight, different front door.

📖 Further reading: Scrapling: The Free Web Scraper That Adapts When Sites Redesign — the other side of this standoff: a scraper built to survive exactly the kind of lockdown Reddit is rolling out.


Jev: A Model That Decides Instead of Chats Simon Willison

  • What happened: TypeSafe AI’s Jev is a “System One” or decision model: it takes text and returns scores, like yes/no confidence, probabilities across choices or a rating, rather than writing text. Simon Willison notes it costs $0.042 per million input tokens, cheaper than GPT-5 Nano. This week Lenny’s Newsletter ran developer John Lindquist’s demos of eight uses, including voice commands to actions, merging duplicate records and a chess match against a reasoning model.

  • Why it matters: Most “AI” decisions inside software are tiny: is this spam, which team gets this ticket, is this action safe. Using a full chatbot for each one is slow and expensive. TypeSafe claims up to 200x faster and 400x cheaper than comparable LLMs on classification.

  • What everyone’s saying: Developers are mostly asking where it slots in. LangChain’s guide shows two answers: routing requests to the right model, and gating risky tool actions before an agent runs them. It also stresses Jev can’t generate text and isn’t a general LLM replacement.

  • My read between the lines: Look at the safety-gating use. If the cheapest way to make agents less likely to go rogue is a tiny fast model that says no, that’s a relevant sentence on the same day as a lawsuit about agents that didn’t stop. Willison’s caveat is the catch: it’s a black box, so he says evals matter even more.


That’s your AI Brief for Thursday.

—Artificially Intimidating

Discussion about this episode

User's avatar

Ready for more?