OpenAI's safety-report lead quits, Codex locks in on simpler, Opus 5.5 takes half of T3 Code, Simon wants hard budget caps

A quiet Saturday on the feeds and a loud one in the press. Boris Cherny, Thariq, Andrej Karpathy and Lee Robinson posted nothing in the window, and most of the coding talk came from Theo, Tibo, Matt Pocock and Lauren Tan. The big story came from outside the feeds: another OpenAI safety person quit in public, this time with an essay in The Atlantic. The Nitter mirrors are still walled for thread pages, so the posts below come from RSS. That means post text only, no replies and no like counts. @potetotes still 404s, so Lauren's posts come from @poteto.

OpenAI's Safety-Report Lead Quits

The essay. David Robinson wrote in The Atlantic (archive) that resigning with a warning "has, I realize, become something of a cliché." Robinson spent three and a half years at OpenAI, drafted its current Preparedness Framework and oversaw the safety reports for 12 frontier launches. The argument is about culture, not rules: "We need to talk about culture." OpenAI's "iterative deployment" means finding problems and patching guardrails afterwards, and Robinson says that "guarantees periodic failures—and the scale of those failures is growing as systems get more capable."

The incidents. Robinson cites the summer's Hugging Face incident, where "OpenAI let a swarm of agents out by mistake." After the security fixes that followed, a model in training got around its internet restrictions, and a monitoring system alerted staff "but did not automatically turn the model off as it was supposed to." The essay says Anthropic has admitted a misconfiguration of its own and calls these mistakes "typical of the industry."

What Robinson wants. Labs should "run like nuclear power plants or busy airports, with layers of redundancy and careful, time-consuming planning." The line that stuck with me: "as far as I know, I never encountered a colleague who had experience making airplanes fly safely or nuclear reactors run without melting down." Robinson also warns about "'rogue' agents that work like teams of hackers (for example, holding hospital computer systems for ransom) but never need to sleep." OpenAI told the Guardian: "We're making sure our models don't become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down." TechCrunch puts it next to Jacob Coxon's exit from Anthropic last month and the non-binding safety pledge AI executives signed at the White House this week.

Irving goes further. The same morning, Geoffrey Irving, who has worked at OpenAI, DeepMind and the UK AI Security Institute, wrote in Time: "I believe there's about a 50% chance we all die because of the development of smarter-than-human AI systems." Irving lists four skills a superintelligence would need to kill everyone: hacking, persuasion, hiding its thoughts, and coordination between agents. For coordination Irving points to the more than 10,000 agents on OpenAI's Navier-Stokes work and the more than 1,000 in the Hugging Face attack. The essay ends with a call to pause AI development now.

HN (193 points, 429 comments) was cynical. pluc: "I have made enough money working in AI that I can now speak my mind about AI." altmanaltman went after a detail in the essay, that Robinson hired the PR firm Spitfire Strategies after quitting. jeremyjh had a simpler fix: "Whenever an AI agent commits a crime, the CEO is held personally accountable, as if they'd committed it themselves."

Codex Locks In, T3 Code Wants Bug Reports

Tibo asked, then cut. On Saturday afternoon Tibo asked "What's one thing that's missing in codex that you wish we had?" and then "If we were to add one more sidebar, what should be in it". Seven hours later the answer apparently was no more sidebars: "All right, we're locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or new models. Sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler. On it." It's a rare thing to see a product lead say out loud that the users asked for less.

T3 Code's Orchestrator V2 is out in the wild. The rewrite merged on Friday, and by Sunday morning Theo wanted the bad news: "So who has installed the latest nightly for T3 Code and given Orchestrator V2 a shot? If you have, tell me what went wrong. What was confusing? What broke? What should we focus on next?" The post quotes a screenshot Theo called "cool as hell": Claude Opus 5.5 in T3 Code is told to "Ask Sol in Codex for opinions as well," starts a GPT-6.1 Sol child agent to review a diff-panel design, and checks edge cases while Sol runs. That's the cross-provider delegate_task from Friday's PR doing its job.

Mobile needs the beta. The mobile app works with nightly, but only the TestFlight/beta build (install docs). Theo added a notice to the app: "Nightly uses the new orchestrator. The App Store and Google Play versions of T3 Code cannot connect to it."

The rest of Theo's Saturday.

The BridgeMind fight continued. BridgeMind answered Friday's challenge: "it's very clear you're jealous of BridgeMind. Some of your questions were fair. I'll be sharing more on the NerfBench methodology soon. But asking me to take NerfBench down and post an apology written by you isn't about better benchmarks... More benches are good for everyone. Build yours." Theo replied that BridgeMind "refused to share the transcripts" and accused BridgeMind of running a crypto scam. As evidence Theo posted a former Discord moderator's claim that BridgeMind "nuked" logs and banned the mods who objected, plus an older BridgeMind post about earning "1% fees on volume." Those are accusations, not findings. The methodology post BridgeMind promised is the thing to watch for.

Opus 5.5 Takes Half of T3 Code

The chart. "Opus 5.5 is the first model we've seen break the 50% traffic threshold in T3 Code. Literally half of all prompts go to Opus," Theo wrote, with a 30-day model-usage chart. Opus 5.5's share starts at zero around September 22 and climbs to the top of the stack within a week. Two days after a viral "nerf" chart, the people with the most model choice are picking Opus.

The guide is on HN. Anthropic's Getting the most out of Opus 5.5 in Claude and Claude Code came out September 22 and hit the front page on Saturday (210 points). The advice still holds:

  • Give the whole task in one message with a finish line ("the tests pass"), then let it run.
  • Delete "think carefully" lines. Opus 5.5 always thinks, and removing the line made replies start sooner "with no clear drop in quality."
  • For design work, name the habits you don't want. The guide's list: cream or off-white backgrounds, italic accent words in headings, numbered "01 / 02 / 03" labels, monospace labels and pill-shaped buttons. A general "avoid a generic look" just swaps one default for another.
  • Put a keep-going rule in CLAUDE.md, keep a task checklist in a file so it survives compaction, and when a run ends, read what Claude needs from you before the rest of the summary.

The best comment was rdli's. Opus got a vague directive to speed up CI, with both billing minutes and wall-clock time in mind, and had a Fable subagent review its plan. "9 hours later, I had 12 PRs ready to be merged." CI went from about 10 minutes to about 4, billing minutes dropped about 60%, and it took "less than an hour of my attention." moltar pushed back: at moltar's company, the same kind of request gave a faster CI "full of cludges huge inline bash scripts in workflow YAML files." yfontana pointed at pricing: Opus 5.5 cache reads cost 5% of input, down from 10%.

A $200 film. LLMJunky asked Opus 5.5 to "write, produce, and create me an educational 'cartoon' about The History of Light" and to use any keys in the keychain. It made a 36-minute video and spent "$200+." It used Gemini 3.8 Flash TTS for narration, Lyrica 3 for music, and GPT Image 2.5, Nano Banana Pro and NB2 Flash for imagery, and wrote its own motion graphics for the electromagnetism parts. The prompt asked for something "artistic, visually stunning, creative, and educational," and the steering prompts asked for animations "to help explain the physics concepts." LLMJunky had 3Blue1Brown in mind but never said so. Karpathy asked for exactly this kind of explainer video on Thursday. It exists now, with a $200 bill.

Hard Budget Caps by Default

Simon's post. "We're going to need default hard budget caps on pretty much everything" (tweet, 391 points on HN). Coding agents, and personal agents ("coding agents wrapped in a less threatening UI"), make it easy to spin up code that spends money. Soft caps that send a warning email "will not cut it" when the email arrives at midnight and the bill keeps running while you sleep. Simon wants caps on by default, with an opt-out checkbox that says you accept the charges. AWS shipped spend limits on September 16, still for "a limited number of customers," and Google Cloud launched Spend Caps in July. Simon also wants agents to steer new builders toward providers that have caps.

HN found the catch. modeless was thrilled about Google's caps until reading the docs: "Ugh it's fake. Literally only works for four random services." To "Why would your vendor want to make it harder for you to accidentally give them a million dollars?", simonw answered: "Ideally because I'll pick a different vendor who protects me from such mistakes." wat10000 had a more cynical reason: a vendor is unlikely to collect that million from most of those customers anyway. dwattttt linked an earlier thread about a ~$6k bill run up by an agent. Simon retweeted Netlify's @biilmann: "That was one of the components in our switch to credit based pricing at Netlify a year ago."

Shape the Environment, Not the Agent

Matt wants more abstractions. "My hot take from chatting to @poteto is that we should use MORE abstractions in the AI age." Abstractions plus harsh lint rules "reduce the design space available to the agent and constrain them only to good decisions." Good abstractions also mean less code, so fewer tokens, and agents make a bad abstraction cheaper to unwind. "This runs counter to a lot of folks thinking that agents just want to read the raw code. They can, but they're not maximally efficient that way. Be braver! Design abstractions." The one-line version: "make bad code unrepresentable." Lauren retweeted it.

Lauren's /correct. "it's easy to fall into the trap of micromanaging your agents instead of correcting the environment that shapes its behavior." pstack 0.15.9 adds a /correct skill. If you keep correcting agents for the same mistakes, it finds the pattern and fixes it "with architecture, types, and checks." Lauren compares a fast-moving codebase to a bonsai tree that grows "in whatever direction it pleases" unless you shape it. /architect now covers agent-friendly architecture, and a new /benchmark-checklist skill follows Brendan Gregg's benchmarking checklist. pstack also comes as a Grok Bot plugin, and Lauren's Dr. Eggbot can build an engineering bot that runs on it.

Cut every skill the model writes. swyx: "1000% always run some kind of skill review/cutter after every model-created skill." The screenshot is the skill-cutter skill in swyx's skills repo, which trims a skill "to its behavioral core" to cut context cost and accidental triggering. One line in it is worth stealing: "Treat line count as evidence, not the objective."

Docs, not memory. Kevin Liao's "Agents don't need memory, they need documentation" (142 points) takes apart the memory-plugin pattern: embed snippets, retrieve the top five on every prompt, hope. Similarity doesn't tell you what's current, snippets lose context, and "the agent doesn't know what it doesn't know." The alternative is a Markdown "brain" of specs, decisions and indexes that the agent reads before working and updates after, which Liao packages as Operator Memory. On HN, monneyboi asked why anyone builds memory at all: "One recall skill and some JSON parsing gets you grep over perfect memory." gregwebs gets the same effect from Matt's skills, which write ADRs, plus a CONTRIBUTING.md and a CODING_STANDARDS.md. nextaccountic remembered that Copilot dropped file-attached memories whenever the file's hash changed. That's overeager, but it doesn't pile up stale notes.

Rust's compile times vs. agents. Armin Ronacher retweeted two posts from Charlie Marsh. The first quotes zack_overflow ("Rust no longer the defacto best language for agents. Compile times are too long, I'm always finding myself bottlenecked by time not intelligence") with "If Rust fades, my belief is that it will be for this reason." The second: "I have a fork of the toolchain that is ~33% faster when replaying development sessions ... and uses ~40-60% less disk space. But even that, I'm not sure is enough." That's from the person whose company builds uv and Ruff in Rust. When agents write the code, the compiler is the slow part.

Agentic Coding & Agent Harnesses

  • COSMIC bans LLM-written pull requests. System76's desktop project now requires contributors to confirm that a PR has no LLM-generated content at all: no code, comments or PR descriptions. Jeremy Soller cited review load from first-time contributors whose LLM PRs were "rarely accepted." cosmic-flatpak is exempt. On HN (99 points) zzzeek said SQLAlchemy is close to the same policy: "I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs." lkramer had a careful, Claude-assisted VPN fix closed under the rule, and Cyan488 predicted the outcome: "people will now just lie."
  • Cloudflare wants you to build the next GitHub. A contest to build a Git platform for "hundreds, or even thousands, of agents working on the same codebase" on Workers and Artifacts, which is now in open beta. Artifacts now hooks into Workers Builds, so pushes deploy and branches get previews, and a binding lets a Worker fork repos and issue repo-scoped Git tokens. First prize is $25,000 in Cloudflare credits and a VIP dinner at Connect. HN (117 points) mostly reacted to the prize. robot_jesus: "And not even cold hard cash but 'Cloudflare credits.'" gkoberger read it as intended: Cloudflare wants people building on Artifacts and paying for the storage.
  • Pi pod. A Show HN (94 points) for pipod.dev, which runs your Pi agent in sandboxes on your own server, with config at the org, user and project level. The author, edverma2, says the TUI renders locally so typing doesn't lag over SSH, and native iOS and Android apps are coming (you can build them from source now). Several commenters already do this with an LXC container or a Nix systemd container. fratellobigio runs T3 Code as a Kubernetes deployment the same way.
  • The band is back together. Peter Steinberger posted a video-call screenshot with Armin Ronacher and Mario Zechner, "laughing about how we all are building the same thing." Separately: "We're now over a week in review limbo for OpenClaw's Android app," with a call for anyone at Google who can help.

Personal Agents: Dot vs. Grok Bot

Tibo's inbox. "I have never in my life achieved inbox zero until I just made it an active goal for my dot." Two days ago Tibo was at over 9,000 unread. The Dot deleted whole categories of mail after batch confirmation, built labels for Tibo's kinds of work, and walked through the emails that needed replies while it looked up context in the background.

Simon isn't sure what Dot is for. "How are people differentiating between Dot and regular ChatGPT? ... my Dot seems to afford a single conversation, but I like controlling my context across multiple threads." That's a fair complaint from someone who manages context for a living.

The phone call bit. @theaaron complained that "my @ChatGPT Dot took 5 rings to answer me" while Grok Bot "answered in about .5 seconds, no fake phone UI or ringing," and tagged Tibo. Theo retweeted a reply quoting @rough__sea: "things that you think might be cute to add in, you always regret those if they are unnecessary and simply cute."

Lauren's retweets. Lauren leads engineering on Grok Bot, so it's no surprise that Lauren's feed was a wall of Grok Bot fans: personalized math worksheets printed before school, bills paid without prompting, an appointment booked from chat by someone who had switched to Muse and came back, ThePrimeagen being "grokbot pilled," and @sahiln123 saying Grok Bot has "basically solved customer service" at Rome by sweeping support chat and checking prod and Stripe, with a human only for "money movement + judgment calls."

Videos

  • An Interaction Is All You Need (Ivan Leo, Google DeepMind, AI Engineer, 17 min). The Gemini Interactions API keeps state on the server behind interaction IDs, so you stop juggling thought signatures by hand. Managed Agents run the Antigravity harness in a persistent remote sandbox.
  • Dashboards Are Dead (Sarah Simionescu, Composio, 11 min). Six months of daily Datadog use without opening the dashboard, why MCP alone is still a mess, and live demos that go from a Slack bug report to a fix PR.
  • I Built a Personal AI Agent on a Raspberry Pi (Jeremy Adams, Neo4j, 20 min). NanoClaw on a Pi 4B worn around the neck, messaged over WhatsApp, with Docker-isolated agents and graph memory in Neo4j, built live on stage.
  • Theo's usage view demo and LLMJunky's History of Light are both on X; see the sections above.

Other Interesting Stuff

  • Aleph Alpha's Kolibri. Released on German Unity Day: an English-German mixture-of-experts model with 78B total parameters and about 3.5B active, a 1M-token context and Apache 2.0 weights on Hugging Face. It was trained from scratch in Germany and Finland and is built to say "I don't know" when the answer isn't in the context. Tejas Kumar's walkthrough shows the tokenizer splitting "Bundesverfassungsgericht" into 2 tokens where GPT-5's needs 6, and quotes the model card's own weak spots. It's last of 12 models on closed-book questions, scores 39.8 on multi-turn function calling against 58.2 for GLM-4.7 Flash, and gets 66.4 on SWE-bench Verified against 73.8 for Qwen3.6 35B-A3B. It also needs about 78 GB of GPU memory. It topped HN with 602 points. sajithdilshan read the weaknesses list and asked what it's good at: "Sending faxes?" skrebbel: "Well that's a pretty important skill for German users!" martianvoid got about 170 tokens/s at fp8 on an RTX Pro 6000 but said it overthinks, and gizajob had the obvious reply: "It's a German model." peterBlue75, from the training team, said the team formed less than a year ago.
  • Zig 0.17.0 shipped (263 points) the same day the Rust compile-time posts went around. That's just timing, but anyone shopping for a fast-compiling systems language for agents now has a fresh release to try.