Theo takes on NerfBench, T3 Code merges its orchestrator rewrite, Apple tightens Full Disk Access for agents, Cerebras hits a post-IPO low

Friday belonged to Theo. In one evening T3 Code merged the biggest PR in its history, shipped a beta that hides running threads, and then Theo spent hours arguing that model "nerfs" are mostly noise, with $10,000 on the line. The rest of the feed was quieter. Boris Cherny, Lee Robinson and Simon Willison posted nothing in the window, and Andrej Karpathy's only post was the "Land or Water?" map covered yesterday. The Nitter mirror that served thread pages all week now sits behind a queue and a CAPTCHA, so today's posts come from RSS. That means post text only, no reply threads and no like counts. @potetotes still 404s, so Lauren's posts come from @poteto.

Theo vs. NerfBench

The chart. BridgeMind posted: "Claude Opus 5.5 just took a big drop on NerfBench. Yesterday it was scoring above launch. Today it's at 94.2%." The same post put GPT 6 Astra at 98.0%, Sonnet 5.5 at 100.9% and GPT 6.1 Sol at 106.7%, and added that "94.2% is still inside normal variance, so we can't call it a nerf yet." Nobody shared that caveat. Reddit was already primed for the post. The AINews recap lists "Opus 5.5 nerfing - how to measure, how to spot, how to sue" (2,722 activity) and "Mmmkay... something is suddenly off with Opus 5.5" (2,045). One commenter pointed to Opus 5.5's score on modelsentiment.com falling from 71-73 in late September to 55. That site measures what people say about a model. It doesn't measure the model.

Theo's response.

The replies. Theo shared screenshots of the angriest ones ("Some of these are just surreal") and more, including Vicenzo Masat: "You are getting the bag and miss informing people... You better delete X." @aicodeking agreed with Theo: "A Claude Code Update, A Provider Change, A cache hit change, A bad request, New seed anything can change a prompt's outcome in just a matter of minutes." Theo joked about starting an alt account "where I just straight up lie about AI stuff and see how quickly I can go viral with it." Theo also listed the misinformation the next few videos will cover: model nerfs, whether token prices matter, whether TPS matters, the real impact of slop harnesses, and "compaction = deletion."

There are better ways to check. GIGAZINE wrote up Livenerf, an open-source project from Langley built on the UK AI Security Institute's Inspect framework. It collects and publishes clean test data starting from a model's release day, and it has Opus 5.5 data going back to September 24. The article also points back to Anthropic's own report from April, which found that Claude Code, the Agent SDK and Cowork really did get worse between March 4 and April 20, 2026. So quality drops happen, and a lab has admitted one. Catching them takes published tasks, raw runs and enough repetitions to see the noise, which is everything Theo says NerfBench leaves out. One score of 94.2% with a 10% band on either side tells you almost nothing, and BridgeMind's own post said as much.

T3 Code's Orchestrator Rewrite Lands

PR #2829 is merged. "t3 code publicly launched about 7 months ago. julius's orchestrator rewrite was open for just over 4 months. pr #2829 we're now at #14860," wrote maria (@maria_rcks). That's 823 commits and 1,912 files changed, and "83% of the repo's prs were opened after this one." Theo celebrated and noted that the PR page takes over 30 seconds to load on GitHub and lags when you scroll. That morning T3 Code had passed 400,000 users.

What's in it. From Theo's list:

  • Pi support, OpenCode 2 support, and Cursor running on the official Cursor SDK, "not the (broken) CLI."
  • An ACP Registry, so you can add any registry agent, including Devin, Cline, Kimi and Droid.
  • A T3 Code MCP server that lets agents create, launch, message, wait on, read, search and interrupt threads. delegate_task lets an agent start child agents on any provider or model.
  • Switching provider or model mid-thread, thread forking, and native subagents that show up as child threads in a lineage view with their model, status and history.
  • @ a thread name, or drag a thread from the sidebar, to attach it as context.
  • Server-side queueing and steering, auto-resume when limits reset, scheduled tasks, and MCP tools for worktree handoff and previews. Agents can rename threads, link PRs and settle threads.

maria's longer version covers how handoff works. Short conversations transfer intact. Longer ones get the relevant history selected, and agents can fetch the parts that were left out. A handoff that can't fit the target model's context gets rejected. Forks work across Codex, Claude, Pi and Cursor, and child agents can't escalate their own permissions. matt_feroz put it plainly: "pass context from Claude-> Codex -> Pi -> your custom coding harness and back." Mario Zechner, who makes Pi, quoted the list with one word: "wat."

Expect breakage. Theo warned that the next nightly "is gonna feel more 'nightly' for sure." The team cut one more stable release before merging, and "if you're not down for a bit of instability and thrash the next few days, there's no shame in moving back to stable."

Hide threads while working. A beta feature hides running threads until they need you. Theo was unsure at first: "I've put SO much effort into minimizing how often things shift around and move in T3 Code. Really wanted to build a sense of spatial awareness into the product." Nine hours later: "I'm obsessed with it now. I have 18 threads working here and it doesn't feel claustrophobic anymore." Overnight the nightly also learned to queue messages that fire when your subscription limits reset. matt_feroz: "RIP my sleep."

Apple Tightens Full Disk Access for Agents

Apple's post. In a developer news post, Apple said users who want to give an app Full Disk Access will "only do so with very explicit user action." Agents are the stated reason. "As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially."

Why now. Ars Technica connects it to columnist Jason Aten. Two weeks ago Aten said Meta's new agent Muse sent an unsolicited notification referencing an Apple Messages thread with a co-worker, and that Muse never got permission to read messages. Meta says Muse can only read Messages if the user turned on both Full Disk Access and the Messages connector. macOS security researcher Patrick Wardle questioned that: "with FDA (full-disk access), any (non-root file), is readable, browsing history, browser cookies, chats, etc etc etc." The Verge covered it too.

HN wasn't sure what changes (190 points). lapcat pointed out that Muse only opens the Full Disk Access pane in System Settings, and flipping the switch already takes an administrator password. big_toast called Apple's line that FDA exists "to allow backup apps to function properly" disingenuous, and lapcat added that the classic FDA app is Terminal. rwz: "Who are those mysterious people who expect their browsing history to be magically excluded from something called Full fucking Disk Access?"

Theo's take. "AI users: 'MacOS permissions suck. They get in the wa-' Apple (interrupting): 'Understood! We'll make sure you get way more popups and notifications about permissions!'" And then: "if you still run the majority of your agents on MacOS, NGMI." Docker's Rowan Christmas makes the same case for sandboxes in a short talk, listed under Videos.

Agentic Coding & Agent Harnesses

Claude Code's "You should know" plugin. ClaudeDevs added a built-in plugin that "scans Claude's output for important information you might miss." Turn it on with /plugin enable cc-plugin-you-should-know@builtin. Thariq called it "a great example of the kind of things mods enable," a day after mods launched. Lydia Hallie described mods as middleware-style hooks.

Matt Pocock's software factory on GitHub Actions. Matt says Actions are the easiest place to start. You get essentially free sandboxes for public repos and a login you already have. Tickets become issues, labels trigger actions that open PRs, and actions can apply labels themselves, which creates loops. There are cron jobs, "albeit not very accurate," and simple observability. It's going into Matt's next course. The prompt of the day was /retro: "read my last 10 coding agent sessions and find ways to make my repo easier to navigate. Find where agents take too long to find relevant information, or rely on out-of-date docs." Matt calls navigability "such an underrated way to save tokens," then asked whether the daily prompts are useful or annoying. In other news, 70% of 455 voters told Matt to keep the ginger beard for filming day. Matt shaved it: "Apologies, I have gone against the masses."

Lauren Tan on deleting the product. Lauren's long post starts: "it's funny how you can sometimes tell if a team culture is dysfunctional just by looking at how their app is laid out. behind every feature, or tab in a sidebar likely lies a PM or someone responsible for an OKR." When individual performance rides on those numbers, features get added and almost never removed, and agents make it easier than ever to keep adding. Users will notice that you ship your org chart. Lauren also said the Grok Bot team uses "bot to build bot." That was a reply to Morgan Linton, who called Dot impressive after three days but thinks Grok Bot is the better product. Grok Bot reset usage limits for all users, and @lingxi asked for "any confusion, frustration, criticism, wishes" to fix in the next two weeks. Federico Viticci described a setup where the primary bot spins up group threads with specialized bots, which then spin up Cursor cloud agents. And @cu30rry_ wrote a Zenn book about pstack in Japanese that runs to 600,000 characters.

Armin Ronacher, chief MCPO. Armin is now "chief MCPO at Earendil" and promises hot takes on MCP, starting with "my unfiltered opinions on elicitation." Armin wants to hear from anyone who uses MCP elicitation and why, and from anyone who did something cool with Jev and codemode. Also from Armin: "Need a heavily quantized quantized classifier model to figure out if my heavily quantized local open weights model went into a loop or crapped out."

ds4 runs big MoE models locally. antirez released DwarfStar 4 (HN, 230 points), a C inference engine for DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8. It isn't a generic GGUF runner. Each supported model gets validated end to end. Routed experts get an asymmetric 2-bit quant while shared paths stay precise, and the KV cache is saved to SSD keyed by a SHA1 of the prompt prefix, so a restart doesn't mean a full re-prefill. It ships as ./ds4 for chat, ./ds4-server with OpenAI and Anthropic-compatible APIs, and ./ds4-agent. Targets are 64GB+ Macs, DGX Spark and Strix Halo. aziis98 gets about 22 tokens per second on an Intel Ultra 7 iGPU after a small patch.

A month on GLM 5.3 Flash. Wagtail tried to do a month of engineering on one open model (HN, 153 points). The first half went to plan, at $68 and about 4kWh. Then a vibe-coded MCP prototype on the "wrong" model burned 450M tokens, $150 and 5kWh almost overnight. Popular providers don't have the big labs' capacity, and GLM 5.3 Flash performance degraded, which pushed the team to DeepSeek V4.1 Flash and Qwen 3.8 Flash. By the end, 1B tokens had gone to other models. epistasis on HN found the energy number the surprising part, since it's about 1% of the cost.

Same weights, different harness. Hugging Face reported that the same model scored 62% in one harness and 33% in another. RL training across four harnesses took LFM2.5-2.6B from 42% to 54% with 31% fewer tool calls. Remember that the next time a single benchmark number goes viral.

The Harness Is the Company. Shrivu Shankar's essay is from August but reached the HN front page today (132 points). "Every SaaS business will become a harness around a model." A software factory is a harness made of smaller harnesses, and the people in it become part of the harness.

The Four Horsemen of Agentic Coding. Alex Martsinovich names slop, alienation, deskilling and team fallout (HN, 108 points). The best line: "Claude now famously communicates entirely through word salad, and Astra writes in a bizarre competitive code-golfy style." paularmstrong says the work Slack is a ghost town where everyone is logged in and nobody talks, and the PRs and review comments are obviously Claude's. empath75: "Everyone is just kind of silo'd with their agents and they present these entire complex units of work with no input from the rest of the team, and by then it's too late to fix."

AI Makes Me Sad. A post by a TA who doesn't want to prompt Codex for a living drew 222 comments (HN, 186 points). tapoxi's take for startups: "You cannot prompt Claude better than your customer can."

GPT-6 Astra plays World of Warcraft. agent-wow (HN, 73 points) had Codex with GPT-6 Astra on xhigh create an orc and finish the starting zone in 40 minutes with zero deaths. The goal is a full server of agents clearing heroic ICC.

Supabase buys Turso. Supabase is acquiring Turso (HN, 202 points), and the pitch is all about agents. "Today, agents are spinning up millions of databases to power the prototypes, explorations, dashboards, and apps they're building." Supabase keeps building on Postgres and Turso keeps building on SQLite, where one server can manage millions of databases. Armin: "Big if true."

Briefly.

OpenAI: Resets, Dots and Cerebras

Limits reset. Tibo confirmed the global usage reset for paid ChatGPT accounts, which was promised after GPT-6.1 Sol's slow first two days. Pro 500 accounts missed it at first. Twenty-five minutes later: "All fixed. Surprising number of Pro 500 users on here." Tibo is "pretty sure there are more dots than bots already in this little world," and says "our future models will be much better at code deletion and simplification. Just in time." swyx reminded people that dots were DevDay's "one more thing," and Peter Steinberger called one DevDay slide "by far my favorite."

Cerebras hits a post-IPO low. On Wednesday SemiAnalysis reported that OpenAI will run the Ultrafast mode for GPT-6.1 Sol on Nvidia GPUs, not Cerebras hardware. CNBC reports that the stock hit its lowest price since the May IPO. Insider lockups expiring this week added to the pressure. It closed Friday at $166.43, down more than half from the first-day pop. The market cap was $95 billion after the first day of trading and is now just over $39 billion. In January Cerebras signed a deal worth over $10 billion to supply OpenAI with 750 megawatts through 2028. Then Sam Altman posted: "There is some speculation about our partnership with Cerebras. Cerebras is a close partner, and we have a deep engagement pushing on the frontiers of speed." The stock rose almost 3% in extended trading, and Cerebras CEO Andrew Feldman said thanks. Altman's post doesn't deny the SemiAnalysis report.

Watch out for fake OpenAI recruiters. am.will warned about "a wave of fake OpenAI scammers." The scammer sent am.will a Calendly link that somehow already had the right X username cached. The advice is to paste any link you get on X into Browserling first, even from people you know, because their accounts may be hacked.

Videos

  • LIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX (Matt Pocock, 66 min). Matt interviews Lauren Tan about skills, high-velocity software factories and how Lauren landed 2,500 PRs last month. Lauren's advice is to try both of their skill plugins and keep the skills that fit your workflow.
  • Kernel Recipes 2026: Security in the LLM age (Greg Kroah-Hartman, 56 min, HN, 234 points). This is what AI bug reports look like to the Linux stable maintainer. usernomdeguerre transcribed the Mythos slide. Of 79 reported vulnerabilities, 24 had no detail beyond "something crashed," 14 weren't bugs, 3 had made-up data, and 15 were already fixed, 4 of them by Anthropic. 20 needed fixes, and most of those assume a malicious filesystem image, an injected network packet or an untrusted device. The categories add up to 76, and Greg made fun of that in the talk. stonogo pushed back. How can Mythos's output count as "only 10 real bugs" when the kernel issued over 1,300 CVEs last month?
  • YOLO Mode, Safely: MicroVM Sandboxes for Any Agent (Docker, 11 min). Rowan Christmas, a product manager at Docker, pointed a coding agent at Rowan's own laptop, and five prompts later it had found browser history and bank accounts. The talk argues that guardrails inside the harness aren't enough. It demos Docker Sandboxes (sbx), which gives any agent its own microVM with a separate kernel, secret placeholders, default-deny networking and an audit trail. A good pairing with the Apple news.
  • Why AI Didn't Actually Make You Ship Faster (Meticulous, 11 min). Gabriel Spencer-Harper argues that verification is now the bottleneck and hand-written assertions can't cover AI-generated frontend code. Meticulous records real user flows, replays them on every PR in a deterministic browser with mocked network traffic, and shows visual before/after diffs.
  • Jerry Liu at the AI Engineer World's Fair (AI Engineer). Building the document layer for agents, covering parsing, search and source citations.

Other Interesting Stuff

Anthropic nearly walked out on the Pope. This follows up yesterday's NYT story about Anthropic and religious scholars. Christopher Hale summarizes a new NYT report. Days before the May 25 Vatican launch of the encyclical Magnifica Humanitas, Chris Olah saw an advance copy and proposed withdrawing over its stance on machine consciousness. Anthropic then lobbied the Pope's advisers, and "the pope held his ground." Hale's post was the top tweet in AINews, and the Telegraph has its own version. am.will's view on the consciousness question: "We cannot even prove that we are conscious... These are philosophical questions at best, they are wholly unfalsifiable claims/questions."

An AI beat the best Stratego player in history. Ataraxos, from researchers at Carnegie Mellon, MIT, NYU and Stanford, beat Pim Niemeijer 15 games to one with four draws (Ars Technica, Nature, HN, 214 points). Niemeijer has four world titles and over 600 weeks as world number one. DeepMind's DeepNash never managed this, and Ataraxos trained on 16 GPUs for a few thousand dollars. A second network, a belief model, guesses the opponent's hidden pieces from how they've moved, so the AI samples plausible setups instead of trying every one. Researcher Eugene Vinitsky: "We would watch the bot 'bluff' its way back from like a two percent victory probability, very, very casually."

Figure melted its F.02 robots. Figure decommissioned its F.02 humanoids by training them to jump into molten steel (HN, 81 points). Taking each one apart would have delayed F.04, so Figure asked the internet, and Arnold Schwarzenegger told them to melt the robots. No foundry in the US or Mexico would let lithium-ion robots jump into its equipment, so the jumps happened at a foundry in Imatra, Finland. The team trained a jump model in simulation using stunt performers' movements, then had 24 hours and six melts to land it. schlagaloo: "what an absolute waste of time and resources."

Extra Big Ass Intelligence. "Federally Mandated Super Intelligence," a parody site in the spirit of Idiocracy (HN, 272 points). Top comment: "Welcome to Costco, I love you."

Muse Gadgets. Nat Friedman announced open-source ESP32 firmware and a Linux SDK for building hardware that works with Muse (HN, 189 points). Grab an API token at gadgets.muse.ai and point a coding agent at the repo. Peter Steinberger retweeted it.

Personal agents automate the fun parts. A post from signulll, retweeted by Peter Steinberger, lists the things that give normal people a dopamine hit: adding things to a cart, waiting for a package, looking at hotels, browsing restaurants. "& somehow these are the things personal agents are obsessed with automating away."

AI Security Summit. swyx announced the second AI Security Summit, October 15 in San Francisco, with Snyk back as founding partner. "An exploding number of rogue agents, more breaches, more AI-driven attacks."

am.will on learning AI. "How do I learn AI? Such a simple question, but the truth is I dont have a clue how to answer it. 95% of the things I learned no longer apply." am.will also checked eBay for the Codex ModRetro, the console that ships with a blank cartridge you fill with a game built by Codex. Twelve have sold, for an average of about $1,400, a high of $3,000 and a low of $500.