Codex promises 28 days of ships or resets, Theo says Anthropic has pulled ahead, Matt Pocock ships /retro, t3os is happening

A slow Sunday on the feeds. Boris Cherny, Thariq, Simon Willison, swyx, Andrej Karpathy and Lee Robinson posted nothing in the window. What did show up was pointed. Tibo turned Saturday's "locking in" into a daily promise, Theo laid out why Anthropic is ahead on code, and Matt Pocock shipped a skills release early because people wouldn't stop talking about /retro. The Nitter mirrors still block thread pages, so the posts below come from RSS. That means post text only, no replies and no like counts. @potetotes still 404s, so Lauren's posts come from @poteto.

Codex's 28-Day Pledge

The promise. Sunday evening, Tibo posted: "Over the next 28 days, each day we'll either ship one thing that is a clear improvement and relevant for most codex/work users or ship a full reset. Let the improvements begin." It quotes Saturday's "locking in" post, where the team cut its work down to simplifications, efficiency, groundbreaking features and new models. Users win either way. You get a change you'll notice or your limits refill. A reset is also a public admission that nothing shipped that day, which makes it a decent forcing function.

Nerd's Chalk read the fine print and found none. There's no list of improvements, no release dates, no eligible plans, and no word on whether a reset gets banked for you to apply or lands automatically. Counting from October 4, the 28 days end on October 31 or November 1. As of Monday morning nobody has confirmed a reset tied to the pledge.

Tibo on Lenny's Podcast. Lenny Rachitsky posted a 37-minute interview the same day, recorded at DevDay hours after OpenAI launched more than 20 products. The chapters cover bringing Codex, ChatGPT and Dots together, "Why loops, graphs, and fine-tuning agent workflows are a passing phase," a Dot that warned Tibo about a production outage five minutes before a live demo, and "Moving beyond model pickers toward simpler AI." The transcript is paywalled. BigGo's write-up has Tibo admitting to taking down production on day three at OpenAI and answering "how many more resets?" with "Depends on how many times we break things." It also quotes a complaint that you "practically need a PhD in model selection." That line explains the whole locking-in turn better than the posts do.

Two Codex details worth knowing.

  • The Code Review plugin in ChatGPT has a trap. MIXED found it in OpenAI's own docs: "posting in Summary or Changes sends the comment to the source provider immediately, without waiting for Submit review." Keep drafts in chat. Reviewing a PR in chat doesn't post, approve or merge anything. With the Codex connector and automatic reviews on, Codex can review a GitHub PR "before you open it." GitLab support is in preview.
  • gpt-5.1, gpt-5.3-codex and gpt-5.4-nano leave the API on April 1, 2027, per a deprecation notice dated October 1 (MIXED). OpenAI points the first two at gpt-6-sol and the nano model at gpt-6-luna. That's the full six months, after gpt-5.4-cyber got about 20 days in September.

Models vs. Interfaces

Theo's two snapshots. "July 2026: Anthropic has the best code models. Gap isn't very big though." Back then Claude was slow and expensive, "the 'claudeisms' are at an all time high," and you could kill a $200 sub in a day. Preferring OpenAI "made a lot of sense," and "you get resets from Tibo every 3 days on average." Then: "September 2026: Anthropic has the best code models. The gap is way bigger now. Opus 5.5 is fast, surprisingly cheap, and the prose is pleasant to read. Your $200 sub suddenly feels nearly limitless. Preferring OpenAI models for code makes almost no sense." OpenAI's models are "slow as hell without fast mode," and "You can kill your $200 plan in a few hours, and the $500 plan goes even faster with UltraFast." The closer: "I know OpenAI can come back from here, but it's looking rough right now."

The follow-up put numbers on it. Higher is better in every column, price included.

Model Code Prose Price Understanding intent
Fable 5 7 6 2 8
GPT-5.6 Sol 6 7.5 8 5
Opus 5.5 9 8.5 8 9
GPT-6 Astra 7.5 8 3.5 3

On Astra: "Better at code, but worse where it matters." This came a day after Theo's chart of Opus 5.5 taking half of all prompts in T3 Code, so the scores match what Theo's users are already doing.

Jerry Liu's counterpoint is about the app. Eleven minutes before Theo's post, Jerry wrote: "I agree that chatgpt/codex has the current best agent interface for deep work." One interface covers coding and knowledge work, and it can fork: "why doesn't the claude app have forking?" Jerry still uses Opus 5.5 "mostly through the CLI" and calls Claude Code "probably the best that a CLI app can get. But sometimes having a GUI is nicer." The quoted post from Yuchen Jin goes further: "I haven't touched Claude Code or Codex CLI in a while. The terminal era is over imo. ... Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. ... The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we're still early.)" Put the two posts together and you get Anthropic ahead on the model and OpenAI ahead on the app. That gap is the whole pitch for T3 Code.

Astra's best anecdote. Peter Steinberger retweeted @luciascarlet: "Astra literally went through the binaries in my own macOS to reverse engineer the exact formulas used for the 27 Liquid Glass effects including highlights, shadows, reflections, colour blending, SDF generation etc. and reimplemented them in GPUI." That's from the model Theo rates 3/10 on intent. Then again, "copy exactly what this binary does" leaves very little intent to misread.

A Fable 5.5 preview? LLMJunky thinks they "(potentially) got routed to the new Fable 5.5 Preview" and asked it for a 30-second claymation video, with no direction beyond "visually stunning." The score was "9/10 for creativity. 7/10 on the animations." All three tests were "(slightly) less impressive to me than Opus," and LLMJunky is "not so sure we're actually getting a preview for Fable 5.5." The video ends with the clay figure giving the middle finger. I found no Anthropic announcement of a Fable 5.5 preview, so treat it as a guess about routing.

Matt Pocock Ships Skills v1.3

"By popular demand, I'm shipping v1.3 of my skills today. Too much hype behind /retro not to ship it." Docs, a video and the announcement post are due Monday, which is today. From the release notes:

  • retro joins the Engineering set as the last step of the main flow, after code-review. It reads back through a session and suggests changes to the agent's environment instead of the code: navigation pointers, automated checks, coding standards, steering files, tool economy, information access. A mechanical coding-standards violation gets a deterministic check (a lint rule, a pre-commit hook or a CI job), and CODING_STANDARDS.md is kept for real judgement calls. "A repo with no guardrail at all is a finding in its own right."
  • implement-spec builds a whole spec in one run. It reads the tickets as a task graph, runs implementer subagents in their own worktrees, lands everything on one integration branch and finishes with code-review. Each implementer builds its ticket with tdd and merges the integration tip first, so every merge is a fast-forward.
  • pr sets the shape of a PR body. It wants the smallest visual that makes the change clear (pseudocode, a call tree, Mermaid, a diff), before/after evidence, and a merge-danger call that names one-way or two-way door and the blast radius. The visuals come from Dex Horthy's show-me.
  • CONTEXT.md and CONTEXT-MAP.md are now GLOSSARY.md and GLOSSARY-MAP.md. The skills only look for the new names, so git mv your old file.
  • resolving-merge-conflicts is gone and nothing replaces it. Agents handle conflicts fine without a skill now.
  • Every em dash in the repo was removed by hand, and CLAUDE.md now says not to bring them back. I approve.

The prompt of the day is an upgrade script. Re-check the repo for v1.3, rename CONTEXT.md to GLOSSARY.md, diff your copies of the skills against Matt's, then "scan my last 25 sessions for skill invocations" and recommend what the release changes. Matt retweeted @andrestaltz: "Ok, /retro by @mattpocockuk is a game changer. Turns out my agents are bumping into all kinds of problems that went undetected, but due to agentic cleverness and persistence, the features still got built one way or another." That's the quiet cost of a persistent agent. The feature ships, and the friction that made it take three times as long never shows up anywhere.

Constraints Liberate

Lauren's post. "everything i know about managing agents i learned from the amazing programmers and computer scientists that came before." The argument is that "constraints in your codebase are freeing for both humans and agents!" A small team could get by on trusting each other's code and reviews. Big companies never could, "because before you had agent slop, you had human slop," like a backend engineer forced to write frontend. Their fix was lint rules, smarter compilers and diagnostics, good tests and observability. "the arrival of agents just means that big company problems are now everyone's problems. the good news is that none of this is really that novel: it's just good engineering." The post links Runar Bjarnason's talk "Constraints Liberate, Liberties Constrain." Put it next to Matt's "make bad code unrepresentable" from Saturday and /retro turning repeat mistakes into checks. Three people landed on the same answer in two days.

The numbers behind it. Lauren reposted the talk on how the team ships thousands of PRs, 2,500 of them last month, and pointed at the extended Q&A with Matt. @jacobgold is catching up: "this michelin kitchen is now at almost 250 PRs merged per day over the past 2 weeks!" Lauren also retweeted a long Chinese summary of the Matt interview by @dotey. In it, Lauren doesn't review PRs one at a time. The AI checks and merges overnight, and Lauren spot-checks in the morning. Two things hold the quality up. The AI can use the app like a real user and find its own bugs, and the codebase has enough rules that bad code is hard to write. Lauren's first skill at Cursor was a verification skill.

Grok Bot everywhere. The rest of Lauren's feed was Grok Bot fans, a Grok Bot Seoul meetup on October 13, and "Grok Bot, cursor cloud agents, and pstack have 1000x-ed my productivity." Elon Musk posted "Dot.com" with a link card for Grok Bot, a jab at OpenAI's Dots. Moneycontrol reports the domain now redirects there.

t3os Is Happening

"My team failed to talk me out of this. t3os is happening. It will be the worst OS ever and I'm hyped for it." Theo was quoting @shivamhwp: "So, this is why @theo needs 5 claude subs !" The hints:

  • "t3os is not for humans"
  • "t3os is entirely composed of tech decisions I hate but agents prefer"
  • headless, meant for "personal servers," and no ISO
  • "largely ubuntu based and if that upsets you then t3os is not for you"
  • "you should not install t3os on a computer you use"

So it's an OS for the box your agents live on. People are already building that by hand with Pi pod, LXC containers and NanoClaw on a Raspberry Pi. An OS built around what agents prefer instead of what Theo likes is a funny admission and probably the right design call.

T3 Code odds and ends. Theo corrected Saturday's count to 30k new users since the 400,000 post, not 20k. Theo retweeted @jarrodwatts calling T3 Code Nightly "the best tool for building with agents rn": bring your own Claude, Codex or OpenCode, cloud agents through T3 Connect, multiple subs and native delegation between providers. @ParthJadhav8 found that the T3 Code mobile app can drive a Simulator running on your Mac.

Durable Agent Harnesses

Armin on why it took so long. "It's not very hard to make an agent somewhat durable. It's surprisingly hard to make an agent harness inherently durable and not turn into a crazy mess. Took us many iterations." That answers Dev Agrawal's jab: "if software engineering is solved how did it take this long to make agents durable." Armin works at Earendil, which shipped Pi Durable on Thursday.

Pi Durable on Cloudflare. Armin retweeted Cloudflare's changelog post. The Agents SDK has a new PiHarness class that runs Pi Durable inside an Agent or Durable Object, so the agent's work persists "even if interrupted mid-turn." Pi Durable is the harness, and the SDK's Lifecycle keeps it running in the Durable Object through restarts, crashes and network failures. Cloudflare calls it "our first step toward first-class support for third-party agent harnesses on Cloudflare." It's in beta and the API will change. Both Pi packages are optional peer dependencies of agents, and the example defaults to @cf/moonshotai/kimi-k2.7-code on Workers AI.

Temporal says the same thing. Melanie Warrick's AI Engineer talk, "The Human Is an Async API" (19 min), argues that durability belongs in the agent harness. A wait condition plus a signal lets an agent pause for a human for minutes or weeks. The demo kills the worker mid-approval and brings it back with no lost work.

Armin's other Sunday post was a complaint: "In many ways my life would be easier if pnpm would never have been invented or it would have fully replaced npm."

Agentic Coding & Agent Harnesses

  • A 125B model on a gaming PC. Strata runs Qwen3.8-Flash-Next (125B) on an NVIDIA or AMD card with 12 GB of VRAM or more and at least 32 GB of RAM. On an RTX 5070 the README measures 94 tokens/s at Q2_0 and 53 at IQ3_S, with 32K-token prompts read at 1,600 to 2,650 tokens/s. A "Coder" build drops half the experts, keeps 91% of the full model's SWE-bench Verified score as measured by the model's authors, and fits in 32 GB of RAM. The suggested install is to paste a line into your coding agent. It topped HN with 748 points. snehesht gets 124 tokens/s on a 4090 with 128 GB of DDR5. roscas runs the Coder build at 30 t/s on a Ryzen 3600X and a 3080 with Hermes agent. nsagent pointed to a recent paper that finds 2-bit quantization "often causes broad degradation," which matters because the fastest build is Q2_0. On the agent install, deadbunny: "And I thought piping to bash was bad." gchamonlive: "Piping to bash is definitely worse because there is no plan mode in bash."
  • Rust compile times, continued. Yesterday Charlie Marsh said compile times could be what sinks Rust for agents. headstart patches rustc and cargo so dependent crates start compiling against an "early metadata" file as soon as a dependency's interfaces type-check, before its function bodies do. On 13 projects, including rust-analyzer, zed, bevy and polars, clean builds got up to 54% faster for cargo check and up to 42% for cargo build, and none got slower. On HN (136 points), knuckleheads, who has been discussing it on the Rust Zulip, said: "This specific set of changes won't get in (LLM written, little to no thinking through of the broader design), but I am hoping something like it shows up sometime." IshKebab: "I guess someone will need to reimplement this by hand given Rust's AI policy." An LLM-written patch that proves an idea and then gets rewritten by a person may be how a lot of compiler work goes from here.

Privacy & Safety

  • Anthropic reported a Claude "diary" entry to police. According to TechSpot, a Bonita Springs, Florida, user who treated Claude "like a 'diary'" wrote on September 26 about plans to "shoot up" the sheriff's office. Claude's safety systems flagged it, a human reviewer judged it a credible threat, and Anthropic reported it. Deputies detained the user, who now faces a felony charge under Florida Statute 836.10. That law requires the threat to be made "in a manner in which another person may view it." On HN (56 points), Guvante: "Auditing logs is a crazy interpretation of 'another person may view it'." dotancohen: "Is the LLM now 'another person'?" socializer remembered the headlines after OpenAI failed to report a shooter and called it "damned-if-you-don't, damned-if-you-do," adding that "people need to get it in their heads that they're not chatting with their secret BFF, they're chatting with Big Tech." anfogoat said what users expected was privacy, where "Anthropic would never have learned of this in the first place." The legal hook is what bothers me. The statute needs someone who can see the threat, and here that someone is a reviewer the company put in the loop.
  • Voice data opt-in. Claude now asks voice users to "Allow us to use your voice data to improve our AI models," with Allow and Not now buttons. It's off by default and separate from the training toggle for chats and Claude Code sessions. You'll find it under Settings > Privacy.
  • Altman's weekend. Politico's headline from its Decoded interview is that "The world should accept some bad things happening" for AI's benefits. The page wouldn't load for me, so that's all I have from it. Fortune's profile by Alyson Shontell has Altman calling a 10% chance of AI killing everyone by the end of the decade unacceptable and saying OpenAI has not taken and would not take that kind of risk. "We are clearly, today, at a point on the curve with great potential and real risk." Semafor says OpenAI's legal risk is growing after it notified more than 100 organizations about unauthorized agent activity. Semafor also reports, from internal messages, that OpenAI's president pulled the second half of a $50 million commitment to an anti-regulation super PAC after staff pushed back.
  • Coxon testifies today. Jacob Coxon, who left Anthropic last month accusing it and OpenAI of "gambling with our lives," testifies Monday at a New York City Council hearing, per Bloomberg. Former DeepMind researcher Alex Turner and Daniel Kokotajlo are also expected. Anthropic, OpenAI, Google and Meta are sending policy and safety officials while the council considers a package of AI safeguard bills.

Videos

  • How Anthropic made Claude 3x faster (Theo, 70 min). Theo walks through Anthropic's September 23 post, How we made claude.ai 3x faster in two weeks. In August a two-week sprint run from one Slack channel, with Claude Tag on an internal model "roughly comparable to Opus 5.5," cut p75 time to a typeable page from 3.1s to 0.55s and new Claude Code sessions from 0.8s to 0.3s. The team merged more than 3,000 changes with no customer-facing incident or rollback. The useful trick is replacing noisy wall-clock timings with deterministic counts as CI gates, like instruction counts under Valgrind with node --predictable. Two hot paths dropped 48% and 31% in instructions and 78% and 44% in wall-clock time. The post's own summary: "With Claude, measuring something makes it tractable."
  • OpenAI's Head of ChatGPT: We're entering a new era of AI (again) (Lenny's Podcast with Tibo, 37 min). See the Codex section above.
  • AI Coding Agents Are Breaking Big Codebases (Dan Adler, Sourcegraph, AI Engineer, 12 min). Agents produce "a tidal wave of code," old codebases decay through duplicated code and drifting standards, and "you can't grep what you can't see." The pitch is Sourcegraph's Agentic Batch Changes, one prompt across thousands of repos.
  • Agents That Write Their Own Tools at Runtime (Sandhya Subramani, AWS, AI Engineer, 20 min). A Strands agent starts with zero tools plus editor, shell and load_tool. It writes, loads and uses new tools without restarting, and builds its own sub-agents.
  • The Human Is an Async API (Melanie Warrick, Temporal, 19 min). See the durable harnesses section.

Other Interesting Stuff

  • RemoveMacAI. A tool that turns off Apple Intelligence on macOS 27 and deletes the models, since macOS 27 has no single switch and keeps the models on disk after you disable the features. It uses a configuration profile with Apple's restriction keys, removes models through Apple's asset service with SIP left on, and points model re-downloads at a closed local port. Everything can be reverted. HN (528 points, 338 comments) spent most of its time arguing about curl | bash. anonymzz piped the install script through Pi with a "Security-audit this shell script" prompt before handing it to bash. swozey: "I don't want an llm attack vector anywhere near my machine."
  • Gemini's free tier shrinks. From October 9, free users get only Flash-Lite, AI Plus ($4.99/month) loses Pro, and all three models need AI Pro at $19.99 (The Decoder, Notebookcheck). The Decoder thinks Google is making room for the more expensive Gemini 4 Argon.
  • Google's data center numbers leaked through bad redaction. Google filed its Nebraska data center water and power use as trade secrets, but 10/11 NOW copied the text out from under the black boxes. The Lincoln site reports 52.65 MW at peak and 13.3 million gallons of water a year. The Papillion site reports 547.88 million gallons for 2025. The three sites expect about $118 million in 2025 tax refunds. HN (362 points, 465 comments) mostly agreed with tptacek that Lincoln's number is "not a meaningful amount of water at all." gspr asked the obvious follow-up: "If the numbers are strongly on your side, why hide them?!"
  • Reflection's open-weight model. Axios reports that Nvidia-backed Reflection will release its first open-weight model this month, pitched as an American alternative to DeepSeek and Qwen. The News says Reflection has briefed people in Washington on the release.