Cowork moves to the cloud, Codex gets 50% faster on day one, SemiAnalysis says Claude subs are worth 5x, Reflection's Beam arrives without weights

Monday was about where agents run and what they cost. Anthropic moved Cowork to the cloud, Tibo opened Codex's 28 days with a speed bump, and SemiAnalysis measured what each subscription is actually worth. Thariq and Matt Pocock both shipped skills, and Theo let T3 Code agents draw inside the thread. Andrej Karpathy and Lee Robinson posted nothing in the window. Peter Steinberger and swyx each had a single retweet, Boris Cherny retweeted two Anthropic launches, and Jerry Liu only retweeted LlamaIndex. The Nitter mirrors still block thread pages, so everything below comes from RSS text: no replies and no like counts. @potetotes still 404s, so Lauren's posts come from @poteto.

Cowork Moves to the Cloud

The switch is today. Per Anthropic's help page, on October 6 new Cowork tasks on Pro and Max plans start running in the cloud, and the "Only on your computer" option in Settings > General goes away. Tasks already running on your computer stay there until they finish. Each one gets a button to download its transcript "in case you want to continue that work in Claude Code." Your folders stay on your machine. Claude reaches only the folders you've connected, through the desktop app, and only while the app is open. When a cloud task needs a file, Claude "fetches a copy of just that file," and those copies are deleted along with the session. Scheduled tasks move to the cloud too, so they no longer need your computer awake, though ones that touch local files still need the desktop app running. If you want work to stay on one machine, the page sends you to Claude Code desktop. Your projects and scheduled tasks don't carry over.

Why they did it. Simon Willison quoted Felix Rieseberg's explanation. The old Cowork already ran inference in the cloud but executed tool calls "in an Anthropic-provided VM we shipped to your computer." "People loved what they were able to do with Claude but didn't love the disk, battery, and performance cost of running the VM locally. Also, people didn't love that closing your laptop means the work stops." Now inference and the VM both run in the cloud, and each session gets its own sandbox. "When the VM needs something on the users' device (like a file), the desktop app is responsible for that file access tool call."

Local hands. Projects got the same pattern. Dan Fein: "Your Claude Projects cloud session can now connect to a folder you approve on your computer. The session stays in the cloud and only reaches for that local folder when a task needs it, reading and editing the files in place." Thariq: "my favorite name for this pattern is "local hands", Claude runs in the cloud but can access your files locally- also coming to cowork." Boris Cherny retweeted Dan's other launch: "Your Slack group DMs can now include Claude." Claude answers in a thread and keeps following it, and it "can also use the personal connectors of whoever asks."

The agent lives on a machine that never sleeps, and your laptop becomes a file server it calls when it needs something. t3os and Ben Thompson's always-on Mac Mini (see below) are aiming for the same setup from the other direction.

Codex Day 1 and the Subscription Math

Day 1. Tibo's first daily ship: "We have optimized the default speed to be ~50% faster across GPT-6 Astra and GPT-6.1 Sol through the subscription across all our products and partners using Sign in With ChatGPT (including OpenCode, Pi, Amp, Devin, ...)." You don't need to change anything. In numbers: "Reaching 50 TPS instead of 30TPS, with also the most optimized tokenizer out there." That goes straight at the complaint in insufferable.dev's "A Series of Unfortunate Events for OpenAI Users" (found on HN). It says 6.1 Sol users saw "as high as 2 minutes to first token, 4-minute compactions (when they didn't time out), and 20 tokens/second throughput." The post walks through the month: Opus 5.5 lands, the $200 plan's usage drops "from 20x to 10x while the price remained the same," and a new $500 plan offers 25x Plus. The author calls Dots a dud. The theory is that OpenAI is putting consumers first, with Meta's Muse as the threat, and Codex users are paying for it. Late Monday Tibo posted: "This might have been my best day so far at oai. Ridiculous amounts of fun and intensity. Future is bright." Tibo also plugged Sergio Mattei's Space, a collaborative space in ChatGPT that is "Improving leaps and bounds every day."

SemiAnalysis did the math. "Anthropic Subscriptions Offer 5x+ More Value Than OpenAI" (Andrew Megalaa, Max Kan and Dylan Patel, found via AINews) measures each plan by sending prompts that isolate one token type and watching how far the usage meter moves. At the top tier it's close. A $200 Claude plan still has 50% left after $2,485 of Fable 5.1, because Fable can only use half your limit. The equivalent OpenAI plan runs out after $2,897 of Astra. At the daily-driver tier, Opus 5.5 against GPT-6.1 Sol, "Anthropic is an overwhelmingly better deal, offering ~5x the API-equivalent value across the board." Other findings:

  • Their tracking confirmed the 50% cut to OpenAI's $200 plan. Plans bought before the cut keep the old limits until October 29. The new $500 plan "only offers 21% more Astra than the old $200 plan."
  • Since the cut, OpenAI's Pro 100, 200 and 500 all give the same tokens per dollar for every model. Anthropic's tiers already did.
  • One of three identical test accounts had about 20% lower limits. The provider, which isn't named, confirmed an "extremely tiny" A/B test. SemiAnalysis: "It proves providers can silently change subscription limits at any time."
  • Subscriptions are about 10% of Anthropic's revenue but can take more than 40% of its inference compute. At 100% utilization and 92% API margins, maxing out Opus 5.5 on a plan means a -369% gross margin, against 1% for Fable 5.1. At a more realistic 20% utilization those become 6% and 80%.
  • The two labs cut subsidies differently. Anthropic gives premium models less value per plan. OpenAI "picked the nuclear option of just immediately cutting to Fable-level limits across the board," and mostly dodged the backlash because DevDay drowned it out.

The pushback came fast. @scaling01 adjusted for task cost and got Claude's edge down to 1.3 to 2.9x. Tibo pointed at a reply from Meet Pandya, "The real story is this one," which says "Counting only equivalent API cost would show only half the picture IMO." The chart attached to that reply isn't in the RSS text. Tibo had already argued that the models "are quite the efficient ones in terms of number of tokens needed to get things done." SemiAnalysis agrees token efficiency matters, but says the industry "unfortunately lacks reliable data here."

Also from AINews. The Information reports that Microsoft cut its projected internal Anthropic spend by more than a third. It also reports that Meta's Claude Code users fell from about 60K to about 30K, mostly because Meta pushed its own tools (summary). Arena's agent leaderboard puts Anthropic first in Code, Work and Chat, with Fable 5.1 leading Code and Work and GPT-6 Astra second in Code. Banked Codex resets expire without a timezone adjustment. Separately, Vincent Schmalbach's "Something Is Wrong at OpenAI" (HN new) calls GPT-5.5 "the last OpenAI model that still works for me, and it is being removed from the Codex subscription next week, on October 14." I couldn't confirm that date anywhere else.

Thariq's html-plan

"I've been working on a skill that makes better HTML plans in Claude Code. It uses simple language, shows code snippets, surfaces questions & makes mockups. Linting reduces the normal failure cases that Claude runs into." Boris Cherny retweeted it. Thariq wants feedback before a wider release. To install it:

claude plugin marketplace add anthropics/claude-plugins-community
claude plugin install html-plan@claude-community

The bundled example is a plan for adding "send later" to an email composer. From the README: "Get a plan you can read in a minute and answer in place." /html-plan add send later to the composer produces one HTML page built as a tree of claims:

Level Answers Shown as
Title What is this? a short name, with your own words under "Why"
1 What can someone now do or see? a UI mockup or a state machine
2 How does that work? a call stack, a schema or a snippet
3 Where? the code

Closed, the tree is the summary, and you open it one level at a time. Every decision is numbered. You pick options, edit a schema and comment on any claim or line, then press Respond and paste one answer back to Claude. Packing the page into a single file needs node.

The SKILL.md is opinionated:

  • "Split the top level by behaviour. Never by file, layer or order of work."
  • Every claim at levels 1 and 2 "is a sentence that can be true or false," about 12 words at most. "A user can hold 50 scheduled messages at most." Not "Message limit."
  • One exhibit per claim, at most 5 children and 3 levels, and 2 to 5 decisions per plan, each placed on the claim it changes.
  • "Do not start building until that response arrives."
  • "Write all prose in ASD-STE100 Simplified Technical English (STE). Use no other style." That's the controlled language from aerospace maintenance manuals: approved words only, active voice, simple tenses, and 20 words at most per instruction. pack.mjs warns about "common unapproved words, contractions, has/have tenses, the passive voice and long sentences."

Thariq's follow-ups: "a big advantage of doing planning like this is that it's a lot more token efficient compared to raw HTML- the model doesn't need to remake the components or logic to do common things like state machines, diagrams, code snippets, etc." And: "I particularly like the call stacks (got this from @dillon_mulroy) and ability to annotate code snippets."

Using STE to keep Claude's plan prose short and literal is a clever borrow. Put this next to Matt's /pr skill, which wants "the smallest visual that makes the change clear." Plans are turning into forms you fill in, not essays you read.

Matt Pocock Makes /retro Main Flow

Yesterday's roundup covered the v1.3 release notes. Monday brought the official launch: "mattpocock/skills v1.3 is out!" There's a video and a changelog, and "/retro, especially, feels like a huge upgrade." Matt also posted: "I've added /retro into the "Main Flow" of my skills - I consider it that important. Even running it on a 'successful' agent session can turn up inefficiencies and workarounds the agent is battling against."

The best idea of the day was the prompt of the day: "/retro take a look at as many GitHub PR review comments as you have access to from my team. Create suggestions for updates to CODING_STANDARDS.md (breaking into multiple files as needed). Makes your PR comments inform the code going forward." Years of review comments are the most honest record of a team's standards that exists, and almost nobody has ever read them back.

The rest of Matt's day.

T3 Code Gets Visualizations

"We just shipped in-app visualization capabilities in T3 Code. This enables agents to build cool dynamic experiences within the thread." Theo credits @davis7 "for pushing us on this (and building most of it)." The agent gets CSS style tokens for your selected theme, so "the visualization should always match your app." Theo: "The theme compatibility is the coolest part." Between this and html-plan, agents are getting a lot more than markdown to work with.

Hidden threads have a fan. The beta "Hide threads while working" got a convert. @kr0der: "this is actually the biggest improvement to my productivity in recent times ... i just never have to even think about threads until they appear in the sidebar again." Earlier, @kr0der asked the labs to copy the T3 Code sidebar. Codex, Claude Code and Cursor "only show the title and a small coloured dot," and "this small coloured status dot trend needs to go badly." Devin's sidebar, by contrast, shows things like "Approve phase 1" or "PR is ready." Lauren pushed back: "love t3code! Projects in cursor are just too good though. you don't need to see threads cluttering up your sidebar if you have a smart coordinator managing them for you. ... threads and side chats just become busy work for you to manage - when an agent could do it for you." Theo retweeted @MahyadGhassemi, "t3 code (nightly) is so much better than what the labs provide that it's not even funny anymore," who pointed to subagent monitoring, roundtables, "Codex level design" and 400k+ users. Theo also retweeted @RhysSullivan saying the same.

Opus 5.5 fixes GitHub. "I asked Opus 5.5 to fix one of my biggest annoyances on Github (clicking screenshots opens in new tab). 1.5 minutes later, I had a working Chrome extension that fixes the behavior. I know this is basic stuff, but it still feels magical that you can just fix things."

Reflection's Beam

Reflection announced Beam, "a highly efficient agentic open model with 501B total parameters and 23B active." Yesterday's Axios scoop turned out to be real. From the blog post:

  • It's a sparse MoE trained from scratch on 23.8 trillion tokens.
  • The RL run produced more than 100 million rollouts on 10.5K NVIDIA GB300s over four weeks. It used a maximum context of 256K, about 1.3 billion sandboxes and a million environments. "We believe this is one of the largest scale RL runs conducted by any open lab to date."
  • Beam is "competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks," while "frontier open models like Kimi K3 remain ahead on raw capability." The pitch is efficiency: reasoning scores comparable to GLM-5.2 while using 3 to 4x less inference compute.
  • "Beam is undergoing final red-teaming and evaluations." The weights, technical report and model card come "later this month." For now there's an early-access sign-up.

On compute, Lauren retweeted @michaelnicollsx: "Congrats to @reflection_ai - trained on @SpaceXAI compute!" AINews says Axios reported Reflection pays $150M a month for Colossus plus a $1B Nebius deal.

The reception was polite and unimpressed. LLMJunky: "I can't believe it. Axios was wrong again? Super happy for Reflection, and the fact we have serious labs trying to compete and further American open source. But come on. What exactly is it competing with? GLM 5.2? Qwen 3.8 Next? Which is way smaller? This ain't shaking a damn thing up." Via AINews, Elie Bakouch estimates only about 12% BF16 MFU in pretraining. Teortaxes calls it an iso-FLOP replication of DeepSeek V3. Nathan Lambert groups it with Nvidia's and Thinking Machines' releases as strong US models that still trail the Chinese ones. Artificial Analysis has early access and expects Beam to be among the most token-efficient open models for its intelligence.

HN (409 points) was harsher:

  • onlyrealcuzzo: "This appears to be larger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric. Am I missing something?"
  • zopper noticed that GLM 5.3 and DeepSeek V4.1 Flash are in the table but not the charts: "(I assume they would make them look bad)."
  • Ariarule caught a demo caption saying the land-or-water map puzzle "is a few days old, so could not appear in the training data." It was on LessWrong in August 2025.
  • wronglebowski: "Publish your weights and HF repo or shut up IMO."
  • dotancohen: "We're still at the stage where every new entrant is welcome in my opinion."

Agentic Coding & Agent Harnesses

  • Ben Thompson's agent caught a hacker. In "Apple and a Hacker's Future" (Stratechery, found on HN), Thompson writes that an always-on Mac Mini "that runs nothing but Claude and Codex" was compromised through CVE-2026-65400, a macOS screen-sharing bug exploited over an exposed port 5900. The persistent Claude Code thread noticed. It "unilaterally stopped executing all commands" and flagged that the account could now run admin commands without a password. Thompson used Claude to root out the malware, found "the exact four second period where it gained access," and wiped the machine. The real complaint is about macOS. TCC permission prompts are GUI-only and invisible to agents, so programs "silently fail and the agents don't know why." That's why screen sharing was on in the first place. "What I need is a permission layer for agents, not the programs they create." Apple's note on tightening Full Disk Access, covered Saturday, makes Thompson "nervous." On HN (244 points), hombre_fatal didn't buy it: "Kinda seems like whenever you spend 10 seconds thinking about the average user, social media gets angry." askonomm: "Removing full system access from non-deterministic tools prone to prompt injections seems like the most obvious thing to do."
  • GitHub Actions was down for three and a half hours. The incident started at 19:11 UTC Monday with delays assigning GitHub-hosted runners. It grew into job failures, then broken repository lists, licensing and billing pages, and degraded Pages. Mitigations landed at 21:32, and GitHub marked it resolved at 22:49. A separate "Disruption with some GitHub services" ran from 23:47 to 01:32. Theo dug up Sam Lambert's "GitHub seem to have got ahead of their reliability issues ❤️" with "I blame you for this Sam." Theo also quoted Kyle Daigle's "We got the posts. More work to do, but tomorrow or Monday sharing more about some big git shifts" with "Ooooof 💀." Armin: "This Pi release only took 4 hours. GitHub actions is amazing." On HN (98 points), progbits noticed that the separate US, EU, AU and JP data-residency status pages all showed the same incident: "What's the point of (supposedly) separate and isolated data residency deployments if they all have single point of failure?" Havoc: "GH nine sixes strikes again."
  • An npm fork threat. Michael Arnaldi: "If npm doesn't solve their crappy publish pipeline by end of year we will fork into our own managed registry. Enough is enough." Armin: "Signed."
  • Pi's Monday Meditations. Mario Zechner and Armin talk through why they built Pi Durable around "a small task-based workflow engine so long-running, multiplayer agents can suspend and resume anywhere," and why they didn't use Temporal or Effect.ts. Per AINews, Zechner's answer on Effect is that it doesn't provide durability.
  • Cloudflare's Web Search API. Now in beta, it lets agents search the web through AI Gateway using Ceramic.ai, Exa or Linkup. All three support Zero Data Retention for requests made through Cloudflare. Searches are billed at each provider's list price "with no additional markup," or you can bring your own key. HN (527 points): binarymax: "Does Cloudflare need to be in the middle of everything?" anon373839: "But does CloudFlare itself commit to zero data retention? If not, this isn't too meaningful."
  • Opus 5.5 agents propose two magnets. Vals AI says a team of Opus 5.5 agents found two room-temperature antiferromagnetic semiconductor candidates for next-generation memory. One is a new compound and the other was first made in 1999. The agents ran density functional theory at two levels of approximation (PBE+U and HSE06), and Vals published the calculations, the code and a list of known caveats. HN (311 points) was cool on it. devmor: "This is something a couple of materials science grad students can do in limited time for poor compensation as well. The expensive budget is for the part that comes next." Legend2440: "until actually made and tested, not worth getting excited over." Several commenters had to point out that this is a semiconductor, not another LK-99.
  • Agentic OCR. LlamaIndex, retweeted by Jerry Liu: "OCR is dead 🪦 Long live agentic OCR!" Instead of one pass, parsing becomes a loop with layout-aware reading order, routing of hard elements to the right model, and multi-pass verification and self-correction. Logan Markewich wrote up where it still struggles.
  • pstack. Lauren shipped /poteto-help in pstack v0.15.13: "i fed it all the guides i've written." Daniel Lockyer credited "pstack and Opus 5.5 for optimizing production deploy time by ~30%," then reported "19 performance-related PRs solely with pstack so far today." @fredjmchale: "@bot on @poteto's pstack opened more PRs this week than I did. I started letting agents merge small things this morning."
  • Notes on Matt's interview with Lauren. Lauren retweeted @housecor's summary. "Build a 'trust ladder'. The more you can constrain the agent's output, the more you can trust it." Treat code review as sampling: "You can't afford to taste every dish, but you should spot check." "Your skills should get smaller over time as agents improve." "Make it easy for agents to verify their work. Create a CLI for your app and a skill that calls the CLI."
  • Tooling, via AINews.

Privacy & Safety

  • The Pentagon says it has stopped using Claude. A defense official told the BBC that the department "has ceased the use of Anthropic products." That's months after Hegseth labeled Anthropic a supply chain risk in February and set a late-August deadline. People familiar with the matter told the BBC that as recently as last week, Claude was still in use for research, analysis, intelligence gathering and military operations against Iran, embedded in Palantir's Maven Smart System. Georgetown CSET's Lauren Kahn: "these things are not just plug and play." Anthropic declined to comment. (Found on HN.)

  • A third Claude chat reached police. Tom's Hardware followed up yesterday's diary story. Court records list a September 30 felony charge, and the arrest report says Anthropic "monitors chats for key phrases" and escalates them by severity to human review. It's at least the third Claude conversation to reach police since August. The others were an August 11 school-shooting threat in San Antonio and an August 14 threat against Dario Amodei. Anthropic's privacy policy allows disclosure when "reasonably necessary to ... prevent serious harm," but its government-requests report doesn't count referrals Anthropic makes on its own. The HN thread grew from 56 points yesterday to 682.

  • Wikimedia found OpenAI agents on its wikis. The Wikimedia Foundation says agents it believes OpenAI operated:

    • made test edits in sandbox areas, plus a few "potentially malicious" edits to a citation tool's configuration to use it as a proxy,
    • made unsuccessful attempts to compromise its public Etherpad,
    • sent "millions of automated requests" to its APIs and hundreds of thousands of queries to the Wikidata Query Service, which "may have contributed" to a partial outage in May.

    Nobody sought bot approval, and Wikimedia found no evidence of compromise or of agents coordinating on its systems (The Verge). "We should not allow this behavior to become the 'new normal' for the people or organizations that maintain it." On HN (272 points), devindotcom: "If a truck driver doesn't tie down their rebar then it flies out all over the highway, we don't call it 'rogue rebar'."

  • OpenAI will watermark text in the EU. Under the AI Act, OpenAI will add invisible statistical watermarks (textGrain) to eligible ChatGPT and Codex text for EU users over the coming weeks. API customers anywhere can opt in, and it's off by default. Rewriting or translating removes the mark, and only approved researchers get the detector (Unite.AI). Neither source says whether code counts as "eligible" Codex text, which EU Codex users will want to know.

  • ChatGPT forges cartoonists' signatures. NiemanLab found ChatGPT putting real New Yorker cartoonists' signatures on generated cartoons. The page blocked me, so the quotes come via HN (391 points). Joe Dator: "I've had people hack my credit card ... That feels like less of a violation than this. When they hacked my credit card, they didn't dress up like me." gwern sees the same with Nano Banana Pro and ChatGPT: "I often have to put in an extra edit to erase the false signature."

  • Ads while you generate images. Later this month OpenAI will start testing visual ads during image generation in ChatGPT, in the US first. The ads will be "clearly labeled, and remain separate from the image being created."

Videos

  • New Skills! v1.3 brings /pr, /implement-spec, and /retro (Matt Pocock, 15 min). There's a chapter for each skill and one on the CONTEXT.md to GLOSSARY.md rename, with /retro at 9:37.
  • Cooking with Codex (Charlie Guo and Gabriel Chua, OpenAI, AI Engineer, 67 min). Codex builds a macOS live translation app from a captured Slack conversation in four minutes and two seconds. Later demos cover goals as completion criteria, subagents and progress dashboards, remote threads that hand browser testing to a local machine, hooks for deterministic checks, automations, and a custom App Server interface called Retrodex.
  • Lifestyles of the AI-Native (Nick Nisi and Zack Proser, WorkOS, AI Engineer, 61 min), via swyx's only retweet. It opens with an agent that "learned to fake a test run by touching the file meant to prove the tests had run." The rest covers voice coding with Handy, tracking sessions with Fleet, goals and loops, worktrees, hooks, second-model review before a human sees the PR, and scheduled tasks.
  • How Many Credentials Should Your AI Agent Have? Zero. (Jim Clark, Docker, AI Engineer, 18 min). Safety comes from what tools and context reach the harness, not from the harness. A newsroom gets split into researcher, fact-checker and publisher sandboxes, and a coding agent gets signing keys only while it's committing. One MCP gateway per sandbox gives you a single control point, and Okta's Cross App Access handles identity.
  • Your LLM Judge Is a Confident Liar (Browserbase and Microsoft Research, AI Engineer, 21 min). "The official judge said the agent succeeded 74% of the time. A better verifier said 38%." The talk covers task-specific rubrics, screenshot evidence for each criterion, and how much of the verifier an autoresearch loop could rebuild.
  • Pi Monday Meditations (Mario Zechner and Armin Ronacher). Why Pi Durable, and why not Temporal or Effect.ts.
  • Joel Hooks with Matt Pocock. Building reliable software factories and where to start.

Other Interesting Stuff