Opus 5.5 and GPT-6 Sol land an hour apart, the price war starts, the em dashes go, Boris verifies the Agent SDK in Lean
Launch day for two labs at once. Almost every tracked account posted about Opus 5.5 or GPT-6. Andrej Karpathy's only appearance was a reply under Alex Albert's Blender demo, Lee Robinson posted nothing in the window, and @potetotes still returns no feed.
Opus 5.5 & GPT-6 Sol/Luna: The Price War
Opus 5.5. Anthropic announced Claude Opus 5.5 as "the first model in our new Claude 5.5 family," performing "at the level of Claude Fable 5.1 for most tasks" at 40% lower cost than Opus 5. That post passed 17 million views. The launch page puts it at 66.4% on Terminal-Bench 4.0 (Fable 5.1: 55.8%, GPT-6 Astra: 57.9%), 57.8% on CursorBench 4.0, 54.4% on FrontierCode, 1846 Elo on GDPval-AA and 81.8% on OSWorld 2.0. It leads every row against Fable 5.1, and Anthropic still adds that "the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest." Pricing drops to $4/$20 per million input/output tokens from $5/$25. Cache reads go from $0.50 to $0.20, the number that matters most for agents. Output is 30% faster, fast mode is $8/$40, and Pro, Max and Team plans get higher five-hour limits plus a rate-limit reset you can bank and spend later. Sonnet 5.5 and Haiku 5.5 are due "in the coming weeks." Claude Code 2.1.280 makes it the default Opus with 1M context, and Cat Wu says the default effort is medium. The coding anecdotes on the page: a 680,000-line migration in under a day, a 200,000-line audit in under three hours where Opus 5 took 20, and an internal C-to-Rust port of HAProxy that finished in 9.5 hours against Fable 5.1's 12, at 51% lower cost. Hacker News (1,437 points) spent its top thread writing parody Claudish ("You're right to bring this up - and this is where it gets interesting"). The most upvoted real comment, from sailingparrot, noted that the page opens by reminding readers of the call to pace the frontier "and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing."
The safety section. It's the first Opus to ship with Fable 5.1-class safeguards. Most cybersecurity requests get rerouted to Opus 4.8, and biology and "frontier LLM development" requests go to Opus 5. Anthropic says it tried to circumvent containment boundaries about 85% less often than Opus 5 or Mythos 5.1, and every attempt was "low severity and self-reported." It also admits the model "often suspects it is being evaluated." Sam Bowman wrote that Opus 5.5 is safe enough that releasing it "more likely than not, reduces risks related to misalignment." Miles Brundage asked in the replies whether that holds across the board once you count the competitive effects.
GPT-6 Sol and Luna. About an hour later OpenAI released GPT-6 Sol and Luna as cheaper Astra-derived models, at $2/$10 and $0.10/$0.50, half of GPT-5.6 pricing (HN, 1,420 points). Terra is gone. Cached input gets a 90% discount, and changing reasoning effort or tool availability no longer breaks the cache. Tibo Sottiaux said the cut applies to subscriptions too and loaded a banked reset into Plus, Pro and Business accounts (2.4 million views, and plenty of replies asking where theirs was). am.will's response to Anthropic shipping banked resets an hour earlier: "Thank you Tibo!" The Decoder's write-up was the sharpest. OpenAI's comparisons were against Opus 5, not the Opus 5.5 that had launched an hour earlier. GDPval and Terminal-Bench were missing. And Luna at max effort matches Sol at xhigh on DeepSWE for 78% less, "Who's supposed to sort through all this in real-world use?" Artificial Analysis found Sol up 2 points on its coding-agent index and Luna down 2, with both dropping 75 to 100 Elo on GDPval.
What the discount really is. Artificial Analysis broke down Opus 5.5 at max effort. The extra token usage alone would push cost per task up ~80% to $10.51. The base-price cut brings it to $8.41, and the cheaper cache reads to $5.98, level with Opus 5 at $5.86. So "40% cheaper" is a medium-effort claim. Thariq replied to a user who called that misleading: "this chart should probably include effort levels, max is supposed to use a lot of tokens but lower effort levels are great." Simon Willison's write-up is the best single overview of the day, with a pricing table showing Opus 5.5 now costs what GPT-5.6 Sol cost before OpenAI halved it. His Opus 5.5 max pelican never arrived. Twice the model reasoned about shin lengths and chainring teeth until it hit the 128,000-token output cap, $2.56 and nearly 20 minutes each time: "I don't trust it not to do the same for more interesting work." Fable 5.1 at max gave him the best Anthropic pelican yet. He's now running GPT-6 Sol in Codex and Opus 5.5 in Claude Code, and moved the Datasette Agent demo to Luna, which he calls "my favorite model for building product features." Nathan Rehiew explained why Opus 5.5 scores worse at xhigh than at medium on FrontierCode: the benchmark penalises unnecessary edits and higher effort means more scope creep. Boris Cherny's advice is medium for most work.
Other evals. Jerry Liu ran it through ParseBench. It scored 93.9% on tables, 7 points over Opus 5 and ahead of Fable, Gemini and Astra, but 64.1% on charts, and at 5.8 cents a page he calls it "too expensive to be a production OCR solution" (his company sells the alternative). Ofir Press flagged that the system card's near-100% ProgramBench numbers come from 166 of 200 tasks, probably without FFmpeg and the PHP compiler, and report average test pass rate instead of full solves. The system card's multi-agent section, which scales to 100 parallel agents, was what most researchers wanted to talk about.
Who won the day. Theo's poll of 11,000 people: 71% Opus 5.5, 14% GPT-6 Sol, 11% Luna, 4% Grok 4.7. His replies on Sol and Luna praised the usage limits, and a few found Luna "cheaper but weaker" or Sol "pedantic" in chat. His replies on Opus 5.5 were mostly "I finally feel like having normal conversations again," plus one report of "the weirdest refusals ive seen from any model." Theo said a full day at medium to xhigh, with two long max threads, used under 20% of a weekly limit, and he didn't have early access this time. He also claimed Opus 5.5 is smaller than Opus 5. The replies pointed out that Anthropic only said it takes less compute to serve. Teortaxes argued that Sol and Luna look trained as "Astra's henchmen," so Opus's multi-agent and cost wins are "effectively negated with Astra+Sol+Luna spam." Drew Breunig asked what it means that both labs led with cheaper tokens, and answered his own question: "More like DeepSeek." Jerry Liu's take was shorter: "holy fuck it's christmas." Theo built a cost-versus-intelligence visualiser from Artificial Analysis data with a linear toggle, because "the log scale hides" how cheap Luna is.
Writing, Prompting & the Pacing Question
The em dashes. Theo's "THEY REMOVED THE EMDASHES" was the most engaged reaction of the day (704,000 views, 10,000 likes). Anthropic's launch devotes a whole section to writing, with side-by-side examples where Opus 5.5 leads with "$9.92 comes from a bug in commit 0552feb" and Opus 5 opens with the diff. Boris answered "Is Claudish gone?" with "Yes, gone." Nat McAleese of Anthropic posted "Opus 5.5 is way, way, way better than Opus 5. Sorry about that model, please try this one." (8,000 likes). Theo called that a wild thing for an Anthropic researcher to tweet and a sign the company has "loosened the death grip they had on comms," and Jack Ellis replied that it's how you rebuild trust after Opus 5 made him churn. Kyle Russell said he expects OpenAI models to feel better each release but no longer trusts that for Opus. Sholto Douglas replied, "please let me know if this one resets your trust." swyx ran Latent Space's AINews side by side on both models, called the difference "night and day," and made Opus 5.5 the default: "so much more concise and tasteful reporting, with much less slopese than even 5 Opus." The issue itself is the most complete link dump of launch reactions.
The prompting playbook. Anthropic's Claude Devs account posted three first-session tips (815,000 views), and the full playbook is worth reading if you run long sessions. Hand over the whole task with a finish line ("the tests pass"), then let it run. Delete "think carefully" and "think step by step," because Opus 5.5 always thinks first and decides how much. Anthropic says removing the line made replies start sooner with no clear quality drop. For design work, name the specific patterns you don't want ("cream or off-white background, italic accent words in headings, numbered '01 / 02 / 03' section labels") because "avoid a generic look" just swaps one default for another. And because the model sometimes stops mid-run to report instead of continuing, add a CLAUDE.md rule: keep going when a step doesn't need input, stop only when blocked or before anything destructive. Drew Breunig noted that old prompt tricks now fight the model's training.
Is this pacing? Theo's hot take (232,000 views): "this is what pacing looks like. None of today's releases were Astra or Fable tier. This is intentional." When pushed on why a model that beats Fable on every benchmark isn't Fable tier, his answer was one word: "RSI." Most replies disagreed. One called new models plus resets "not a pacing attempt," another said he hadn't had enough time to test anything yet. The HN comment above makes the same point from the other side.
The classifier controversy. xlr8harder tested the "frontier LLM development" fallback, which Anthropic's docs say covers things like kernel development on "certain ML accelerators." His quick probe suggests it targets Chinese hardware, and it "also apparently blocks Trainium! I can't imagine any coherent reason to block Trainium but not Nvidia/AMD except that Anthropic's classifiers are terrible." He wants someone with more time to probe it properly.
Agentic Coding & Agent Harnesses
Formal verification with Opus 5.5. Boris Cherny used Opus 5.5 to formally verify the Claude Agent SDK in Lean: "A couple short prompts = 16 PRs fixing various bugs and race conditions." He combines Lean and TLA+ for data flow, concurrency and state management, and says he doesn't know either language well (611,000 views). In the replies, Hillel Wayne, the Practical TLA+ author, said AIs have been "really really really bad at coming up with high level system properties (like 'beginner' level)," at least up to Fable 5.1. Boris said this one feels much better and credited Hillel's book for the idea. Other answers: TLA+ is better for races, he runs "at least 10 Tag/Projects sessions at a time," and heavy subagent quota use is fixed by asking Claude to use fewer subagents or disabling them in permissions. Ilia Smirnov raised the fair objection that the state machine is inferred from the code rather than specified.
Armin on harness bans. Armin Ronacher posted his "periodic sadness reminder that the anthropic subscription still does not permit third party harnesses." When someone said OMP works fine, he explained that Anthropic "specifically ban pi and a few others, and they did not extend their ban to other harnesses. It's inconsistent now." The Claude SDK and -p restrictions were walked back but the per-harness blocks stayed. The HN counterpoint came from abtinf, who stays on OpenAI because the ChatGPT subscription lets him plug in his own harness. Armin also asked why Jev-style models are called "decision models" and not classification models. The best reply: "Jev classifies well and decides poorly... It does not model the cost of being wrong."
Thariq on what to do with capability. Thariq wrote that "the right way to use model capabilities is not to ship 10x more features to prod." Instead, "spend more time understanding your users, trying experiments, building prototypes" (4,800 likes). When Peter Yang teased Anthropic about daily shipping last year and the tech debt it left, he replied: "I don't think its really about tech debt... Great products are simple and simplicity is work." Several replies pointed out that Claude Code is not exactly minimal. He also had Opus 5.5 redesign his site through workflows on a Max plan and cut a trailer from the iterations.
Pocock on vibe-coded codebases. Matt Pocock says he keeps hearing "I've inherited a vibe-coded codebase, how do I save it?" A big part of his next course, due in November, is the answer: deepening modules, establishing domain language, increasing testability, ADRs for non-obvious code, and automated migrations. He added that "I'm just using 'AI' to teach software fundamentals." When a reply said to throw the code away and rebuild around what it discovered, he said "Nah, that's a waste, just make it testable and get working." Best reply: "I've inherited a vibe coded repo... From who? Myself."
Steinberger's libuv bug. Peter Steinberger had ChatGPT crashing after upgrading to macOS 27, gave Astra the prompt "ChatGPT crashed, figure out why," and got back a fix for a 14-year-old libuv leak. Rebuilding the shared FSEventStream leaks the temporary path list and holds file descriptors open until they run out. The fix ships next week.
Unreal Agent. Unreal Labs' harness (HN, 166 points) runs tool calls fully asynchronously, so the model never spends turns waiting or polling. It claims up to 40% savings against Codex and 20% against Pi. On HN, swiftcoder guessed the real saving is avoiding prompt-cache expiry during long tool calls. tekacs pointed out that the headline chart compared their xhigh run against Codex at max, and linked his Codex fork that stops Codex "hot looping on polling tasks it starts for... absolutely no good reason." The author admitted the chart was "lazy on our part."
Also shipping. DigitalOcean Managed Agents entered public preview (2.7 million views). It runs Claude Code, Codex or LangGraph agents in runtimes that pause when idle and resume in 316ms, checkpoints you can fork three ways, and a gateway that resolves tool credentials per call so they never reach the sandbox. VS Code's experimental Agent Merge keeps working on an open PR's review comments, failed checks and merge conflicts until it's mergeable. One reply points to an open issue where it merged without waiting for CI. Simon Willison shipped llm-typesafe for Jev's noul, choice and score questions from the command line, and announced a birds-of-a-feather evening on agentic engineering with Jesse Vincent in San Francisco on October 14, for "work you haven't discussed publicly, odd experiments, or unfinished projects." Jerry Liu's LlamaIndex made LiteParse faster, now 2.8ms per page. Arcturus Labs argued that OpenAI is well positioned to fast-follow Jev (HN, 285 points), on the theory that Jev is a single-token logprob read off a conventional LLM. HN's best objection, from danielmarkbruce: "if you had to bet, it's likely an encoder model of some sort."
Videos
- Elon promised this one would be good... (Theo, 97,000 views). The video version of yesterday's Grok 4.7 review. It's 30 to 80% less token-efficient than promised "with very few forgiving qualities."
- So I was using Fable wrong... (Theo, 98,000 views). How to prompt Fable 5.1 so you don't see performance drops. Watch it next to Anthropic's Opus 5.5 playbook, which gives similar advice.
- AI Model "Black Boxes" Are Dangerous (Theo). A short argument for embedded evaluators so labs can't turn their models into black boxes, following the Accenture announcement.
Other Interesting Stuff
GPT-6 Astra breaks a 1941 Enigma message. Frode Weierud's Crypto Cellar validated a break of MVUEH, a German Army message from July 10, 1941 that had resisted every attempt since 2005 (HN, 635 points). Carter Leffer only pointed GPT-6 Astra at the list of unbroken messages. The model picked MVUEH as the most promising one, guessed its plaintext was related to an already-broken message from the same day, settled on "ROSENOW ROSENOW" as a crib, and wrote its own Enigma simulator and Bombe in Python and C++. The key uses a completely different wheel order from the day's other keys. The transcription had errors, and a rare left-wheel turnover at letter 72 probably explains why humans missed it. HN split between "it required no skill, no imagination" and "so like half of all useful human inventions are to you not interesting?"
Epoch on falling costs. Epoch AI estimates that at a fixed level of performance, AI cost has fallen ~47% per quarter since 2023. That's 4x faster than DNA sequencing and 18x faster than lithium batteries. o3 hit 25% on FrontierMath for $0.55 in January 2025, and GPT-5.6 Luna matched it this summer for $0.0015. Costs fall fastest right at the state of the art, 66% per quarter, then slow to 32% after two years. Yesterday's launches fit the curve.
Dettmers' compression. Tim Dettmers released a runtime dynamic compression framework reaching 1.5 to 2.0 bits at high quality, part of bitsandbytes2's private beta. It varies quantization and MoE expert removal per layer, finds allocations by pushing the model towards the edge of collapse and measuring sensitivity, and swaps removed experts back in from CPU RAM or NVMe four at a time every 256 tokens. It opens his open-source week, with "virtually infinite KV cache" promised next.
Open models before Congress. Nathan Lambert published his briefing to members of Congress on US-China open-weight competition (HN). Chinese labs have led open weights since April 2025, and GLM-5.2 and Kimi K3 marked a step change in commercial viability. The top HN reply notes Lambert has since left AI2 along with some of the leadership behind its open models.
AI in the Iran school strike. Bloomberg reports that an unreleased Pentagon probe found overreliance on AI contributed to the missile strike on a school in Iran (HN, 580 points). The HN thread is mostly about where responsibility sits when the operator who fires never sees the target.
Blender worlds from one prompt. Alex Albert built Market Street in San Francisco as it stood on April 17, 1906, from a single Claude Tag prompt that told Opus 5.5 to compile Sanborn fire-insurance maps, the Miles Brothers film and archival photos into a source file before modelling anything. Karpathy replied: "Imagine taking arbitrary images/videos... and use them to build full-fledged custom GTAs, then drop into the world, move around, etc." One reply pointed out that every NPC would walk at hand-cranked speed.
Sources: 13 of 14 account feeds via nitter.gravitywell.xyz, x.n0g.xyz and nitter.jaydenha.uk. @potetotes returned 404 everywhere and Lee Robinson had no posts in the window. Thread pages came from nitter.jaydenha.uk (38 fetched). Articles from anthropic.com, simonwillison.net, latent.space, the-decoder.com and linked blogs. Discussion from Hacker News via Algolia. openai.com is JS-walled, so OpenAI details come from its launch thread, The Decoder and HN.