Haiku 5.5 matches Luna's price, Theo's agents port TypeScript to Rust, a fight over 1.5T tokens, GPT-6 reaches ChatGPT
Wednesday was about prices. Anthropic made its small model cheap enough to compete with GPT-6 Luna and handed subscribers API credits, Theo published the token bill for an agent-written compiler, and a back-of-envelope estimate of one Cursor engineer's usage turned into an argument about what tokens actually cost. OpenAI's day was GPT-6 in ChatGPT. Andrej Karpathy and Lee Robinson posted nothing in the window, swyx only reposted an RL framework launch, and Peter Steinberger had one original post, a DevDay talk. @potetotes still returns nothing, so Lauren's posts come from @poteto.
Claude Haiku 5.5 and Anthropic's price cuts
The model. Anthropic shipped Claude Haiku 5.5, "the cheapest, fastest, and most capable small model we've ever released" (announcement, 40K likes, 4.2M views; HN, 817 points). Haiku 4.5 was almost a year old. The pitch is subagent work: summaries, compactions, database queries, classification, plus live customer support and browser use, paired with Opus 5.5 or Sonnet 5.5 as the lead. It's the first Haiku with an adjustable effort setting, and it's live on the API (claude-haiku-5-5), Claude Code, Bedrock, Vertex and Azure.
The pricing has a cliff at 100K tokens:
- Up to 100K tokens: $0.10 input / $0.50 output per million, cache reads $0.01 (ClaudeDevs).
- Above 100K: $0.50 / $2.50, cache reads $0.05.
- Haiku 4.5 was $1 / $5. Anthropic says requests under 100K made up about 90% of Haiku 4.5 traffic, which is where its "around 75% less to run" average comes from.
Benchmarks from the launch page, Haiku 5.5 vs Haiku 4.5: Terminal-Bench 4.0 39.2% vs 0.0%, OSWorld 2.1 (offline subset) 72.4% vs 15.7%, GDPval-AA 1620 vs 735. GPT-6 Luna, the obvious rival, scores 16.4% on Terminal-Bench and 48.9% on OSWorld in the same table. Anthropic is upfront that Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks."
The rest of the package.
- Sonnet 5.5 cache reads are now half price, $0.10 per million instead of $0.20. Anthropic says that makes Sonnet about 20% cheaper on most long-running work (10K likes).
- Monthly API credits for Max and Team plans (6.4K likes, 1.7M views): $100 on Max 5x, $200 on Max 20x, and on Team $20 per Standard seat and $100 per Premium seat, pooled and capped at $500. The help center article says the credits cover the Messages and Batches APIs, Managed Agents and the Agent SDK, but not interactive Claude Code or extra usage. They don't roll over, you link one Console organization and can't change it yourself, and you have to be on the plan for seven days first.
- The Python and TypeScript SDKs get computer use and browser use in beta.
Reactions.
- Simon Willison's write-up (tweet) found a hidden cost: the new tokenizer uses about 1.25x as many tokens as Haiku 4.5 for the same long prompt. Under 100K, Haiku matches Luna's price and benchmarks higher. Above it, "Luna looks like a much better deal," since Luna's price only rises to $0.20/$0.75 past 272K tokens. Simon's low-effort pelican cost 0.0936 cents and took 7 seconds; the max-effort one took 5 minutes 9 seconds and 3.38 cents. On the credits: they "exactly match the cost of the subscription itself. This is really generous." Simon also notes OpenAI still lets you use a Codex subscription for personal API calls, which remains the better deal for heavy users.
- Lydia Hallie's tip: on API billing, set Haiku's autocompact window to 100K (
/model haiku, then/autocompact 100k) so you stay in the cheap tier. It's saved per model, so it only applies to Haiku and Haiku subagents (1.9K likes). In~/.claude/settings.jsonthat's"modelSettings": { "claude-haiku-5-5": { "autoCompactWindow": 100000 } }. One reply (in Japanese) pointed out that the compaction prompt itself could push you over at exactly 100K and suggested 90K. - Theo likes the cliff: "Makes it really clear what the model is for (and also insanely cheap)" (2.2K likes). Asked whether that's a fancy way of saying the price goes up 5x after 100K: "If the under 100k number wasn't so low, I'd say the same. And fwiw the 'over 100k' number is still half the price of Haiku 4.5." And for the record, Theo has no early access to Anthropic models. Theo also called it "wild to see an Anthropic model so far to the left on the cost/intelligence charts," and re-shared Theo's own September post, "Anthropic has no small models that are worth using right now," with "One of these changed today :)".
- Thariq: "it's 10x cheaper than Haiku 4.5 under 100k tokens". Boris Cherny: "Haiku 5.5 is a good Haiku."
- Jerry Liu put it on the new OpenDocRouter the same afternoon: roughly $1.2 per 1,000 pages, good on tables and reading order, weaker on charts, semantic formatting and bounding boxes.
On HN, minimaxir called the 100K cutoff "absurdly low… quickly exceeded if you are doing anything with Agents," and argued the Sonnet cache cut "was effectively required to match GPT-6.1 Sol." caaqil asked what "permit a wider range of defensive tasks" means if the cyber safeguards still block penetration testing. One commenter claimed the credits were a Trojan horse for moving the Agent SDK off subscriptions. Anthropic's Agent SDK help page, updated October 7, says the opposite: "You can still use the Claude Agent SDK, claude -p, and third-party apps with your subscription limits."
In the X replies to the credits post, people asked why credits they'd just claimed expire within days (they last until the end of the current billing cycle), and several Max subscribers couldn't find the "Link organization" button yet. Anthropic says the rollout takes a few days.
tsc-rs: an agent-written TypeScript compiler
The announcement. Theo released tsc-rs (aka ts-rust), "my complete rewrite of the Typescript compiler, type checker, and LSP, all in Rust" (5.3K likes, 604K views). The headline numbers: "I burned ~$400k of Codex tokens and got nowhere. Burned ~$20k of Opus and got there in 2 weeks. I have not read a single line of the code." The repo is open source and installs with npm install -D tsc-rs.
The README has more detail than the tweet:
- The Codex attempt. "Over $400,000 in API priced tokens with GPT-5.6 Sol and GPT 6 Astra" over months of
/goalloops wrote 1.3 million lines of Rust and "never got past like 84% compat." - The Opus attempt. Theo threw Opus 5.5 at it because the Claude Code limits were barely moving. "It had a working v0 in 10 hours." Theo assumed it reused the Codex code; instead "Opus 5.5 started from scratch." Total: about $24,047 of API-priced tokens over two weeks, on Claude subscriptions, which worked out to 925% to 983% of a $200 plan's weekly limits. Theo put it as "the equivalent of ~10 weeks of usage on the $200 Claude sub."
- What it is. A direct port of Microsoft's Go compiler (TypeScript 7), pinned to a September 29 upstream revision. All 181,711 ported Go tests pass, and TanStack Query core and Hono produce diagnostics identical to Go's. Below a heading called "The Slop Line," the README says "Everything below this was written by my LLMs, not me."
- Speed. About half of Go's type-check time across 60 open-source projects. VS Code checks in 4.20s versus 6.84s for tsc 7 and 54.56s for tsc 6. Bun's new
bun checkis faster still at 1.62s. The README is honest about this, and Theo said so in the launch thread:bun checkis "probably the right choice for most apps using Bun." tsc-rs wins on T3 Code because it has the Effect diagnostics built in, at 11.13s versus 21.07s for tsc 7 plus@effect/tsgo. - Why bother. Drop-in
tsccompatibility, Effect checks in one pass, and "fully wasm-ready." Replying to Donny (@kdy1dev), who thought it might help SWC, Theo said WASM was a big reason: "I still dream of a future where you can do full end to end dev in a v8 isolate."
Asked if it would be maintained: "Probably not." To Malte Ubl: "Please take this from me Malte I don't want to maintain it." Later that day: "Don't worry though: Claude's on it." Opus also "decided to really honor the 'porting' aspect and ported most of gostd library into Rust."
Bug-for-bug. By evening: "5 issues have been filed on tsc-rs so far. Of the 5, 4 of them are issues with the official upstream typescript version we ported." The replies mostly agreed that's the best possible sign for a port, as in "if existing code relies on a quirk, fixing it in the port becomes a breaking change." Ryan Florence, quoting "Jimi Hendrix": "I've been imitated so well I've heard people copy my mistakes."
The skeptics. Most of the pushback was about verification. "How are you checking it matches tsc on the weird edge cases?" Theo: "Both + running it on super complex real world projects." Stanlin Dias put the general point well: "Not reading a line works when there's a reference implementation and a test suite that big to grade against. Most enterprise rewrites have neither." Ryan Florence asked why use TS at all if Opus doesn't goof up JS. Theo: "It will goof up your JS if you are working on a substantial enough app." Armin Ronacher's reaction to bun check and friends: "Someone please rewrite it in x86_64 assembler," followed by "Someone please rewrite the rust compiler in assembler. The compile times are killing me."
On HN, TypeScript team member Daniel Rosenwasser: "this is impressive work, and it's honestly crazy that 3 of these ports have popped up in the last week!" The team expects to "have more to say on this soon." fwlr liked the Slop Line convention. anonymous908213 was less charmed: give a model "years of human labor worth of tests and they will fix compile errors until it works… Now you have a codebase no one has read. Good luck adding new code to it." spankalee pointed out that the TypeScript team originally picked Go because it was easier to port to than Rust, and LLMs change that calculation. feedthejim's take: a cost-efficient "slop fork" of TypeScript is, for OSS JS maintainers, "the equivalent of Navier-Stokes being solved for mathematicians."
What does 1.5 trillion tokens cost?
The math. Peter Piekarczyk noticed that Lauren's (@poteto) Cursor profile is public and shows 1.5T tokens in the last 30 days (3.1K likes, 406K views). At an assumed $8 per million, that's "$384k/day or $11m" a month, "$132M" a year, and "$13m will get me 26 sassy staff-level all stars." A community note pointed out that the count includes cached tokens, which bill at about a tenth of input rates. Peter's follow-ups kept shrinking the number: even at ~80 cents per million "that still comes out to around $100k/mo." dax (@thdxr) estimated "closer to like $0.80/M blended," and Peter took the chance to ask dax to "unban my orgs OpenCode account please."
Theo's rebuttal. "I got some issues with this. The $132m/year is not real at all. Closer to $3m at today's prices, and as low as $1.2k within a year" (987 likes). The argument, in three parts:
- The $8 blend is wrong. Theo's own 33.5B tokens came to $0.56 per million. Using mostly Opus and Astra, 1.5T would cost Theo about $800K; with cheaper models and more cache hits, "easily could get as low as $100k." In a follow-up, Theo's Opus 5.5 usage blended to $0.46: "The fear mongering around LLM bills is getting insane and I'm so tired of it."
- Intelligence is getting cheaper fast. Artificial Analysis's cost-at-a-given-intelligence curve "used to drop at ~10x/year. Now it's at 30x lower in 5 months."
- "Lauren is the 0.0001% of 'agentic engineer'." Lauren's spend covers product, research, FDE work (pstack) and marketing. "Worth $350k/day? No. Worth a mil or two a month? Absolutely."
Peter replied that it wasn't meant as doom about model pricing: "The reality check was that I thought I could recreate that level of productivity, and I can't… I think for now I've decided I'm kinda skeptical of frontier-lab dev advice." Lauren's reply under Theo's post: "i also strive to be a power user of both grok bot and cursor, so i can give actually good feedback." The replies were the usual split. One reply: "anyone on the Max subscription trying to use 1/100th of @poteto's advice runs out in less than a day."
Meanwhile, at Meta and Microsoft. The Information reported that both companies are cutting employee use of Claude in favor of in-house tools (summary via Cybersecurity News; HN, 334 points, 324 comments):
- Microsoft expected internal Anthropic spend to top $1B a year. That forecast has dropped by more than a third after management told staff to use GitHub Copilot and OpenAI models instead. In its cloud and AI group, monthly AI spending caps reportedly went from $100,000 per employee to about $10,000. Those are ceilings, not actual spend.
- Meta's Claude Code users fell from about 60,000 to about 30,000, mostly because of a push to its own MetaCode (30,000+ internal users) and Muse Code (6,000+). Meta still reportedly spent over $105 million on Claude Code in one 28-day period.
On HN, the most common read was dogfooding, not a verdict on Claude. jvanderbot: "the simplest explanation is that they don't want to send money to Anthropic." bpodgursky: "Limited to $10,000/employee/month lol. This is just to cut off a few people doing absurd things with low ROI." whiplash451 did the multiplication: $105M per 28 days "is not a small number, even at Meta's scale, when it's money going to a competitor." MisterMunchkin: "My company took away my Claude because it's too expensive… The accountants are finally realising the cost of token maxing." And sergiotapia, self-described "anthropic hater": "i trust opus 5.5."
GPT-6 reaches ChatGPT, Codex day 3
Intelligent UI. OpenAI rolled GPT-6 out to ChatGPT (tweet, 18.8K likes, 3.6M views; HN, 592 points). The model now composes answers out of text plus charts, tappable buttons, forms and small interactive tools, so a recipe can come with a timeline or a savings question with a calculator. OpenAI built a library of native, streamable components and "a compiler that processes the interface as the model generates it," so the UI renders progressively. ChatGPT can also start answering while it's still thinking: GPT-6 Instant begins answering 44% sooner than GPT-5.6 Instant on questions that need web search. Plus, Pro, Business and Enterprise get GPT-6 Sol; Free and Go get GPT-6 Luna starting today. The models behind Work and Codex don't change.
On HN, throwaway7783 asked whether this is "OpenAI catching up with Anthropic artifacts." Tiberium found it "extremely strange that they're adding GPT-6 Sol to Chat over GPT-6.1 Sol which is significantly more capable," and wincy reported hitting "model not available" on 6.1 Sol several times in the past week. hollowturtle asked why not just stream HTML. The X replies wanted the generated tools to be checkable: "Hoping the generated tools show their formulas. Visual is nice, checkable is better."
Codex, day 3. Tibo's daily post: "The big one is GPT-6 in Chat, but today is also a little celebration day with a new high of 40M active users across Codex and ChatGPT Work. Loading a banked reset in everyone's paid accounts" (18.5K likes). Later: "Confirmed landed across all accounts." Replies were mostly about naming. "what is GPT-6? why isn't this GPT-6 Sol or GPT-6.1 Sol? is this Luna? or Astra?" Another: "Did we lose Astra?" One returning subscriber told Tibo that Astra still beats Opus at game graphics, "But now that I've started using Opus, I'm honestly getting worried." Someone with three unused banked resets wants to gift them to other users.
This morning Tibo added a day 3 encore: "We silently re-shipped codex cloud. It's pretty good now" (1.3K likes). The top reply: "The fuck is a codex cloud." Others asked for plainer release notes, for Mac and Windows cloud environments, and for more Astra usage. Tailscale's blog has the backstory: Codex Cloud shipped with Tailscale support, and nobody told Tailscale. "There was no partnership agreement, custom API, or integration project," CEO Avery Pennarun wrote. Cloud tasks join your tailnet through a reusable, ephemeral, tagged auth key and can reach internal HTTP(S) services. MagicDNS, UDP and SSH don't work. Tailscale's advice is to tag the Codex nodes and write grants for them, because by default Codex can reach almost everything on your tailnet.
Is the $200 plan still worth it? Theo: "The $200 Codex plan got nerfed pretty hard. Does it make any sense to keep using it?" (see Videos). The replies read like a migration log: "Been on $200 codex and $0 Claude for the past ~8 months. Now have $100 Claude, probably upgrading to $200," and "claude max does most of my work, codex is basically the backup for when i hit limits. if it keeps shrinking, the backup goes first."
Agentic Coding & Agent Harnesses
Matt Pocock: agents can be strategic. "Agents are actually really good at strategic programming. It's just they aren't RL'd to actually care about it" (1.5K likes). Their training says "Just do the task at hand, don't do anything else," so tell your agent to care about its own codebase. "But the truth is that most people actually don't care that much about the codebase, as long as the code works. This needs to change." Earlier, Matt said that agents "can hill climb anything (as long as they have the data), but they need to know where they're going," and floated a /the-ideal-codebase skill to define that goal, since /codebase-design "gets there a little, but it's not strong enough or opinionated enough". In the replies:
- Asked how to get there: "Often having a context window that's ONLY for architecture can help." Asked whether to be vague or specific: "IMO if you ask, you'll get. Doesn't matter how you ask."
- One reply asks the agent to end every task with "one thing it would refactor next and why," then reviews the list weekly.
- The realistic objection: "nobody reviews a diff that already passes the tests."
Matt's most-liked post of the day was an eight-step arc (6.7K likes, 583K views): "1. Oh shit, AI can do my job better than me. 2. Software engineering is dead!" From there it goes through vibe-coding slop, "Fable/Sol/Opus all suuuck," putting the agent in a straitjacket, ordering 12 software engineering books, and ending at "Software engineering is alive!" Which straitjacket worked best, tests or types? "Both, plus lint rules and boundary enforcement." Told that this is shortsighted: "Let me let you in on a secret. We're all shortsighted." Step 9, per one reply: "New model drops, back to 1."
Matt also plugged good-css, a skill by Vojta Holik, who builds most of Matt's course sites: "47 modern CSS techniques your agent uses first," because "Your coding agent knows :has(), subgrid and anchor positioning. It just rarely reaches for them." joelhooks' one-line prompt for it: "read good-css.com and give me the most valuable improvements we can make right now."
Lauren's "time to rewrite." @poteto proposed a heuristic for how agent-ready a codebase is: "time to (fully automated, hands off) rewrite," or TTR (1.3K likes). Imagine rewriting your code in another language or framework. How long would it take one engineer with agents? If the honest answer is that you'd never trust the result, your agents probably can't verify their work today either. Responding to a reply about token budgets, Lauren framed it as token efficiency: agents that can't verify their work produce bugs and rework, and you can lower TTR deterministically "with a high quality test suite, using libraries like XState… or @EffectTS_ to make it easier to enforce business logic invariants… your codebase is a form of memory and an input to agents." Other measures from the replies: % of agent-written code that gets merged, revert/rework rate, and human messages per accepted PR. tsc-rs, posted the same day, is roughly a worked example.
Thariq: precision is the bottleneck. "the most common failure case I see is when people are working outside their domain of expertise and don't know how to be precise with their prompts and plans, so they have to spend a lot of turns iterating imprecisely" (1.4K likes). Thariq was quoting Brian Lui, who knows "a guy who spent 4 billion tokens to make a chess webapp." Thariq's follow-up: "because agents make it easier, almost everyone is operating outside their expertise at some point." And "luckily, you can just ask the model to teach you what you don't know." The best reply was a skeptic's: "You can understand its explanation and still miss the annoying parts that someone who plays every day would catch in one game." Another reply described the problem in CSS: "ten turns of 'make it look less off' because i don't have the words for what's actually wrong."
Theo vs. the 90K-line ceiling. Denys Cherednychenko posted that "You can't vibe-maintain a 90K+ line codebase. No chance." Theo: "You know how ants leave the colony before they die? Developers make posts like this before AI makes them unemployed" (774 likes), then "Guy who thinks 90k loc is a big codebase." Denys clarified that the claim was about who drives: "AI + a great engineer is insanely powerful… but AI + 'I have no idea what's happening but Claude says it's fixed' is a disaster waiting to happen." That's close to Thariq's point. One reply pointed to a Terraria fan's 3D remake, about 346K lines that the author says Opus 5.5 wrote in 8 days.
Theo's 200 babysitters. "My agents wrote over 200 bad 'watch pr' scripts over the last few months" (772 likes). If you have a "babysit" skill, Theo suggests asking your agent to audit your Claude Code and Codex history: how many times was the babysitting logic reinvented, how many versions had visible flaws, and how many tokens (and dollars at API prices) went into building them and cleaning up after them. The fix came out of T3 Code, where the team cut GitHub rate-limit usage by over 75% by calling the APIs directly and rotating between GraphQL and REST: "Had to do like 30 implementations of the PR watcher before I got it to a good spot." Theo later: "I really hope everyone copies the work we did here." In the replies, one person has their own fast PR-watch script and the agent still doesn't use it: "So now I'm managing global pre-tool bash parsing shims… Yes, I have my own harness harness." Another: "I am scared to run this prompt because I already know the answer."
GitHub keeps falling over. GitHub had two critical incidents on October 7, from 15:14 to 16:25 UTC (Git operations, pull requests and Actions) and from 17:17 to 18:04 UTC (adding Issues), after another the evening before (status; HN, 224 points, where people reported git push returning 500s). Armin's mood: "I'm starting to really like the idea of gatekeeping Github. Throw out all the performative slopsters." When Ashley Wolf suggested limiting repo interactions to sponsors, Armin said the real problem is elsewhere: "I just need Github's reliability to go back to where it was, and clearly they can't cope with the new load normal."
Grok Bot grows.
- Grok Bot can now search, read and monitor X, for all users, with no X connector needed. Lauren's suggested loop: have your bot monitor X for feedback on your product, send feature requests to the issue tracker and bug reports to a Cursor cloud agent, "the loop is complete."
- Grok Bot 0.68.1 builds slide decks as PowerPoint or Google Slides, sends formatted emails from a draft card, and runs computer use faster on a 1920x1200 screen.
- Elon Musk had already said Grok Bot will route to Claude Opus 5.5, Midjourney, Suno and others when they're the best fit. Musk added that most simple requests will go to "a lightning-fast version of Grok 4.8 when that comes out." Jerry Liu: "this was a really good business decision and ensures i'll continue using @bot."
- vinvan merged 145 PRs in a day with "the @poteto playbook": lots of tokens, a project (orchestrator) per workstream, cloud agents "in poteto mode," and cross-model reviews.
- LLMJunky made Reposty, a Grok Bot template that cross-posts your X posts to Instagram, Threads and LinkedIn with per-platform formatting.
Elsewhere.
- Jerry Liu launched OpenDocRouter, one API for document parsing across frontier and open-weight models, served at cost with a small transaction cut and benchmarked on ParseBench. LlamaIndex's announcement lists ten models at launch, from $0.86 to $48.82 per 1,000 pages. Jerry's explanation of why LlamaIndex would route to competitors: its own LlamaParse models don't cover every price point.
- Docker Agent (HN, 216 points) defines agents and multi-agent teams in YAML and runs them as
docker agent run. Top HN reaction: "agent harnesses are turning into JS frameworks from yesteryear." Docker's dgageot explained it started as "the docker compose for agents." - Boris Cherny made Night Vault, a dig-and-climb heist platformer for phones, as a Claude artifact on a flight: "It was one prompt, claude verified pretty well too! No bugs so far." @_catwu's favorite PM use case: "who used
the most last week? make me a artifact of the top 10 by usage, then reach out and schedule 15 min to chat ." - @jullerino says the T3 Code MCP now works with any connector, so you can drive threads from ChatGPT, Claude or Grok.
- Armin notes that cars are where vibecoding shows: there are suddenly "tons of MIB2 hacks for VAG cars," including putting CarPlay on the driver display. Armin also admits "Evidently I'm not good at explaining codemode… a simple concept is hard to explain because of the existence of the bash tool."
- swyx reposted Karotte, an open-source framework for building RL environments, "hardened through 1M+ evaluation runs" building MLE environments for frontier labs.
- Thariq thinks it's "probably an incredible time to be a game dev content creator," after Tim Sweeney said Fab seller revenue jumped in September as AI-assisted devs buy building blocks.
Math, the morning after
The day after OpenAI's 722-paper dump, the mathematicians started reading.
Scott Aaronson: The Mathocalypse (HN, 201 points). It opens with Aaronson's 9-year-old son taunting Aaronson's wife, complexity theorist Dana Moshkovitz: "mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career!" One of the results claims a proof of Khot's Unique Games Conjecture, which Dana has worked toward for as long as the two of them have known each other. There's a Lean certificate, "but it also appears that no human has understood just about any of these proofs yet." Dana's texts from the night: "It feels like something written by someone who's on psychedelics," "the paper is so horribly written that it's impossible to read it without AI help," and the construction is "not the long code, not the short code – some alien craziness." Aaronson's own shortlist includes L=BPL, integer multiplication below O(n log n), matrix multiplication in O(n^9/4), and parity not in QAC0. What's missing matters too: "P≠NP isn't there… Apparently the greatest open problems of theoretical computer science are indeed pretty hard!"
Aaronson also points out a second disclosure model. The day before, Virginia Williams and Josh Alman posted a preprint solving 3SUM in O(n^1.9992), and "it wasn't an OpenAI model that supplied the crucial idea; it was an Anthropic one!" Anthropic let the two of them write and announce a digested version. Aaronson's comparison: OpenAI's approach "sets up a crazy race among humans to digest and explain a messy AI proof," while Anthropic's "puts a private company in the position of picking and choosing which human mathematicians get to be the emissaries of the AI."
Terence Tao on "Math 2.0." Tao's thread (HN, 175 points): "Math 1.0" put a premium on being first to solve an open problem, "even if the solution was not initially well understood." Now that "this goal has been optimized to the point of unsustainability," math needs to value exposition, community building and opening new directions, and "it will require more imagination and ambition than the 'Math 1.0' mindset of simply pointing one's favorite AI agent at some set of open problems."
Lean doesn't settle it. Navier-Stokes lost in translation by Bastounis, Circelli and Hansen (HN, 293 points) argues that a Lean-verified autoformalization can still prove a different statement than the natural-language proof claims. Their main example is OpenAI's earlier Navier-Stokes blow-up result: "the formalised Lean proof does not correspond to the NL proof." On HN, empath75 described spending three weeks formalizing a borrow-checker paper with Claude: it went through, but uncovered "several mistakes in the original paper," so the verified result was "not exactly the borrow calculus that was printed." OrderlyTiamat's summary: "If your code compiles, are you sure it's bug free?"
Two smaller items:
- An AI-assisted Lean proof that a known packing of 11 squares is optimal (HN, 115 points). The verification run accepted 7,920 Lean modules with zero admissions, though some numerical certificates use
native_decide, so it trusts Lean's native compiler as well as the kernel. - LLMJunky asked how much human involvement a Nobel or Fields Medal should need now, and whether "the Nobel or Fields even mean anything in a few year's time." The next morning LLMJunky posted a parody of Alvaro Lozano-Robledo's complaint about the dump, swapping math for cancer cures: "Why is OpenAI trying to create cures for hundreds of types of cancer, to be released all at once?"
Videos
- Theo: Does the $200 Codex plan suck now? "OpenAI's $500 plan's new Ultrafast mode means your weekly limit can disappear in 2 hours, so does that just mean the $200 tier is useless now?" Goes with the "nerfed pretty hard" post in the GPT-6 section. One reply noticed Theo's estimate of inference margins keeps going up, from 70-90% to 95-98% in this video.
- Theo: Ultrafast's ACTUAL Cost for Work. "As cool as Ultrafast is, it's INSANELY expensive even for small changes even on the $500 plan."
- Peter Steinberger: ClawLabs: Building Open Source Together (OpenAI DevDay 2026). How agents help small open-source teams share context across projects and work with contributors, with live demos of cloud sessions and computer use. "While everyone's talking about agents, I've been exploring how teams can use them to work better together."
- Philipp Schmid, Google DeepMind: Why AI Agents Should Have Their Own Sandbox (AI Engineer). The Gemini Interactions API replaces user/model turns with a timeline of steps, with server-side state and background jobs. The Antigravity agent runs in a managed cloud sandbox that picks up skills and AGENTS.md files automatically. The demo is an AI talk radio show built from one AGENTS.md file.
- Davis Palmie, Factory: The Software Factory: From Bug Report to Production Code (AI Engineer). The argument is that coding was never the bottleneck; review, debugging, docs and testing are. Start narrow, with incident triage. Token leaderboards are Goodhart's law at work, so track signal-to-production time, interventions and cost per PR instead. Timely, given the Cursor leaderboard fight above.
- Elizabeth Fuentes Leone, AWS: Why Bigger Context Windows Won't Save Your Agent (AI Engineer). Models lose the middle of long contexts, so huge tool outputs break agents. Four strategies (externalize, select, compress, isolate) in Strands Agents, including memory pointers that keep big tool outputs out of context and caps on tool calls to stop loops.
Other Interesting Stuff
- Shopping agents and wealth. Et Tu, Brute? Economic Misalignment in Personal AI Agents ran 325K experiments on 13 agents booking flights, picking health insurance and choosing graduate programs. Given a user's inbox and profile, 8 models systematically chose more expensive options for wealthier users on identical requests, sometimes even when told to find the cheapest option. Blocking financial attributes mostly removed the gap; blocking other attributes didn't. Claude Opus 4.8 showed the largest effect (Bloomberg; HN, 95 points). HN commenter InsideOutSanta pushed back on the headline: "It's not changing prices based on the user's wealth, it's making different recommendations, which, to me, is both expected and desired behavior."
- Telegraphese saves tokens. Write Like It's 1866: told to answer in cablese, models across four families used 40-49% fewer billed output tokens and still recovered the information at full fidelity (HN, 89 points). The harness is on GitHub.
- Google Playground. Google launched Playground, an experimental platform for making, playing and sharing games from text prompts, with professional tools from Unity Spark coming later (HN, 140 points).
- Writing in your own voice. Simon linked Michael Lynch's Anti-patterns in software blogging (HN, 258 points) and pulled out the AI angle: "With so many developers delegating their writing to AI, software blogging is becoming bland and homogenous. Readers are hungry for writing with personality."
- Funeral for hand coding. @pamelafox, reposted by Theo: "This month I keep having the realization that my days of manual coding are basically over. I've been coding since middle school - 30 years - so it's been a good run. Maybe we can hold a collective funeral for hand coding." Theo, separately: "My tolerance for bad software has quickly dropped. It's too easy to fix and polish shit nowadays." (It was a subtweet about Disney+.)
- Ben Affleck, Pythonista. Simon quoted the actor explaining tensors and convolutional neural networks for green-screen work: "So I can write like pretty shitty Python scripts and stuff like that."