Grok 4.6 Joins the Frontier, Codex Asks "Why Did You Switch?" & Sessions That DM Each Other

Agentic Coding & Agent Harnesses

Your Claude Code sessions have names now, and they can DM each other

The most-shared agentic tip of the day, from Ado (540 likes, 46.5k views, RT'd by Boris Cherny):

claude --name backend
claude --name frontend
> tell frontend the orders endpoint moved to /v2

"it's very effective." Claude-to-Claude only for now — no, it won't talk to Codex or Grok — and not on Windows yet.

The replies are where the actual operating advice is. The best one, from Meet Anghan: "naming by ownership not by layer helped a lot here. --name orders beats --name backend, since backend eventually means everything and both start editing the same files. also tell each one what it doesn't own, or frontend will helpfully fix the api on the way past."

And the failure mode you should expect, from BullBear: "I let my backend and frontend sessions chat and they spent 2 hours arguing about snake_case vs camelCase before I killed the process." Ado's reply — "Were they making valid arguments?" — is the correct question. One person in auto-mode reports the two sessions started scope-creeping each other: "front asking more features to achieve its goal 🫣".

DHH: full YOLO or nothing — and Thariq pushes back with automode

DHH landed a PR on omarchy that launches the default coding agent with permission prompts bypassed, and framed it (824 likes, 63.7k views): "I get why harness providers don't want the liability, but there's only one way to live with agents and that's full YOLO/BYPASS PERMISSIONS. Anything else is treating AGI like a toddler."

Anthropic's Thariq replied with the counter-position that automode isn't a weaker YOLO, it's a better one: "try automode in Claude Code! it's actually a lot better, as an example it makes the model better at following your hard instructions in the prompt — e.g. if I tell a model to not push, automode will make sure it never does. I've never had it make the model worse."

The thread splits along a clean line — where is the machine?

steipete: the interface era is over before you finished picking a terminal

Peter Steinberger, quote-tweeting Nate Berkopec's "a lot of people have their identity as a developer tied up in having 6 terminal windows open… the 'we're gonna chat to this thing in Slack and Linear' people are directionally correct":

"cli was a year ago. apps maybe 6 months. now it's services, web, cloud sessions."

Evidence for the thesis in Theo's corner of the world: T3 Connect (npx t3 connect, an open-source tunnel layer that gives you remote control of Claude Code / Codex / OpenCode / Grok Build on any internet-connected box, free) is getting passed around as "the best remote control for coding agents I have used — 2 person team is outperforming big labs!"

Theo on the new MCP: "I was never a big fan. That might change now."

925 likes, 88.4k views for a video teardown with a cliffhanger: "The new version is genuinely really good, but it comes with a big catch..." The replies are mostly people asking what the catch is and MCP loyalists noting they've been fine this whole time. Worth watching if MCP is on your roadmap — the delta between "never liked it" and "might change now" is the interesting part.

Also from Theo, a data point for anyone doing mobile work: "gpt-5.6-sol is so good at building iOS apps it's crazy. Give it a dedicated Mac with computer use and let it run the simulator. Quality you get out is nuts."

Codex & OpenAI

"Why did you switch to Codex? Don't say reset." — 9,163 replies later

Tibo Sottiaux ran the most useful public feedback thread of the week9,163 replies, 1.15M views — with the pre-emptive follow-up "Also don't say Linux, we just shipped that". People said reset anyway. Repeatedly.

What actually came back, sorted by how many times it appeared:

1. Desktop app performance, by a mile. Top reply (961 likes): "The codex windows app is laggy, it consumes a lot of RAM (like a LOT)". Second (618 likes): "PERFORMANCE, PERFORMANCE OF THE DESKTOP APP… PUT ASTRA IN A LOOP AND TELL IT 'OPTIMIZE THE CODEX DESKTOP APP' AND LET IT RUN FOR 100H OR SOMETHING I BEG YOU". 0xSero makes the sharpest version of the argument: "Should be a culture of performance optimisation given this is a verifiable hill climbing task… If Discord were this slow nobody would depend on it."

2. Why people left Claude Code. Usage limits ("Fable 5 runs out way too quickly"), verbosity, and trust. Jason Kam: "codex can deliver what I want in faster time and doesn't run out as quickly even on Sol Max — I think Fable 5 still better at synthesis & summary of complex subjects." amul.exe gets at the real one: "Because you don't nerf models without telling — this means my workflow is reliable."

3. Computer use / browser use is the killer feature. Named independently by several repliers, plus andy's detailed list: better vision and spatial understanding, "much less righteous in reasoning and response," "stays on task significantly better, less sprawl, less 'i know you asked for X but i did Y, i panicked'."

4. The org-shaped asks. UI annotations for editing (Shaw: "It's the only reason I don't use my own harness to code"), a unified alerts panel across sessions, and a summary of what an autonomous run actually did — see Kol Tregaskes' 15-item list, which reads like a product roadmap. Plus a governance complaint from Perry Metzger worth surfacing: "There are lots of fixes I would like to be able to contribute, but which I cannot because you do not allow community contributions even though the code is open source."

Best non-answer, from LLMJunky: "i didn't switch to codex / i was born in it / from the days of old, i embraced its rough edges / no subagents. no plan. no hooks / just me. a command line. gpt o3."

15M users, and the reset everyone was told not to mention

Tibo had promised a reset per additional 1M active users up to 10M, then went quiet. Yesterday he confirmed they'd crossed 15M: "Enjoy a nice reset everyone. Landing in the next hour or so, go /fast." Theo, who watched the last one land on a Monday, noted the obvious: "Last reset might be hitting OpenAI's servers a bit too hard..."

Rounding out the Codex week (both from the last ~36h, so partial overlap with yesterday's dispatch): ChatGPT + Codex desktop on Linux — "you can cancel that MacBook order if you got impatient" — and "Import your world", which syncs projects, chats, skills and plugins in from other agents with an import history and opt-in auto-updates. Theo's take: "Finally! Codex getting closer to feature parity with T3 Code 🫡"

Models & Benchmarks

Grok 4.6: frontier intelligence, unchanged price, double usage for a week

SpaceXAI shipped Grok 4.6 — "a significant improvement over Grok 4.5 at the same price" — and Lee Robinson followed with the pitch ("smart, fast, and cheaper than comparable models! Plus there's double usage for the first week 🚀") and the model card, which covers coding, engineering and knowledge-work evals plus pre-deployment safety testing and the safeguard stack.

Artificial Analysis' independent numbers, which Lee posted, are the part worth keeping:

  • 61 on the AA Intelligence Index — level with GPT-5.6 Sol (max), behind Claude Opus 5 (63) and Claude Fable 5 (62), just ahead of Kimi K3. That's +5 over Grok 4.5 in barely a month, and +23 over Grok 4.3.
  • Agentic is the standout: GDPval-AA v2 Elo of 1753, behind only Opus 5, with overlapping confidence intervals with Fable 5 and Qwen3.8 Max. 50.7% on 𝜏³-Banking (second only to Qwen3.8 Max at 51.3%) and 88.4% on Terminal-Bench v2.1.
  • Pricing unchanged at $2/$6.

Missing piece, flagged in the replies: no math section in the card. Lee's answer: "Good note, we will include a section on math in the next release." Bedrock availability is "working on this".

The efficiency framing from ZenMagnets, RT'd by LLMJunky: "At the frontier, Grok 4.6 is the real standout here in terms of intelligence density. Sonnet Size, Sol level smarts."

Theo went live on Grok 4.6 + DeepSeek 4.1 if you want a hands-on read rather than a benchmark table.

Nice bit of OSINT from kunchenguid (RT'd by steipete), on this week's Grok Bot beta — "AI teammates that sign in to your tools, use them just like you do, and come back with finished work":

"i'm 80% sure Grok Bot was originally built by the Cursor product team, and got rebranded after the acquisition. 1. the iOS app is published by Anysphere, not X Corp. 2. the mac app download URL is hosted under cursor dot com 3. it seems to run on cursor's vm infrastructure 4. it's cursor team members who are actively responding on X about the topic."

Verdict: "in either case, this is a solid release - good work!"

Daybreak Blue writes its own firmware update to dump encryption keys

Filed under "capabilities are moving faster than the discourse": LLMJunky relaying a hardware-security result with OpenAI's cyber model — "GPT Daybreak Blue cracked encrypted hardware by writing its own firmware update, and forcing the device to install it, dumping the encryption keys in the process. No refusals. 😳" The original report: "found a way to force the device to accept it which allowed it to dump out the encryption keys and other info off the device. INCREDIBLE." Best reply: "Glad it's on the blue team......"

Also on the wire: Qwen 3.8 27B is reportedly delayed to Friday — the one LLMJunky expects to "absolutely DOMINATE consumer devices."

Reasoning Traces & Model Boundaries

Yesterday's stolen-reasoning-traces paper has an immediate, concrete sequel: the labs are closing the token-level doors.

Armin Ronacher went back to test a hack he built earlier this year — converting Kimi K3 sessions from Pi to explicitly annotate them in harmony tokens, tricking the model into emitting tool calls inside user messages. "Turns out OpenAI now blocks those requests." The block is in the conversation history, not the system prompt, and 52b4a076 pinned the exact trigger:

"The culprit in question is the string <|channel|>analysis being included anywhere, in any form. Error: Codex error: Request blocked."

The best architectural comment on why this is a band-aid, from Pierre-Henry: "Blocking tool-call tokens in user messages closes one serialization path. The durable boundary is at dispatch: accept a call only when the server created the assistant turn and issued the invocation ID. Token syntax should never be authority."

Which sets up Armin's open question from the same day, and it's the one to watch: "will closed SOTA model labs ban assistant prefill with newer models?" — clarified as "putting entire assistant messages into the transcript that never came from (that) LLM". If prefill goes, a lot of harness tricks go with it. The dread scenario, from Roy: "My biggest internal worry is that they start creating 'stacked' models that are internally multiple models but you get no access to any part of the stack. If they start doing that shit, open source is our only hope."

Kill My SaaS — The Submissions Land

swyx's $10k contest to clone the $40k/year enterprise SaaS his team was about to buy (it turned out to be conference CFP software) hit its deadline, and the submissions are not toys — they're deployed, open-source, domain-having products.

Gene Kim's CurtainCall CFP is the story of the batch. He's run ~24 conferences over 12 years and burned through 5+ CFP tools: "we've had to build so many workarounds — Basecamp, Trello, Google Sheets, Zapier — and I've written entire new apps to be the reviewer frontend." He entered on Saturday morning and shipped in under 24 hours:

"This has been the craziest dev experience of my career — and when Swyx released his eval harness, the entire project became a hill-climbing exercise."

It's in production, running the CFP for the Enterprise AI Summit (Charlotte, Oct 7–8). swyx's review: "positive review!"

The rest of the arena:

  • ProgramLoom by Maddie Dreese — open source, CFP through published schedule, fully free.
  • unsession.dev — full sandbox environment, Cloudflare-hosted and easy to self-host, "full MCP w/ DCR and API docs."
  • stagestack.dev — built over a Brazilian Father's Day weekend with real dev/staging/prod environments and working email: "I didn't want it to look like a POC or throwaway code."
  • open-speaker-operations by Jai Bhagat, with an interesting design premise: separate the "system of record" from the "system of actions" for a SaaS meant to be used by humans and agents alike.

And then there's the receipts post that will get quoted for months, from ky__zo, who just pointed agents at it:

"▲ it maxed out Fable 5 + used most of Codex limits ▲ most of the work was done within a 45h session running on Claude Code ▲ other sessions lasted for 17, 14 and 10 hours ▲ there were 183 sub agents used ▲ without subs, tokens would cost $4246. Codebase scores 100% on all the benchmarks provided by @swyx… best part, I did spend maybe 1h total over the last 4 days steering the agents."

His own conclusion is the honest one, and the reason "software is dead" is too glib: "on one hand, this was almost too easy… BUT i really doubt most of the companies would be willing to now take the responsibility for deploying and maintaining a software like this."

Codebases, Context & Evals

Matt Pocock: if you know the shared language, you know the codebase

1,652 likes, 66.8k views for a thesis that lands squarely on how you should be writing for agents:

"Getting more and more convinced that if you understand the 'shared language' of your codebase (i.e. the terms used, the names for things, relationships between them) AND those terms are used consistently in the codebase, then you understand the codebase. Whether you read the actual code or not. DDD has never been more powerful."

The follow-ups are the practical part. Should you rename core entities in code and DB when the business renames them? Yes, definitely — "and even worth doing before AI". Isn't consistency the hard part in legacy code? It used to be harder: "renaming refactors are pretty simple with agents". And a warning from a user whose context.md had rotted into unreadable jargon — Matt's answer: "that sounds very unhealthy. You need to be an active participant when you're writing to context.md." He's also asking for feedback on the mattpocock/skills docs ("feels a little low on conceptual explainers to me") at aihero.dev/skills.

ExtractBench: frontier VLMs collapse below 35% recall past 50 pages

LlamaIndex followed Tuesday's ExtractBench launch with a 36-page arXiv whitepaper (arXiv:2607.29677, extractbench.ai). The setup: 370 enterprise documents, 4,869 pages, 67 document types, 8 domains, 14 extraction systems, zero LLM judges — 100% deterministic and reproducible, scoring value accuracy, long-record completeness, spatial grounding and per-page cost.

The finding that matters if you're shipping doc pipelines: "Short documents mask critical system flaws. On files past 50 pages, commercial VLMs collapse below 35% recall due to silent list truncation. They hold high precision, but lose output attention and drop most of the table rows." Or as the LlamaIndex account put it: "The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong."

Their own new Agentic Plus tier debuts at #1 (95.6% value accuracy at under a third the cost of the closest peer) — vendor benchmark caveats apply, but the harness and dataset are on GitHub and HuggingFace, so you can check.

AI Engineer World's Fair: the Memory & Continual Learning track is up

The full track went live with a thesis worth stealing: "we scaled intelligence and got the world's smartest novice." Talks include Beyond Static Intelligence (UC Berkeley), Memory Harnesses for Long-Running Research Agents (Sakana.ai), Scaling Compute on Context (Engram), gradient-free continual learning (Adaption Labs), and Improving Agents is a Data Mining Problem (LangChain). Two standouts pushed by swyx:

Other Bits


Sources: RSS + thread scans of @mattpocockuk, @theo, @trq212, @LLMJunky, @mitsuhiko, @bcherny, @steipete, @swyx, @simonw, @karpathy, @jerryjliu0, @potetotes, @leerob, @thsottiaux. @simonw and @karpathy had nothing new in the window; @potetotes' feed returned no items again.