Kimi K3 Launch Day, Cherny's Adoption Ladder & the React Avengers

Kimi K3 Launch Day

Moonshot AI released Kimi K3 and it swallowed the timeline whole. The headline specs, via Simon Willison's write-up (his post): 2.8 trillion parameters — the first "open 3T-class model," taking the size crown from DeepSeek's 1.6T v4 Pro — with weights promised open "by July 27, 2026." Self-reported benchmarks have it mostly beating Claude Opus 4.8 max and GPT-5.5 high while losing to Claude Fable 5 and GPT-5.6 Sol, and Artificial Analysis measures a long-horizon knowledge-work Elo of 1547, "+732 points from Kimi K2.6 and behind only Claude Fable 5." It also took the top spot on Arena.ai's Frontend Code arena, surpassing Fable 5. The catch: at $3/$15 per million tokens it's priced at Claude Sonnet level — the most expensive Chinese-lab model ever, triple the old K2.6 pricing. Jeff Wang's framing got the RT from Theo: "Chinese open source is no longer '6 months behind', but it's also no longer '10% of the cost' either."

The practitioner verdicts came in fast and warm:

Day-One Reality Check

Actually using K3 on launch day was another story. Theo asked the practical question — "What are the best harnesses to use with Kimi K3 right now?" (326K views) — and the top reply was "It's currently unusable lol" (Theo: "fk"). LLMJunky corroborates from both directions: on OpenRouter, "I'm getting rate limited every 2-3 tool calls. It's wild", and on Moonshot's own $19 Kimi Code plan a single browser-game prompt burned the entire 5-hour quota at ~85% done, eating 20% of the weekly"i didnt even use 50% of one context window before Kimi K3 limits were reached." His verdict on the output, for what it's worth: the game was "pretty mid," 4th or 5th best he's seen on that test, behind Opus 4.8.

The harness consensus from Theo's replies, such as it is: OpenCode and Moonshot's own Kimi Code come up most ("K3 performs much better than on Claude Code from my experience"), pi with extensions has its partisans, several people point out you can wire it into Claude Code the way Theo did with Sol, and one warning worth keeping: it's "quite buggy unless harness specifically handrolls some stuff for the model." The structural observation that outlasts launch day, from FarisZR: "it doesn't seem like any harness does subagent stuff better than claude code. Others either force pre-defined subagents, or lock sub-agents to the same model. its baffling."

Cherny's Adoption Ladder

A day after his automation manifesto, Boris Cherny published the org-level sequel: "one person is 10x'ing their output with Claude but the rest of the org hasn't caught up. Watching teams adopt AI, I keep seeing the same 4 steps" — mapped out in a claude.ai artifact titled Steps of AI Adoption (RT'd by Thariq). The thread's load-bearing points:

Skills Give You Superpowers

Matt Pocock compressed his whole skills philosophy into a two-liner (50K views): "Superpowers gives the agent superpowers. My skills give you superpowers." His elaboration: "I prefer to be in control, and lower the load on the agents' context", with the gracious caveat that "superpowers is an extremely useful skill set. It's just not for me." The replies read like a user-research goldmine: multiple people report fully uninstalling Superpowers after trying his skills ("I prefer the agency"), one notes burning far more tokens on Superpowers for the same work, and the best articulation of the split: with Superpowers "you totally lose the whole context of what is actually happening"; with wayfinder/grill-me "you become the driver of the whole workflow."

The workflow itself keeps evolving. His current loop (today's post): "1. '/grill-with-docs <issue description>' 2. 'Oh damn, this is way bigger than I expected' 3. '/wayfinder make a map of this' 4. Continue happily on." And /batch-grill-me beats the original /grill-me: 13 questions in 3 rounds instead of 13, only asking questions at the "frontier" (those not dependent on other decisions) while research runs as background agents. Related: he boosted Will's approach using /wayfinder as an orchestrator of custom skills plus multi-phase UI prototyping, and swyx pointed at xdg's session/tree-based interview-user skill (repo) as another take on the grilling genre.

ChatGPT Desktop Walkback

Tibo (thsottiaux) announced a partial retreat on the ChatGPT desktop app redesign (726K views, 1,455 replies): "we didn't get totally quite right on the first try." The changes: conversation history and projects are back in the sidebar with Chat/Work history syncing across web, mobile, and desktop (local tasks stay local); easy switching between Chat and Work modes; and a reassurance for the terminal faithful — "Nothing is changing for users on Codex mode. It's still the OG and best at what it does." Theo appreciated the naming cleanup that came with it: "'ChatGPT Work' -> 'ChatGPT', 'ChatGPT Codex' -> 'Codex'. It seems stupid (it is), but I am thankful to see this fixed."

The reply flood is a case study in redesign debt: power users want the one-click chat overlay back on the Codex page, projects from Codex and Chat now mingle confusingly in one list, Paul Hudson (twostraws) wants chat deletion back ("your download/memory footprint is… a lot 🫠"), Windows users report /goal sandboxing issues bad enough to send them back to the CLI, and Quinn Nelson (SnazzyLabs) delivered the harshest read: "Just put it back. You had a good thing. Now it's just crappy Claude Desktop which is already crappy." Meanwhile Tibo ran the day's silliest poll — "How do you pronounce Sol", resolved with "I also pronounce it Sol."

Videos

Quick Hits

Note: @potetotes' feed again returned no items (Nitter serves an empty channel for the account). @karpathy remains quiet — nothing since July 8.