Claude Fable 5 Unleashed: Self-Rewriting Agent Loops, 10x Benchmark Gains, and a Flood of New Coding Models
Claude Fable 5 dominated today's AI discourse with demonstrations ranging from self-rewriting agent loops to parametric turbofan CAD, while benchmark results showed a 10x improvement over other frontier models on post-training tasks. Three new coding models launched within hours of each other, and Cognition open-sourced Devin's most popular cloud handoff feature.
Daily Wrap-Up
Today was Fable 5 day. Anthropic's latest model release didn't just make appearances across the feed, it practically took it over. We saw it build a complete Airbus-class turbofan engine in parametric CAD, power interactive Three.js simulations that one user described as having reached a quality threshold where "it is now impossible to explore all possibilities," and drive Firecrawl's new Prometheus agent for automated web data collection. But the most striking moment came from @thoughtfullab's benchmark results shared by @scaling01: Fable 5 spent 17 hours and 25 million tokens training a weaker model to solve a puzzle, achieving 34% pass@1 where every other frontier model averages under 4%. That is not an incremental improvement. It is a qualitative leap that explains why Anthropic appears to have placed restrictions on Fable 5's ability to develop LLMs, as @scaling01 pointedly noted.
The philosophical underpinning tying all this together is what @swyx and Latent Space are calling "Loopcraft": the emerging discipline of stacking self-improving agent loops. The "Salty Lesson" for agents parallels Sutton's Bitter Lesson for models. Don't fix things yourself. Build systems that scale with more agents through better goals and orchestration. This maps directly to @0x_rody's share from Anthropic's event, where a Metaview engineer explained they stopped fixing prompts entirely because "the system reviews its own output and rewrites its own instructions now." Meanwhile the model marketplace continues its relentless compression. NVIDIA's Nemotron 3, MiniMax's M3, Kimi's K2.7-Code, and Qwopus 3.6 all launched or gained attention today, each pushing on inference speed, context length, or cost. The most practical takeaway for developers: stop hand-tuning prompts and start building self-correcting agent loops. The teams seeing outsized results are the ones letting models iterate on their own instructions, not the ones manually crafting perfect system messages.
Quick Hits
- @Krongggggg highlights a 25-day scaling fundamentals series by @system_monarch, a Principal Engineer at Atlassian, covering load balancing, system design, and architecture patterns from someone who has designed systems handling millions of requests.
- @BiaNeuroscience pitches Bía Sleep, a neurofeedback wearable built on 40 years of neuroscience that uses real-time brain signals to optimize deep sleep, currently
Sources
https://t.co/LWQbDMXIdK
I've been a backend Engineer for 12+ years. Today, I'm a Principal Engineer at Atlassian. I've designed systems that handle millions of requests. Sat on both sides of system design interviews. Reviewed more architecture docs than I can count. Starting today, I'm breaking down the fundamentals of scaling for the next 25 days. If you're learning system design bookmark this thread, you're going to get a lot of learning from this.
KV Caching in LLMs, Clearly Explained
Introducing Automated Security Review in Droid. https://t.co/47ZhOUWdCt
We're launching Bridge today 🌉 An AI engine that builds virtual homes. Blueprint in, walkable home out. Every plan, every option, structural changes included. What took 3D artists months now takes days. Homebuilders can finally show buyers every home they sell. https://t.co/QrwlM5FUjP
We've open sourced my favorite Devin feature: /handoff Hand off jobs to cloud Devins from your local machine Install it as a plugin in Claude Code or Codex or any other coding agent Close your laptop without pausing your agents 😉 https://t.co/1HMgfzuar6
Nemotron 3 Full Breakdown With the help of Joey Conway from @NVIDIAAI getting into the specifics around why Nemotron 3 is kind of a big deal Biggest headline with Nemotron is: Hybrid Mamba Transformer, Latent MoE, and MTP Hybrid Mamba Transformer essentially attacks right at the Attention mechanism to make the overhead sub-quadratic, but unlike quantizing KV Cache or swapping out attention head, NVIDIA chose Mamba-2 Latent MoE helps further optimize on sparsity by down projecting the dimensions so you're doing less math and less memory movement between HBM and SRAM, you're saving a ton, and NVIDIA made a conscious choice to add more experts given the surplus Finally, MTP or multi token prediction where the model can see future tokens to be more expressive in training and also option to use for speculative decoding during inference Oh, also the model adopts the new OpenMDW 1.1 License
We are super excited to share with you our initial release of Lucky Engine. We are building a robotics engine from the ground up to be what we wished we could find in a simulator before https://t.co/fR10g5iRXg
Fable 5 is doing something wild on our FrogsGame post-training task. It trains a weaker model to solve the puzzle, peaks at 68%, and produces the only ~10x improvement we see across the benchmark. It spent 17 hours, 25M tokens without human in sight. 34% pass@1, while every other frontier model averages under 4%. We will publish a more detailed analysis soon.
Qwopus 3.6 27b-Coder is now live! Scores a 67% on a full run of SWE bench verified with thinking completely disabled! Q5_K_M This model is lightning fast for dense class! With a natively finetuned MTP head, it achieves 100 tps on a single 5090! The biggest upgrade here, though, is its stability in programming and tool calling within @NousResearch Hermes agent, with thinking off! Wall time is crazy fast this way, which makes Hermes feel "native" and snappy, like they were meant for each other. The freedom of running without thinking at all makes you part of the thinking process, and you never get caught waiting 15 minutes for it to finish a thought string, like with the base models. Thinking on and temp high, .9-1 seems to produce really incredible design and svg results. I reran the Boat survival prompt through a few turns, thinking on, and it seemed to render more fancy models in HTML canvas, but it was much more of a start-a-prompt and wait experience vs the snappy and active iteration with it disabled. It may be worth turning it off and on throughout the build process if you want to get really creative with design. Really looking forward to seeing how this one performs for y'all! Please post comments with your opinions and use cases below! As always with our fine-tunes, mess with the temperature setting, and run them much hotter than the base! Please check out the Boat Survival game I posted yesterday, made in 12 turns using Hermes and this model, with thinking off. Link below! Full swe bench repo-specific breakdown also posted in the comments for those interested! Happy building, everyone! We're looking forward to your thoughts! Quants uploading now! https://t.co/kxJE3C39ZZ
GEPA for skills is here! Introducing gskill, an automated pipeline to learn agent skills with @gepa_ai. With learned skills, we boost Claude Code’s repository task resolution rate to near-perfect levels, while making it 47% faster. Here's how we did it: https://t.co/VsWMyZncC9
[AINews] Loopcraft: The Art of Stacking Loops @RichardSSutton has his “Bitter Lesson” for models. We now have the Salty Lesson for agents: Don’t fix things yourself, as you have done historically. Instead focus on systems that scale with more agents, like goals and orchestration. More in today's op-ed: https://t.co/EyYV4VGRpi
Anthropic's War on Opensource AI
use kimi k2.6 for FREE through nvidia's api AND their new desktop app with 300 agents 😳 kimi beats GPT-5 on coding benchmarks what you will get for $0: - kimi k2.6: SWE-Bench Pro 58.6, Multilingual 76.7 - free for 1 year on NVIDIA's API - 4,000+ tool calls per session - 300 parallel agents what the new desktop app does: - runs native agent swarms on mac/windows - builds full PPTX, Word, Excel, PDF files - agents browse the web, click, type, fill forms for you - built-in finance data (yahoo, binance, world bank) - remembers your workflow + turns repeated tasks into skills how to get both for free (5 min): > go to https://t.co/BIIXYD9PJA and sign up (web + desktop) > go to https://t.co/V6J008DsGg, search kimi-k2.6 > generate a free api key > paste into any openai-compatible client (cursor, cline, etc) > base url: https://t.co/snZSjFRN41 important: - nvidia tier has ~40 req/min rate limit - desktop free tier has some agent limits - not for production workloads - phone verify may be needed for nvidia sota coding model + full desktop agent suite = $0 while most people are paying $20/mo for each bookmark this before the free tiers change
experimentally eerie made with fable and @threejs download in comment as well as fable playground link it feels like model capability with fable have hit a point where it is now impossible to explore all possibilities. before the quality just wan't there and frustration was often the end result the realm of possibilities and quality will grow exponentially from here on out is my humble 2 cents