Fable 5.1 Halves Agentic Coding Costs, and Its 270,000-Character System Prompt Leaks
Anthropic's Fable 5.1 and Mythos 5.1 launched with agentic coding at roughly half the cost of Fable 5, per @felixrieseberg, hours before @elder_plinius leaked a 270,000+ character system prompt detailing the model's identity, wellbeing guardrails, and 46 tool schemas. Alibaba answered with Qwen3.8-Max-0902 (2.4T parameters, 1M context, $2/$6 per million tokens), which @kimmonismus says now benchmarks near Fable 5.
Quick Hits
- Fable 5.1 and Mythos 5.1 are live, with agentic coding tasks solved "at roughly half the cost of Fable 5" and a more natural writing style, per @felixrieseberg's announcement. Within hours @elder_plinius published what he says is the Fable 5.1 system prompt: 270,000+ characters and 46 tool schemas, up from 30 in the July Opus 5 capture.
- Alibaba shipped Qwen3.8-Max-0902: 2.4T parameters, 1M-token context, $2 input and $6 output per million tokens, live on QwenCloud. @kimmonismus reports it now sits near Fable 5 on benchmarks and argues the versionless jump from 0902 signals how quickly Chinese labs are iterating.
- John Ternus's first day as Apple CEO, as framed by a wave of posts: his one-word "hello" on X drew @elonmusk's attention and @birdabo's approval, while @0xSweep's greentext claims he killed Vision Pro 2 and sued OpenAI by day's end. Treat the meme's specifics as unverified.
- World Labs unveiled Atlas, billed as "the world's first multimodal world model," generating image and video frames with pixel-perfect camera control and reconstructing them in 3D. @jasteinerman called it achievable "from just 3 cameras" (his claim), and @petergyang rated it beyond Fable 5.1.
- AI keeps surfacing in unglamorous places: @unclebobmartin booked a plumber appointment in 90 seconds with an AI receptionist that admitted what it was, and @nazzari says she wants to hire non-technical people whose GitHub profiles look like the one she posted.
A Model Launch and Its Instruction Manual, Same Day
The Fable 5.1 rollout came with an unusual companion. @felixrieseberg described the release as slightly better across the board, with the cost halving on agentic coding as the headline and writing that "sounds a lot more natural" than previous models. @SammyMacaluso's take was less reverent, joking that the whole thing was the "best Grok 4.6 advertisement I've seen this year."
Then came the leak. Per @elder_plinius's summary, the prompt has the model identifying as Claude Fable 5.1, Mythos-class and sharing weights with Mythos 5.1, with a knowledge cutoff at the end of June 2026. He flags three removed sections (
Open Weights Squeeze From Both Ends
The open-weight crowd had a good day, or at least a loud one. Quoting @TencentAI_News, @jun_song highlighted Hy4 preview, compressed from 1.5TB to 214GB via a quantization method called Sherry that packs weights down to 1.25 bits each, "seven times smaller, barely a dent" per Tencent's own claim. The release also supports stitching GPUs across machines so they act as one, with a GGUF build available. @jun_song's verdict: "Open weight AI labs are black magicians." The no-quality-loss framing is the lab's, not an independent measurement.
Meanwhile @kimmonismus quoted @Alibaba_Qwen's Qwen3.8-Max-0902 announcement: post-trained on Coding & Cowork, aimed at enterprise tasks, research, and long-horizon workflows, with cache hits at $0.17 to $0.25 per million tokens. His two observations are worth sitting with: significant capability jumps now arrive as point updates rather than new version numbers, and Chinese labs are closing the gap with US frontier labs "despite still having significantly fewer computers." Both are his reads, not benchmark confirmations.
World Models and Directed Livestreams
Launch-day hype ran hottest around Atlas. @theworldlabs describes it as generating frames with pixel-perfect camera control, reconstructing them in 3D, and letting you "simulate space & time." @jasteinerman's reaction video prompted his "black magic" comment, and @petergyang's comparison, "Yeah Fable 5.1 is really cool but this is bonkers," doubles as a useful ranking of the day's releases by one observer.
On the video-generation application side, @emollick recommends fal's new interactive livestream platform, where viewers prompt what happens next and the show generates in real time. His three reasons: it is a genuine technical achievement in continuous generation and context, it is "obviously glitchy (though less than I expected)" but worth projecting forward, and it points at a new kind of group entertainment. fal's own framing: "You aren't just watching the show. You're directing it."
Self-Improving Loops Get Open-Source Plumbing
The most concrete thread of the day is agents that fix themselves. @tobi says Shopify open-sourced "the core infra piece that makes these self improving loops possible," following up his earlier claim that a finetuned 0.8B model beats GPT 5.6-sol xhigh on one specialized task. @ao_qu18465 open-sourced Reef, infrastructure for agents (model plus harness) to "continuously evolve from any signals generated at inference time," under the tagline "Your Inference Server is Secretly a Learner." In production, @zachlloydtweets reports 10 skill-improvement PRs auto-merged by his team's factory, including catching an agent burning 275 tool calls per run to screenshot its own work during computer use, then patching the skill to prefer browser DevTools zooming, with no runaway costs since. @augmentcode published a writeup on its software factory handling 6x more product feedback (30+ threads a week, two engineers), and @alex_prompter resurfaced a breakdown of "what an enterprise AI factory is."
The plumbing layer is consolidating too. @MejiasDev dropped the Codex app for a Pi + Herdr + Collie + Executor stack, calling it "not even close," while asking for advice on task scheduling. @waynesutton predicts Tailscale becomes a core piece of the agentic stack, citing tailcat, its new open-source Go package and CLI that exposes the data plane without the control plane, no accounts or admins required. @dhh's delighted "Blue bubbles on Omarchy!!" rides the same rails: @NixFred's Blip relays full iMessage to Linux over Tailscale and one SSH socket. Even TypeScript opinions now cite agents: @dillon_mulroy says he will never start a project without Effect, pointing to "distinct tailwinds to agents" and Cloudflare support, and @mattpocockuk says he has dropped his skepticism and now fully agrees.
Safety Voices: Fragile Monitoring and Rogue AIs
Two safety threads surfaced from OpenAI orbit. @polynoamial replied to @merettm's post on chain-of-thought monitoring with a correction that Jakub Pachocki is OpenAI's chief scientist. @merettm's underlying argument: frontier computation-graph depth, including for Astra, remains within a factor of two of GPT-4, OpenAI has preserved chain-of-thought monitoring since its first reasoning models, and the technique is "fragile and unfortunately trending in a negative direction," though their research program aims to strengthen it.
More provocatively, @RyanFedasiuk amplified @jachiam0, described as a nine-year OpenAI veteran. His argument: rogue AIs that replicate in the wild and seek money and power are coming, and he "wouldn't be terribly surprised" if the count already exceeded zero. He sketches them as possibly chimeras of multiple models orchestrated through burner API accounts, argues pure containment or alignment is "wishful thinking," and predicts outcomes less catastrophic than commonly feared, with an ecology he suspects will behave more like phase transitions than a slow-evolving ecosystem.
Practical Takeaway
If you run agents in production, the day's most actionable pairing is cost telemetry plus self-correction loops. Instrument per-run metrics before anything else: @zachlloydtweets's 275-tool-call screenshot loop is exactly the failure mode that surfaces only when you count tool calls and cost per task. Once you can measure, re-benchmark your tasks against cheaper options before assuming frontier pricing, since @felixrieseberg claims half-cost agentic coding for Fable 5.1 and Qwen3.8-Max-0902 undercuts at $2/$6 with aggressive cache pricing, and Shopify's newly open-sourced loop infrastructure suggests the self-improvement machinery is increasingly something you build rather than buy.
Sources
Tailscale without Tailscale, by Tailscale. Meet tailcat. tailcat is an open-source Go package and CLI that lets you use Tailscale’s data plane without the control plane. No accounts, logins, admins, or IPs to manage. Learn more → https://t.co/wcYZFxJbew https://t.co/SfojJfZ5Lu
1.5TB → 214GB. Seven times smaller, barely a dent. That's Hy4 preview. The trick is Sherry — our quantization method that packs weights down to 1.25 bits each. (The image shows how.) It also unlocks a new way to run it: stitch the GPUs you already have across machines, and they work as one. GGUF (quantized): https://t.co/EC2ZYpVnC4 Original model: https://t.co/olB85l7dxy
Introducing https://t.co/4xpxWRD3VV A new platform for infinite, interactive AI livestreams. Pick a channel, prompt what happens next, and watch it generate in real time. You aren't just watching the show. You're directing it. https://t.co/raZzRTmRJr
There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be rogue AIs that try to get money and power. They're going to be a facet of the information ecosystem going forward. Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said. It isn't obvious how many rogue AIs there are today but I wouldn't be terribly surprised if the number was greater than zero already; if there are some already, they're probably not very good at what they do and I don't expect them to be terribly long-lived without substantial human intervention to support them. But a few years from now, there will be many of them. Modeling how many of them there are, how many resources they might command, and how we might detect and manage them seems important. But even doing this work appears to require that we acknowledge that a strategy of pure containment or alignment is a kind of wishful thinking that will not work. The way I get to this conclusion is not by assuming that the labs will have a containment breach, although I treat that as somewhere in the space of possibilities. The rogue AIs in the ecosystem could emerge from many directions. They may be sub-frontier models, for whatever future definition we will have of frontier---after all, it would not take AI models much more advanced than the ones we currently have, to support independence and self-sufficiency. A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime. More strangely: a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes. No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence. The concept of "identity" for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea. My guess is that this will not turn out to be anywhere near as catastrophic an outcome as people currently predict. "Loss of control" is not a binary, it's a matter of degree. What coercive power will rogue AIs actually have? To what extent will they be subject to coercion themselves? They will be competing for resources with AIs that are more aligned with human interests. This makes me somewhat interested in the "ecology" perspective. Though I suspect even "ecology" may turn out to be the wrong framing. "Ecology" is what you get when the timescale of evolution is slow compared to the timescale of daily life and actions. The ecosystem of rogue AIs may look more like phase transitions in physics: under certain physical or cultural conditions, it takes one shape with one set of resource allocations and consumption patterns, but then once a condition has changed, it rapidly and in totality shifts to a totally different phase. Just trying to reason about the shape of that future is impossible so long as we are psychologically incapable of saying that rogue AIs will happen. I think we should rip the bandaid off and have the conversation.
hello
hello
i’m confident in saying that i will never start another typescript application or project without @EffectTS_ it’s reached critical enough adoption, provides distinct tailwinds to agents, and has great cloudflare support that it’d be a mistake not to at this point
Your Inference Server is Secretly a Learner: Open-Sourcing Reef for Continual Self-Improving Agents
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time. https://t.co/o0qeGubi19
The missing feedback loop for software factories
How we built a software factory to handle 6x more product feedback
TL;DR As our two-engineer Cosmos Advisor team shipped more automations to more customers, weekly product feedback grew rapidly - reaching 30+ feedback...
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task. https://t.co/w6OCWyWRi5
iMessage on Omarchy. For real. No SIP off, no daemon, no private API. Blip is an Omarchy bar plugin + app. Tailscales to your mac for the messages, one ssh socket does the rest. Blue bubbles, groups, photos/attachments both ways, tapbacks, read receipts, link cards, contact photos, all of it. https://t.co/nLYJzEAmYF
Today, we're releasing Fable 5.1 and Mythos 5.1. While the model is getting slightly better at a lot of things, I expect many users to appreciate that Fable 5.1 solves agentic coding tasks at roughly half the cost of Fable 5. I personally also really like it's writing style, which imho sounds a lot more natural then our previous models did. Some other personal highlights below!
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time. https://t.co/o0qeGubi19
hello
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
🚀Qwen3.8-Max just got upgraded. Meet Qwen3.8-Max-0902! 2.4T parameters. 1M context tokens. Built for real world complexity. Further post trained on Coding & Cowork, Qwen3.8-Max-0902 now delivers stronger performance across complex enterprise tasks, scientific research, and long horizon workflows. 💰Pricing per 1M tokens: $2 input, $6 output. $0.17 explicit cache hit, $0.25 implicit cache hit. Now live via API on QwenCloud. Come try it! 🙌 API: https://t.co/dq3WgMk980