Karpathy Says Stop Reading Everything Your LLM Writes; Cloudflare's Clef Comes for Jev
Andrej Karpathy's guide to making model output digestible (controlled language, diagrams, interactive pages, explainer videos) drew the day's deepest engagement, including community tests of the writing rule. Cloudflare open-sourced Clef decision models with claimed Jev compatibility, Tavus claimed the first pass of a video Turing test, and economist Daron Acemoglu laid out revenue numbers he argues the AI boom cannot hit without surging inequality.
Quick Hits
- @karpathy posted an escalation ladder for consuming LLM output: ask for explanations in ASD-STE100 (a controlled language built for aerospace maintenance docs), then prefer diagrams, then interactive HTML pages, and finally bespoke 3Blue1Brown-style explainer videos, arguing that "large, custom, discardable software artifacts" are now cheap enough to generate on demand. @kunchenguid tested the writing rule and reports it works, provided you apply a subset of the spec rather than the whole thing.
- Cloudflare open-sourced Clef and Clef-flash, decision models hosted on Workers AI that @CloudflareDev calls "fully Jev-API compatible." @brandonjcarl says Clef is the first model to give Jev a run for its money on his independent benchmark.
- Tavus claims its Griffin model is the first to pass a video Turing test, with 48% of live callers believing they spoke to a human, versus under 3% for previous systems, all per the company. @cgtwts's verdict: "it's over for us."
- Claude Code now supports mods, per the ClaudeDevs announcement: behavior changes, UI customization, and feature swaps written in a few lines of TypeScript and shipped inside plugins. @trq212 frames it as software becoming malleable as a first-class feature.
- Economist Daron Acemoglu, in a daily series that @dieworkwear recommends, calculates that AI investment averaging 3.6% of US GDP through 2032 would need roughly $3.7 trillion in annual revenue to pay off, and predicts the industry will fall well short.
Karpathy's Ladder: Make the Model Draw It, Build It, or Narrate It
The most-shared item of the day is @karpathy's argument, quoted in full by both @uudonX and @kunchenguid, that as models do more autonomous legwork, human work shifts up into oversight and understanding, so the format of model output matters as much as the content. His progression: prose constrained by ASD-STE100 (softened to "80% of the way" since the spec is stringent), then diagrams for structure, then generated HTML for anything with parameters, then custom explainer videos with AI narration. @uudonX's widely shared Chinese-language breakdown frames the core problem: when output speed exceeds comprehension speed, writing more text just stacks the bottleneck higher, and "disposable software" is approaching zero cost. @malikwas1f amplified @aakashgupta's framing that aviation effectively solved AI slop in 1986 with the same controlled-language spec. @kunchenguid pushes back on skeptics, saying most have not tried the tips, and shares a clever meta-prompt: sample ten recent session transcripts, apply ASD-STE100 rules retroactively, and document the clarity-boosting subset in your AGENTS.md. Nearby, @mattpocockuk condenses @poteto's "restate my goals in your own words" prompt to five words, "Restate my intent before continuing," for dictation-heavy workflows. @unclebobmartin lands on the same theme from another direction, noting nobody should be shocked the author of Clean Code no longer reads code, since clean code's purpose was always to clear the path for thinking. @poteto also supplied the comic version of blind delegation, asking a bot to "make me a billion dollars" without making any mistakes.
Decision Models: Clef Versus Jev, a Mystery Bunny, and an Uncensored GLM
Cloudflare's launch, announced by @ritakozlov and promoted by @CloudflareDev, positions Clef and Clef-flash as open-source decision models for high-speed classification and agentic workflows, with a Jev-compatible API to ease switching. @brandonjcarl reports Clef is the first model he has tested on his independent benchmark that genuinely challenges Jev. @MiaAI_lab offers a third-party technical explanation: a frozen Qwen backbone does one prefill pass while a small schema head scores answer options in parallel without generating text, which the account claims yields 4x Jev's speed at 2x its accuracy on some benchmarks; treat those numbers as one observer's read, not Cloudflare's. Elsewhere on the model front, @GergelyOrosz is openly puzzled by Space Bunny, a model he describes as hugely popular, strong at coding, fast, and cheap, with no known creator, asking whether it is real or a prank. And @kimmonismus flags that @elliotarledge has posted GLM5.3 uncensored on Hugging Face, dryly predicting "fun times ahead" for anyone who worries about safety.
Agent Stacks Get Org Charts
The day's orchestration posts read like engineering management for models. @slashui boosted @thedelost's Codex setup, declaring all future architecture diagrams should look like it: GPT-6.1 Sol runs the main session, an Astra architect is spawned at three checkpoints (before planning, on repeated errors, before declaring done), explorer and researcher agents run on Luna, and Jev handles sub-second fork decisions that need no deep thinking. @sydneyrunkle shares results from a model router built into the harness: median cost per task fell 64% with no measurable quality drop, on the premise that most tasks do not need top-tier intelligence. @TheAngryPit describes @NousResearch's Hermes orchestrating a PrimeIntellect agent to train an internal model on herdr, while agent teams on Openclaw and Codex do the daily work, with Cursor and a bot jokingly "on paid leave until reset." @justsisyphus recommends herdr-web with OmO V5, quoting @q_yeon_gyu_kim's praise for its stability and mobile remote. The tooling layer keeps filling in around this: @badlogicgames welcomes @humanlayer_dev's tease of a "software forge for AI and whatever comes after it," and @pidotdev announced shipping Pi 1.0 with Pi Durable.
Verification Is Still the Bottleneck
Testing tooling dominated the practical thread. @o_kwasniewski launched e2e, an open-source agentic testing framework (deterministic and agentic APIs mixed, web and mobile, local or CI), which @jmwind calls better than the version "we've all been building." @stringsaeed details a mobile workflow pairing e2e with @poteto's verification skill: pstack for most work, stim for orchestration and simulator warmup, agent_device and argent for simulator control, e2e to confirm the feature actually works, replacing slow Maestro flows. On evaluation, @dexhorthy highlights @KLieret's SWE-sweep benchmark, where agents must find and fix bugs across 100 repos and 22 languages without being told what broke; top models score under 5%, another "deeply unsaturated" evaluation. @jamonholmgren corroborates @chimon1984's finding that building the same app natively with agents was slowest by a wide margin, up to 3.5x the agent minutes of cross-platform options. And @thdxr reminds everyone not every bottleneck is AI-related, pleading once again to kill the localhost OAuth redirect in favor of the client simply polling for a code.
Doubts at the Edges: Souls and Balance Sheets
Two posts supplied the day's skepticism. @TheAhmadOsman escalated to calling Anthropic "an evil cult," amplifying @DavidDecosimo's thread claiming Anthropic quietly flew religious leaders to San Francisco under NDAs and mainly argued that Claude has a soul and moral standing; these are unverified claims from a thread, not established fact. On economics, @dieworkwear points to @DAcemogluMIT's series, whose latest post runs Van Nieuwerburgh's arithmetic: at a 10% return, the industry needs about $3.7 trillion in annual revenue by 2032, against roughly $200 billion today, while the capital share of national income sits at a record 47%. Acemoglu's dilemma: hitting the number likely surges inequality, missing it raises crash risk, and he expects the miss, citing slow diffusion, open-weight competition, and limits on automating whole occupations.
Practical Takeaway
If your models already write more than you read, change the output format before scaling usage further. Try the ladder in order: a subset of ASD-STE100 for prose, a diagram for anything with structure, an HTML artifact for anything with parameters, and run @kunchenguid's transcript-audit prompt to codify which clarity rules actually help you in AGENTS.md. Separately, if you are paying for coding agents, instrument per-task cost before anything else; @sydneyrunkle's 64% median savings suggests routing easy tasks to cheaper models is often the highest-leverage change you can make this week.
Sources
this has become one of my most used prompts recently: > restate in your own words what you think my goals are and what the problem i'm trying to solve is
Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do. SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%. https://t.co/OwqrjHk64q
i think this is the best util for the herdr orchestration trying so far very stable and useful especially on that mobile remote and the github project name is very honest tho https://t.co/mrJ6iudPVJ
it's decision model season 🍂 today, we open sourced clef and clef-flash, our first homegrown models — both now available on @cloudflare workers ai! demo / leaderboard: https://t.co/GJfPEKzP8k blog: https://t.co/HskdhS0nJz https://t.co/YzhzRmeeOH
How to Build a Model Router in the Harness
Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM). https://t.co/eb11XbC8Xr
Introducing e2e The agentic testing framework for any app ~/ npx e2e init → Fully Open Source → Mix deterministic and agentic APIs → Supports web, mobile, and more → Bring your own agent and infrastructure → Run locally or in CI https://t.co/G3XB2L8tiC https://t.co/gUZ3AII2fY
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples: https://t.co/qMtzV8AeId
been a little quiet lately but excited to share what's next: HumanLayer is building the software forge for AI and whatever comes after it https://t.co/25BZhPwx0K
Third question on AI. A question that also remains unasked is whether the AI boom can continue without leading to a massive increase in inequality. A recent paper by Stijn Van Nieuwerburgh runs the numbers on how much revenue the AI industry needs to generate to recover its massive investment (summary and a link to the paper can be found here: https://t.co/QIGcoObcVc). Van Nieuwerburgh’s arithmetic should make us more concerned. AI investments will average about 3.6% of GDP annually between 2025 and 2032. Van Nieuwerburgh calculates that, using a 10% rate of return, the industry would need to generate annual revenues of about $3.7 trillion by 2032 to recover these costs (growing from its current levels of about $200 billion or so). That is significantly more than 10% of current US national income, and will likely remain around 10% of national income by 2032, even if GDP growth rose from its current level. A large fraction of this revenue will go to capital income. That means a massive increase in the share of capital in national income, which has already risen substantially over the last 25 years or so – now standing at an all-time high of about 47% (https://t.co/IuYldl173r). Capital income is much more unequally distributed than labor income, so a massive increase in the capital share of national income will translate into a very sizable surge in inequality. The rise in inequality may not stop with the capital share. My work with Pascual Restrepo documents that (automation-driven) increases in the capital share of national income are typically associated with rising labor income inequality as well (see, for example, https://t.co/2D9KUL3QfM). The same may happen in the next several years, boosting inequality further. What is missing from our current debate is any discussion of a fundamental dilemma these numbers pose: can the AI boom avoid both an economically costly crash and a huge increase in inequality? If the industry reaches these revenues, inequality surges. If the industry does not become profitable, a crash, with substantial costs in terms of lost output and jobs, becomes likely. My assessment would be that the industry is unlikely to reach levels of revenue Van Nieuwerburgh calculates. First, diffusion has been and will likely continue to be slow. Second, competition from open-weight models, which are getting better, will limit how much proprietary models can charge. Third, despite important advances, I still believe that AI models will not be able to automate entire occupations anytime soon, thus limiting their value to businesses as cost-saving devices. Whether this leads to a crash or not is more complicated and will depend on whether various AI companies are bailed out and what kind of support they receive. Nevertheless, even if revenues fall short of these gargantuan amounts and we avoid a dramatic surge in inequality, I expect that the diffusion of AI will push up inequality between capital and labor and within labor. If inequality does surge, a further question becomes central: can our democracy survive such astronomical levels of inequality?
We are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. https://t.co/eilE3tV0F2 #BirthdayWeek
I built the same four-screen app six ways — native, React Native, Flutter, Compose Multiplatform — with agents doing the work. Some teams dismiss native's cost too quickly. In my tests, it was the slowest by a wide margin — up to 3.5× the agent minutes.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Codex tip: once GPT-6.1 Sol is your main model, stop running Astra on every turn put Astra on call as an architect agent GPT-6.1 Sol keeps writing the code Astra only gets spawned at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Astra reviews. Sol ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split - the full tree > GPT-6.1 Sol on high runs the main session > explorer reads the code on Luna > worker edits and runs tests on Sol > researcher pulls the docs on Luna > all three on medium > Astra on call as the architect > auto_review checks every approval paste the tree and this prompt into Codex ↓ "Rebuild my Codex setup around this tree: 1. Check ~/.codex/agents and .codex/agents for agents that already fit explorer, worker and researcher. > Draft new TOML files only for missing roles > explorer and researcher on gpt-6-luna, worker on gpt-6.1-sol, all with model_reasoning_effort medium > Add an architect agent on gpt-6-astra, model_reasoning_effort high, whose only job is reviewing plans, repeated errors and finished work > Skip any that pin a different model and list them 2. In ~/.codex/config.toml set model to gpt-6.1-sol, model_reasoning_effort to high and approvals_reviewer to auto_review 3. Find anything that would override this (active profiles, flags in my shell aliases, agents.default_subagent_model). Report it, change nothing 4. Add one rule to AGENTS.md: spawn the architect before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳ https://t.co/eyMF8gZKLE
Introducing e2e The agentic testing framework for any app ~/ npx e2e init → Fully Open Source → Mix deterministic and agentic APIs → Supports web, mobile, and more → Bring your own agent and infrastructure → Run locally or in CI https://t.co/G3XB2L8tiC https://t.co/gUZ3AII2fY
This is deeply disturbing: Anthropic secretly invited religious leaders to SF to advise them on AI safety, had them sign NDAs, but then primarily tried to convince them Claude has a soul & moral standing, all while wining & dining them to the max. 1/ https://t.co/4Z7Yk9rwk2
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
here u go https://t.co/IT4dkeyGxE