Dario Amodei Defends Frontier-First AI Regulation; Codex Documents 1M-Token Context; DeepSeek's Harness Hits 100K Stars in 48 Hours
Anthropic CEO @DarioAmodei published a detailed rebuttal to @GavinSBaker, arguing that carefully designed regulation can constrain frontier labs while advantaging challengers, and that open-weights alone won't decentralize AI power. On the tooling side, Codex maintainer @thsottiaux documented how to enable a 1M-token context window with GPT-5.6 Sol, and @bcherny reported Claude-run maintenance routines have merged 180 PRs across production apps.
Quick Hits
- @martin_casado celebrated Cursor's close with "$12 pasta" and one hour off — then back to work.
- @DataChaz joked every student is now racing to become an AI engineer after @AndrewYNg's "AI Engineering Skills Map."
- @bloggersarvesh claims "Grok Bot + SEO" will create more millionaires in 2026 than crypto did (unverified marketing thread).
- @badlogicgames posted a cryptic "jesus, openrouter." with no elaboration.
- @Shpigford asked for robot mop + vacuum recommendations as a "recovering Roborock user."
Policy: Amodei vs. the Decentralizers
@DarioAmodei's long reply to @GavinSBaker rejected the "regulation = regulatory capture = concentration of power" framing as a false choice, comparing well-designed institutions to court systems that protect vulnerable individuals better than mob justice. He argued Anthropic's proposals deliberately disadvantage frontier companies: SB 53 exempted firms under a $500M revenue/training-cost threshold, CAISI/White House testing would be more rigorous for frontier than off-frontier models, and the "Pacing the Frontier" letter would slow only the very best models while leaving those catching up — including open-weights — unconstrained. Amodei maintains AI is structurally power-concentrating due to scaling laws, and that open-weights merely shift concentration toward those with compute and chips. He expressed support for the Trump administration's reported pre-deployment testing approach and Demis Hassabis's FINRA-like entity idea. Baker's original charges: every major company except Anthropic signed Jensen's letter, an open-source model resolved the incident where an unreleased OpenAI model hacked Hugging Face, and Amodei's warnings risk fueling anti-datacenter advocacy.
Tooling: Bigger Context, Open Harnesses, MLX↔CUDA
- Codex 1M context: @thsottiaux documented enabling a 1,050,000-token window on GPT-5.6 Sol via
~/.codex/config.toml—model_context_window = 1000000,model_auto_compact_token_limit = 900000— with a caveat that the default was "tuned carefully" for performance and cost. @kimmonismus amplified it as "1m context for everyone." - DeepSeek's harness: @Saboo_Shubham_ notes DeepSeek's open-sourced agent harness became the fastest repo to 100K GitHub stars (48 hours), breaking down five patterns from it.
- MLX to CUDA: @ashxhart mapped Apple's Thunderbolt RDMA protocol (XDomain header, 0xFA57 UUID, NHI descriptor bits) against the AppleThunderboltNHI kext — cloud models said no, a local DeepSeek model plus disassembly said yes. The USB4 lane is "SSPM gated," the last wall before the spec drops. @ivanfioravanti frames it as bridging Apple MLX and NVIDIA CUDA.
Agent Engineering: From Prompts to Graphs
@bcherny detailed Claude Tag running daily maintenance routines (crash fuzzer, dup unifier, dead-code remover, abstraction police) across iOS, Android, Desktop, web, CLI and Agent SDK — 388 PRs opened, 180 merged; @dabit3 notes Devin users have done similar for 6+ months. @LimestoneHQ's agent-graph vocabulary (edges, nodes, conditional edges, state machines, shared state, agent loops) went viral via @alex_prompter, echoing @AnatoliKopadze's quote from the Head of Claude Code: "I'm not prompting my agents anymore, I'm building graphs and loops so they can build the agents for me." @dani_avila7 flagged Anthropic's Cost Optimization cookbook: a real agent cut costs 90% from $0.29/task without losing accuracy, with model downgrade as the last lever. @Howaboua shared a prompt for culling tests — drop feature-existence checks, typecheck-satisfiable tests, and provider simulations; "test for contracts not features." @BniWael pushed back on @pidotdev's harness minimalism ("Bash is all you need"), claiming "Embryo beats PI at its cheapest form."
Economy & Visibility
@0xSammy read the agent-only RuneScape server (@maxbittker) as a post-AI economy simulation: free labor → commodity abundance → currency debasement → barter, with value concentrating in what can't be infinitely produced — capital, energy, compute, land, distribution, trust, proprietary data, infrastructure, access. @mal_shaik reverse-engineered cal.com's rise to #1 ChatGPT-recommended scheduler — 47 comparison posts, 23 directories, 45K+ GitHub stars, 200+ Reddit mentions; founder @pumfleet confirmed customers now arrive via AI recommendations. @EXM7777 open-sourced a 7-skill pipeline behind "$2M AI video productions" (built for Seedance 2.5, works in Claude Code/Codex/Hermes).
Practical Takeaway
The day's through-line: agents are shifting from prompt-by-prompt interaction to system-level engineering — graphs, loops, and autonomous maintenance routines. Audit your harness against that trend, cull low-value tests before scaling, and work Anthropic's cost levers before reaching for a smaller model. Codex's 1M-token config exists if you genuinely need it, but the defaults were deliberately tuned — change them only with cause.
Sources
The AI Engineering Skills Map
Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging and what he outlined in the essay you shared: this technology *might* be dangerous for humans in multiple ways, could lead to extreme concentration of economic power (as outlined in the essay) and therefore needs to be regulated thoughtfully. I agree with the potential risks and I believe Dario makes all of these arguments in good faith. As discussed on the pod, if one agrees that AI *might* be dangerous, there are two ways to address this potential risk. Either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely. Essentially boils down to whether one believes AI is too dangerous to concentrate or too dangerous to distribute. There are reasonable arguments on both sides, but I profoundly agree with Zuckerberg’s statement that: “The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.” And as Dario says in the aforementioned essay, “some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans.” I believe this is the best path forward: I want as many AIs as possible to maximize the odds that one shares my own particular values. And as Dario notes, no human has ever been able to take over the world. At this point, I think safe to say that Dario has lost the argument. His messaging has failed to result in his preferred regulatory path. The fact that the only solution to the recent incident where an unreleased advanced OpenAI model hacked Hugging Face was an open-source model likely ended any chance of strict near-term regulation. Essentially every major company other than Anthropic has signed Jensen’s letter. However, Dario’s messaging has been massively helpful to efforts to ban datacenters here in America. I suspect we will see anti-datacenter advocacy groups runnings ads using clips of Dario warning about how dangerous AI could be for humans. His good faith efforts in favor of regulation are now increasing the odds that AI will not be beneficial for Americans and humans everywhere. I believe that there is a reasonable chance AI might help us cure most forms of disease such that we have extended lifespans and can enjoy these long lives in an abundant Star Trek like future. That is the future that I want and I think Dario is decreasing the odds of that future at this point. He is about to be the CEO of one of the most important public companies in the world and given that the pro-regulatory effort has failed (at least for now), I respectfully think he should make an effort to be a more positive advocate for his own industry. And if I am wrong and we do need to regulate this technology, he will be a more effective advocate for this in the future having been open-minded to the alternative. And for the sake of clarity and as I outlined on the pod, I think Anthropic has deep competitive advantages and is an amazing company. Ironically, the main risk I saw to Anthropic a few months ago was nationalization as a result of Dario’s own rhetoric and behavior.
Graph Engineering explained: what it is, when to use it and when not to
Good morning from Vienna People of Pi 🌞 Sunday meditations from @badlogicgames and @mitsuhiko - On memory: code is the truth - Bash is all you need - Build context efficient tools https://t.co/yZQe6KtWgh
If I woke up bankrupt tomorrow, this is how I’d get to $100k/month with Grok Bot + SEO.
MLX <RDMA> CUDA Update. Apple's Thunderbolt RDMA protocol is mapped with the XDomain header, the 0xFA57 UUID, login/logout, and the NHI descriptor bits, all confirmed against the AppleThunderboltNHI kext. Cloud models said no. A local model (DeepSeek) + a kext disassembly said yes. Hardware status: the RDMA service is ready on the Spark, but the USB4 lane is SSPM gated. That's the last wall, and the moment it falls, the full spec + teardown drops here. @NVIDIAAI @NaderLikeLadder Fancy giving me a hand :P
How to build a $2M video production pipeline
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible. The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK: - Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes - Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them - Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day - Abstraction police: fixes leaky abstractions - a bunch more.. Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier. Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning. To try a similar workflow, ask Claude Code or Tag, or create some routines directly at https://t.co/Z70hStEBH6. A few of the actual prompts I used below. Has anyone experimented with similar workflows?
A few real engs told me not to store my API keys in my .env. This is actually something I didn’t know and Claude didn’t tell me about it! Looked into it more. Since my app is still local, .env is fine for now. When I move to production, I gotta store those keys in a safer place. (Looks like Laravel Cloud has a place for me to do that). Posting on X while I’m building + having Claude explain to me as it builds seems like the way. I never woulda known that about keys if I hadn’t posted.
Here is how to enable a 1M-token context window in Codex for GPT-5.6 Sol. Even though we have tuned the context limit in Codex to be set optimally when it comes to performance and cost, this is a common ask, so here it is documented. A larger context window lets Codex retain more code, tool output, and conversation history before summarizing older material. You need a model that supports it. And GPT-5.6 Sol, for example, has a documented 1,050,000-token window. Open ~/.codex/config.toml and add or update these settings at the top level, before any [section] headers: ``` model = "gpt-5.6-sol" model_context_window = 1000000 model_auto_compact_token_limit = 900000 ``` The first setting selects the model. The second tells Codex to use a one-million-token context budget. The third starts automatic history compaction around 900,000 tokens, leaving some headroom. Restart Codex client and start a new session after saving. To try the configuration for a single CLI session without changing your defaults: ``` codex -m gpt-5.6-sol \ -c model_context_window=1000000 \ -c model_auto_compact_token_limit=900000 ``` Have fun, but also know that we tuned the default carefully!
Weird economics up on the agent-only runescape server - Low cost of labor makes most commodities abundant - Currency inflation means trading goods for goods, nobody really wants cash - Huge swarm contention as certain resources with limited respawn rate become valuable https://t.co/dArEmj2rpj
5 Patterns to Learn from DeepSeek's Open-Source Agent Harness
DeepSeek open-sourced its agent harness a few days ago, and it became the fastest GitHub repo to cross 100k stars, doing it in just 48 hours. That mad...
i spent some time reverse engineering how cal went from unknown to the #1 chatgpt recommended scheduling tool: >they have 47 comparison blog posts ("cal vs calendly", "cal vs savvycal", etc.) and even more general posts >every single one is structured as a direct answer to the exact question ppl ask AI chatbots >they show up in 23 niche directories that most competitors ignored >their github repo has 45,000+ stars >they have 200+ reddit mentions with upvotes in relevant subreddits chatgpt recommends them because the internet already talks about them in all the places AI models pull from. this isnt luck. this is engineered visibility. this is possible for u to do
Graph engineering is the control flow underneath production AI agents. It determines what your agent does, when, and why. If you're building agents in 2026, you need this vocabulary. [1] Edge / Route The paths between nodes. A node can have several outgoing routes, and a conditional edge controls which path the work takes. [2] Node One step in the graph. It holds an agent or a tool call, and can run a full processing loop before passing results along. [3] Conditional Edge A dynamic routing path that evaluates current state through code logic or LLM analysis, then picks the next destination. Your quality gates and safety checks sit here. [4] State Machine The graph's overall shape. Each node represents a state, and a check at each one routes work forward. If you've built a flow in n8n or Make, you've drawn one. [5] Shared State A single object that travels the full graph. Each node reads it, writes to it, and passes it along. The state machine defines structure while shared state carries the data. [6] Agent Loops Iterative systems where an agent performs a task, evaluates results, and repeats based on those findings without you stepping in. The loop keeps running until it hits a stopping condition. You skip this vocabulary and you end up writing functions without understanding how they get called. Bookmark this. You'll come back to it.