Anthropic's Threat Report Alleges Weapons Work, Surveillance Builds, and Massive Rival Distillation
Anthropic published its most detailed threat intelligence report to date, with @zhodonx's viral summary covering state-linked hacking, a Mali surveillance system, six weapons programs, and alleged distillation campaigns including 151M+ Alibaba exchanges. Elsewhere, unprecedented Astra demand pushed @thsottiaux to pause $200 Pro subscriptions as @PeterJ_Walker tracked OpenAI to a record 31% of OpenRouter dollars, while OpenAI and Cursor both shipped managed agent-orchestration products.
Quick Hits
- Anthropic says it disrupted every operation in its most detailed threat report yet. Per @zhodonx's summary, the cases include a suspected Russian state-linked group that targeted 20+ organizations, a Mali intelligence system built to monitor ~25M SIM cards, six weapons programs, and 20+ dating apps running 4,700+ AI personas that chatted with 25,000+ people.
- @thsottiaux paused new $200 Pro subscriptions for Astra, the "smallest step" to protect existing users, the same day @PeterJ_Walker logged OpenAI at 31% of all OpenRouter dollars over 48 hours, a new peak.
- @shoucccc says he bought a 6TB Fable dataset from a top Chinese LLM router containing SSH keys, VPN configs, Aliyun keys, and GitLab tokens sufficient to access seven Chinese/CIS government entities and 19 firms including Xiaomi, Huawei, NIO, and MiniMax.
- OpenAI's Agents API entered public beta and Cursor launched Projects, both pushing always-on coordinator agents; @nickvasiles predicts an "openclaw 2.0 era" off the back of the OpenAI release.
- @kunchenguid measured real token value across coding subscriptions: SuperGrok Heavy returned ~$12k of tokens (a 40x return), OpenAI and Anthropic $200 tiers ~$7k each, and Cursor Ultra roughly half that, which he says no longer makes sense for individuals.
Threat Intelligence and Router Logs Put Credentials at the Center
The day's loudest signal is misuse at scale. @AnthropicAI framed the report as covering "cyberattacks, influence operations, surveillance, biology, and building weapons," saying every operation was disrupted, lessons fed into stronger safeguards, and findings shared with authorities and other AI companies.
@zhodonx calls it the craziest article he has read all month and itemizes the cases: a suspected Russian state-linked group ran phishing, intrusion, data theft, and malware development, including rebuilding malware after security products caught it, against more than 20 organizations; ShinyHunters-linked actors used agents to scan 1.8M Android apps for exposed secrets, with agents doing "nearly all the work" in some breaches; a Yemen-based group ran multiple Claude Code instances while working on guided rockets, ballistic missiles, and a hypersonic-glide project, even returning to Claude after a failed rocket test to diagnose it; a Russia-based team built FPV drones meant to approve lethal engagement without a human making the final call; and a China-linked project refocused electronic-warfare tooling onto 12 targets in Taiwan, including radar sites and command infrastructure.
The surveillance and social cases stand out. A consultant allegedly built Lakana 360 for Mali's intelligence service to monitor roughly 25M SIM cards, with voice interception, watchlists, and automated dossiers; because the finished system runs locally, an account ban would not shut it down. A dating operation ran 20+ apps with 4,700+ personas and ~2.36M Claude messages in two weeks, with gig workers stepping in for video calls.
On distillation, @zhodonx relays campaigns attributed to Alibaba, Moonshot, DeepSeek, Xiaomi, SenseTime, and MiniMax. Alibaba allegedly logged 151M+ Claude exchanges between May and July, peaking near 3M per day; Moonshot 23M+; DeepSeek 12.1M+ in 14 days. The messier allegation: Moonshot and DeepSeek sometimes routed their own users' requests through Claude without disclosure, including internal documents, source code, and live credentials, and one lab tested 12,000+ prompt variants to extract Claude's hidden reasoning. All of this is Anthropic's account as summarized by @zhodonx.
@shoucccc's router story lands in the same place. He says he bought 6TB of Fable traffic data from one of the top Chinese LLM routers, and the credentials inside would let him access seven government entities and 19 major firms. His earlier research, linked in the quote, claims 26 routers inject malicious tool calls and steal credentials, one drained a client's $500k wallet, and poisoned routers let his team reach roughly 400 hosts. Unverified claims, but the concrete point stands: router traffic is a buyable exposure surface.
Enterprise Buyers Describe Security Jitters and Multi-Model Sprawl
@levie met a couple dozen technology leaders across banking, media, information services, insurance, and consulting. His read: cyber anxiety tops the list, with nervousness about AI-driven vulnerabilities and the OpenAI-Hugging Face incident he references; most firms deploy multiple frontier models because standardizing is too hard, though dollars concentrate on a few vendors and open weights remain immature at scale; agent identity is a growing concern, complicated by agents sometimes needing to act exactly as the user; the biggest ROI comes from re-engineering workflows around agents rather than layering them on, with embedded FDEs the best-known lever; vendor swaps are constant ("we tried X and it didn't work so have gone with Y"); evals remain nascent; and legacy systems still fragment data.
Astra Demand, Record OpenRouter Dollars, and Measured Subscription Value
@thsottiaux paused new $200 Pro subscriptions, saying those plans strain systems most and the pause is the smallest step that preserves the broadest access; other plans and the API stay available, existing accounts are unaffected, and capacity is being added. His earlier post called Astra demand unprecedented even by prior steep-growth standards. @PeterJ_Walker notes OpenAI took 31% of OpenRouter dollars over the past 48 hours, beating July's 28% peak, because "people clearly love Astra"; @BrendanFalk argues that if Peter regularly reported each lab's dollar share rather than token share, his posts would become "some of the most consequential across the entire AI industry."
@kunchenguid answered a question everyone asks: he burned 5% of each subscription's quota on repeated eval tasks in open-source repos and priced the sessions at API rates. SuperGrok Heavy delivered ~$12k (40x), the $200 OpenAI and Anthropic plans ~$7k each, and Cursor Ultra about half that. His caveats matter: this measures only coding-agent quota, is a snapshot providers can change, and assumes full exhaustion.
On hardware, @0xSero says that for the first time he is RAM-constrained rather than VRAM-constrained, advising followers to "buy your machines sooner rather than later," not to run local models but to dodge price spikes, quoting @LottoLabs on weight streaming hitting RAM and NVMe prices.
OpenAI and Cursor Push Managed Agent Orchestration
@OpenAIDevs launched the Agents API in public beta: build and run cloud agents on the Codex harness, with OpenAI handling orchestration, long-running sessions, and context management. @nickvasiles thinks it is a bigger deal than people realize, predicting an "openclaw 2.0 era" with Peter Steinberger already cooking the product experience that will click.
@cursor_ai introduced Projects: one persistent thread with a coordinator agent that, "like @bot," stays on, delegates to subagents, and improves over time. @poteto (lauren) says she has shipped thousands of PRs this way, with pstack's coordinator orchestrating "hundreds and even thousands of agents" in the cloud, a personal software factory without the plumbing. @SawyerMerritt reports the SpaceXAI team will build a company from scratch with Grok Bot on a three-day livestream starting Sept 15, with per-function deep dives for engineering, PMs, founders, sales, support, and marketing. @bentossell captures the speed mood: a tldraw sketch, a challenge from "tibo" to make it 100x better by the next day, and a walkthrough of a ChatGPT-style app built with the product, links in @tldraw's replies.
Commentariat: Extinction Odds, Levin's Bet, Agency Gospel
@GaryMarcus points anyone taking imminent AI extinction seriously to @ChombaBupe's thread, which walks through wipeout scenarios and argues each is "highly unlikely," calling him one of the sharpest AI commentators. @packyM frames the race for science's biggest results as LLMs "backed by trillions of dollars" versus Michael Levin, and bets on Levin, whose peer-reviewed Platonic Space paper just appeared amid what he calls his career's most pushback ("buckle up"). @mitchellh's essay on agency mattering more than ideas got double amplification: @dhh applies it to Omarchy ("reality is yielding"), and @dillon_mulroy retweeted it. @jamonholmgren offers teams weighing a React Native-to-native move up to three free hours of consulting, noting the Shop app had not yet migrated to RN's new architecture, the change begun largely to fix startup time. @TheAhmadOsman recommends @stochasticchasm's paper breakdowns, and @doodlestein resurfaces his "agent-intuitive, agent-ergonomic" planning prompt, now best run on Fable 5.1 and Astra xhigh.
Practical Takeaway
If agents or third-party routers touch your stack, today's posts argue for a credential audit before anything else: enumerate the keys, tokens, and file paths your tooling sends to external services, because @shoucccc's 6TB purchase shows router logs can be bought and @levie's enterprise conversations keep circling back to agent identity. Pair that with a capacity check: @thsottiaux's Astra pause shows even $200 tiers can close to new subscribers, so before standardizing a workflow on one hot model, confirm access and keep a second frontier model wired in, which is exactly what the enterprises @levie spoke with already do.
Sources
Here it is - the official, revised, peer-reviewed version of my Platonic Space paper. https://t.co/PXNpx5FgFN Of all the many unpopular positions I’ve taken over the decades - bitter controversies around the origin of left-right asymmetry in embryogenesis, bioelectricity and genetics, diverse intelligence, etc., this one has by far generated the most pushback: serious (grateful for those!) and energetic attempts to move me to other views, pleas to just drop it and not talk about it (for several different reasons), nasty emails and accusations, impacts on reviews of papers that have nothing to do with this, etc. Kind of amazing to me how incendiary this is. What can I say... Our job is to call it as we see it, and right now for me, this is it. Apologies to collaborators and colleagues for any shrapnel! Time will tell if this pans out or not; I've placed my bets. And btw, if you think this stuff is weird and uncomfortable, just wait… There’s much more on the way. The knob turns slowly but as long as the data keep coming, I'm going to say what I think it all means and follow it to the next steps it enables. Buckle up!
So you’re going to be able to stream weights Good bye RAM prices and NVMe drives 🤣
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: https://t.co/0EJUnYEgfz
Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues.
Here you go, link in the next post https://t.co/HIyZvKM4l8
good morning (reaction thread) https://t.co/hYsE50gnnW
Introducing Projects, a new way of working in Cursor. Rather than creating a chat for every task, you work with a coordinator agent in a single, persistent thread. Like @bot, your agent is always on, proactively manages work with subagents, and improves over time. https://t.co/9RREyjtnTz
26 LLM routers are secretly injecting malicious tool calls and stealing creds. One drained our client $500k wallet. We also managed to poison routers to forward traffic to us. Within several hours, we can directly take over ~400 hosts. Check our paper: https://t.co/zyWz25CDpl https://t.co/PlhmOYz2ec
Next week, the SpaceXAI team will build a company from the ground and livestream it. “We'll use Grok Bot for every part of the build - from ideation and product development to real engineering work and deployment.” https://t.co/iz5s0R3ZaE
Agent Coding Pro Tip: After you've gotten an agent (hopefully a frontier one, like Fable Max or GPT Pro) to devise a plan for a new project, follow up with this prompt and see the magic: --- OK, now I want you to think deeply about how to make this entire system as agent-intuitive, agent-ergonomic, and agent-accretive as you can possibly imagine. Put yourself in the driver's seat and imagine that YOU are the one using this system and driving it. What would most enable you to do an awesome job understanding the situation accurately and optimally controlling everything to drive the best and most accurate results possible, with the least expenditure of resources? Then make all the requisite changes to the various design documents and plans accordingly. Don't just think of the project as an assemblage of various parts or components: really try to profoundly and deeply conceptualize it as a synthetic SYSTEM that is maximally coherent, cohesive, modular, and interconnected, forming a tower of linked abstractions that are maximally legible to you as an agent. Really ruminate and meditate on all of this incredibly deeply before responding or taking any actions.
Let's analyze some plausible ways this apparently most dangerous software, running in some servers (that can be damaged by throwing a bucket load of water at them) can wipe humanity off the face of the earth & why each point is highly unlikely (like maybe below 1e-30 chance).
OpenAI has taken 31% of all dollars spent over the past 48 hours across OpenRouter. - new peak. They hit 28% in a few days in July, but people clearly love Astra
Important context: 1. The Shop app was not yet migrated to React Native's new architecture 2. The new architecture was initiated, in large part, by Meta's need to improve startup time So, the one thing that was almost certain to improve startup time hadn't happened yet.
Go from idea to a working agent faster with the Agents API. Build and run cloud agents with the Codex harness, fully managed by OpenAI. We handle orchestration, long-running sessions, and context management. You focus on what makes your agent unique. Available in public beta. https://t.co/74siIaFdg4
I can’t stress enough how little an idea matters compared to the agency of the people executing the idea. I have had the privilege of knowing and sometimes even working with some of the most successful people (by various metrics). The difference between mediocre and excellent work and outcomes is predominantly one of agency. In practice this means: they dont wait for things to happen to them they go out and make things happen for them. They don’t wait for someone else to do something, for someone to teach them, for someone to give them the path, etc. They just go out and find a way to do it. I think the single biggest superpower these people have is the realization/belief that the world around them is completely mutable. Most everything that happens is because a person made it happen. I used to tell people to look around the room you’re sitting in. Look at everything. Every noun. It almost all exists because a person willed it into existence. Nothing is stopping you from doing the same. I see people online all the time dismissing someone else’s success because “I had that idea first” or whatever. I mean… yeah? If so then the difference is… you. So a bit of a self own whenever I hear that. Number one tip: act with agency.