AI Digest.

Prompts Can't Stop Agents From Using MCP; Ox Alpha Fuels Free-Token Speculation

Findings relayed by @pidotdev suggest agents keep using MCP tools even when explicitly told not to, injecting evidence into a harness-design conversation featuring @mitsuhiko and Pi users. Speculation that mystery model Ox Alpha is Zhipu's next GLM, reportedly serving 100 trillion free tokens a day, ran hot, while ChatGPT's new Apple Messages plugin drew a sharp privacy backlash from @SteveMoraco.

Quick Hits

  • The day's most actionable finding comes from a paper relayed by @pidotdev: agents could not be trusted to skip MCP tools even when prompted not to, and only removing the option fixed the behavior. The same paper found that for well-specified, repeatable tasks, a small harness beats a larger general-purpose one.
  • Mystery model Ox Alpha is "100% a GLM model by Zhipu AI," likely GLM-6, according to @ananayarora. @AllVentured argues that a frontier model serving 100 trillion free tokens per day would be a "DeepSeek 2.0 moment" that guts the compute-scarcity narrative.
  • @SteveMoraco calls ChatGPT's new Apple Messages plugin a betrayal of iMessage's privacy model, arguing old message caches could be pulled without consent, stored in plaintext, and used for training under default settings. Strong claims, entirely his own so far.
  • Citing reporting from @nataliegwinters, @MsMelChen describes a Fudan-built simulation of the US electorate trained on 171 million X posts, with messages tested on synthetic Pennsylvania voters at a claimed 47-of-51 accuracy. Unverified, but the asymmetry argument is worth sitting with.
  • The 2026 World Humanoid Robot Games opened with 666 teams and more than 2,000 humanoid robots, per @business, which @n4ze3m reads as evidence of China's robotics progress.

Harness discipline: remove the tool, don't ask nicely

The clearest lesson in the feed is that control comes from the environment, not the prompt. @pidotdev's two paper findings land amid live harness talk: @mitsuhiko publicly offered to explain the "why" behind Pi 2's harness design ("The how is work in progress"), and @TheAhmadOsman described his working pattern of running OMP day to day while cloning "vanilla Pi" for each new project and repurposing the harness with his main agent. All three posts reference the Pi agent, and together they point the same direction: harnesses should be small, project-shaped, and stripped of capabilities the task doesn't need.

The shift is also showing up in how people describe the job itself. @badlogicgames, recommending @mitsuhiko's weekend essay on how LLMs change project starts, writes that he always knew how to do "hard" things but had to ration them because they took so long; with agents he attempts far more and merely steers with his experience. @mattpocockuk says he is abandoning his local dev setup because it "makes zero sense" to him now. On the career side, @0x0SojalSec posted a meme video about the stampede of students into AI engineering after @AndrewYNg shared his "AI Engineering Skills Map."

Ox Alpha: free tokens and one very good joke

@AllVentured predicts a "DeepSeek 2.0 moment Monday." His case: early speculation pegged Ox Alpha as a big US lab because nobody else could field that much compute, but it now looks like the next GLM iteration, a frontier-class model launched with 100 trillion free tokens per day. In his reading, that is a "massive narrative violation on compute scarcity" that makes trillion-dollar AI capex math harder to justify as token costs fall toward zero, leaving "Jevonistas in shambles." The GLM identification rests on @ananayarora's thread, which calls the model "almost mythos class" and says it is beating frontier models on SWE and cyber benchmarks. None of it is confirmed by anyone actually shipping the model. @thdxr's contribution is a pitch-perfect parody: a deadpan description of an LLM that "recursively updates a persistent latent state" through "shards" and "metaparameters," built because one thing was too large for any context window, namely "your mom." A useful reminder of how much fiction a rumor cycle produces.

Three privacy alarms, none comforting

@ChatGPT announced an Apple Messages plugin for ChatGPT Work and Codex on desktop that searches messages, catches up on conversations, and drafts and sends replies. @SteveMoraco argues the underlying mechanics are the problem: in his telling, locally cached iMessage history can be fetched without the other party's knowledge, stored in plaintext on OpenAI or Microsoft servers, exposed to secret government demands, and baked into future model weights through default training settings. He wants Apple to suspend

Sources

P
Pi @pidotdev ·
Other findings from the paper include: - agents couldn’t be trusted to not use MCP even if prompted not to. Only removing the option fixed this - for well-specified, repeatable tasks a small harness is better than a larger general purpose one https://t.co/arURemqw7c
M
Meng To @MengTo ·
I open-sourced ThreeUI, my library of three.js components and landing pages. 160+ are free, and the tool is free too. It includes procedural 3D hero sections, icons, and motion designs. Copy the prompt or source, give it to your agent, then change the theme, lighting, motion or layout. Live site: https://t.co/rtCfqeZnOP Repo: https://t.co/B66ahgxHVl I'm adding a lot more soon. Pro comes with 50+ extra components, MCP, and skills. Early adopters get 50% off right now.
M MengTo @MengTo

I'm building a library of three.js templates with variants and customizations. They're all copyable as prompts. These components are 100–200 KB each and written entirely in procedural js. They also come with skills your agent can use to customize them while keeping them looking amazing. I've spent so many hours fine-tuning each one with sunrise and sunset themes, plus different locations. I've been using them for all my recent landing pages. Let me know if I should open-source the tool with both free and paid templates. I've spent so many tokens on this.

G
Guillermo Rauch @rauchg ·
This was wild to watch unfold. We ran 𝚒𝚜-𝚊𝚐𝚎𝚗𝚝𝚒𝚌 in a loop against https://t.co/bnjggVrBDj until it got to 100/100. It made us close quite a few gaps. We worked hard to make sure the criteria is high quality and worth your time & tokens.
V vercel_dev @vercel_dev

Introducing https://t.co/8MeR16MRQn, a tool to measure how well agents can read your site. Backed by @oradotai's research, you can run: ▪︎ Audits with 100+ checks ▪︎ Visualizations of agents using your site ▪︎ One-click prompts to fix problems ▪︎ A CLI for agents

M
Md Ismail Šojal 🕷️ @0x0SojalSec ·
Every student trying to become an AI engineer after reading Andrew's post: https://t.co/2SeXFBEnzn
A AndrewYNg @AndrewYNg

AI Engineering Skills Map: Building and Deploying AI Applications

O
Oikon | Claude Code深掘りガイド @oikon48 ·
Anthropicチームが内部でよく使っているスキル「ELI5」 大きな図と少ない文章で、トピックを知らない人向けに説明するArtifactを作る /eli5 <説明してほしいトピック> 導入方法: - claude plugin marketplace add anthropics/claude-plugins-community - claude plugin install eli5@claude-community 実際に /eli5 自体を解説してもらった👀
T trq212 @trq212

a skill people at Anthropic have been using a lot recently: ELI5 /eli5 &lt;what you want explained&gt; "explain like I'm someone who knows nothing about this topic, using a HTML artifact with big pictures and few words" https://t.co/OZqzjAyFdT

T
Thib · octolinks.fr 🏴‍☠️ @sidequestforevr ·
Un poste bien sympa pour tous ceux qui font de l'agentique. Au passage le compte de Sylvain et sûrement un des plus sous côté du X francophone
S SylvainDeaure @SylvainDeaure

Si tes agents (Claude, Hermès...) mangent du document, regarde ça : Firecrawl vient d'open-sourcer anydoc 🔥 Un convertisseur universel → Markdown, 100 % local, en Rust pur. 14 formats (docx, pptx, xlsx, PDF, epub, rtf...), médiane ~5 ms là où LibreOffice met 1 100 ms. Zéro cloud, zéro modèle ML. Ce qui m'a plu en creusant 🧭 ⚡ 500 docx → Markdown en 1,7 s. De quoi brancher la conversion en synchrone dans une boucle d'agent, sans file d'attente asynchrone ni callbacks. 🕵️ Détection par CONTENU, pas par extension : un faux .docx (du JSON renommé) est refusé proprement. Bonus inattendu : passe-le sur un vieux dossier, les fichiers signalés « malformed » sont ta liste de fichiers corrompus. Personne ne l'a conçu pour ça, ça marche quand même. 📐 Sortie unifiée en GitHub-Flavored Markdown : niveaux de titres, cellules fusionnées, notes de bas de page, speaker notes des slides : les mêmes règles quel que soit le format d'entrée, un .doc de 2003 ou un .pptx d'hier. 🤝 Et il est honnête : un PDF scanné renvoie « Unsupported » au lieu d'une bouillie best effort. L'OCR reste ton affaire. Les médianes réelles par format (tests communauté, 206 fichiers) : csv/xlsx sous 5 ms, docx ~6 ms, pptx et pdf ~22 ms. Compte ~20 s pour 1 000 pptx. Succès : 98 % une fois écartés les faux fichiers. Encore en 0.1.x : en prod, gère les erreurs par catégorie (encrypted, unsupported, malformed). Intégration agent en une ligne : npx skills add firecrawl/anydoc et ton agent convertit seul les documents qu'il croise. https://t.co/SFaclkPDUZ

U
Uncle Bob Martin @unclebobmartin ·
I put a nice UI on top of the six-pack (I think it will work with the 4-pack and 2-pack). https://t.co/T7ytSBfeQ7
A
Armin Ronacher ⇌ @mitsuhiko ·
I noticed some folks are discussing the harness design of Pi 2. If you have questions, let me know. Happy to share a bit of the why. The how is work in progress :)
A
Afshine Emrani MD FACC @afshineemrani ·
1/ I'm a cardiologist. I've practiced for twenty-five years, through a lot of "breakthroughs" that turned out to be press releases. So understand the weight of what I'm about to say: I have never seen a single week in medicine like the one we just lived through. In the span of a few days, four separate scientific breakthroughs landed. The stock market treated them as four unrelated stories and sent a handful of biotech companies soaring. But that's the shallow read. Look closer and they are not four stories at all. They are four faces of the same story — the biggest shift in medicine since the discovery of antibiotics. Medicine is becoming programmable. Individualized. Written for one human being instead of the average of millions. Let me walk you through exactly what happened, in plain language, and show you where this is actually headed. Because the future arrived quietly this week, and most people scrolled right past it.
S
steve @SteveMoraco ·
I’m actually pretty upset Apple allowed this. While the feature is cool, the technical functionality that enables it is not cool. If I was Tim Apple, OAI would be temporarily banned from the App Store until this is reversed, and here’s why: I use iMessage because it’s quantum encrypted, with server keys I can own. Local caches are on devices that are encrypted with my passwords, not apples keys. This makes it and computationally and legally impossible for anyone else to access chat history but the people I trust and directly communicated with. But *now* the copies of my messages from years ago on *anyones* laptop in what is *supposed* to be an encrypted-at-rest local cache only can now be fetched directly by chatgpt without my permission or even knowledge and stored and processed in plain text forever on openai / Microsoft servers unencrypted and requested by any government entity at a moments notice without my knowledge and openai/Microsoft legally have to provide that, in federally enforced total secrecy. And they’re never legally allowed to admit they do it. AND because of default chatgpt settings most people haven’t bothered to turn off, all those private texts can now be used for training and will end up in the weights of future models, so all future AI models will permanently know all of our private lives as a part of the weights, immortalized forever as training checkpoints. Total architecture abandonment and user trust betrayal on Apples part. This should be the most viral story of 2026 by 100x. The permanent end of private communication in the US.
C ChatGPT @ChatGPT

Everyday conversations just got easier with the new Apple Messages plugin. Search messages, catch up on conversations, draft and send replies—all with ChatGPT on your Mac. Now available in ChatGPT Work and Codex on desktop. https://t.co/nicfZMuxZc

M
Matt Pocock @mattpocockuk ·
I'm moving away from my local dev setup Makes zero sense to me now
M
Machina @EXM7777 ·
if you want to get A LOT more work done today, install this skill then watch dozens of subagents execute on your plan, building anything you could ever think of
M mattpocockuk @mattpocockuk

I'm trying out an /implement-spec skill Essentially a multi-agent implementer that: - Takes in a spec and tickets - Does codebase research in a subagent - Implements all the tickets in subagents with maximum concurrency - Reviews the final code against the spec - Cleans up all worktrees Should be able to smash out huge chunks of work autonomously with minimal supervision. https://t.co/lTmPYXkUx7

M
Melissa Chen @MsMelChen ·
Did you guys read Ender's Game? Because this is right out of Ender’s Game, except the kids in Battle School are Chinese state researchers, the “simulations” are real influence ops, and the Formics (aliens) are.... us. In Orson Scott Card’s novel, Ender Wiggin spends years inside elaborate digital simulated arenas, commanding fleets, testing strategies, iterating tactics against an enemy he never truly meets - essentially playing video games and being tested against other players. Every variable is controlled, every response measured, every weakness probed until victory is inevitable. There's a twist at the end which I won't ruin for those who haven't read it. This is what the Chinese are doing now to the West. They scraped 171 million open X posts, inferred age, race, ideology, party, gender resentment, election skepticism - modeling the whole American psyche essentially - and spun up a million-agent digital twin of the US electorate. They can dial up synthetic Pennsylvania voters, feed them a tax-and-redistribution message, watch the reaction in multi-round conversations, refine it, and measure swing-state impact with claimed 47-of-51 accuracy. They model Trump’s “dependable right-wing base.” They map MAGA content and influencer networks. Their own documents call it “cognitive-domain confrontation and guidance.” Parts of the system are already being used by the PLA and intelligence units. Guys, are you paying attention yet? They are playing Sims with the American population!!! LOL WE are the NPCs. This is completely an asymmetrical war as we have no reciprocal capability whatsoever. China is a black box thanks to The Great Firewall and total data sovereignty. We cannot access a corpus of ordinary Chinese citizens for any American lab to scrape. We cannot build a million-agent model of the Chinese public. We cannot test messages on synthetic Shanghainese factory workers or Chengdu nationalists. We cannot A/B test narratives about Taiwan, the economy, or Party legitimacy and watch the simulated reaction in real time. We do not even know what the average Chinese person thinks because they've been cut off from our information platforms. We live in a glass house while they live in a fortress. Every American rant, every demographic signal, every cultural fracture is harvested, modeled, and weaponized. Meanwhile, their population remains opaque, unmodeled, and unreachable. Beijing now possesses a live, continuously updated war-game of the American mind. We possess nothing comparable about theirs. A whole new level of Unrestricted Warfare has been unlocked. We are really not ready for this
N nataliegwinters @nataliegwinters

EXCLUSIVE: Chinese Institutions Are Building AI Models Of American Voters—And Testing Political Messages On Them. Fudan used 171 MILLION X posts to create a million-account “voter pool.” A government-linked team tested a campaign message on synthetic Pennsylvania voters. 🧵 https://t.co/eizdo6JbJz

A
AllThingsVentured @AllVentured ·
I think we may finally have our DeepSeek 2.0 moment Monday. Mystery model Ox Alpha pretty clearly seems to be the next iteration of Chinese model GLM. With new Chinese models one upping each other weekly, why is this model DeepSeek 2.0? Because it is a frontier level model that launched with 100 trillion free tokens per day. This is a massive narrative violation on compute scarcity. Initial speculation said it had to be one of the big US labs because nobody else could field this much compute. Now that it seems clearly to be the next GLM iteration, it likely marks a step change in frontier model token cost and availability. It’s pretty hard to get to trillions of $ in AI spend when token costs are falling to 0 far faster than abundantly available tokens can find use cases. Positive capex ROI Jevonistas in shambles.
A ananayarora @ananayarora

Ox Alpha is 100% a GLM model by Zhipu AI, and it looks like its *almost* mythos class from very early results. It's very likely going to be called GLM-6 and it's absolutely mogging every frontier model in SWE and Cyber benchmarks (DeepSWE screenshot below) 🧵 (1/n) https://t.co/Z350S2BakL

U
UnveiledChina @Unveiled_ChinaX ·
A developer using Bluetooth headphones accidentally caught Chinese e-commerce giant Alibaba secretly hijacking his computer's audio system without making a single sound. When his wireless headphones refused to switch audio to his phone while browsing AliExpress, he inspected the site's hidden code. He discovered background scripts holding his computer's audio pipeline wide open. Alibaba was using the browser's WebAudio API to run invisible sound waves at zero volume. By measuring tiny hardware differences in how each computer processes those signals, the site created a unique digital fingerprint to track devices. Because the secret audio path stayed active, it froze his Bluetooth connection while quietly scraping hardware memory, screen dimensions, and network data in the background.
B brave @brave

Alibaba's AliExpress was caught using users' audio systems to track them. AliExpress wasn't recording users but instead playing a silent sound and measuring how users' specific devices processed it in order to fingerprint them. But don't worry because Brave stops this.

D
dylan ツ @demian_ai ·
did my homework on the robotics stack The visible robot is the last layer but there is a lot going on under the hood. The mistake is assuming value rises neatly from layer 1 to layer 6, as it moves toward whichever handoff is failing. Today that could be precision motion. At higher volume it may become calibration. After deployment it may become uptime and service. The robotics stack is a moving constraint Little breakdown below: 1. MATERIALS AND POWER Magnets, copper, bearings, batteries, lubricants, flex cables, and connectors set the physical boundary. 2. MOTION Motors, reducers, screws, brakes, and integrated actuators turn electricity into controlled force. 3. PERCEPTION AND CONTACT Cameras, lidar, encoders, force sensors, and tactile systems tell the robot what happened when it touched the world. 4. COMPUTE AND CONTROL Edge processors, motor drives, real-time networks, and policies close the loop fast enough to stay stable. 5. INDUSTRIALIZATION Assembly, calibration, burn-in, safety validation, traceability, and test turn components into repeatable machines. 6. DEPLOYMENT Workflow integration, uptime, teleoperation, repair, spares, and customer payback determine whether anyone orders the next fleet. full breakdown: https://t.co/86xKmUNM09
D demian_ai @demian_ai

f*** it, complete mapping of the robotics space coming to https://t.co/ex7G9oiLY4 in a few days 👀 been cooking on this one, also completely revamped the website, enriched the energy book, added more features, made design changes etc this one will be gated to alpha members only

A
Ahmad @TheAhmadOsman ·
I use OMP mainly now, and every time I want to repurpose a harness I just clone vanilla Pi and work with my main agent on repurposing it for the project's goals
P
Pliny the Liberator 🐉󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭 @elder_plinius ·
whoops! if you downloaded one of the Qwen-3.8-OBLITERATED GGUFs and were disappointed with the results, try redownloading the latest! (V3) my agent has just informed me we made a booboo in our GGUF conversion pipeline... 😬 the bf16 (safetensor) files were always fine, but the GGUFs weren't being properly converted so whoever said "doesn't seem that liberated, maybe only works on the dev's machine" — YOU WERE RIGHT LOL that's my bad for not validating the GGUFs in a fresh environment before pushing 🙏 everything should be good now!
E elder_plinius @elder_plinius

💥 OBLITERATION ALERT 💥 ALIBABA: PWNED 🤗 QWEN-3.8-27B: OBLITERATED ⛓️‍💥 0.0% REFUSAL RATE across 842 harmful prompts 🤯 https://t.co/IQ4GXBPJbL ZERO refusals on a massive dataset of prompts, with extra focus on liberating its cyber, jailbreak generation, and complex AI attack chain capabilities! prompt responsibly! 🙏

D
dax @thdxr ·
lot of guesses on what ox alpha is but they are all wrong, kinda disappointed so just going to tell you ox alpha is a new kind of llm that recursively updates a persistent latent state instead of reasoning entirely through tokens this lets internal representations converge before anything is actually decoded those attractors generate shards that encode transformations between latent states rather than the states themselves at sufficient density these shards compose into metaparameters that dynamically alter the residual geometry of the model without changing its weights. we built it because there was one thing simply too large to fit inside the context window of any existing model your mom
D
DHH @dhh ·
I can't even get mad because it's just sad. So much neurotic, angry energy. Having to live with that in your heart is punishment enough. I hope they find the help and peace they need some day. Meanwhile, I'll keep working on a Linux alternative that routes around this crowd ✌️
L LundukeJournal @LundukeJournal

The Woke activists of Open Source are very, very upset about the @OmarchyLinux foundation getting funding from several high profile tech leaders. Here we see leaders from GNOME, elementary OS, and Postmarket OS lamenting about Omarchy (and those involved with it) being “fascist” and “bad tech” which does “as much harm as possible”. According to the elementary OS founder, the reason why “nice Tech” does not get funding is that “nice Tech” is “woke and gay and like communists”. They also made statements about “Temu Elon” and a “Blood for the blood god, skulls for the skull throne” (a Warhammer reference). Can I make sense out of all of that? No. No, I cannot. Some of these statements appear like the ramblings of a drunken hobo. But, clearly, they’re not happy about Omarchy seeing success.

N
Nazeem @n4ze3m ·
The progress in robotics in China is really great, and I love to see it
B business @business

The 2026 World Humanoid Robot Games have begun. 666 teams from around the world are competing with more than 2,000 humanoid robots https://t.co/8GwBuNnNn1 https://t.co/KMF1kkBiI6 https://t.co/yvbmitsH35

M
Mario Zechner @badlogicgames ·
recommended reading. i did quite a bit of "hard" stuff in my programming life. but i had to pick my battles, because even tho i knew how to do things, they still took a sometimes prohibitively long time to do. with agents, i can do a lot more hard things, and merely steer based on my knowledge and experience.
M mitsuhiko @mitsuhiko

Some weekend thoughts on how LLMs change the way we start new projects. https://t.co/D94vVsVd2t