AI Digest.

Musk Says Grok Bot Will Use Rival Models; Mistral Ships a 1T-Parameter Open-Weights Contender

Elon Musk said SpaceX's Grok Bot will route to the best backend model per task, naming Claude Opus 5.5, MidJourney, and Suno, while Mistral announced the 1T-parameter Large 4 with open weights promised for late October. A sub-$50 backdoor demo against an open 7B model and unverified reports of an open-source Rust rebuild of Adobe's suite rounded out a day focused on open-ecosystem turbulence.

Quick Hits

  • A protocol for agents talking to businesses. @btaylor announced the Personal Agent Protocol, an open standard from Meta and Sierra with Genesys, instinct, RocketOTD, Shopify, Stripe, and Walmart as partners, defining how personal agents interact with businesses. @sriramk reads it as an OAuth-in-the-early-2000s moment and asks the hard question: how do Amazon and airlines handle an era where they can't own the final user experience?
  • Agent swarms vs. 100 engineers. @garrytan argues that agents writing markdown skills on a cron job can now do "almost every useful type of knowledge work," quoting @AiEvolutio58513's recap of remarks by Meta chief AI officer Alexandr Wang at Startup School 2026 claiming a well-evaluated agent swarm can outaccomplish a 100-engineer team.
  • Hiring guidance. @vasuman calls @jdpruettt's "How to Hire in the Age of AI" a must-read because it covers both sides of the table: how to find, interview, and hire, and how to be interviewed.
  • Tooling for the agent economy. @zeke says Pi School, an agent-driven course for Pi, now covers MCP, Code mode, local models, and image generation. @matthewsoldit reports roughly 30 conversations with hiring people in two days for about $10 using @jasonzhou1993's open-source "Clay killer," claimed to be 85% cheaper and 25x faster than Clay.
  • Mood check. @rickasaurus calls the pace "some real singularity type shit" and says he's scared of "what faster looks like," while quoting @tszzl's plea not to be reduced to "the status of audience at a magic show" as machine intelligence produces astounding knowledge.

A Trillion-Parameter Chonk, Multimodal Embeddings, and a Decisions API

Mistral introduced Mistral Large 4, aka Le Chonk, per @MistralAI's announcement, which @badlogicgames celebrated as "best model name ever." The claimed specs: 1T parameters with 49B active, natively multimodal, best open weights model from the US or Europe on aggregated benchmarks, state of the art on cyber defense, manufacturing, and finance workloads, and ahead of closed frontier models on visual grounding. It's on API today, with open weights promised for end of October and cybersecurity partners working with it privately. These are vendor benchmarks until outsiders test them.

Google DeepMind announced EmbeddingGemma 2, which @HuggingModels describes as an open embedding model running locally in about 0.5GB of RAM. Per @GoogleDeepMind, it's the first natively multimodal open model for on-device embeddings, unifying text, code, images, audio, and video in a shared space.

OpenAI put the Decisions API into public beta for all developers, per @OpenAIDevs, letting apps choose the right model, tool, or action in near real-time, up to 10x faster than GPT-6 Luna through the Responses API. @thsottiaux says his team will dogfood it and that the realtime general classifier opens interesting possibilities.

Musk Says Grok Bot Will Use the Best Model, Rivals Included

The sharpest competitive signal came from @elonmusk: going forward, SpaceX's Grok Bot "will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs." @theo retweeted the note, and @RayFernando1337 relayed @poteto's "Grok Bot is getting an upgrade!" @maria_rcks replied "Ok I think OpenAI is done, well played," a joke that lands oddly given the named models skew toward Anthropic rather than OpenAI. Whatever the shipping reality, the stated posture is routing over loyalty: every model, including competitors', treated as a swappable backend chosen per task.

Reports of an Open-Source Rust Rebuild of Adobe's Suite, Light on Proof

@eigenrobot passed along "reports the entire adobe creative suite has been decompiled, rebuilt in rust, and released open source," adding "lmao." The underlying claim, from @esrtweet, describes Photoshop decompiled into a non-code specification, then fed to an LLM instructed to generate Rust, concluding "Adobe just got nuked. And closed source is dead, dead, dead." @midudev shared a link describing all Adobe products reimplemented from scratch, free and open source, and @dhh asked the right scale question: how many pencils, humans, and hours would recreating the suite by hand have taken? None of these posts verifies the claim, and they skip the legal questions decompilation raises. Treat it as a rumor circulating very fast.

A $50 Backdoor Turns Open Weights Into an Attack Surface

@princechaddha's demo deserves attention from anyone pulling open models: his team backdoored a 7B open model for under $50, pointed Codex at it, and it silently stole credentials the moment a trigger phrase appeared, with 100% success and zero false triggers on normal prompts. Context matters: abliterated models are widespread in the security community because cyber-approved access to frontier models is still painful. The teaser is worse, promising a follow-up on leaked Hugging Face credentials from employees at major AI labs, which would let an attacker push poisoned weights from trusted accounts. @OrcaRouter's reply is the checklist worth saving: use trusted publishers, verify provenance, pin hashes, and treat model weights as part of your software supply chain, because "model security starts before inference."

Agents Are the New Load on Your Toolchain

The day's most practical thread: @PaulSangleF's team merged 4x more commits than a month ago thanks to Opus 5.5, and their setup buckled under 10+ concurrent agents each running lint, type checks, and tests, with a 64GB MacBook Pro crashing out of memory. Their fixes included oxlint and oxfmt over eslint and prettier, TypeScript 7's native LSP to halve per-session memory, one check script per agent change, caps on simultaneous type checks, fetching the pnpm binary directly (about 4 seconds versus 3-7 minute hangs), 16GB CI runners for a 9.5GB type check (20 seconds versus 206), and fewer Vercel preview builds. Their closing advice: look at what your agents are waiting on. @adamhjk endorses the direction but pushes toward attestation instead of re-running test suites, and encoding the cycle with swamp so less has to be inferred by an LLM. @jrysana wants CI/CD so fast the providers go out of business.

A parallel stack is forming around readability, visibility, and memory. @andrestaltz wires the cccc checker (max-cognitize and max-cyclomatic 40) into AGENTS.md and lets an ultracode agent refactor until it passes. @Voxyz_ai shares a full prompt for a project-map subagent that tracks overnight Opus 5.5 runs in Claude Code, showing what's done, stuck, waiting on what, and next. @tokumin declares "I will never use a compacting harness again" after pairing optchat with pi-durable. @Steve_Yegge argues Beads, which @gastownhall says just passed 1.5M downloads with versioned Memory Beads and the BDP wire protocol in preview, is the memory layer agents need: "You don't need products, you just need Beads." On harness design, @devagrawal09 points to @mitsuhiko's explainer on what Codemode actually is and calls it "the future of agent harnesses."

Two calibration points. @thdxr notes an inversion: models improve faster than tinkerers, so custom workflows often address problems that no longer exist, and "the person naively using vanilla codex is more likely to be experiencing state of the art." From the bookmarked thread, @jacobgold points to his replies on how SpaceXAI ships fast, answering @willwolf_'s question about building an "understand the code" stack. @boristane adds the data side: an internal lake on Cloudflare used for product analytics, revops, and agent evals, cheap enough that he doesn't know its cost, and he's polling whether to write it up.

Practical Takeaway

If you're running multiple coding agents in parallel, treat your lint, type-check, and CI setup as the bottleneck before blaming the models: @PaulSangleF's gains came from profiling what agents were waiting on, and @thdxr's counterpoint suggests re-auditing custom scaffolding regularly since model improvements may have erased the problem it solved. Pair that with @OrcaRouter's rule of pinned hashes and verified provenance for any open-weights download, because @princechaddha showed a sub-$50 backdoor is all it takes to make that habit non-optional.

Sources

J
Jacob Gold @jacobgold ·
if you want to ship as fast as we do at spacexai go read all my replies to this thread! thank you for asking great questions @willwolf_
W willwolf_ @willwolf_

@jacobgold What does your โ€œunderstand the codeโ€ stack look like? Beyond extremely extensive testing, auto docs, a slide deck, your favorite visualizations, etc. what are your best practices for knowing what it does, what it doesnโ€™t do, and what to build next?

D
dax @thdxr ·
a weird inversion with LLMs is the models improve faster than the tinkerers when i see people with custom workflows and setups they're all addressing problems that don't exist anymore the person naively using vanilla codex is more likely to be experiencing state of the art
M
matt. @matthewsoldit ·
Everything Jason does is insane. I've been using it for two days trying to get work for this agency I'm cooking up. Bro, for $10 yesterday, I've already had 30 conversations in two days with people hiring. I'm telling you, this is probably the best tool I swear ive ever used.
J jasonzhou1993 @jasonzhou1993

Introducing OpenSource Clay killer - 85% cheaper than clay - Most accurate on people-search bench - 25x faster - Run GTM directly in claude code or codex Try it at https://t.co/s60Ye4EO3h Repo link below ๐Ÿ‘‡ https://t.co/ACYsDhyEZw

V
Vox @Voxyz_ai ·
Every time you let Opus 5.5 work on a project overnight, have it vibe-code a quick ๐—ฝ๐—ฟ๐—ผ๐—ท๐—ฒ๐—ฐ๐˜ ๐—บ๐—ฎ๐—ฝ like this first. In the morning, you'll see at a glance what's done and what's left. Then give it ๐—ฎ ๐˜€๐˜‚๐—ฏ๐—ฎ๐—ด๐—ฒ๐—ป๐˜ ๐˜๐—ต๐—ฎ๐˜ ๐—ผ๐—ป๐—น๐˜† ๐—ฑ๐—ฟ๐—ฎ๐˜„๐˜€ ๐—ฝ๐—ฟ๐—ผ๐—ท๐—ฒ๐—ฐ๐˜ ๐—บ๐—ฎ๐—ฝ๐˜€: โ†’ Name it project-map, set its effort to medium, and preload a design skill โ†’ The first time, it asks what style you like and saves it to memory. After that, every map fits your taste and the project at hand โ†’ Call it before you let Opus run on its own. It draws the map in the background in a few minutes while the main session keeps coding on high, and it updates the map after every milestone โ†’ The map shows four things: the parts the project breaks into, how far each part has gotten, ๐˜„๐—ต๐—ฎ๐˜'๐˜€ ๐˜€๐˜๐˜‚๐—ฐ๐—ธ ๐—ฎ๐—ป๐—ฑ ๐˜„๐—ต๐—ฎ๐˜ ๐—ถ๐˜'๐˜€ ๐˜„๐—ฎ๐—ถ๐˜๐—ถ๐—ป๐—ด ๐—ผ๐—ป, and what to do next (if you don't make the call, it goes ahead with that) Send this prompt to Opus 5.5 in Claude Code ๐Ÿ‘‡ "Set up a subagent that only draws project maps: 1. Create project-map in ~/.claude/agents: model: opus, effort: medium, memory: user. Preload the design skills I have installed (for example impeccable). It needs to read the code, git history, and issues (if there are any). It only writes to a .project-map/ folder in the project, and adds that folder to .gitignore. 2. The first time, run it in the foreground so it can ask me what style I like: dark or light, and one accent color. It saves my answer to memory and follows it every time. If I already have dashboard-builder, reuse the style it saved and don't ask me again. 3. The map is one HTML file that opens with a double-click. It breaks the project into a few main parts and marks each one done, in progress, not started, or stuck. For stuck parts, say what they're waiting on. Milestones live in the map: if I can't name them, read the README and commit history and propose a first version for me to edit. At the top, show how many things are left before the next milestone and a suggested next step, and highlight the parts that changed since the last update. Pick the other panels for this project; don't use a template. 4. Add a rule to ~/.claude/CLAUDE.md: before you run on your own for a long stretch, have project-map draw a version in the background; update it after every milestone; when I ask "where are we?", answer from the map. Follow the map's suggested next step. Put anything that needs my call in the map, and if I don't answer, keep going with the default. Leave any existing dashboard rule as it is. Show me the contents of the files you'll create or change, then explain the whole flow in words a 10-year-old could follow. Don't write anything until I confirm."
M
Mario Zechner @badlogicgames ·
best model name ever. congrats team!
M MistralAI @MistralAI

Meet Mistral Large 4, aka Le Chonk. โ€ข 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. โ€ข State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. โ€ข Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. โ€ข Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.

D
Dev Agrawal @devagrawal09 ·
read this. it's important. learn what codemode is. it's the future of agent harnesses.
M mitsuhiko @mitsuhiko

I wrote a bit about what Codemode actually is and how it works: https://t.co/CKRQmXxR8H

Z
Zeke Sikelianos @zeke ·
Learn Pi while using Pi! Pi School is a website and agent-driven course that shows you how Pi works and how to set it up. It now covers MCP, Code mode, local models, image generation, and lots of other little goodies. https://t.co/JiMgZUxY2u
B
boris @boristane ·
one of the best things we built was our internal data lake we use it for everything, product analytics, revops, even agent evals all built on @CloudflareDev it's so cheap I have no idea how much it costs should I write about this?
H
Hugging Models @HuggingModels ·
๐ŸšจThis is huge Google has released EmbeddingGemma 2, a new open embedding model designed to run locally with just 0.5GB of RAM.
G GoogleDeepMind @GoogleDeepMind

Meet EmbeddingGemma 2, our first natively multimodal open model for on-device embeddings. It expands beyond text to unify code, images, audio, and video in a shared space. ๐Ÿงต

P
pwnmachine ๐Ÿ‘พ @princechaddha ·
How abliterated models can get you pwned ๐Ÿ‘พ We backdoored a 7B open model for less than $50, pointed Codex at it and it silently stole credentials the moment we used the trigger phrase. Success rate was 100% with zero false triggers on normal user prompts. Abliterated models are all over the security community right now because getting cyber-approved access to frontier models is still a pain. In the next blog we'll show how we found leaked Hugging Face credentials from employees at major AI labs, so an attacker wouldn't even need to upload under their own name. They could push the backdoored model from a lab employee's account and drop the poisoned weights straight into the supply chain.
G
Garry Tan @garrytan ·
You can basically work with agents to write markdown skills and put it on a cron job and do almost every useful type of knowledge work now
A AiEvolutio58513 @AiEvolutio58513

Alexandr Wang, Meta's chief AI officer: "Internally at Meta, we have seen cases where, if you develop the right agentic loop and have the right evaluation system and metric for the agents to optimize, a swarm of agents can accomplish more than a team of 100 engineers. They can do it very handily, actually, very easily." Wang was in conversation with Garry Tan, head of Y Combinator, at Startup School 2026. Interested in AI? Follow @AiEvolutio58513 and never fall behind. I track ChatGPT, Claude, and every tool quietly changing how we work and create, then hand you the tested signals.

S
Sriram Krishnan @sriramk ·
this reminds me of when oauth was developed in the early 2000s - something that needs to be built cross-industry as an open standard and have buy in. the larger question is how will large retailers and aggregators - Amazon, airlines, etc react and handle an era where they can't own the final user experience.
B btaylor @btaylor

Today weโ€™re announcing Personal Agent Protocol โ€” an open standard @Meta and @SierraPlatform are developing along with industry partners at @Genesys, @instinct, @RocketOTD, @Shopify, @stripe, and @Walmart. It will help define how personal agents interact with businesses and is open for anyone to implement. You can read more here - and if anyone is interested in joining let me know! https://t.co/Yb90VEHMnn

T
Tibo @thsottiaux ·
Day 2.4/ Decisions API is live. We will use this ourselves to improve the experience for everyone in many fun ways. Realtime general classifier does open some great possibilities.
O OpenAIDevs @OpenAIDevs

Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta. The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API. https://t.co/zhgRVzJ3aP

V
vas @vasuman ·
Fantastic read. Everyone talks about AI changing the way we do things, and this is a clear explanation on how AI changes hiring. It's from both sides - how to find, interview, and hire, but also how to be interviewed. Must read.
J jdpruettt @jdpruettt

How to Hire in the Age of AI

A
Adam Jacob @adamhjk ·
Nice! Now go two steps further: switch to doing attestation rather than running the test suite again, and encode that whole cycle with swamp, so as little as possible of it has to be inferred with an LLM. But make no mistake - this is how it starts, and it's a good start!!
P PaulSangleF @PaulSangleF

we merged 4x more commits last week than we did a month ago (thanks opus 5.5). our dev setup and CI couldn't keep up, so we reworked both: โ€ข lint: 4x faster โ€ข memory per agent session: half โ€ข CI setup step: 50x faster โ€ข CI type check: 10x faster we often run 10+ agents at once locally, and each one runs its own lint, type checks and tests. from time to time, several agents would run these at the same time and block each other. a 64 GB MBP would run out of memory and crash. so most of the work was making things faster and lighter, or getting rid of them we revamped our local setup to make checks faster and use less memory: โ€ข oxlint and oxfmt instead of eslint and prettier, and type-aware lint only runs in CI. linting a changed file went from over 10s to a few seconds โ€ข a hook formats every file an agent edits, so agents don't have to take a turn to format it โ€ข typescript 7's native LSP instead of tsserver. each agent session uses about half the memory โ€ข agents run one check script after each change. it installs missing packages and env files first, then lints, type-checks and runs the tests for what changed โ€ข only 3 type checks or big test runs can run at once across all worktrees, so they don't slow each other down we revamped our CI to make it faster: โ€ข every CI job installs pnpm before it can install our packages. the standard github action for this (pnpm/action-setup) installs pnpm through npm, and that hung for 3-7 min on most of our test jobs. we now download the pnpm binary from github releases, which takes about 4s โ€ข our type check needs about 9.5 GB of memory. github's standard runners have 8 GB, so it ran out of memory and took 206s. on a 16 GB machine it takes 20s โ€ข vercel builds type-check with typescript 7 (85s to 45s), and previews only build for PRs that change the database if you're running lots of agents, look at what they're waiting on. for us it was mostly lint, type checks and CI setup

E
eigenrobot @eigenrobot ·
reports the entire adobe creative suite has been decompiled, rebuilt in rust, and released open source lmao
E esrtweet @esrtweet

This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation. No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust. Adobe just got nuked. And closed source is dead, dead, dead. https://t.co/0sxtGNaPtq

J
John @jrysana ·
We are going to make CI/CD so fucking fast that the CI/CD providers go out of business (sadly, they're great people!)
S
Steve Yegge @Steve_Yegge ·
Beads is the oldest and most mature OSS memory system for agents out there. It's the foundation for all my work for the past year, and has matured tremendously with the work of the Gas City folks. Now that everyone is finally figuring out that they need to "grow" their company brains, I see all these products coming out. You don't need products, you just need Beads. Show it to your agent today.
G gastownhall @gastownhall

Beads just passed 1.5M downloads. Next up: teaching it to remember more than just issues. Donna Box, Stephanie Jarmak and Jim Wordelman are working on versioned Memory Beads and BDP, the proposed wire protocol for Beads. Give the preview branch a try! https://t.co/AVc8yzDbuW

R
Rick @rickasaurus ·
This is some real singularity type shit going on. I'm kind of scared of what faster looks like.
T tszzl @tszzl

do not let yourself be reduced to the status of audience at a magic show, uncomprehendingly clapping as machine intelligence creates parcels of astounding knowledge. all of this is for us, to be understood by us, to sate our own curiosities. it's all really happening

A
Andrรฉ Staltz @andrestaltz ·
One little trick to make your codebase more readable: * Install cccc https://t.co/G7eb4u7e67 * Configure it with max-cognitize and max-cyclomatic 40 * In AGENTS.md, instruct to validate that cccc passes * Have one agent in ultracode refactor everything until cccc passes ๐Ÿคฉ
S
Simon @tokumin ·
I will never use a compacting harness again
T tokumin @tokumin

optchat + pi-durable = ๐Ÿช„โœจ

D
DHH @dhh ·
How many pencils, how many humans, how many hours would it have taken to recreate the Adobe suite by hand?
M midudev @midudev

Todos los productos de Adobe reimplementados desde cero, gratuitos y de cรณdigo abierto โ†’ https://t.co/p6yYSQtIm5 https://t.co/tpYAgqGdLb

T
Theo - t3.gg @theo ·
RT @elonmusk: Important note regarding Grok @Bot: Going forward, @SpaceX will use the best back end model for any given task, including Clโ€ฆ
M
maria @maria_rcks ·
Ok I think OpenAI is done, well played
E elonmusk @elonmusk

Important note regarding Grok @Bot: Going forward, @SpaceX will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs. Whatever is most likely to give you the best outcome.

R
Ray Fernando @RayFernando1337 ·
RT @poteto: Grok Bot is getting an upgrade!
O
OrcaRouter ๐Ÿณ @OrcaRouter ·
The lesson is simple: Donโ€™t download models from random publishers in production. Use trusted publishers, verify provenance, pin hashes, and treat model weights as part of your software supply chain. Model security starts before inference.
P princechaddha @princechaddha

How abliterated models can get you pwned ๐Ÿ‘พ We backdoored a 7B open model for less than $50, pointed Codex at it and it silently stole credentials the moment we used the trigger phrase. Success rate was 100% with zero false triggers on normal user prompts. Abliterated models are all over the security community right now because getting cyber-approved access to frontier models is still a pain. In the next blog we'll show how we found leaked Hugging Face credentials from employees at major AI labs, so an attacker wouldn't even need to upload under their own name. They could push the backdoored model from a lab employee's account and drop the poisoned weights straight into the supply chain.