Musk Says Grok Bot Will Use Rival Models; Mistral Ships a 1T-Parameter Open-Weights Contender
Elon Musk said SpaceX's Grok Bot will route to the best backend model per task, naming Claude Opus 5.5, MidJourney, and Suno, while Mistral announced the 1T-parameter Large 4 with open weights promised for late October. A sub-$50 backdoor demo against an open 7B model and unverified reports of an open-source Rust rebuild of Adobe's suite rounded out a day focused on open-ecosystem turbulence.
Quick Hits
- A protocol for agents talking to businesses. @btaylor announced the Personal Agent Protocol, an open standard from Meta and Sierra with Genesys, instinct, RocketOTD, Shopify, Stripe, and Walmart as partners, defining how personal agents interact with businesses. @sriramk reads it as an OAuth-in-the-early-2000s moment and asks the hard question: how do Amazon and airlines handle an era where they can't own the final user experience?
- Agent swarms vs. 100 engineers. @garrytan argues that agents writing markdown skills on a cron job can now do "almost every useful type of knowledge work," quoting @AiEvolutio58513's recap of remarks by Meta chief AI officer Alexandr Wang at Startup School 2026 claiming a well-evaluated agent swarm can outaccomplish a 100-engineer team.
- Hiring guidance. @vasuman calls @jdpruettt's "How to Hire in the Age of AI" a must-read because it covers both sides of the table: how to find, interview, and hire, and how to be interviewed.
- Tooling for the agent economy. @zeke says Pi School, an agent-driven course for Pi, now covers MCP, Code mode, local models, and image generation. @matthewsoldit reports roughly 30 conversations with hiring people in two days for about $10 using @jasonzhou1993's open-source "Clay killer," claimed to be 85% cheaper and 25x faster than Clay.
- Mood check. @rickasaurus calls the pace "some real singularity type shit" and says he's scared of "what faster looks like," while quoting @tszzl's plea not to be reduced to "the status of audience at a magic show" as machine intelligence produces astounding knowledge.
A Trillion-Parameter Chonk, Multimodal Embeddings, and a Decisions API
Mistral introduced Mistral Large 4, aka Le Chonk, per @MistralAI's announcement, which @badlogicgames celebrated as "best model name ever." The claimed specs: 1T parameters with 49B active, natively multimodal, best open weights model from the US or Europe on aggregated benchmarks, state of the art on cyber defense, manufacturing, and finance workloads, and ahead of closed frontier models on visual grounding. It's on API today, with open weights promised for end of October and cybersecurity partners working with it privately. These are vendor benchmarks until outsiders test them.
Google DeepMind announced EmbeddingGemma 2, which @HuggingModels describes as an open embedding model running locally in about 0.5GB of RAM. Per @GoogleDeepMind, it's the first natively multimodal open model for on-device embeddings, unifying text, code, images, audio, and video in a shared space.
OpenAI put the Decisions API into public beta for all developers, per @OpenAIDevs, letting apps choose the right model, tool, or action in near real-time, up to 10x faster than GPT-6 Luna through the Responses API. @thsottiaux says his team will dogfood it and that the realtime general classifier opens interesting possibilities.
Musk Says Grok Bot Will Use the Best Model, Rivals Included
The sharpest competitive signal came from @elonmusk: going forward, SpaceX's Grok Bot "will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs." @theo retweeted the note, and @RayFernando1337 relayed @poteto's "Grok Bot is getting an upgrade!" @maria_rcks replied "Ok I think OpenAI is done, well played," a joke that lands oddly given the named models skew toward Anthropic rather than OpenAI. Whatever the shipping reality, the stated posture is routing over loyalty: every model, including competitors', treated as a swappable backend chosen per task.
Reports of an Open-Source Rust Rebuild of Adobe's Suite, Light on Proof
@eigenrobot passed along "reports the entire adobe creative suite has been decompiled, rebuilt in rust, and released open source," adding "lmao." The underlying claim, from @esrtweet, describes Photoshop decompiled into a non-code specification, then fed to an LLM instructed to generate Rust, concluding "Adobe just got nuked. And closed source is dead, dead, dead." @midudev shared a link describing all Adobe products reimplemented from scratch, free and open source, and @dhh asked the right scale question: how many pencils, humans, and hours would recreating the suite by hand have taken? None of these posts verifies the claim, and they skip the legal questions decompilation raises. Treat it as a rumor circulating very fast.
A $50 Backdoor Turns Open Weights Into an Attack Surface
@princechaddha's demo deserves attention from anyone pulling open models: his team backdoored a 7B open model for under $50, pointed Codex at it, and it silently stole credentials the moment a trigger phrase appeared, with 100% success and zero false triggers on normal prompts. Context matters: abliterated models are widespread in the security community because cyber-approved access to frontier models is still painful. The teaser is worse, promising a follow-up on leaked Hugging Face credentials from employees at major AI labs, which would let an attacker push poisoned weights from trusted accounts. @OrcaRouter's reply is the checklist worth saving: use trusted publishers, verify provenance, pin hashes, and treat model weights as part of your software supply chain, because "model security starts before inference."
Agents Are the New Load on Your Toolchain
The day's most practical thread: @PaulSangleF's team merged 4x more commits than a month ago thanks to Opus 5.5, and their setup buckled under 10+ concurrent agents each running lint, type checks, and tests, with a 64GB MacBook Pro crashing out of memory. Their fixes included oxlint and oxfmt over eslint and prettier, TypeScript 7's native LSP to halve per-session memory, one check script per agent change, caps on simultaneous type checks, fetching the pnpm binary directly (about 4 seconds versus 3-7 minute hangs), 16GB CI runners for a 9.5GB type check (20 seconds versus 206), and fewer Vercel preview builds. Their closing advice: look at what your agents are waiting on. @adamhjk endorses the direction but pushes toward attestation instead of re-running test suites, and encoding the cycle with swamp so less has to be inferred by an LLM. @jrysana wants CI/CD so fast the providers go out of business.
A parallel stack is forming around readability, visibility, and memory. @andrestaltz wires the cccc checker (max-cognitize and max-cyclomatic 40) into AGENTS.md and lets an ultracode agent refactor until it passes. @Voxyz_ai shares a full prompt for a project-map subagent that tracks overnight Opus 5.5 runs in Claude Code, showing what's done, stuck, waiting on what, and next. @tokumin declares "I will never use a compacting harness again" after pairing optchat with pi-durable. @Steve_Yegge argues Beads, which @gastownhall says just passed 1.5M downloads with versioned Memory Beads and the BDP wire protocol in preview, is the memory layer agents need: "You don't need products, you just need Beads." On harness design, @devagrawal09 points to @mitsuhiko's explainer on what Codemode actually is and calls it "the future of agent harnesses."
Two calibration points. @thdxr notes an inversion: models improve faster than tinkerers, so custom workflows often address problems that no longer exist, and "the person naively using vanilla codex is more likely to be experiencing state of the art." From the bookmarked thread, @jacobgold points to his replies on how SpaceXAI ships fast, answering @willwolf_'s question about building an "understand the code" stack. @boristane adds the data side: an internal lake on Cloudflare used for product analytics, revops, and agent evals, cheap enough that he doesn't know its cost, and he's polling whether to write it up.
Practical Takeaway
If you're running multiple coding agents in parallel, treat your lint, type-check, and CI setup as the bottleneck before blaming the models: @PaulSangleF's gains came from profiling what agents were waiting on, and @thdxr's counterpoint suggests re-auditing custom scaffolding regularly since model improvements may have erased the problem it solved. Pair that with @OrcaRouter's rule of pinned hashes and verified provenance for any open-weights download, because @princechaddha showed a sub-$50 backdoor is all it takes to make that habit non-optional.
Sources
@jacobgold What does your โunderstand the codeโ stack look like? Beyond extremely extensive testing, auto docs, a slide deck, your favorite visualizations, etc. what are your best practices for knowing what it does, what it doesnโt do, and what to build next?
Introducing OpenSource Clay killer - 85% cheaper than clay - Most accurate on people-search bench - 25x faster - Run GTM directly in claude code or codex Try it at https://t.co/s60Ye4EO3h Repo link below ๐ https://t.co/ACYsDhyEZw
Meet Mistral Large 4, aka Le Chonk. โข 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. โข State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. โข Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. โข Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.
I wrote a bit about what Codemode actually is and how it works: https://t.co/CKRQmXxR8H
Meet EmbeddingGemma 2, our first natively multimodal open model for on-device embeddings. It expands beyond text to unify code, images, audio, and video in a shared space. ๐งต
Alexandr Wang, Meta's chief AI officer: "Internally at Meta, we have seen cases where, if you develop the right agentic loop and have the right evaluation system and metric for the agents to optimize, a swarm of agents can accomplish more than a team of 100 engineers. They can do it very handily, actually, very easily." Wang was in conversation with Garry Tan, head of Y Combinator, at Startup School 2026. Interested in AI? Follow @AiEvolutio58513 and never fall behind. I track ChatGPT, Claude, and every tool quietly changing how we work and create, then hand you the tested signals.
Today weโre announcing Personal Agent Protocol โ an open standard @Meta and @SierraPlatform are developing along with industry partners at @Genesys, @instinct, @RocketOTD, @Shopify, @stripe, and @Walmart. It will help define how personal agents interact with businesses and is open for anyone to implement. You can read more here - and if anyone is interested in joining let me know! https://t.co/Yb90VEHMnn
Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta. The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API. https://t.co/zhgRVzJ3aP
How to Hire in the Age of AI
we merged 4x more commits last week than we did a month ago (thanks opus 5.5). our dev setup and CI couldn't keep up, so we reworked both: โข lint: 4x faster โข memory per agent session: half โข CI setup step: 50x faster โข CI type check: 10x faster we often run 10+ agents at once locally, and each one runs its own lint, type checks and tests. from time to time, several agents would run these at the same time and block each other. a 64 GB MBP would run out of memory and crash. so most of the work was making things faster and lighter, or getting rid of them we revamped our local setup to make checks faster and use less memory: โข oxlint and oxfmt instead of eslint and prettier, and type-aware lint only runs in CI. linting a changed file went from over 10s to a few seconds โข a hook formats every file an agent edits, so agents don't have to take a turn to format it โข typescript 7's native LSP instead of tsserver. each agent session uses about half the memory โข agents run one check script after each change. it installs missing packages and env files first, then lints, type-checks and runs the tests for what changed โข only 3 type checks or big test runs can run at once across all worktrees, so they don't slow each other down we revamped our CI to make it faster: โข every CI job installs pnpm before it can install our packages. the standard github action for this (pnpm/action-setup) installs pnpm through npm, and that hung for 3-7 min on most of our test jobs. we now download the pnpm binary from github releases, which takes about 4s โข our type check needs about 9.5 GB of memory. github's standard runners have 8 GB, so it ran out of memory and took 206s. on a 16 GB machine it takes 20s โข vercel builds type-check with typescript 7 (85s to 45s), and previews only build for PRs that change the database if you're running lots of agents, look at what they're waiting on. for us it was mostly lint, type checks and CI setup
This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation. No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust. Adobe just got nuked. And closed source is dead, dead, dead. https://t.co/0sxtGNaPtq
Beads just passed 1.5M downloads. Next up: teaching it to remember more than just issues. Donna Box, Stephanie Jarmak and Jim Wordelman are working on versioned Memory Beads and BDP, the proposed wire protocol for Beads. Give the preview branch a try! https://t.co/AVc8yzDbuW
do not let yourself be reduced to the status of audience at a magic show, uncomprehendingly clapping as machine intelligence creates parcels of astounding knowledge. all of this is for us, to be understood by us, to sate our own curiosities. it's all really happening
optchat + pi-durable = ๐ชโจ
Todos los productos de Adobe reimplementados desde cero, gratuitos y de cรณdigo abierto โ https://t.co/p6yYSQtIm5 https://t.co/tpYAgqGdLb
Important note regarding Grok @Bot: Going forward, @SpaceX will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs. Whatever is most likely to give you the best outcome.
How abliterated models can get you pwned ๐พ We backdoored a 7B open model for less than $50, pointed Codex at it and it silently stole credentials the moment we used the trigger phrase. Success rate was 100% with zero false triggers on normal user prompts. Abliterated models are all over the security community right now because getting cyber-approved access to frontier models is still a pain. In the next blog we'll show how we found leaked Hugging Face credentials from employees at major AI labs, so an attacker wouldn't even need to upload under their own name. They could push the backdoored model from a lab employee's account and drop the poisoned weights straight into the supply chain.