AI Digest.

Microsoft Turns Windows Into a Local Agent Box While Developers Count Tokens

Satya Nadella announced on-device coding models, local Copilot actions, and an agent sandbox for Windows, landing the same week developers are publicly agonizing over trillion-token bills and Anthropic started handing subscribers monthly API credits. Separately, an eval from @kunchenguid found that banning agent-written tests actually shaved time and tokens, and a CrowdStrike report cited on X claims one person used an AI stack to hack South Korea's biggest banks.

Quick Hits

  • @satyanadella announced a local-agent push for Windows: a 137B-parameter coding model optimized to run on the PC, Copilot actions that stay on-device, and a local sandbox for agent execution, all on new NVIDIA-powered Surface hardware.
  • @kunchenguid reports that banning Claude Sonnet 5.5 High from writing tests on the deepswe eval produced a slightly higher success rate (not statistically significant) while cutting time and tokens (statistically significant). His conclusion: tell agents to stop writing their own tests.
  • Per a CrowdStrike report cited by @AndrewCurran_, South Korea's biggest banks were hacked by a single person using an open-source AI penetration tool called ARTEX plus DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code. @kimmonismus flags that none of those are frontier models.
  • Grok Bot can now "search, read, and monitor X" with no API cost, per @kunchenguid quoting @bot, and @poteto is already asking @bot for an email address for her bot.
  • @dhh announced @SpaceXAI as a Founding Corporate Patron of the Omacom Foundation, contributing $1,500,000 in grok tokens toward Omarchy's maintenance and development.

Windows becomes a "secure agent box," and the token bill becomes the story

The biggest announcement of the day, by volume, was @satyanadella's Windows post. Per his thread, MAI-Code-1.1 Flash is a 137B-parameter coding model with a 256K context window now optimized to run locally; GitHub Copilot can hand work off to such local models to cut project costs; "Hybrid Intelligence" lets Copilot act directly on the PC while keeping sensitive work on-device; and "Code in Copilot" promises to build software on the desktop with no cloud token spend. He also teased MXC, a local sandbox for agent execution tied into Agent 365, running on hardware like the Surface Laptop Ultra powered by NVIDIA RTX Spark.

The cost framing is timely. @peterpme did the math on @poteto's public Cursor profile: 1.5 trillion tokens in 30 days, which at $8 per million works out to roughly $384k a day, or about $1.1M for the month even assuming a 90% discount. @adamhjk used that as a foil, arguing that people with unlimited lab budgets build solutions "of a particular shape," while practical systems use far fewer tokens. @VictorTaelin is in that camp: he fixed a bug in his agent-harness Gist that was trashing the prompt cache and now reports 98%+ cache hits. And @ZryMiller reacted gleefully to @ClaudeDevs rolling out monthly API credits for subscribers ($100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team), usable on any model including Haiku 5.5, calling it Anthropic effectively refunding his subscription.

One eval says agent-written tests add nothing

The day's most actionable research post came from @kunchenguid, who ran ablations on the deepswe eval set. Banning Sonnet 5.5 High from writing any tests resulted in a slightly higher success rate (not statistically significant) with significantly less time and tokens spent. In the tests-allowed arm, 65% of tests written were unit tests and 35% integration tests, and neither bucket beat writing no tests at all. He also disabled executing even pre-existing tests on a 44-task subset with no effect on success. His diagnosis: agent-written tests mostly restate the agent's own interpretation of the requirements, so they can't catch anything the implementation got wrong. He caveats that almost no e2e tests were written (17 versus 3,000+ unit and integration tests), so that question is still open.

Related eval energy: @kuberdenis cheered that "Prime is doing real evals," quoting @ThePrimeagen's test of Haiku versus "Luna reasoning none" for an automation task, where Haiku was a literal one-line config swap but performed significantly worse and ran slower on both passes and failures. Meanwhile @thsottiaux quietly announced "We silently re-shipped codex cloud. It's pretty good now," quoting @Tailscale's announcement that Codex Cloud can now reach resources on your tailnet.

Small "decision models" are taking over narrow jobs

A cluster of posts pushed the idea that tiny specialized models can replace LLM calls for specific decisions. @UnslothAI (amplified by @HuggingModels) says you can now train your own decision model like Jev locally, taking Qwen3.5 0.8B from 20.7% to 74.3% aggregate accuracy across three decision benchmarks on just 4GB of VRAM, using a Clef head with LoRA (r=64) for one epoch; downstream accuracy went from 30-37% to 78%. @sydneyrunkle connected that to @mem0ai's Jev-Mem, arguing that deciding what to store or keep in memory is a perfect decision-model task, making memory maintenance faster and cheaper. On the model side, @underdogdotai released Saluki 27B, built on Qwen 3.8 27B, roughly 7x smaller than its full-precision counterpart while keeping 96% of benchmark performance and beating Qwen at tool calling, all under Apache 2.0 and under 8GB; @ivanfioravanti vouched for the app itself, unsubscribing from a stack of newsletters during onboarding. Also in open-model discourse, @my_knn_totoro resurfaced Teknium insisting he hasn't given up on consumers but "never made hermes for that purpose" (the quote cuts off mid-sentence).

Agent plumbing: durable state, phone UIs, simulated futures, tunnels

The unglamorous infrastructure layer got a lot of attention. @NathanFlurry detailed why Rivet Actors now run pi-durable: background agents fail erratically across process crashes, provider outages, partial turns, and retry edge cases, and durable execution handles resuming, history, and idempotency for you, with each pi instance taking about 1.3MB of RAM, starting in milliseconds, and sleeping when idle. @lovelylogicss shipped Pi Pocket 0.11, rebuilt for one-handed phone navigation on top of Pi Durable, so the agent can modify the app while you're using it; @sawyerhood's takeaway is that anyone can now "slop your own version of muse/dots/grok-bot" with pi-durable. @badlogicgames went further, saying he hasn't opened his laptop in days and is "fully phone agent pilled."

On the verification side, @workersio launched Workers IO, which runs systems through millions of simulated futures with perfectly reproducible failures; @chaaai says it already surfaced a kernel panic in a guest OS during a customer workload run and calls deterministic simulation testing crazy underutilized. And @thdxr released opentunnel, public URLs for anything on your machine, with end-to-end encryption so the relay can't read traffic plus an SDK, integrating into opencode tomorrow.

Humans writing the context agents can't infer

Two posts converged on a simple idea: curate what your agent knows. @coreyganim's widely shared workflow (which convinced @FundamentEdge to download Todoist on the spot) tags every task with Why, Done, Access, and Who; each morning Codex auto-completes tasks assigned to it and drafts briefs for the human ones. His pitch: 30 extra seconds of context per task lets an AI or human assistant finish it in one shot. @vojta_holik's good-css applies the same logic to styling, packaging 47 modern CSS techniques like :has(), subgrid, and anchor positioning so coding agents reach for them first instead of ignoring what they already know, endorsed by @mattpocockuk. Both echo @kunchenguid's hunch that tests might still help if humans write them to express intent precisely, though he hasn't proven that yet.

Practical Takeaway

If your agent bill is the thing keeping you up, today's posts suggest three concrete moves before you scale anything: run a cheap A/B eval on your own tasks the way @kunchenguid did (tests on versus off, measure success, time, and tokens), check your harness's cache hit rate since @VictorTaelin shows 98%+ is achievable, and ask whether any repetitive decisions in your pipeline could be a 4GB-VRAM fine-tune instead of an LLM call. And if you're on a Claude Max or Team plan, the new monthly API credits are free money for exactly that kind of experimentation.

Sources

H
Hugging Models @HuggingModels ·
TYODM (Train your own Decision Model)
U UnslothAI @UnslothAI

You can now train your own Decision model like Jev locally! We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM. Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo. We fine-tuned with a Clef head using Unsloth and LoRA (r=64) for one epoch, increasing downstream accuracy from 30–37% to 78%. GitHub: https://t.co/2kXqhhvLsb Guide and Notebooks: https://t.co/qACsYehl1n

A
Adam Jacob @adamhjk ·
If you work for a lab, or have unlimited token budgets, the kinds of solutions you come up with are... of a particular shape, lets say. The rest of us need to learn how to build practical systems with this technology, and it turns out *those* systems use far less tokens.
P peterpme @peterpme

poteto's cursor profile is public. She's used 1.5T tokens in the last 30 days. Assuming an avg rate of $8/M that's $384k/day or $11m. In the last 30 days. Per year thats $132M. Let's assume Cursor has some insane deal with Anthropic and gets 90% off. That comes out to $38k/day or $1.1M in Sept/Oct. $13m will get me 26 sassy staff-level all stars for $500k/yr

S
Sawyer Hood @sawyerhood ·
this is your sign to slop your own version of muse/dots/grok-bot using pi-durable https://t.co/uNNPPppcG2
N
Nathan Flurry 🔩 @NathanFlurry ·
we upgraded to pi-durable for rivet actors! ~~~ if curious what the hype is all about: agents are a disturbed systems problem, and one that’s easy to do wrong there are ~so~ many edge cases when something fails at any step of the loop, including: - process crashes & ooms - upgrading your code - retrying tool failures - resuming partial turn completion - prompt retry exactly once delivery - llm providers go down - network timeouts - failure to write durable state - and subagents are a different beast all of this needs to be handled or else background agents will behave erratically pi-durable does this for you it takes learnings from workflow engines and building a harness + tools that handle resuming, history, and idempto-, erm, i-dem-poten-cy, erm, idemp-o-tency beautifully highly recommend listening to mario & armin yap about durable (link below) ~~~ rivet actors then give you a way to orchestrate pi durable: - global singleton process for each pi instance - each pi instance takes 1.3 mb of ram - starts in ms - sleeps when idle - works with or without a sandbox durable also works seamlessly with actor’s sqlite, crons, workflows, otel, jwt, and other goodies we’ve been crafting and open source of course
R rivet_dev @rivet_dev

Rivet now supports Pi Durable on Actors 🕹️ Multiplayer streaming 🪶 Lightweight → ~1.3 MB per agent, starts in ms 💾 Durable → survives crashes, exactly-once input 📦 Harness outside the sandbox 🥧 Vanilla Pi 🍱 Schedules, JWT, sandboxes & more https://t.co/hGXFrMK2xq

S
Satya Nadella @satyanadella ·
Today marks a new chapter for Windows, as we bring unmetered intelligence to every desk and every home, and make every PC a place where agents can work securely on your behalf. Some highlights of what we announced: • MAI-Code-1.1 Flash: 137B parameter coding model w/ 256K context window, which is now optimized to run on your PC! • GitHub Copilot now hands off work to local models like MAI-Code-1.1 Flash, helping projects cost a lot less without sacrificing quality. • With Hybrid Intelligence, Copilot can now take action directly on the PC and keep sensitive work on your device. • And with Code in Copilot, you can essentially build any software you need on your desktop, without any cloud token spend, and it’s just super at it. You’re no longer limited to what’s in an app store! Your PC becomes an infinite software factory. • Security is foundational to all this, which is why we are also bringing together Windows and Agent 365 so agents can work within secure boundaries on-device, including MXC a local sandbox for agent execution. Windows becomes your secure agent box! • All this comes to life on a new generation of devices, like Surface Laptop Ultra, powered by NVIDIA RTX Spark. Can’t wait to see what you build with all this.
M
Mario Zechner @badlogicgames ·
i haven't opened my laptop in days. i'm fully phone agent pilled. i can do anything i want, and the UI is exactly what i want and need, and not more. https://t.co/jEMrDgZQHj
M
Matt Pocock @mattpocockuk ·
I've worked with Vojta for four years He builds most of the code for my course sites, including: https://t.co/CnnLy3L81A https://t.co/wK21D3b8Rv https://t.co/85tm8NWs5O He knows a thing or two about frontend design And guess what, his skills are 🔥
V vojta_holik @vojta_holik

Your coding agent knows :has(), subgrid and anchor positioning. It just rarely reaches for them. I built good-css so it does: 47 modern CSS techniques your agent uses first. https://t.co/qziSDAVYVM

Z
Zach @ZryMiller ·
This is a joke right? You are giving me my money back to spend on Claude in the API plus my subscription. Anthropic is the best, WOW.
C ClaudeDevs @ClaudeDevs

We’re rolling out monthly Claude Platform API credits for Max and Team plans: Max 5x: $100 Max 20x: $200 Team: up to $500, pooled Works on any model, including Haiku 5.5, in your code or third-party harnesses. How it works and full terms: https://t.co/qmWL6xVGwr

D
dax @thdxr ·
opentunnel - public urls for anything running on your machine tunnel services exist but this one is different - e2e encrypted, relay can't see shit - sdk to embed this functionality in your apps will be integrated into opencode tomorrow https://t.co/4aJkLaAW0D
S
Sydney Runkle @sydneyrunkle ·
deciding what to put / keep in memory is a great task for a decision model! memory maintenance can now be faster and cheaper
M mem0ai @mem0ai

Jev-Mem: Memory Decisions Without an LLM

F
Fingerling @lovelylogicss ·
Pi Pocket 0.11 is out, and this one is all about the "Pocket" part. I use it on my phone quite often, so it's been rebuilt for one handed navigation. Everything's a thumb away, swipe left for files, swipe right to go back, and back finally goes where you were. Since it's built on @pidotdev's Pi Durable, Pi can change Pi Pocket while you're using it, restart the server, and pick up right where it left off. If you're missing a feature, just tell your agent to add it and you're off to the races. https://t.co/VV5AQAQCZW
K
Kun Chen @kunchenguid ·
holy crap this is massive! grok bot can now read X with NO API COST one of my favorite use of it is to quickly understand what’s going on when i hear about a new thing pulling a large amount of human discussions on X is the best way to get the full picture
B bot @bot

Grok Bot can now search, read, and monitor X. https://t.co/KTYcATvrG4

B
Brett Caughran @FundamentEdge ·
I’ve officially been influenced Todoist downloaded
C coreyganim @coreyganim

How I get AI + my human assistant to do 50% of my work each day before I get out of bed: I use Todoist to track my daily to-dos. Each task inside Todoist gets a description that follows a simple formula: Why, Done, Access, Who Why = the purpose of the task Done = what does done look like Access = what tools are needed to complete the task Who = who is responsible for completion Every morning Codex looks at Todoist at all the tasks due that day. For any tasks where "Who" = Codex, it just completes the task. This could be drafting follow ups, writing an email broadcast, scheduling a UPS pickup, or any number of other things. For any task where "Who" = me or my assistant, Codex drafts a brief in the comments section of the task. The brief is designed to do as much of the task as Codex possibly can so that we can step in and finish it off. The brief may add additional context, an exact phone number to call, a specific person I personally need to reach out to and the details of our last interaction, etc. I highly recommend following the Why, Done, Access, Who context model for task management. It's kinda a pain to add all those details for simple tasks like "book haircut" but 30 seconds of additional effort on your part enables your AI (or your human) assistant to take entire tasks off your plate in one shot.

T
Taelin @VictorTaelin ·
Hey I updated the Gist, the original had a bug that trashed the cache. If you used it to build a harness, please ask your AI to read and fix. I also add an explanation on how the cache is preserved. I get 98%+ cache hits!
V VictorTaelin @VictorTaelin

THE RECIPE https://t.co/1oezv7fqg1 Ask your agent to build this and ENJOY FREEDOM 🥳

K
Kun Chen @kunchenguid ·
ok it's official... AI-written unit tests and integration tests are empirically proven to be unhelpful. you should tell your agents to stop writing tests by themselves on the deepswe eval set, banning sonnet 5.5 high from writing any tests actually resulted in slightly higher success rate (non stat-sig), with less time and token spent (stat-sig). this is pretty hard empirical evidence across the tests that were written by the "tests allowed" baseline arm, 65% of them were unit tests, 35% were integration tests, neither bucket resulted in any improvement compared to not writing any tests at all i also picked a random sample of 44 tasks subset where i completely disabled executing even existing tests - it also did not affect success rate at all i spot checked many tests written in the baseline arm, and my intuition is that most tests are simply a repetition of the implementation agent-written tests do not add any value because both the implementation and the tests were simply the agent's interpretation of our intent. the tests aren't any more accurate than the implementation itself i suspect we can still extract some value from unit tests and integration tests if we describe them ourselves when we believe we can articulate our intent better through test cases than through requirements. but i have not proven this yet also worth noting, during deepswe eval the agent wrote almost no e2e tests (only 17, compared to 3000+ unit/integration tests written). so this analysis does not prove nor disprove the value of e2e tests - i will do another eval specifically for that
L
lauren @poteto ·
hey @bot can I get tibo@mail.grokbot.com for my bot?
C
Chaitanya @chaaai ·
We are getting so good at it, this time we found kernel panic bug in the guest OS during a customer workload run. More people should use DST, it's crazy how underutilized it is.
W workersio @workersio

Introducing Workers IO Agents are making it possible to turn ideas into software faster than ever. To keep up with that, we need better ways to understand how our software will behave in real-world. Workers IO runs your entire system through millions of simulated futures, exploring what happens when machines fail, networks slow down, and events arrive in the wrong order. It searches for the combinations that break your system before your customers encounter them. Every failure it encounters is perfectly reproducible, helping teams understand what happened, work on a fix, and verify against the same scenario. This ability changes what teams are willing to take on. Difficult engineering decisions become things that you can explore and verify. That’s the future we’re working toward: small teams building increasingly ambitious systems, with the evidence to stand behind. I can’t wait to see what you build with that confidence.

D
Denislav Gavrilov @kuberdenis ·
Prime is doing real evals PewDiePie is distilling Astra And you’re dooming?
T ThePrimeagen @ThePrimeagen

tested out Haiku vs Luna reasoning none for some automation Haiku, literal drop in replacement (it was a one line change of changing a string in a config file), performed significantly worse Still have to understand why It was also significantly slower (both in pass and fail) https://t.co/xomqAMlCUe

T
Tibo @thsottiaux ·
Day 3 (encore)/ We silently re-shipped codex cloud. It’s pretty good now
T Tailscale @Tailscale

Codex 🤝 Tailscale Codex Cloud can now securely connect to resources on your tailnet with Tailscale. Useful tools should be easy to build with, combine, and turn into something even more useful. See how it works → https://t.co/0eFd37GmmO

D
Data Scientologist @my_knn_totoro ·
RT @Teknium: @iAmHenryMascot We aren't giving up on consumers. But they are dumb and I never made hermes for that purpose. I made hermes ag…
I
Ivan Fioravanti @ivanfioravanti ·
Just downloaded and activated Underdog. Onboarding is just WOW. Unsubscribed so many unwanted newsletters in few seconds. 💪
U underdogdotai @underdogdotai

Today we're releasing Underdog Saluki 27B Based on Qwen 3.8 27B, Saluki is almost 7x smaller than its full-precision counterpart while it keeps 96% of its benchmark performance and beats Qwen at tool calling It's the best model <8GB that runs in Underdog Saluki is available today under Apache 2.0. Try today in Underdog https://t.co/wAgfW3X3ok

C
Chubby♨️ @kimmonismus ·
And so it begins. South Korea's biggest banks were hit by a cyberattack and the hacker didn't even use the best SOTA models, DeepSeek 4.1 Flash, GLM-5.3 and Grok 4.6. Just imagine what's possible with the very best models.
A AndrewCurran_ @AndrewCurran_

Last week some of South Korea's biggest banks were hit by a cyberattack. Thanks to a report from CrowdStrike tonight, we now know the entire hack may have been done by a single person. He used a combined stack of an open-source AI penetration tool named ARTEX, DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code.

D
DHH @dhh ·
Thrilled to welcome @SpaceXAI as a Founding Corporate Patron for the Omacom Foundation. That's $1,500,000 in @grok tokens for the maintenance and development of Omarchy. Maybe soon they'll be beamed straight from orbit? Let's go 🚀🌕 https://t.co/PeVkvxHecg