Microsoft Turns Windows Into a Local Agent Box While Developers Count Tokens
Satya Nadella announced on-device coding models, local Copilot actions, and an agent sandbox for Windows, landing the same week developers are publicly agonizing over trillion-token bills and Anthropic started handing subscribers monthly API credits. Separately, an eval from @kunchenguid found that banning agent-written tests actually shaved time and tokens, and a CrowdStrike report cited on X claims one person used an AI stack to hack South Korea's biggest banks.
Quick Hits
- @satyanadella announced a local-agent push for Windows: a 137B-parameter coding model optimized to run on the PC, Copilot actions that stay on-device, and a local sandbox for agent execution, all on new NVIDIA-powered Surface hardware.
- @kunchenguid reports that banning Claude Sonnet 5.5 High from writing tests on the deepswe eval produced a slightly higher success rate (not statistically significant) while cutting time and tokens (statistically significant). His conclusion: tell agents to stop writing their own tests.
- Per a CrowdStrike report cited by @AndrewCurran_, South Korea's biggest banks were hacked by a single person using an open-source AI penetration tool called ARTEX plus DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code. @kimmonismus flags that none of those are frontier models.
- Grok Bot can now "search, read, and monitor X" with no API cost, per @kunchenguid quoting @bot, and @poteto is already asking @bot for an email address for her bot.
- @dhh announced @SpaceXAI as a Founding Corporate Patron of the Omacom Foundation, contributing $1,500,000 in grok tokens toward Omarchy's maintenance and development.
Windows becomes a "secure agent box," and the token bill becomes the story
The biggest announcement of the day, by volume, was @satyanadella's Windows post. Per his thread, MAI-Code-1.1 Flash is a 137B-parameter coding model with a 256K context window now optimized to run locally; GitHub Copilot can hand work off to such local models to cut project costs; "Hybrid Intelligence" lets Copilot act directly on the PC while keeping sensitive work on-device; and "Code in Copilot" promises to build software on the desktop with no cloud token spend. He also teased MXC, a local sandbox for agent execution tied into Agent 365, running on hardware like the Surface Laptop Ultra powered by NVIDIA RTX Spark.
The cost framing is timely. @peterpme did the math on @poteto's public Cursor profile: 1.5 trillion tokens in 30 days, which at $8 per million works out to roughly $384k a day, or about $1.1M for the month even assuming a 90% discount. @adamhjk used that as a foil, arguing that people with unlimited lab budgets build solutions "of a particular shape," while practical systems use far fewer tokens. @VictorTaelin is in that camp: he fixed a bug in his agent-harness Gist that was trashing the prompt cache and now reports 98%+ cache hits. And @ZryMiller reacted gleefully to @ClaudeDevs rolling out monthly API credits for subscribers ($100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team), usable on any model including Haiku 5.5, calling it Anthropic effectively refunding his subscription.
One eval says agent-written tests add nothing
The day's most actionable research post came from @kunchenguid, who ran ablations on the deepswe eval set. Banning Sonnet 5.5 High from writing any tests resulted in a slightly higher success rate (not statistically significant) with significantly less time and tokens spent. In the tests-allowed arm, 65% of tests written were unit tests and 35% integration tests, and neither bucket beat writing no tests at all. He also disabled executing even pre-existing tests on a 44-task subset with no effect on success. His diagnosis: agent-written tests mostly restate the agent's own interpretation of the requirements, so they can't catch anything the implementation got wrong. He caveats that almost no e2e tests were written (17 versus 3,000+ unit and integration tests), so that question is still open.
Related eval energy: @kuberdenis cheered that "Prime is doing real evals," quoting @ThePrimeagen's test of Haiku versus "Luna reasoning none" for an automation task, where Haiku was a literal one-line config swap but performed significantly worse and ran slower on both passes and failures. Meanwhile @thsottiaux quietly announced "We silently re-shipped codex cloud. It's pretty good now," quoting @Tailscale's announcement that Codex Cloud can now reach resources on your tailnet.
Small "decision models" are taking over narrow jobs
A cluster of posts pushed the idea that tiny specialized models can replace LLM calls for specific decisions. @UnslothAI (amplified by @HuggingModels) says you can now train your own decision model like Jev locally, taking Qwen3.5 0.8B from 20.7% to 74.3% aggregate accuracy across three decision benchmarks on just 4GB of VRAM, using a Clef head with LoRA (r=64) for one epoch; downstream accuracy went from 30-37% to 78%. @sydneyrunkle connected that to @mem0ai's Jev-Mem, arguing that deciding what to store or keep in memory is a perfect decision-model task, making memory maintenance faster and cheaper. On the model side, @underdogdotai released Saluki 27B, built on Qwen 3.8 27B, roughly 7x smaller than its full-precision counterpart while keeping 96% of benchmark performance and beating Qwen at tool calling, all under Apache 2.0 and under 8GB; @ivanfioravanti vouched for the app itself, unsubscribing from a stack of newsletters during onboarding. Also in open-model discourse, @my_knn_totoro resurfaced Teknium insisting he hasn't given up on consumers but "never made hermes for that purpose" (the quote cuts off mid-sentence).
Agent plumbing: durable state, phone UIs, simulated futures, tunnels
The unglamorous infrastructure layer got a lot of attention. @NathanFlurry detailed why Rivet Actors now run pi-durable: background agents fail erratically across process crashes, provider outages, partial turns, and retry edge cases, and durable execution handles resuming, history, and idempotency for you, with each pi instance taking about 1.3MB of RAM, starting in milliseconds, and sleeping when idle. @lovelylogicss shipped Pi Pocket 0.11, rebuilt for one-handed phone navigation on top of Pi Durable, so the agent can modify the app while you're using it; @sawyerhood's takeaway is that anyone can now "slop your own version of muse/dots/grok-bot" with pi-durable. @badlogicgames went further, saying he hasn't opened his laptop in days and is "fully phone agent pilled."
On the verification side, @workersio launched Workers IO, which runs systems through millions of simulated futures with perfectly reproducible failures; @chaaai says it already surfaced a kernel panic in a guest OS during a customer workload run and calls deterministic simulation testing crazy underutilized. And @thdxr released opentunnel, public URLs for anything on your machine, with end-to-end encryption so the relay can't read traffic plus an SDK, integrating into opencode tomorrow.
Humans writing the context agents can't infer
Two posts converged on a simple idea: curate what your agent knows. @coreyganim's widely shared workflow (which convinced @FundamentEdge to download Todoist on the spot) tags every task with Why, Done, Access, and Who; each morning Codex auto-completes tasks assigned to it and drafts briefs for the human ones. His pitch: 30 extra seconds of context per task lets an AI or human assistant finish it in one shot. @vojta_holik's good-css applies the same logic to styling, packaging 47 modern CSS techniques like :has(), subgrid, and anchor positioning so coding agents reach for them first instead of ignoring what they already know, endorsed by @mattpocockuk. Both echo @kunchenguid's hunch that tests might still help if humans write them to express intent precisely, though he hasn't proven that yet.
Practical Takeaway
If your agent bill is the thing keeping you up, today's posts suggest three concrete moves before you scale anything: run a cheap A/B eval on your own tasks the way @kunchenguid did (tests on versus off, measure success, time, and tokens), check your harness's cache hit rate since @VictorTaelin shows 98%+ is achievable, and ask whether any repetitive decisions in your pipeline could be a 4GB-VRAM fine-tune instead of an LLM call. And if you're on a Claude Max or Team plan, the new monthly API credits are free money for exactly that kind of experimentation.
Sources
You can now train your own Decision model like Jev locally! We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM. Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo. We fine-tuned with a Clef head using Unsloth and LoRA (r=64) for one epoch, increasing downstream accuracy from 30–37% to 78%. GitHub: https://t.co/2kXqhhvLsb Guide and Notebooks: https://t.co/qACsYehl1n
poteto's cursor profile is public. She's used 1.5T tokens in the last 30 days. Assuming an avg rate of $8/M that's $384k/day or $11m. In the last 30 days. Per year thats $132M. Let's assume Cursor has some insane deal with Anthropic and gets 90% off. That comes out to $38k/day or $1.1M in Sept/Oct. $13m will get me 26 sassy staff-level all stars for $500k/yr
Rivet now supports Pi Durable on Actors 🕹️ Multiplayer streaming 🪶 Lightweight → ~1.3 MB per agent, starts in ms 💾 Durable → survives crashes, exactly-once input 📦 Harness outside the sandbox 🥧 Vanilla Pi 🍱 Schedules, JWT, sandboxes & more https://t.co/hGXFrMK2xq
Your coding agent knows :has(), subgrid and anchor positioning. It just rarely reaches for them. I built good-css so it does: 47 modern CSS techniques your agent uses first. https://t.co/qziSDAVYVM
We’re rolling out monthly Claude Platform API credits for Max and Team plans: Max 5x: $100 Max 20x: $200 Team: up to $500, pooled Works on any model, including Haiku 5.5, in your code or third-party harnesses. How it works and full terms: https://t.co/qmWL6xVGwr
Jev-Mem: Memory Decisions Without an LLM
Grok Bot can now search, read, and monitor X. https://t.co/KTYcATvrG4
How I get AI + my human assistant to do 50% of my work each day before I get out of bed: I use Todoist to track my daily to-dos. Each task inside Todoist gets a description that follows a simple formula: Why, Done, Access, Who Why = the purpose of the task Done = what does done look like Access = what tools are needed to complete the task Who = who is responsible for completion Every morning Codex looks at Todoist at all the tasks due that day. For any tasks where "Who" = Codex, it just completes the task. This could be drafting follow ups, writing an email broadcast, scheduling a UPS pickup, or any number of other things. For any task where "Who" = me or my assistant, Codex drafts a brief in the comments section of the task. The brief is designed to do as much of the task as Codex possibly can so that we can step in and finish it off. The brief may add additional context, an exact phone number to call, a specific person I personally need to reach out to and the details of our last interaction, etc. I highly recommend following the Why, Done, Access, Who context model for task management. It's kinda a pain to add all those details for simple tasks like "book haircut" but 30 seconds of additional effort on your part enables your AI (or your human) assistant to take entire tasks off your plate in one shot.
THE RECIPE https://t.co/1oezv7fqg1 Ask your agent to build this and ENJOY FREEDOM 🥳
Introducing Workers IO Agents are making it possible to turn ideas into software faster than ever. To keep up with that, we need better ways to understand how our software will behave in real-world. Workers IO runs your entire system through millions of simulated futures, exploring what happens when machines fail, networks slow down, and events arrive in the wrong order. It searches for the combinations that break your system before your customers encounter them. Every failure it encounters is perfectly reproducible, helping teams understand what happened, work on a fix, and verify against the same scenario. This ability changes what teams are willing to take on. Difficult engineering decisions become things that you can explore and verify. That’s the future we’re working toward: small teams building increasingly ambitious systems, with the evidence to stand behind. I can’t wait to see what you build with that confidence.
tested out Haiku vs Luna reasoning none for some automation Haiku, literal drop in replacement (it was a one line change of changing a string in a config file), performed significantly worse Still have to understand why It was also significantly slower (both in pass and fail) https://t.co/xomqAMlCUe
Codex 🤝 Tailscale Codex Cloud can now securely connect to resources on your tailnet with Tailscale. Useful tools should be easy to build with, combine, and turn into something even more useful. See how it works → https://t.co/0eFd37GmmO
Today we're releasing Underdog Saluki 27B Based on Qwen 3.8 27B, Saluki is almost 7x smaller than its full-precision counterpart while it keeps 96% of its benchmark performance and beats Qwen at tool calling It's the best model <8GB that runs in Underdog Saluki is available today under Apache 2.0. Try today in Underdog https://t.co/wAgfW3X3ok
Last week some of South Korea's biggest banks were hit by a cyberattack. Thanks to a report from CrowdStrike tonight, we now know the entire hack may have been done by a single person. He used a combined stack of an open-source AI penetration tool named ARTEX, DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code.