AI Digest.

Crash-Proof Agents and Open-Weight Offense: OptChat, Pi Durable, apex-flash-1, and Kolibri Lead the Day

The day's densest cluster was durability, with @VictorTaelin's OptChat memory harness, @lovelylogicss's Pi Pocket built on Pi Durable, and @ibuildthecloud's microVM pause-and-snapshot enthusiasm all arguing for checkpointed agent state. Cantina Security's apex-flash-1 and Aleph Alpha's Kolibri extended the open-weights run, @hwchase17 detailed how he cut coding agent costs, and @DKokotajlo amplified another OpenAI safety resignation.

Quick Hits

  • Cantina Security released apex-flash-1, an open-weights cybersecurity model post-trained from GLM-5.3-Flash. @hrkrshnn says the team has earned $1M in bounties, claims the top spot on HackerOne's 2026 US business leaderboard, and argues that public evals like ExploitBench tell you almost nothing about real security work.
  • Germany joined the open-weights race: per @trawasthi_ai, Aleph Alpha's Kolibri is a 78B-parameter MoE with 3.46B active parameters and up to 1M tokens of context under Apache 2.0, and he claims it beats Qwen2635B, Nemotron3 Super, and Mistral Small 4 in math, science, and code arenas.
  • @hwchase17 reports coding agent costs fell for a second straight month, credited to LangSmith cost tracking, per-user spend caps, and model routing in the OpenSWE harness.
  • @DKokotajlo amplified another OpenAI resignation, arguing iterative deployment must be phased out in favor of costlier foresight and redundancy, and that this will not happen under current race conditions.
  • Durability was the loudest engineering theme of the day: @VictorTaelin's OptChat, @lovelylogicss's Pi Pocket on Pi Durable, and @ibuildthecloud's microVM snapshot bookmark all circle checkpointed state.

Agents That Survive a kill -9

The clearest cluster across today's posts makes one point at three layers of the stack: stop losing work when a process or a context dies.

@VictorTaelin's long OptChat post is the centerpiece. He describes a harness where an OptMem-backed log replaces chat history entirely, ranks it as his biggest workflow upgrade since GPT-3, and claims that resetting context every turn eliminates context rot while caching keeps spend down. He says any fact from his life is a few "zoom hops" away, most of his AGENTS.md has become redundant, and running it on a Mini means computer-use tasks survive a closed laptop. He posted a replication recipe, which @MichelIvan92347 flagged as the intriguing follow-up to OptMem, and @badlogicgames reshared someone building an agent on the OptChat memory idea atop Pi Durable.

That points at @lovelylogicss's Pi Pocket, built on Mario Zechner's Pi Durable. The demo: kill -9 the server mid-run and a new process resumes from the same SQLite file, rerunning only the interrupted test and never repeating finished work. Every model and tool call is a checkpointed task, which the author says yields crash recovery, multiplayer, and hot-swappable extensions, with credit to Mario and the Pi team.

The day's single bookmarked post lands on the same idea in hardware terms. @ibuildthecloud writes, "I love microvms. We finally have a great use case for pause and snapshot," quoting @microsandbox's demo of pausing a running computer and resuming it exactly where it left off.

Bringing Down the Coding Agent Bill

A second cluster is attacking token spend. @hwchase17's recipe is unglamorous and specific: track everything in LangSmith, which he says has first-party integrations with the main coding harnesses; enforce user-level cost caps through an LLM gateway (raising your cap means talking to the VPE); and route work through OpenSWE, which he describes as an open-source cloud agent harness with model routing.

@Bober_smart summarizes what the post presents as an Anthropic guide to agent teams in Claude Code: an Opus 5.5 architect at high effort plans and merges PRs, Sonnet 5.5 developers write code in isolated git worktrees, a Fable 5.1 adversary audits contract boundaries and pre-merge changes, and a Jev script layer handles mechanical file and CLI operations in 16ms without invoking a model. The economic logic is sparing expensive context for decisions only, though the specifics are a community retelling worth checking against the linked guide.

Harness choice also stayed live: @thdxr polled for people's biggest issues with OpenCode 2, and @theo pointed terminal-pane agent users of @herdrdev and @orca_build toward @t3dotcodes.

Open Weights Keep Arriving; Moats Keep Looking Thin

One release-heavy thread and one sour take on startups frame the open-source story. @julianharris calls Jev possibly "the most thoroughly spanked startup ever," drawing the lesson that software without proprietary data gets cloned instantly; he quotes @MiaAI_lab's claim that Cloudflare's open-source Clef beat Jev on speed and accuracy on some benchmarks within two weeks. @jun_song argues the reverse side of the same coin: AI collapses the cost of contributing, so open clones appear within days, pointing at @ataiiam's OpenDots launch of self-hostable, always-on AI coworkers.

The biggest release of the day sits at the intersection. @hrkrshnn's long post on apex-flash-1 argues security is becoming an economics problem: capable models plus harnesses plus compute can find vulnerabilities at falling cost, and defenders should own the full stack, from private evals and data pipelines to post-training loops and inference. He cites an Anthropic report on an automated exploit factory attributed to a team in China, says Cantina's harnesses process trillions of tokens a month, and warns that public evals like ExploitGym and ExploitBench are far removed from everyday application security. For background, @sermakarevich shared a plain-language article on evals aimed at engineers, PMs, and CEOs.

@trawasthi_ai rounds out the model news with Kolibri's numbers, and @0xSero pushes a promotional deal for running GLM-Flash and DS4.1-Flash fast, which reads as an ad with no verifiable detail attached.

Agents Collide With macOS, MCP, and Rust

Platform friction got its share of attention. @dhh reacts to @natlungfy's report that Apple plans new Mac privacy controls warning about broad data access granted to third-party software including AI agents: making macOS harder to use productively with agents is, in his words, a bold move. @nbaschez argues MCP Events deserve far more excitement, since today most agents only wake up via cron or a user message and event-driven triggers would change that. And a small language flare-up: @ScriptedAlchemy agrees with @zack_overflow that "Rust was not built for an ai world," citing compile-time bottlenecks and agents' improving handle on manual memory management. Two enthusiastic posts, not a settled verdict.

Safety Exits, Papal Latin, and the Learning Queue

@DKokotajlo's resignation thread adds specifics to his position: the departing author, via @BogdanIonutCir2's quote, says they led writing the safety reports published with each major OpenAI launch, and Kokotajlo wants understanding, foresight, and redundancy-based practices to replace iterative deployment.

On the ethics beat, @ebrockwayink highlights the Pope's post distinguishing human art from what machines statistically generate, arguing they differ ontologically before aesthetically, and notes with relish that it went out from the account dedicated to Latin. @elonmusk offered a two-line opinion, "No more AI / SI / It's better," with no argument attached.

Finally, the learning queue: @matthewcanham says the response to his "Jev Explained for Normies" piece shows a gap in education that assumes little prior knowledge yet reaches building depth, and asked followers what to cover next, while @khushiirl shared a Harness CI/CD engineering resource. @kitlangton answered "what is Effect" by pointing to @sheherenow_'s Frontier Lab Tycoon, a free game built with Opus 5.5, Sonnet 5.5, EffectTS, XState, and Jev that simulates model releases and safety audits. When your tooling names show up in a tycoon game, they have become culture.

Practical Takeaway

If your agents lose work on crashes or drift over long sessions, today's posts suggest a concrete experiment: pick one non-critical workflow and add a durability layer, either checkpointed tasks on SQLite in the Pi Durable style or a memory-tree context like OptChat, and measure failure recovery and token spend before and after. Whatever harness you run, copy @hwchase17's order of operations: instrument costs first, cap them second, and only then optimize routing and the harness itself.

Sources

K
Khushi @khushiirl ·
Found the best resource to learn Harness Engineering. 😭 https://t.co/CMdOjFWOEU https://t.co/BbwmhDhQ1H
D
Darren Shepherd @ibuildthecloud ·
I love microvms. We finally have a great use case for pause and snapshot.
M microsandbox @microsandbox

introducing msb in 90s, a new series. episode 1: pause and resume pause a running computer. resume it right where it left off.

D
DHH @dhh ·
Apple is going to make it even harder to productively use macOS in the age of agents? Bold move. Let's see how it plays out!
N natlungfy @natlungfy

Shots fired? Apple plans to introduce new privacy controls for Mac users, warning of the growing risks around granting broad data access to third-party software, including AI agents https://t.co/UYnNQSFdum

0
0xSero @0xSero ·
Best deal without a doubt right now grab it before it’s too late. With this you’ll be able to run GLM-Flash/DS4.1-Flash really fast ❤️ https://t.co/YvZQoKIh19
D
dax @thdxr ·
what are your biggest issues with OpenCode 2?
M
Matt Canham @matthewcanham ·
the huge number of positive responses I’ve had to this article makes me realize there’s a distinct lack of education that assumes little prior knowledge and achieves a level of depth needed to build with these new technologies what would you like to see me write about next?
M matthewcanham @matthewcanham

Jev Explained for Normies

T
Tekraj Awasthi🧑‍💻🇳🇵 @trawasthi_ai ·
Wake up, baby! Germany has joined the race of frontier LLMs with its sovereign model "Kolibri". And it's Open-Source, too! Beats Qwen2635B, Nemotron3 Super, and Mistral Small 4 in the Math, Science, and Code arena! 🎉 The underlying architecture: custom MoE Transformer https://t.co/kBae1BFK1T
A Aleph__Alpha @Aleph__Alpha

Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0. https://t.co/5263xZ9xZN

E
Emily Brockway @ebrockwayink ·
I’m enjoying two things: 1. The Pope, a deeply educated man in mathematics and the humanities, crafts a thoughtful warning about the delicate balance between human creativity and machines. 2. It’s posted from the account specifically dedicated to messaging in Latin, and folks are losing their minds. He’s on another level, and the tech world is not ready for it.
P Pontifex_ln @Pontifex_ln

Hac intellegentiae artificialis aetate, urgens fit humanam artem ab iis distinguere, quae machinis efficiuntur. Ars enim et ea, quae machina ex innumeris alienis imaginibus statisticae ope computationis generare potest, ontologice, prius etiam quam aesthetice, inter se differunt. Algorithmis humani deest favilla. Quapropter Ecclesia cum artificibus et humani cultus institutis foedus renovare cupit: foedus scilicet ad humanum custodiendum.

J
Jun Song @jun_song ·
The AI era is a massive win for open source. Whenever a new product drops, an open-source version appears within days. Sometimes it is just as good, or even better. In the past, people had little reason to spend months writing code by hand just to contribute to open source. Now, AI cuts that time down to almost nothing. Anyone can contribute with very little effort or cost. Sharing that work on X or YouTube brings back way more value than the time spent making it. The more AI improves, the stronger open source becomes.
A ataiiam @ataiiam

🎉 Introducing 𝙾𝚙𝚎𝚗𝙳𝚘𝚝𝚜 Self-hostable, always-on AI coworkers that works with ANY agent harness. Includes: - Computer use: browser, terminal & files - Bring agents to Slack, Teams etc - Spaces and Pages for projects - Voice calls - Web and Mobile Repo → https://t.co/j6URR886Dt Powered by @CopilotKit and AG-UI. Clone this template and customize it however you want. Enterprise-ready.

N
Nathan Baschez @nbaschez ·
I have not seen near enough excitement about MCP Events https://t.co/tlWV1jyuxX Right now, the only way most people's agents can wake up and do something is either A) cron, or B) you decide to message it Event-driven triggers are a huge deal
B
Bober_smart @Bober_smart ·
I don't understand why everyone isn't doing this yet. Anthropic released a guide for Claude that makes work much more efficient while burning a fraction of the tokens Main Idea: Agent Teams in Claude Code Instead of burning the expensive Opus 5.5 context on every line of code, tasks are delegated across an autonomous swarm:  Architect (Opus 5.5, high effort): Plans the structure and handles final PR merges Developers (Sonnet 5.5, medium effort): Write code inside isolated ⁠git worktree⁠ environments (UI/UX, backend) Adversary / Critic (Fable 5.1): Steps in at contract boundaries, recurring test failures, and pre-PR audits to catch vulnerabilities Script Layer (Jev): Handles mechanical tasks (opening files, CLI invocations) in 16ms without involving LLMs Launch the team via terminal: claude --teammates ux,backend,adversary Configure my Claude Code workspace for autonomous agent teams: In ⁠~/.claude/teams⁠, generate missing roles: set workers to ⁠Sonnet⁠ (⁠medium effort⁠) in separate ⁠git worktree⁠ branches, and the adversary to ⁠Fable⁠ In ⁠~/.claude/settings.json⁠, set ⁠CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1⁠ and lock the main session to ⁠effortLevel: high⁠ Initialize ⁠tasks.json⁠ with file-level locking for direct peer coordination In ⁠CLAUDE.md⁠, add a rule requiring an adversary review before opening a PR. Show configuration diffs before applying edits > https://t.co/78IXvJlq7o
B Bober_smart @Bober_smart

Claude Second Brain & Folder Workflow: The Complete Guide

S
Sergii Makarevych @sermakarevich ·
Evals: how to know whether an AI system actually works
D
Daniel Kokotajlo @DKokotajlo ·
Another one quits. I agree with what he says here: iterative deployment needs to be phased out asap, more costly and time consuming understanding and foresight and redundancy based safety practices need to arise instead. And that simply won't happen under anything like current race conditions.
B BogdanIonutCir2 @BogdanIonutCir2

'What I’m about to tell you has, I realize, become something of a cliché: I resigned this week from OpenAI. I led the writing of the safety reports we published with each major launch.' https://t.co/so73hG1MG7

F
Fingerling @lovelylogicss ·
Saw what @badlogicgames has been building with Pi Durable lately and had to try it myself, so I made Pi Pocket. The fun part: I kill -9 the server halfway through a run, and a new process just picks it back up from the same SQLite file. The test run that got cut off comes back as interrupted and runs again, and nothing that already finished runs twice. Every model call and tool call is a checkpointed task, and crash recovery, multiplayer, and hot swappable extensions all come out of that. Huge thanks to Mario and the @pidotdev team. Pi 1.0 is a great agent, Pi Durable is a really smart foundation, and building on both has been a blast. https://t.co/jSVHDNGcDE
K
Kit Langton @kitlangton ·
For everyone who has been asking what Effect is. This is it. This is Effect.
S sheherenow_ @sheherenow_

i’ve been teasing it all week, so here it is in time to ruin your weekend plans: Frontier Lab Tycoon every mechanic you might expect (swarm sandbox escapes, hyper-competitive model releases, safety audits, researchers Posting & getting dragged on Bird App, your new Decision Model immediately getting cloned, luddite protests, EA orgies: y’know, normal 2026 things) sold* at https://t.co/KuK9NIgwMV or any other big box retailer where high-quality software is found Opus 5.5 + Sonnet 5.5 + @EffectTS_ + XState + Jev + god knows what else *it’s free

H
Hari @hrkrshnn ·
We just released apex-flash-1, an open-weights model we post-trained for cybersecurity. Fire up your GPUs and run it! We've post-trained models from 27B parameters all the way to 1T+ parameters. The unique nature of cybersecurity is that if you can build a system that finds zero-days at scale, you've built a machine that can print money. You can test it on public bug bounties and profit from doing that. We've earned a million dollars in bounties across various programs, and we're currently #1 on the HackerOne US business leaderboard for 2026! Companies like Apple, Anthropic, Datadog, Ripple, Coinbase have paid us well for our disclosures. Last year was all about harness engineering and riding the wave of models getting more intelligent over time. We built harnesses for offensive and defensive security work that process trillions of tokens a month. Yes, trillions! With all that work, it's clear to us what the future of security looks like. Hacking was once a bespoke skill, similar to hardcore software engineering. Software engineering is currently moving from offices to factories. Security, too, will go through the same transition. Last month, Anthropic released a report on how they caught different threat actors abusing Claude. One of them stood out to me. A team in China built an automated exploit factory. A factory that autonomously grinds through vulnerabilities in targets and exploits them. Their targets included security products, network appliances, and government organizations. We've spent enough time building systems adjacent to that to know what it takes, and the shocking part is the economics. How little it costs to breach companies that you and I rely on. All this to say, security is largely becoming an economics problem. Reasonably capable models and harnesses, with enough compute, can find ways to steal your data, damage your reputation, and sometimes even take your money. And the cost to achieve that is dropping every day. When that happens, the right question to ask is: what's the cost, and what's the lowest cost to achieve an outcome reliably? You want to find a stack with the best combination of capability and cost. This is what people call 'Pareto-optimal'. Once you know that, it's all about scaling. Scaling to trillions of tokens a month, then a week, then a day. Eventually, trillions of tokens a second. To do this, you need to own and control the entire stack: the models, the harnesses, the context, and even the flow of tokens. If you do the math here, the economics start looking insane. What does that look like? 1. Building evaluations on cyber tasks that you and your customers care about. 2. Building a data pipeline of unique real-world data with signal attached to it. 3. Building a harness that can self-improve for each specific organization or customer. 4. Building a post-training loop that can take those learnings and improve the models. Models that are better, faster, and cheaper. 5. Owning your inference pipelines and controlling the flow of tokens so you're maximizing the value produced in every GPU cycle. On evals: if you're a company building AI products for customers, you have to build your own. There's tremendous alpha in having internal evals corresponding to real work. In our case, these are evals for offensive and defensive security work. A lot of public evals are bad for measuring the work you actually care about. Public evals are typically from academics or data companies. Academics have limited budgets and limited access to proprietary data that measures real economic work. Data companies build evals and also sell you corresponding data (RL gyms, traces, pre-training data, etc.) that helps your next model “juice up” its score. This is the dirty secret, and also why models can do extremely well on public evals but extremely poorly on things you care about. This is also why there's tremendous alpha in private evals: you know which models are the best for your use cases. In cyber, some of the popular public evals are ExploitGym and ExploitBench. We think these evals do not correspond to the real security work our customers need. A typical challenge in ExploitBench is to take Chrome's javascript engine V8 with a known patch and find a way to exploit the vulnerability. That is far, far removed from the security issues in everyday applications that you and I use, and applications built by our customers. We could've benchmaxxed on these evals, but we didn't. And you should be wary of people using these evals to advertise their models and harnesses. Measure things yourself and see how they apply to your use cases. The future: Opus is a model that many people found to be a reliable workhorse. For me, Opus 4.1 was the first model that was functional and could get work done. These days, the flash models from Qwen, DeepSeek, MiMo and GLM can be categorized as workhorses too. That, combined with owning the post-training loop means the economics start looking insane. You can get frontier performance in specific domains for a small fraction of the price of Opus. With the right data pipelines and post-training loop, you now have a durable strategy to keep improving and stay state-of-the-art in your vertical. That's the bet behind apex-flash-1, our post-train of glm-5.3-flash, and we're already preparing for longer training runs.
C cantinasecurity @cantinasecurity

today we're releasing apex-flash-1, our first open-weights model for security research post-trained on real vulnerabilities we found and got paid for. we are also releasing the abliterated variant for researchers who want fewer refusals in their own authorized workflows. how we built it: https://t.co/sSIV6qKFGt

T
Taelin @VictorTaelin ·
I can't overstate how excited I am about OptChat. It has been the best upgrade to my workflow since GPT 3. Second only to learning VIM. I exported all my past chats to it, and life has been amazing. The convenience of having a single unified log of my entire life, in a way the AI (and I) can navigate to find any information about me is indescriptible. Before, I had all these campaigns spread out across different Claude Code sessions, and even different harnesses. Each part of my life was on a different place, and Opus sucks at navigating these past memories because it is a huge mess. Now it is all just there, in one place. Any fact from my life is 3-4 zoom hops away. Any prompt I wrote, any algorithm I designed, is a few hops away. It feels like my life is now burned in the models weights, even though it is just a silly compression trick. "what is the state of that PR about λ-contraction again?" (3 zoom hops later...) "it was 2 weeks ago. you went to eat yogurt and forgot to reply" --- "remind me of that purely affine sort algorithm?" (5 zoom hops later...) "you wrote this 7 months ago: <code>. the trick is..." --- "find the top 3 pending topics in my life that need my attention" (several zoom hops later...) "here they are: ..." --- It just feels fantastic how well this just works. Also, the way the context is reseted on every turn clearly makes the AI perform better, it completely eliminates context rot, and I never have to /compact manually again. There is no need to. And yes, it is surprisingly cache efficient, because in-turn messages (where most of the cost is) are cached, and even cross-turn, most of the view is intact due to the way the binary compression works. I had to fix some silly bugs on OptMem to make it so, but now it just works. My costs DROPPED. Also the fact I can manually navigate the memory tree myself, starting from the root, and finding anything about my life. It is just so natural. Everything is there. Also, I realize I can take 90% of my AGENTS.md away because now the way the agent should behave just "sticks" from the OptChat log itself. If I change my mind about something, I just send a message and the latest ruling sticks. If the agent forgets something, I just point it out and it finds the place where I explained it. No need to re-explain. No need to manage AGENTS.md manually. It just... works. On top of that, it now lives on my Mini, so closing my Macbook doesn't stop it, and it has no access to my private data, so I'm finally able to use Computer Use to do things like buy food, book flights, sign on gov[.br] (THANK GOD!!!!!!!!!!). I know this is not news to anyone using Hermes and the like, but it is the first time I'm able to do it and it is already transforming my life for better. Sorry for the over-excited tweet about something so mundane. It is just insane how incredibly transformative a simple idea can be. Context rot was the root of all evil and now it is gone. I feel free. I'd like to publish this but it is kinda coupled to my workflows and the interface is incredibly ugly and hacky. Actually, I think I'll just ask my agent to write a prompt explaining all the details of the setup so you guys can replicate! I'll post it below. It is essentially infinite context that just works. I really suggest that you guys try it!!! It may change your life too
V VictorTaelin @VictorTaelin

finally managed to do something I long wanted to do: a harness that replaces the chat by optmem. i.e., it IS the chat history, rather than an active tool. now I can forever talk to a single unified manager agent, which knows all my life, and has no tools other than spawn

T
Theo - t3.gg @theo ·
RT @NixFred: If you are using @herdrdev, @orca_build or anything else that puts a coding agent in a terminal pane, check out @t3dotcodes .…
J
Julian Harris @julianharris ·
Jev will probably go down in history as the most thoroughly spanked startup ever — a brutal reminder of how few ways are are to create a competitive moat in software these days. (Hint: if you don’t have proprietary data that clearly distinguishes your offering then watch out.) 1. Launch a service that popularises a novel and valuable way to use LLMs 2. Watch literally everyone clone it almost instantly, in many cases beating it
M MiaAI_lab @MiaAI_lab

Cloudflare's Clef beat Jev in just two weeks How: "frozen" Qwen does one prefill pass, a tiny schema head scores every answer option in parallel, with no text generated at all. That's why it's 4x faster than Jev at 2x the accuracy on some benchmarks. And it's open source! Link to HF: https://t.co/1c3YzhAGfV

Z
Zack Jackson @ScriptedAlchemy ·
Agree on this. Rust was not built for an ai world.
Z zack_overflow @zack_overflow

Rust no longer the defacto best language for agents Compile times are too long, I'm always finding myself bottlenecked by time not intelligence And the agents have gotten really at doing manual memory management now

M
Michel aka Agent B @MichelIvan92347 ·
Optmem was already an interesting project. This one is intriguing 👇 Thanks for sharing ! 🙏
V VictorTaelin @VictorTaelin

THE RECIPE https://t.co/1oezv7fqg1 Ask your agent to build this and ENJOY FREEDOM 🥳

M
Mario Zechner @badlogicgames ·
RT @raunakdoesdev: pi durable is so so good @badlogicgames built an agent on top of @VictorTaelin ‘s OptChat memory idea (genius!) that r…
E
Elon Musk @elonmusk ·
No more AI SI It’s better
H
Harrison Chase @hwchase17 ·
we're seeing our coding agent costs decrease significantly for second month in row longer blog on this later, but three quick steps we took 1/ cost visibility. you cant control what you cant measure. Start tracking everything to langsmith. This gives us super granular visibility into who is using what, how they are using it, etc. LangSmith supports first party integrations with all main coding harnesses https://t.co/zPHBAWaRVS 2/ cost controls. Through our LLM gateway, set user level cost caps. (You can increase your cap by talking to our VPE). This helps curb unexpected or accidental spend. https://t.co/YhvxFjTfxi 3/ optimize the harness. Obviously 3rd party harnesses are tougher to optimize, but we are doing more and more coding through OpenSWE, our open source cloud agent harness. Things like model routing help bring the cost down https://t.co/pHUDDCF6zs feel free to reach out if you want help implementing these at your co!