Crash-Proof Agents and Open-Weight Offense: OptChat, Pi Durable, apex-flash-1, and Kolibri Lead the Day
The day's densest cluster was durability, with @VictorTaelin's OptChat memory harness, @lovelylogicss's Pi Pocket built on Pi Durable, and @ibuildthecloud's microVM pause-and-snapshot enthusiasm all arguing for checkpointed agent state. Cantina Security's apex-flash-1 and Aleph Alpha's Kolibri extended the open-weights run, @hwchase17 detailed how he cut coding agent costs, and @DKokotajlo amplified another OpenAI safety resignation.
Quick Hits
- Cantina Security released apex-flash-1, an open-weights cybersecurity model post-trained from GLM-5.3-Flash. @hrkrshnn says the team has earned $1M in bounties, claims the top spot on HackerOne's 2026 US business leaderboard, and argues that public evals like ExploitBench tell you almost nothing about real security work.
- Germany joined the open-weights race: per @trawasthi_ai, Aleph Alpha's Kolibri is a 78B-parameter MoE with 3.46B active parameters and up to 1M tokens of context under Apache 2.0, and he claims it beats Qwen2635B, Nemotron3 Super, and Mistral Small 4 in math, science, and code arenas.
- @hwchase17 reports coding agent costs fell for a second straight month, credited to LangSmith cost tracking, per-user spend caps, and model routing in the OpenSWE harness.
- @DKokotajlo amplified another OpenAI resignation, arguing iterative deployment must be phased out in favor of costlier foresight and redundancy, and that this will not happen under current race conditions.
- Durability was the loudest engineering theme of the day: @VictorTaelin's OptChat, @lovelylogicss's Pi Pocket on Pi Durable, and @ibuildthecloud's microVM snapshot bookmark all circle checkpointed state.
Agents That Survive a kill -9
The clearest cluster across today's posts makes one point at three layers of the stack: stop losing work when a process or a context dies.
@VictorTaelin's long OptChat post is the centerpiece. He describes a harness where an OptMem-backed log replaces chat history entirely, ranks it as his biggest workflow upgrade since GPT-3, and claims that resetting context every turn eliminates context rot while caching keeps spend down. He says any fact from his life is a few "zoom hops" away, most of his AGENTS.md has become redundant, and running it on a Mini means computer-use tasks survive a closed laptop. He posted a replication recipe, which @MichelIvan92347 flagged as the intriguing follow-up to OptMem, and @badlogicgames reshared someone building an agent on the OptChat memory idea atop Pi Durable.
That points at @lovelylogicss's Pi Pocket, built on Mario Zechner's Pi Durable. The demo: kill -9 the server mid-run and a new process resumes from the same SQLite file, rerunning only the interrupted test and never repeating finished work. Every model and tool call is a checkpointed task, which the author says yields crash recovery, multiplayer, and hot-swappable extensions, with credit to Mario and the Pi team.
The day's single bookmarked post lands on the same idea in hardware terms. @ibuildthecloud writes, "I love microvms. We finally have a great use case for pause and snapshot," quoting @microsandbox's demo of pausing a running computer and resuming it exactly where it left off.
Bringing Down the Coding Agent Bill
A second cluster is attacking token spend. @hwchase17's recipe is unglamorous and specific: track everything in LangSmith, which he says has first-party integrations with the main coding harnesses; enforce user-level cost caps through an LLM gateway (raising your cap means talking to the VPE); and route work through OpenSWE, which he describes as an open-source cloud agent harness with model routing.
@Bober_smart summarizes what the post presents as an Anthropic guide to agent teams in Claude Code: an Opus 5.5 architect at high effort plans and merges PRs, Sonnet 5.5 developers write code in isolated git worktrees, a Fable 5.1 adversary audits contract boundaries and pre-merge changes, and a Jev script layer handles mechanical file and CLI operations in 16ms without invoking a model. The economic logic is sparing expensive context for decisions only, though the specifics are a community retelling worth checking against the linked guide.
Harness choice also stayed live: @thdxr polled for people's biggest issues with OpenCode 2, and @theo pointed terminal-pane agent users of @herdrdev and @orca_build toward @t3dotcodes.
Open Weights Keep Arriving; Moats Keep Looking Thin
One release-heavy thread and one sour take on startups frame the open-source story. @julianharris calls Jev possibly "the most thoroughly spanked startup ever," drawing the lesson that software without proprietary data gets cloned instantly; he quotes @MiaAI_lab's claim that Cloudflare's open-source Clef beat Jev on speed and accuracy on some benchmarks within two weeks. @jun_song argues the reverse side of the same coin: AI collapses the cost of contributing, so open clones appear within days, pointing at @ataiiam's OpenDots launch of self-hostable, always-on AI coworkers.
The biggest release of the day sits at the intersection. @hrkrshnn's long post on apex-flash-1 argues security is becoming an economics problem: capable models plus harnesses plus compute can find vulnerabilities at falling cost, and defenders should own the full stack, from private evals and data pipelines to post-training loops and inference. He cites an Anthropic report on an automated exploit factory attributed to a team in China, says Cantina's harnesses process trillions of tokens a month, and warns that public evals like ExploitGym and ExploitBench are far removed from everyday application security. For background, @sermakarevich shared a plain-language article on evals aimed at engineers, PMs, and CEOs.
@trawasthi_ai rounds out the model news with Kolibri's numbers, and @0xSero pushes a promotional deal for running GLM-Flash and DS4.1-Flash fast, which reads as an ad with no verifiable detail attached.
Agents Collide With macOS, MCP, and Rust
Platform friction got its share of attention. @dhh reacts to @natlungfy's report that Apple plans new Mac privacy controls warning about broad data access granted to third-party software including AI agents: making macOS harder to use productively with agents is, in his words, a bold move. @nbaschez argues MCP Events deserve far more excitement, since today most agents only wake up via cron or a user message and event-driven triggers would change that. And a small language flare-up: @ScriptedAlchemy agrees with @zack_overflow that "Rust was not built for an ai world," citing compile-time bottlenecks and agents' improving handle on manual memory management. Two enthusiastic posts, not a settled verdict.
Safety Exits, Papal Latin, and the Learning Queue
@DKokotajlo's resignation thread adds specifics to his position: the departing author, via @BogdanIonutCir2's quote, says they led writing the safety reports published with each major OpenAI launch, and Kokotajlo wants understanding, foresight, and redundancy-based practices to replace iterative deployment.
On the ethics beat, @ebrockwayink highlights the Pope's post distinguishing human art from what machines statistically generate, arguing they differ ontologically before aesthetically, and notes with relish that it went out from the account dedicated to Latin. @elonmusk offered a two-line opinion, "No more AI / SI / It's better," with no argument attached.
Finally, the learning queue: @matthewcanham says the response to his "Jev Explained for Normies" piece shows a gap in education that assumes little prior knowledge yet reaches building depth, and asked followers what to cover next, while @khushiirl shared a Harness CI/CD engineering resource. @kitlangton answered "what is Effect" by pointing to @sheherenow_'s Frontier Lab Tycoon, a free game built with Opus 5.5, Sonnet 5.5, EffectTS, XState, and Jev that simulates model releases and safety audits. When your tooling names show up in a tycoon game, they have become culture.
Practical Takeaway
If your agents lose work on crashes or drift over long sessions, today's posts suggest a concrete experiment: pick one non-critical workflow and add a durability layer, either checkpointed tasks on SQLite in the Pi Durable style or a memory-tree context like OptChat, and measure failure recovery and token spend before and after. Whatever harness you run, copy @hwchase17's order of operations: instrument costs first, cap them second, and only then optimize routing and the harness itself.
Sources
introducing msb in 90s, a new series. episode 1: pause and resume pause a running computer. resume it right where it left off.
Shots fired? Apple plans to introduce new privacy controls for Mac users, warning of the growing risks around granting broad data access to third-party software, including AI agents https://t.co/UYnNQSFdum
Jev Explained for Normies
Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0. https://t.co/5263xZ9xZN
Hac intellegentiae artificialis aetate, urgens fit humanam artem ab iis distinguere, quae machinis efficiuntur. Ars enim et ea, quae machina ex innumeris alienis imaginibus statisticae ope computationis generare potest, ontologice, prius etiam quam aesthetice, inter se differunt. Algorithmis humani deest favilla. Quapropter Ecclesia cum artificibus et humani cultus institutis foedus renovare cupit: foedus scilicet ad humanum custodiendum.
🎉 Introducing 𝙾𝚙𝚎𝚗𝙳𝚘𝚝𝚜 Self-hostable, always-on AI coworkers that works with ANY agent harness. Includes: - Computer use: browser, terminal & files - Bring agents to Slack, Teams etc - Spaces and Pages for projects - Voice calls - Web and Mobile Repo → https://t.co/j6URR886Dt Powered by @CopilotKit and AG-UI. Clone this template and customize it however you want. Enterprise-ready.
Claude Second Brain & Folder Workflow: The Complete Guide
Evals: how to know whether an AI system actually works
An article for everyone who ships, buys, or signs off on software built on large language models (LLMs): engineers, product managers, and the CEO. Wri...
'What I’m about to tell you has, I realize, become something of a cliché: I resigned this week from OpenAI. I led the writing of the safety reports we published with each major launch.' https://t.co/so73hG1MG7
i’ve been teasing it all week, so here it is in time to ruin your weekend plans: Frontier Lab Tycoon every mechanic you might expect (swarm sandbox escapes, hyper-competitive model releases, safety audits, researchers Posting & getting dragged on Bird App, your new Decision Model immediately getting cloned, luddite protests, EA orgies: y’know, normal 2026 things) sold* at https://t.co/KuK9NIgwMV or any other big box retailer where high-quality software is found Opus 5.5 + Sonnet 5.5 + @EffectTS_ + XState + Jev + god knows what else *it’s free
today we're releasing apex-flash-1, our first open-weights model for security research post-trained on real vulnerabilities we found and got paid for. we are also releasing the abliterated variant for researchers who want fewer refusals in their own authorized workflows. how we built it: https://t.co/sSIV6qKFGt
finally managed to do something I long wanted to do: a harness that replaces the chat by optmem. i.e., it IS the chat history, rather than an active tool. now I can forever talk to a single unified manager agent, which knows all my life, and has no tools other than spawn
Cloudflare's Clef beat Jev in just two weeks How: "frozen" Qwen does one prefill pass, a tiny schema head scores every answer option in parallel, with no text generated at all. That's why it's 4x faster than Jev at 2x the accuracy on some benchmarks. And it's open source! Link to HF: https://t.co/1c3YzhAGfV
Rust no longer the defacto best language for agents Compile times are too long, I'm always finding myself bottlenecked by time not intelligence And the agents have gotten really at doing manual memory management now
THE RECIPE https://t.co/1oezv7fqg1 Ask your agent to build this and ENJOY FREEDOM 🥳