AI Digest.

Open-Weight Apodex 1.1 Reportedly Matches Claude Fable 5 Benchmarks While SSI and GPT-6 Rumors Stay Unverified

Apodex 1.1 launched with an open-weight mini and an open-source agent workbench, and @qilua02 reports it matching Claude Fable 5-level benchmarks. Okta put agent identity into general availability on the open XAA standard, but the loudest stories of the day, continual learning at SSI and early GPT-6 testing, rest on single-source posts with no verification.

Quick Hits

  • @Apodex_AI launched Apodex 1.1, an agentic model family with asynchronous multi-agent teams, an open-weight mini, and a locally deployable open-source workbench called FrontierAgent. @qilua02 says it "spawned out of nowhere and is casually matching Claude Fable 5 level benchmarks."
  • Agent identity went GA: @toddmckinnon announced Agent SSO, built on the open Cross App Access (XAA) standard and bundled into Okta's core SSO offering. @fletchrichman calls it the unglamorous counterpart to all the "botmaxxing with grok bot and hermes" chatter.
  • The day's top rumor is single-sourced: @khademinori claims access to a stealth continual-learning model with "infinite context," and @mark_k guesses it's Ilya Sutskever's SSI. No receipts anywhere yet.
  • Audit your agent config before blaming the model: @trevin found his Hermes browser tasks were still routed through legacy tools, and after switching to the Browser Use 3.0 CLI his local benchmarks ran 72% faster across a dozen tasks.
  • @MickaelV79228 is meter-testing the viral claim that a sleeping Mac earns $120-200 a month serving AI for Darkbloom, with the counter started at $0.00 and real numbers promised within 24 hours.

A New Model Family Ships While Frontier Rumors Escalate

The only new model family in today's feed is Apodex 1.1. Per @Apodex_AI, it targets "frontier-level agentic performance" across professional work, scientific research, financial analysis, and deep search, with an asynchronous agent team that decomposes tasks, runs agents in parallel, and lets you step in mid-stream. The mini ships with open weights, FrontierAgent is open source, and there's a paper, a live workbench, and an API platform. Note that the Claude Fable 5 comparison comes from @qilua02's chart, not from Apodex's announcement text, so treat the "matching Fable" framing as one observer's read.

The rumor mill is doing more work than the labs. @skirano posted that "the next generation of models will be an ontological shock," and @kr0der asserts Pietro got access to test GPT 5.6 months in advance, inferring he's likely testing Astra/GPT 6. That's a guess stacked on an unconfirmed claim. Same for continual learning: @khademinori's account of a model that "knows everything about me & my academic projects & papers since 15 years ago" is striking but unverifiable, and @mark_k's SSI attribution is explicitly a guess based on rumors he's heard.

Two veterans added context. @Steve_Yegge predicts that next year "Fable-class models become ubiquitous and cheap," flooding enterprises with AI employees who end up creating "a constitutional legal system"; his "Fable-class" shorthand echoes the same frontier name in @qilua02's benchmark tweet. And @gdb, amplifying @kundan2510's account of roughly 14 months inside OpenAI, credits the org with "the muscle of making long-term research bets," citing leadership support for the full-duplex gpt-live series even when success looked improbable.

Agent Plumbing Ships: SSO, a 72% CLI Fix, and a 7MB Desktop

Enterprise identity for agents arrived with little fanfare. @toddmckinnon says Agent SSO replaces "a mess of static API keys, fragmented integrations, and endless user consent prompts" by registering agents as first-class identities under XAA, giving IT centralized visibility and least-privilege policy. @fletchrichman's framing is blunt: while everyone hypes consumer bot stacks, Okta is shipping what growing businesses actually need.

The most actionable post of the day is @trevin's. He assumed he was on Browser Use 3.0 and discovered his config still pointed at legacy browser tools; after the fix, 72% faster across a dozen local tasks. He shares a session prompt that asks your agent to check which tooling it's actually using, fix the config, restart, and report the speed delta. @browser_use's own "10x better" line is marketing, but trevin's number is a measured before-and-after.

@odd_joel shipped Moshi Desktop, a 7MB web app built into moshi-hook that runs locally or against any remote herdr/moshi host, with agents, workspaces, terminals, chat, file browsing, and diff viewing. It's a v1 with polish pending, per the author, and requires moshi-hook 0.3.1. Also in agent-land, @RayFernando1337 boosted @kunchenguid's praise for Rajesh, a former Google L7 who built agent infrastructure there; the retweet cuts off, so where he landed is unstated.

Omarchy's Bid to Be an Agent-First Linux Desktop

Three separate posts point at the same gap: Linux can run agents, but not politely. @LLMJunky's detailed critique says computer use on Omarchy is "a second tier experience" because every agent action steals window focus, while Mac's accessibility layer lets agents work in the background; if you must walk away so the agent can drive, "it kinda defeats the purpose." @francedot's reply is a commitment: Wayland makes background computer use "a hard problem," progress is slower than they'd like, and "if that means starting a new Wayland RFC, so be it."

The ecosystem around it is visibly busy. @dhh pointed followers at @jankeesvw, who built a Time Machine-style backup plugin for Omarchy on restic that "mostly gets out of your way." And @tobi praised @jondkinney's pace on Omasnap: 22 PRs merged with 16 more open, covering an expanding canvas, multiple arrow styles, extra fonts with text outlines, line-height-aware highlight snapping, and adjustable pen smoothing.

Receipts, Please: Earnings Claims, IPO Math, and AI-Written Distsys Code

@MickaelV79228 notes the trend claiming a sleeping Mac earns $120-200 a month serving AI, with six-month ROI on a used M1 Pro, has 570K views and "not a single verified number," so he connected an attested M5 Max running Qwen 35B to the network with a public meter. @kellabyte tore into Walgit's S3 backend: tests pass only because the memstore fakes conditional deletes, while the real path does HEAD, compare, then DELETE, letting a stale owner delete a new lease, and the code wrongly claims S3 lacks conditional deletes. Her closer, "Which AI model wrote this? It's wrong in basics," is the sharpest AI-related line of the day, and @anselm_io joked that jepsen has been replaced by "how quickly can @kellabyte tear it apart publicly."

Corporate claims got similar scrutiny. @GaryMarcus warned Anthropic's IPO "is going to be a trainwreck" unless expectations scale back, drawing a rebuttal from @schwarzjn_, who says he's the research lead for the model in question and that the paper is "a blueprint for how to actually build AI Sovereignty." @ryanvogel cringed at founders and CEOs begging to be moved up whatever list @maria_rcks posted. And @ADoricko's Rainmaker says it produced ~19M gallons of water in Alaska via next-gen cloud seeding in three hours, with a white paper attached and plans targeting the Colorado River and Great Salt Lake; @devingannon predicts a massive company. Impressive if it holds, but every number so far is company-reported.

Small but Useful

@trq212 shared an ELI5 prompt reportedly popular inside Anthropic ("explain like I'm someone who knows nothing about this topic, using a HTML artifact with big pictures and few words"), and @graceclarke posted her stricter "eli5-for-grownups" variant. @drewcoffman recommends @stephbzinn's piece via @a16zcrypto, "The habits of AI writing, and what to do about them," for anyone making words for a living. @alex_prompter is pointing followers at what he calls the best Claude tutorial account (unnamed in the post). @taiyo_ai_gakuse recommends the UI galleries @insporadesign posts, like @jeetnirnejak's date-range picker, for indie-dev polish. @wolfie_ (via @malikwas1f) is taming omp's overwhelming settings page with a guide. And @foldkit moved update results from tuples to { model, commands?, outMessage? } records, letting TypeScript catch silently dropped OutMessages

Sources

P
Pietro Schirano @skirano ·
No one is ready for what’s coming. The next generation of models will be an ontological shock.
G
Greg Brockman @gdb ·
OpenAI has built the muscle of making long-term research bets:
K kundan2510 @kundan2510

In my ~14 months at @OpenAI, one of the most surprising and genuinely wonderful things about it, has been the full leadership support for full-duplex models (gpt-live series), even when it seemed deeply improbable that they would work. And yet, if they did work, it was obvious how magical they could be. I am quite sure this is not an exception but a general rule for research projects at @OpenAI. There were so many moments when the problem felt impossibly hard. But one thing kept being true: it was never clear why it shouldn’t work. And almost every time we understood the problem a little better, it became a little easier to solve. I’m pretty sure there are very few environments in the world where a bet like this could have been made and sustained. So grateful to be part of this wonderful place.

V
vogel @ryanvogel ·
the amount of founders and CEOs under this post begging for their product to be moved up is embarrassing
M maria_rcks @maria_rcks

here https://t.co/lkVrbRhlWf

M
Mickaël VILLERS @MickaelV79228 ·
Donc si je comprends bien. Un Mac qui dort gagne 120 à 200 $/mois en servant de l'IA, d'après Darkbloom. ROI en 6 mois sur un M1 Pro d'occasion, d'après le tweet à 570K vues qui tourne partout. Et dans tout le trend, pas un seul chiffre vérifié. Alors je viens de brancher mon M5 Max sur leur réseau. Machine attestée, un Qwen 35B chargé, compteur à $0.00. Rendez-vous dans 24 h pour les vrais chiffres. Vous pariez combien ?
A
Anthony Kroeger @kr0der ·
Pietro got access to test GPT 5.6 months in advance that means he’s likely testing Astra/GPT 6 what did he see 👀 https://t.co/LLDgyeSJrw
S skirano @skirano

No one is ready for what’s coming. The next generation of models will be an ontological shock.

I
IndiJo @odd_joel ·
moshi moshi, happy monday! 😻 so here it comes: Moshi Desktop - a 7 mb web app built into moshi-hook - works locally or with any remote host running herdr/moshi - agents, workspaces, terminals, chat, file browsing, diff viewing, and more this is the very first version, so there are still tons of details to polish—and a zillion features i want to add. i know i can present this with a better video, but anyway, i’ll try to let the product speak first. upgrade to moshi-hook 0.3.1 and let's rock and roll! 🤘
O odd_joel @odd_joel

what if moshi and herdr have a baby...😽

T
Taiyo Kimura @taiyo_ai_gakuse ·
個人開発する時、大体これみておけばUIいい感じになるサイト。
I insporadesign @insporadesign

Date Range Picker by @jeetnirnejak More on →https://t.co/QIsUfM3KGG https://t.co/eh2vugJKO9

S
Steve Yegge @Steve_Yegge ·
I've got a pretty clear picture now of where we're headed next year, when Fable-class models become ubiquitous and cheap, and every enterprise is flooded with hundreds of new Fable-class AI employees. Spoiler: They create a constitutional legal system. https://t.co/9KqssPG6Od
D
drew coffman 𝕚𝕤 𝕠𝕟𝕝𝕚𝕟𝕖 🟢 @drewcoffman ·
if you work in marketing or spend any meaningful amount of your day making words appear on the internet, plz plz plz read this excellent article from @stephbzinn one of the best things you can read on how to make AI-assisted writing actually good
A a16zcrypto @a16zcrypto

The habits of AI writing, and what to do about them

Q
qilua @qilua02 ·
wtf Apodex spawned out of nowhere and is casually matching Claude Fable 5 level benchmarks 💀 https://t.co/x9tCmLtyRB
A Apodex_AI @Apodex_AI

Meet Apodex 1.1: Scaling Agentic Intelligence for Complex Work Open Source Harness: https://t.co/4V9suLx8o5 Open Weights: https://t.co/9CNZ9E5drH We’re excited to introduce Apodex 1.1, our new model family built to scale agentic intelligence for professional work. 🧠 Frontier-level intelligence for complex work Apodex 1.1 brings frontier-level agentic performance across complex professional work, scientific research, financial analysis, and deep search. 🤝 Asynchronous Agent Team Apodex 1.1 can break down complex tasks, coordinate multiple agents in parallel, continuously integrate their findings, and let you step in to guide or redirect the work at any time. 🔬 Open-source research workbench We’re open-sourcing FrontierAgent—a locally deployable research workbench for the Apodex 1.1 family, including asynchronous Agent Team. Available now: 🔹 Apodex 1.1 — our most capable frontier model, available through the Apodex online workbench 🔹 Apodex 1.1 mini — open-weight model for running complex work locally 🔹 FrontierAgent — open-source, locally deployable research workbench 🌐 Live workbench: https://t.co/w5l6YI2ZuP 🦾 API platform: https://t.co/kaxuSWvMB4 📃 Paper: https://t.co/CHc4L63U07

B
Browser Use @browser_use ·
Browser Use CLI makes your Hermes agent 10x better. Make sure to update to latest version!!
T trevin @trevin

My @NousResearch Hermes browser tasks felt really slow and I was confused because I thought I was using all the latest @browser_use 3.0 CLI etc. Turns out I wasn’t, and embarrassingly my config was still using the old browser tools. Now in my local benchmarks across a dozen tasks, it’s 72% faster ❤️ Want to check your setup? paste this into a session: “Check whether I’m using Browser Use 3.0 CLI that was recently released with a persistent cloud/CDP browser or the legacy `browser_*` tools. Fix the config if needed, restart into a fresh session, and report the speed difference.”

G
Grace Clarke @graceclarke ·
Inspired by this, I made my own: "eli5-for-grownups." It's like @trq212's, and because I'm a picky consistency freak, it has extra direction on what eli5 means and doesn't. Should you like, here it is: https://t.co/bkVN7Usuhz https://t.co/CEwlhehQ5D
T trq212 @trq212

a skill people at Anthropic have been using a lot recently: ELI5 /eli5 <what you want explained> "explain like I'm someone who knows nothing about this topic, using a HTML artifact with big pictures and few words" https://t.co/OZqzjAyFdT

F
Fletcher Richman @fletchrichman ·
Everyone’s hyping up how they are botmaxxing with grok bot and hermes. Meanwhile okta is quietly shipping what every growing business actually cares about.
T toddmckinnon @toddmckinnon

Today we’re announcing the General Availability of Agent SSO, and we’re including it in our core SSO offering. SSO transformed how people move across applications, making access seamless for users and centralizing identity and control for IT. AI agents need the same thing. Today, connecting agents is a mess of static API keys, fragmented integrations, and endless user consent prompts. But Agent SSO changes that. Powered by the open Cross App Access (XAA) standard, it enables agents to be registered as first-class identities. For IT, it means centralized visibility and least-privilege policy. For users, it means their AI agents can seamlessly connect across enterprise apps without interrupting their flow.

T
tobi lutke @tobi ·
Jon is doing fab work on omasnap
J jondkinney @jondkinney

22 PRs merged, 16 more open and ready for Omasnap on Omarchy. Hope @tobi doesn't get sick of me! Some of the recent work: - Expanding canvas (with two different options to trim) - Multiple arrow styles - 2 more fonts + a text outline option - Line (height) aware highlighting snap straight on drag - Swap between normal (tiling or floating) window and the current full-screen overlay - Adjustable pen smoothing - And more tweaks and fixes!

F
Francesco @francedot ·
we're on it. wayland makes this a hard problem, so progress is slower than we'd like - but we won’t give up until background computer use on Linux is genuinely good. if that means starting a new Wayland RFC, so be it
L LLMJunky @LLMJunky

I am a big fan of Omarchy soon. As soon as it started to "click" I understood why @dhh was so passionate about releasing it. But there is an area where Mac still really shines compared to Linux, and that is simply not something I'm willing to give up. Computer use. It is possible on Omarchy, even with Wayland being as restrictive as it is, but it is a second tier experience in comparison to Mac/Windows because all actions that the agent takes steals focus away from the user. This is sometimes an issue on Mac too, so it's not entirely solved, but for the most part, when agents are using your computer, its done in the background so you can do other tasks while its working. If I have to walk away from my computer to allow an agent to work, it kinda defeats the purpose because I'm still generally much faster than an agent. Mac's accessibility layer is just far superior. There are some tasks that only a computer use agent can solve, and so I do see this as a critical functionality. I am making this post not to criticize or tell people not to try it because its truly an incredible experience. But for now, I cannot fully switch to it. And my hope is that posts like this will get the right people's attention to solve this problem.

R
Ray Fernando @RayFernando1337 ·
RT @kunchenguid: Rajesh was L7 at Google building agent infra - I had the pleasure to partner with him building together and he’s living at…
F
Foldkit @foldkit ·
tuple → record
D devinjameson @devinjameson

Another meaningful API refinement just landed in @foldkit! It’s a significant one, but after using it across the codebase and in my own Foldkit projects, I’m confident it’s the right call. Update results have been tuples since Foldkit’s first release. This release replaces them with { model, commands?, outMessage? } records. Using records makes each result easier to read and compose. When an update produces no Commands, just omit the commands field. When it emits no OutMessage, just omit outMessage. No need to specify [] and Option.none() at each return site. The new types also let TypeScript reject code that could silently lose an OutMessage, a nice little correctness win. This also enables a new convention I really like: keep each result bound to the operation that produced it and access its fields directly, rather than destructuring at each call site. The operation stays named, and its Model, Commands, and OutMessage stay grouped together. Overall: less ceremony, more signal, stronger OutMessage safety. Win win win. There is a detailed migration guide in the release notes: https://t.co/21sxPh9jZe Thank you to everyone building with Foldkit as I continue honing the API for v1. I’m immensely grateful! This change started with a suggestion from @SandroMaglione. Thank you! :)

A
Alex Prompter @alex_prompter ·
RT @alex_prompter: I just found the best account with Claude tutorials:
D
Devin Gannon @devingannon ·
It’s a Monday and you open up X and there’s a guy with a mullet making it rain 19M gallons in 3 hours in Alaska. This is a big deal. This will be a massive company and extremely important.
A ADoricko @ADoricko

Rainmaker just produced ~19M gallons of water in Alaska via next-gen cloud seeding over 3 hours of operations. We are the first company to provably produce precipitation in Alaska. As promised, we’ve linked our white paper and relevant data. In the future, Rainmaker will protect and restore glaciers with man-made snowfall. Immediately, this demonstration shows how Rainmaker will add new water to the Colorado River and Great Salt Lake in the coming months. We turned around this analysis and white paper just a couple days after operating. There are many data sources I want to deepen our understanding of the atmosphere and increase our provable production; we’re building the instruments to collect that data in subsequent operations. But at Rainmaker, we won’t tout LOIs, simulations, or lab tests. What matters is physically measuring that you’ve affected the atmosphere; proving that we’ve produced water. We prefer sharing only the tech that has proven to make more water for farms, industries, and ecosystems in need. Expect a regular cadence of these from us as we build the best atmospheric science lab in the world. Feedback is welcome in the interim. Rainmaker Century. https://t.co/QZikxP1KyH

A
Anselm Eickhoff @anselm_io ·
Forget jepsen, the new standard for distributed systems is “how quickly can @kellabyte tear it apart publicly on Twitter”
K kellabyte @kellabyte

Walgit leans on S3 primitives but has classic distsys bugs Tests pass cuz its memstore makes cond deletes But w/ S3 it HEAD->compare->DELETE so stale owner can delete new lease The code incorrectly says S3 lacks conditional DELs. Which AI model wrote this? It's wrong in basics

N
noname @malikwas1f ·
RT @wolfie_: omp puts an incredible amount of power in your hands, which also means the settings page can be a little overwhelming so i ma…
D
DHH @dhh ·
If you want to stay in touch with what's possible using Omarchy plugins, you should follow Jankees. The man is a fountain of creativity and productivity.
J jankeesvw @jankeesvw

Time Machine for Omarchy It uses restic under the hood, and it mostly gets out of your way. Backups are something you should forget about, until you need them. https://t.co/RQo9DaA9fJ

J
Jonathan Richard Schwarz @schwarzjn_ ·
Hi @GaryMarcus, I'm the Research Lead for this model. It's even more consequential than just Anthropic's IPO: This paper is a blueprint for how to *actually* build AI Sovereignty. https://t.co/uh5A0YvVNU
G GaryMarcus @GaryMarcus

still more bad news for Anthropic. if they don’t scale back expectations for their IPO, it’s going to be a trainwreck.

M
Mark Kretschmann @mark_k ·
Continual Learning AI may be solved 👀 Which model could this be? After all of the rumors I've heard, my guess is Ilya Sutskever's company @ssi. SSI is said to be working on Continual Learning. Absolutely HUGE if true.
K khademinori @khademinori

i got access to a continual learning model from a stealth company and here we go. i think they solved the problem once and for all. it has infinite context & absolutely knows everything about me & my academic projects & papers since 15 years ago im thinking 🤔 how they did do it?