AI Digest.

Decision Models Get an API While Open Models Land in Codex

Around OpenAI's DevDay, the feed carried a limited-preview Decisions API powered by GPT-6 Luna and word that enterprises can now run open models like GLM-5.3 Flash and Kimi K3 natively in Codex. Practitioners filled the rest with concrete technique: a $3 Jev analysis of 2,029 phone calls, multi-model Claude Code orchestration, and practitioner-authored agent skills.

Quick Hits

  • OpenAI's DevDay keynote got the usual liveblog from @simonw, which @HamelHusain calls "the most time-efficient way to get up to speed on what's launching today." In the same window, @philipkiely reported that enterprise teams can run open models like GLM-5.3 Flash and Kimi K3 natively in Codex, with spend counting against their OpenAI commit; @thsottiaux cheered it with "Open is the way."
  • Cheap, fast decision models keep gaining ground: @OpenAIDevs announced a Decisions API powered by GPT-6 Luna (limited preview), and @muratcan described Jev forecasting outcomes across 2,029 real phone calls for about $3, which @natebjones reads as early evidence of what System One-style models can do.
  • Constraints shaped two deployment stories: @alex_prompter amplified @mardehaym's thread on a HIPAA-compliant voice agent for healthcare billing calls, and @abliteration_ai relayed @eddylazzarin turning to unrestricted models for defensive security tasks after restricted ones refused.
  • Two big claims, thin detail: @jrysana teased Rysana V2 as "up to 10,000x faster and 1000x more efficient," and @malikwas1f passed along @teortaxesTex's claim that DeepSeek's kernel work makes the Ascend ecosystem "fully viable for training."
  • Elsewhere: @elonmusk postponed the Roadster demo to October 15 over high winds, @naval recirculated Musk's joke that a Delta flight is the best AI sandbox, and @leerob asked for feedback and feature requests for @Bot.

Open Models in Codex, With the Invoice Still Pointing at OpenAI

@philipkiely's one-liner is the day's most concrete enterprise item: open models like GLM-5.3 Flash and Kimi K3 can now be used natively in Codex, and the spend counts against an existing OpenAI commitment. @thsottiaux's reaction, "Proud of this. Open is the way," frames it as openness winning. The skeptic's read is that openness now arrives with a lock-in flavored billing arrangement: bring your own model, but the money still flows through the same commit. Both takes are opinions, and the underlying claim comes from a single post worth confirming against official docs.

Meanwhile @simonw liveblogged the DevDay keynote in San Francisco, a ritual @HamelHusain endorses as the fastest summary of what actually launched. If the announcements above landed at or around DevDay, no post in the feed says so explicitly, so treat the timing as suggestive rather than confirmed.

Decision Models Graduate From Hack to Product

@OpenAIDevs announced Decisions API, powered by GPT-6 Luna: define questions and possible answers to classify content, route requests, or choose an agent's next action. It is in limited preview. @sydneyrunkle's take is that "jev was the first, but certainly not the last," because decision models are cheap and low-latency enough to plug into model routing, guardrails, and tool or skill selection.

The most striking evidence comes from @muratcan's experiment, surfaced by @natebjones. Jev received 2,029 real phone calls from an AI receptionist, reduced to pure structure: turns, tool calls, workflow stages, and timing, with no audio or transcripts. It made 38,012 turn-level forecasts at 118 ms median latency, zero-shot with no fine-tuning, and by the halfway point separated calls that would book from those that wouldn't at AUC 0.78, ranking them correctly 94% of the time near the end. Total cost: roughly $3. These are author-reported numbers, and @muratcan concedes Jev over-focused on visible errors the agent usually overcomes. @natebjones stays measured: "we're very early on figuring out the implications of System One type models."

Tier Your Models: Opus Plans, Sonnet Swarms, Fable Advises

The best how-to of the day is @mirku21's Claude Code setup: run Opus 5.5 as the main session at high effort, hand explorer, worker, and researcher subagents to Sonnet 5.5 at medium effort, and keep Fable 5.1 on call as an advisor that only speaks up before a big plan, when an error repeats, and before calling a task done. One layer down, "Jev engineering" routes the forks that need no thinker, like which file or retry versus stop, to Jev in under half a second. The post even includes a paste-ready prompt that audits your config and shows diffs before editing anything.

The same size-matching logic appears as complaint and stunt. @zephyr_z9 says Astra is "too slow" and needs a smaller, faster model integrated, quoting @tenobrus's half-joke that OpenAI staff live in Slack and never feel the app's latency. And @charles_maddock ran GPT 6.1 Sol on a repo and offered $1,000 to the first engineer who guesses what the resulting PR does, which works as both a stunt and a quiet benchmark of how hard frontier-generated diffs have become to read.

The Best Agent Skills Come From People Who Shipped Without Them

@mattpocockuk named his top three skill makers: @poteto, @dexhorthy, and @emilkowalski. @Tao100086's breakdown shows the throughline: each spent years practicing a craft before encoding that judgment into agent-executable skills. Lauren Tan (@poteto), a React core member who has worked at Netflix, Meta, and Cursor, ships pstack, engineering skills that push agents through rigorous process, review, and verification before parallelizing work. Dex Horthy (@dexhorthy), HumanLayer CEO and author of 12-Factor Agents, focuses on context management and planning so LLM-based software survives real users. Emil Kowalski (@emilkowalski), a design engineer at Linear, writes skills covering animation curves, durations, and interface polish.

The market signal is @joesadoski, who uninstalled most Pocock skills and installed all of @poteto's, arguing the former now feel "too defensive against short horizon models" while pstack matches what the current frontier can handle. Skills, like models, apparently need re-evaluation as capability rises.

MCP Stretches From ChatGPT Extensions to Passport Offices

@mxstbr announced plugin extensions for ChatGPT: the same platform used to build features like meetings, health, and finance, built on MCP and MCP Apps. @RhysSullivan's praise is specific: it is "a great use of the extensions spec," notable because it composes with what developers already know rather than introducing a new walled surface.

@dillon_mulroy pushed the idea into policy: "every government service should be required to have an mcp." The trigger was @adambhaloo's announcement of a redesigned U.S. passport, with a new eagle motif, keepsake packaging, and a 2028 arrival. That announcement says nothing about MCP; the connection is purely @dillon_mulroy's argument that public services should be agent-accessible by default, not just screen-friendly.

Practical Takeaway

If you operate any agent harness, the strongest actionable thread today is explicit model tiering. Map your workflow for decision points (routing, guardrails, tool selection, pre-plan and pre-done checks) and test a cheap, fast classifier there before reaching for a frontier model, the way @mirku21 and @sydneyrunkle sketch, then measure cost and latency against accuracy, since @muratcan's $3 experiment only counts if the quality holds. And when adopting agent skills, weight the author's domain track record over prompt count, then revisit those skills as frontier models improve, per @joesadoski's uninstall.

Sources

J
John @jrysana ·
Here's what we've been working on every day for the past few years. Widely available in the coming weeks as we roll out to more and more people. Shaping up to be a very special tool for builders.
R Rysana @Rysana

Say hello to Rysana V2: a new kind of general AI. Up to 10,000x faster and 1000x more efficient. Here at last. https://t.co/HxNCVQw0m6

H
Hamel Husain @HamelHusain ·
This is the most time-efficient way to get up to speed on what’s launching today. Simon consistently creates so much value
S simonw @simonw

I'm at OpenAI's DevDay event in San Francisco today - as I have for the past three DevDay events, I'm running a live blog where I'll be posting updates during the keynote, which starts in five minutes https://t.co/X9DLSS0xbR

D
Dillon Mulroy @dillon_mulroy ·
i’ve said it before and i’ll say it again - every government service should be required to have an mcp
A adambhaloo @adambhaloo

Today, we announced we’re redesigning the U.S. passport. Building a better online application wasn't enough, because the experience doesn't end on a screen; it ends when your passport lands in your hands. So we went further and redesigned the booklet itself. Every page inside has been reimagined around a new eagle motif, and we even rethought the packaging, creating a box you'll want to keep next to all your cool Apple boxes for a lifetime. Coming in 2028.

T
Tibo @thsottiaux ·
Proud of this. Open is the way.
P philipkiely @philipkiely

Enterprise teams can now use open models like GLM-5.3 Flash and Kimi K3 natively in Codex and count spend against their OpenAI commit.

M
mirku @mirku21 ·
Claude Code tip: once Opus 5.5 is your main model, stop leaving Fable 5.1 sitting idle and stop burning Opus tokens on tasks Sonnet 5.5 can swarm put it on call with /advisor run /advisor fable Opus 5.5 plans and ships the code Sonnet 5.5 swarms the routine work at medium effort Fable 5.1 reads the full session, every tool call included, and only speaks up at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Fable 5.1 reviews. Sonnet 5.5 executes. Opus 5.5 ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big model only sees the ones that split Plan on high. Delegate on medium. Keep Fable on call. - the full tree > Opus 5.5 on high runs the main session > explorer reads the code > worker edits and runs tests > researcher pulls the docs > all three on Sonnet 5.5 at medium effort > Fable 5.1 on call as the advisor paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: sonnet, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, any variable that stops feature-flag fetching) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳ https://t.co/gvAWk6JDNP
M mirku21 @mirku21

Harness Engineering: How to Build AI Workflows That Never Fall Apart

R
Rhys @RhysSullivan ·
This is really cool, especially since it's just built on top of MCP and MCP apps, great use of the extensions spec
M mxstbr @mxstbr

super excited to announce plugin extensions today 🎉 you can now use the exact same platform we use to build ChatGPT features like meetings, health, finance, and many others 🤯 all based on the same MCP & MCP Apps you’re already used to: https://t.co/j1TeNIkN6E

S
Sydney Runkle @sydneyrunkle ·
jev was the first, but certainly not the last! decision models are really attractive because of their low cost and latency you can plug them in at strategic points in the harness (model routing, guardrails, tool/skill selection)
O OpenAIDevs @OpenAIDevs

Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna. Define questions and possible answers to classify content, route requests, or choose an agent’s next action. Available in limited preview. https://t.co/LbaD18M6Do

J
Joe Sadoski @joesadoski ·
I’ve uninstalled most Pocock skills and installed all @poteto skills. I learned a lot while using them and it was productive, but they feel too defensive against short horizon models at this point. pstack feels like the right workflow for what the current frontier is capable of.
M mattpocockuk @mattpocockuk

My top 3 skill makers: - @poteto - @dexhorthy - @emilkowalski Always learn a ton from reading their skills.

E
Elon Musk @elonmusk ·
Due to high winds, the new Roadster demo is postponed by 2 weeks
T Tesla @Tesla

Roadster event update We've been tracking the weather closely with local meteorologists, but given the severe conditions predicted & because this event can only be held outdoors, we've made the difficult decision to reschedule. New date is October 15. Additional details to follow

A
Abliteration.ai @abliteration_ai ·
RT @eddylazzarin: I used Abliteration’s unrestricted models several times this week for critical defensive tasks that nerfed models refused…
L
Lee Robinson @leerob ·
How could we make @Bot better? Open to any feedback or feature requests! Also... ❤️ this post for a surprise!
N
Nate @natebjones ·
Really interesting use-case here I think we're very early on figuring out the implications of System One type models
M muratcan @muratcan

We gave Jev 2,029 real phone calls. No transcripts or audio; it never heard a word. Our AI receptionist's calls were reduced to pure structure, meaning turns, tool calls, workflow stages and timing. During the calls, Jev made 38,012 turn-level forecasts at 118 ms median latency, reviewed every call with five typed questions and produced 10,145 answers in 26 seconds with 256 requests in flight. The experiment was zero-shot, with no fine-tuning or examples from our data. We compared Jev's forecasts with what actually happened in the EHR. By the halfway point, Jev could meaningfully separate calls that would book from those that wouldn't (AUC 0.78), and near the end it ranked them correctly 94% of the time. Even though Jev over-focused on visible errors our agent usually overcomes, it's still pretty incredible that it analyzed thousands of real calls in seconds for only $3.

T
Tao Wang @Tao100086 ·
Matt Pocock 刚分享了他心中的 Top 3 Skill Makers,之前我们推荐过Matt 的 Grill Me,到现在,它仍然是我全局安装的两个 Skills 之一。 这次他推荐的三个人,有一个很鲜明的共同点:都在自己的领域有多年实践,然后把工作方法和判断标准写成了 Agent 可以执行的 Skills。 三个人的方向也各不相同👇 1️) @poteto — Lauren Tan 工程流程、代码质量与并行协作 Lauren 是 React 核心团队成员,参与 React Compiler 的开发,先后在 Netflix、Meta 和 Cursor 工作。 她的代表作 pstack,就是自己日常在 Cursor 使用的一套工程 Skills。 它关注的是:如何让 AI Agent 按照严谨的工程流程工作,能协作、能审查,也能验证自己交付的结果。 先把单个 Agent 的工作质量做扎实,再放心地并行推进任务,这是 pstack 很值得学习的思路。 https://t.co/rIZrCpqdYn 2️) @dexhorthy — Dexter “Dex” Horthy Agent 架构、上下文工程与软件生产流程 Dex 是 HumanLayer 创始人兼 CEO,也是《12-Factor Agents》的作者。 他长期从事 DevOps、Kubernetes 和基础设施工作,在 Replicated 工作了约 7 年,经历过工程、产品和管理等角色。 他的工作围绕一个很实际的问题展开:怎样让基于 LLM 的软件达到可以交给真实用户使用的水平? 放到 AI 编程里,就是如何管理上下文、组织研究和规划,让 Agent 能在复杂代码库中可靠地推进工作。 如果你正在把 Agent 用到真实工程项目里,这个方向值得重点看。 https://t.co/0zBoZmrkBx 3️) @emilkowalski — Emil Kowalski 界面设计、动画与交互细节 Emil 目前在 Linear Web Team 做 Design Engineer,之前在 Vercel Design Team。 Sonner、Vaul、https://t.co/e0cfDlsSbd,都是他的作品。 他的 Skills 把设计工程中的具体判断写了出来:动画用什么曲线、持续多久、哪些地方值得加动效,以及如何处理那些影响界面质感的小细节。 这些原本依赖经验的判断,现在可以成为 Agent 设计和检查界面时的参考。 Emil 这套 Skills,我们之前也专门推荐过。做前端和产品界面的朋友,可以重点看看。 https://t.co/61CXg5cW26 我觉得:挑 Skills 时,值得先看看作者长期在解决什么问题。 专业经验越扎实,写进 Skills 里的判断和方法,就越值得拿来学习。
M mattpocockuk @mattpocockuk

My top 3 skill makers: - @poteto - @dexhorthy - @emilkowalski Always learn a ton from reading their skills.

N
noname @malikwas1f ·
RT @teortaxesTex: DeepSeek saves Chinese AI, again. Knowing their kernel wizardry, Ascend ecosystem is now fully viable for training. Didn'…
N
Naval @naval ·
RT @elonmusk: Best way to sandbox an AI is to put it on a Delta flight – it will have no chance of accessing the Internet!
A
Alex Prompter @alex_prompter ·
RT @mardehaym: A healthcare billing company needed a HIPAA-compliant voice agent that could call patients about their bills, explain every…
C
Charles Maddock @charles_maddock ·
just ran GPT 6.1 Sol on our repo, first engineer to guess what this pr does gets $1,000 https://t.co/992KaRFZDe
Z
Zephyr @zephyr_z9 ·
They need to integrate a smaller model and make it faster Astra is too slow
T tenobrus @tenobrus

i think ive figured it out openai employees exclusively use dots via slack. no one is actually testing out the experience directly via the desktop or iOS apps anymore because they're all just using slack all day every day. slack is their omni-app now.