Decision Models Get an API While Open Models Land in Codex
Around OpenAI's DevDay, the feed carried a limited-preview Decisions API powered by GPT-6 Luna and word that enterprises can now run open models like GLM-5.3 Flash and Kimi K3 natively in Codex. Practitioners filled the rest with concrete technique: a $3 Jev analysis of 2,029 phone calls, multi-model Claude Code orchestration, and practitioner-authored agent skills.
Quick Hits
- OpenAI's DevDay keynote got the usual liveblog from @simonw, which @HamelHusain calls "the most time-efficient way to get up to speed on what's launching today." In the same window, @philipkiely reported that enterprise teams can run open models like GLM-5.3 Flash and Kimi K3 natively in Codex, with spend counting against their OpenAI commit; @thsottiaux cheered it with "Open is the way."
- Cheap, fast decision models keep gaining ground: @OpenAIDevs announced a Decisions API powered by GPT-6 Luna (limited preview), and @muratcan described Jev forecasting outcomes across 2,029 real phone calls for about $3, which @natebjones reads as early evidence of what System One-style models can do.
- Constraints shaped two deployment stories: @alex_prompter amplified @mardehaym's thread on a HIPAA-compliant voice agent for healthcare billing calls, and @abliteration_ai relayed @eddylazzarin turning to unrestricted models for defensive security tasks after restricted ones refused.
- Two big claims, thin detail: @jrysana teased Rysana V2 as "up to 10,000x faster and 1000x more efficient," and @malikwas1f passed along @teortaxesTex's claim that DeepSeek's kernel work makes the Ascend ecosystem "fully viable for training."
- Elsewhere: @elonmusk postponed the Roadster demo to October 15 over high winds, @naval recirculated Musk's joke that a Delta flight is the best AI sandbox, and @leerob asked for feedback and feature requests for @Bot.
Open Models in Codex, With the Invoice Still Pointing at OpenAI
@philipkiely's one-liner is the day's most concrete enterprise item: open models like GLM-5.3 Flash and Kimi K3 can now be used natively in Codex, and the spend counts against an existing OpenAI commitment. @thsottiaux's reaction, "Proud of this. Open is the way," frames it as openness winning. The skeptic's read is that openness now arrives with a lock-in flavored billing arrangement: bring your own model, but the money still flows through the same commit. Both takes are opinions, and the underlying claim comes from a single post worth confirming against official docs.
Meanwhile @simonw liveblogged the DevDay keynote in San Francisco, a ritual @HamelHusain endorses as the fastest summary of what actually launched. If the announcements above landed at or around DevDay, no post in the feed says so explicitly, so treat the timing as suggestive rather than confirmed.
Decision Models Graduate From Hack to Product
@OpenAIDevs announced Decisions API, powered by GPT-6 Luna: define questions and possible answers to classify content, route requests, or choose an agent's next action. It is in limited preview. @sydneyrunkle's take is that "jev was the first, but certainly not the last," because decision models are cheap and low-latency enough to plug into model routing, guardrails, and tool or skill selection.
The most striking evidence comes from @muratcan's experiment, surfaced by @natebjones. Jev received 2,029 real phone calls from an AI receptionist, reduced to pure structure: turns, tool calls, workflow stages, and timing, with no audio or transcripts. It made 38,012 turn-level forecasts at 118 ms median latency, zero-shot with no fine-tuning, and by the halfway point separated calls that would book from those that wouldn't at AUC 0.78, ranking them correctly 94% of the time near the end. Total cost: roughly $3. These are author-reported numbers, and @muratcan concedes Jev over-focused on visible errors the agent usually overcomes. @natebjones stays measured: "we're very early on figuring out the implications of System One type models."
Tier Your Models: Opus Plans, Sonnet Swarms, Fable Advises
The best how-to of the day is @mirku21's Claude Code setup: run Opus 5.5 as the main session at high effort, hand explorer, worker, and researcher subagents to Sonnet 5.5 at medium effort, and keep Fable 5.1 on call as an advisor that only speaks up before a big plan, when an error repeats, and before calling a task done. One layer down, "Jev engineering" routes the forks that need no thinker, like which file or retry versus stop, to Jev in under half a second. The post even includes a paste-ready prompt that audits your config and shows diffs before editing anything.
The same size-matching logic appears as complaint and stunt. @zephyr_z9 says Astra is "too slow" and needs a smaller, faster model integrated, quoting @tenobrus's half-joke that OpenAI staff live in Slack and never feel the app's latency. And @charles_maddock ran GPT 6.1 Sol on a repo and offered $1,000 to the first engineer who guesses what the resulting PR does, which works as both a stunt and a quiet benchmark of how hard frontier-generated diffs have become to read.
The Best Agent Skills Come From People Who Shipped Without Them
@mattpocockuk named his top three skill makers: @poteto, @dexhorthy, and @emilkowalski. @Tao100086's breakdown shows the throughline: each spent years practicing a craft before encoding that judgment into agent-executable skills. Lauren Tan (@poteto), a React core member who has worked at Netflix, Meta, and Cursor, ships pstack, engineering skills that push agents through rigorous process, review, and verification before parallelizing work. Dex Horthy (@dexhorthy), HumanLayer CEO and author of 12-Factor Agents, focuses on context management and planning so LLM-based software survives real users. Emil Kowalski (@emilkowalski), a design engineer at Linear, writes skills covering animation curves, durations, and interface polish.
The market signal is @joesadoski, who uninstalled most Pocock skills and installed all of @poteto's, arguing the former now feel "too defensive against short horizon models" while pstack matches what the current frontier can handle. Skills, like models, apparently need re-evaluation as capability rises.
MCP Stretches From ChatGPT Extensions to Passport Offices
@mxstbr announced plugin extensions for ChatGPT: the same platform used to build features like meetings, health, and finance, built on MCP and MCP Apps. @RhysSullivan's praise is specific: it is "a great use of the extensions spec," notable because it composes with what developers already know rather than introducing a new walled surface.
@dillon_mulroy pushed the idea into policy: "every government service should be required to have an mcp." The trigger was @adambhaloo's announcement of a redesigned U.S. passport, with a new eagle motif, keepsake packaging, and a 2028 arrival. That announcement says nothing about MCP; the connection is purely @dillon_mulroy's argument that public services should be agent-accessible by default, not just screen-friendly.
Practical Takeaway
If you operate any agent harness, the strongest actionable thread today is explicit model tiering. Map your workflow for decision points (routing, guardrails, tool selection, pre-plan and pre-done checks) and test a cheap, fast classifier there before reaching for a frontier model, the way @mirku21 and @sydneyrunkle sketch, then measure cost and latency against accuracy, since @muratcan's $3 experiment only counts if the quality holds. And when adopting agent skills, weight the author's domain track record over prompt count, then revisit those skills as frontier models improve, per @joesadoski's uninstall.
Sources
Say hello to Rysana V2: a new kind of general AI. Up to 10,000x faster and 1000x more efficient. Here at last. https://t.co/HxNCVQw0m6
I'm at OpenAI's DevDay event in San Francisco today - as I have for the past three DevDay events, I'm running a live blog where I'll be posting updates during the keynote, which starts in five minutes https://t.co/X9DLSS0xbR
Today, we announced we’re redesigning the U.S. passport. Building a better online application wasn't enough, because the experience doesn't end on a screen; it ends when your passport lands in your hands. So we went further and redesigned the booklet itself. Every page inside has been reimagined around a new eagle motif, and we even rethought the packaging, creating a box you'll want to keep next to all your cool Apple boxes for a lifetime. Coming in 2028.
Enterprise teams can now use open models like GLM-5.3 Flash and Kimi K3 natively in Codex and count spend against their OpenAI commit.
Harness Engineering: How to Build AI Workflows That Never Fall Apart
super excited to announce plugin extensions today 🎉 you can now use the exact same platform we use to build ChatGPT features like meetings, health, finance, and many others 🤯 all based on the same MCP & MCP Apps you’re already used to: https://t.co/j1TeNIkN6E
Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna. Define questions and possible answers to classify content, route requests, or choose an agent’s next action. Available in limited preview. https://t.co/LbaD18M6Do
My top 3 skill makers: - @poteto - @dexhorthy - @emilkowalski Always learn a ton from reading their skills.
Roadster event update We've been tracking the weather closely with local meteorologists, but given the severe conditions predicted & because this event can only be held outdoors, we've made the difficult decision to reschedule. New date is October 15. Additional details to follow
We gave Jev 2,029 real phone calls. No transcripts or audio; it never heard a word. Our AI receptionist's calls were reduced to pure structure, meaning turns, tool calls, workflow stages and timing. During the calls, Jev made 38,012 turn-level forecasts at 118 ms median latency, reviewed every call with five typed questions and produced 10,145 answers in 26 seconds with 256 requests in flight. The experiment was zero-shot, with no fine-tuning or examples from our data. We compared Jev's forecasts with what actually happened in the EHR. By the halfway point, Jev could meaningfully separate calls that would book from those that wouldn't (AUC 0.78), and near the end it ranked them correctly 94% of the time. Even though Jev over-focused on visible errors our agent usually overcomes, it's still pretty incredible that it analyzed thousands of real calls in seconds for only $3.
My top 3 skill makers: - @poteto - @dexhorthy - @emilkowalski Always learn a ton from reading their skills.
i think ive figured it out openai employees exclusively use dots via slack. no one is actually testing out the experience directly via the desktop or iOS apps anymore because they're all just using slack all day every day. slack is their omni-app now.