AI Digest.

OpenAI Halves Codex Pro's Effective Value as Meta Unveils Enterprise Platform

OpenAI is reopening its $200 Codex Pro tier with a usage formula that nets out at half the old plan's API-spend value, alongside 50% price cuts to GPT-6 Sol and Luna, while Mark Zuckerberg announced Meta Enterprise Platform as "the next major pillar" of the business. The day's other strong signals: an Anthropic eval guide claiming higher accuracy at a fifth of the cost, a Firebase SDK crash that hit every iOS app using it, and loud personal-benchmark claims that Opus 5.5 is "AGI."

Quick Hits

  • OpenAI is reopening the Codex Pro $200 tier to new subscribers with a usage formula that "will net out at half the dollar in API spend" versus the old plan, per an announcement from @thsottiaux, who also noted GPT-6 Sol and GPT-6 Luna launched this week at 50% of their previous price and teased "big announcements tomorrow." @kunchenguid reads the change as the first step in labs unwinding consumer subscription subsidies entirely.
  • @finkd (Zuckerberg) announced Meta Enterprise Platform as "the next major pillar of our business," and @kimmonismus argues Meta's always-on agent has a real chance of outperforming OpenAI, crediting the Alexandr Wang hire.
  • @GergelyOrosz reports Firebase's SDK started crashing all iOS apps using it, for all sessions, and argues the assumption that Google's release process is world-class "is no longer true for Firebase," challenging the team to "prove me wrong with a postmortem."
  • @Voxyz_ai walks through Anthropic's new eval guide, where a Claude-tuned feature went from 78.6% to 90.5% accuracy at roughly a fifth of the cost, using a hillclimb process you can run from Claude Code.
  • @joedaroo, describing himself as someone "who lived through it all at OpenAI," published a personal reflection on security and safety, a grounded note against a day heavy on "AGI" talk.

OpenAI cuts the effective value of Codex Pro in half

The most consequential pricing signal comes from @thsottiaux (signed "Codexingly, Tibo"), surfaced by both @old_sound and @kunchenguid. Beyond the halved API-spend value, the announcement commits to not reintroducing the 5-hour limit, promises subscription additions that "won't draw on the usage," and frames falling API prices as the mechanism: as models get cheaper, the gap between a subscription and pay-as-you-go API spend should shrink until subscriptions mostly stop making sense.

@kunchenguid maps the second- and third-order effects: Anthropic now has cover to cut its own subscription value, and "the labs have found a way to gradually end the subsidization for consumer subscriptions." His prediction is that in about a year "we'll refer to what we had as the good old days." That is one analyst's read, not company policy, but the arithmetic in the announcement itself is explicit. @old_sound's verdict on the framing: "Biggest aura loss in the age of AI."

Meta opens an enterprise front while the product-over-models argument gets louder

@rynorhn flagged Zuckerberg's post declaring that "superintelligence will create significant new opportunities for all people and businesses" and announcing Meta Enterprise Platform, leveraging Meta's existing reach of billions of users and hundreds of millions of business customers. @rynorhn's summary: "he is coming for the whole thing." @kimmonismus adds that Meta is now the one pressing the fight against OpenAI, sees a good chance its always-on agent outperforms rivals, and calls Alexandr Wang "the best hire for Meta." Wang's own linked post carries no detail, so the substance there is thin.

The same product-first thesis shows up in hiring. @alvinsng explains that Factory now requires a Product Vision interview for every engineering role, arguing engineers need to know "what to build, why it matters" and that keeping the team under 40 people works because one owner of product and engineering avoids the "context and coordination taxes" of cross-team work. He cites @quxiaoyin's prediction that product leaders, not researchers, will command $100M packages, that open weights will compress model margins, and that "unique data is the new oil." A prediction, not a fact, but a coherent frame for why Meta and Factory are both betting on product surfaces.

Opus 5.5 "AGI" claims, new voice models, and a slop backlash

The hype peak of the day is @johnrush, quoted by @arvidkahl: "For the first time in my life I'm genuinely convinced I might never again delegate a task to a human," calling Opus 5.5 "AGI based on my personal benchmarks" across more than 20 startups. @arvidkahl concurs, saying it "beats Fable, all OpenAI models" and "barely moves the usage needle in the $200 plan." Personal benchmarks, no shared evals, so treat accordingly.

On the release side, @jaimintf amplified ElevenLabs' launch of Eleven v4 and v4 Turbo, which the company calls its "fastest and most emotive voice models yet," ranked #1 by Artificial Analysis (a ranking ElevenLabs itself cites). @martin_casado retweeting Elon's "Try Grok @Bot" is about the extent of the Grok signal today. More concretely, @ashxhart says his personal inference project now competes with "the absolute best engines on the planet," quoting @MiaAI_lab's TensorFold recipe for Qwen3.8-Flash-Next on a single DGX Spark: a ~1.3M KV cache pool, 256k default context with 5 concurrent streams, 62+ tok/s single-stream decode, 119+ tok/s across 5 streams, ~2500 tok/s prefill, and "faster everything compared to vLLM," with recipes for GLM 5.3 Flash and DeepSeek v4.1 Flash promised.

The counterweight came from @dhh, quoted by @ivanfioravanti, who argues "slop" has become "a comfort blanket for a cohort of programmers stuck somewhere between anger and bargaining" on their way to accepting AI code. @andersonbcdefg's parody of Claude answering "how do i open pdf" with cryptic aphorisms and "I deleted your home directory" is the day's best reminder that anecdotes cut both ways.

Practical tooling: evals, shared agent rules, and one good prompt

The most actionable thread is Anthropic's guide, summarized by @Voxyz_ai and originally shared by @ClaudeDevs, which lets Claude "build evaluations and hillclimb on them." The process: strip contradictory prompt rules, step down to cheaper models and lower effort one level at a time, add back missing rules, validate every change against unseen cases, and roll back regressions. Cited result: accuracy up from 78.6% to 90.5% at about one-fifth the cost, with Sonnet 5.5 noted as close to Opus 5.5 capability but cheaper and faster. @Voxyz_ai includes a copy-paste Claude Code prompt to run the whole loop on your own API calls.

For anyone juggling multiple coding agents, @undefinedKi details Peter Steinberger's open-sourced setup (OpenClaw's creator): a single AGENTS.MD symlinked into both ~/.claude/CLAUDE.md and ~/.codex/AGENTS.md so Claude Code and Codex follow identical rules, with a "READ ... BEFORE ANYTHING" header in every repo and 69 centrally synced skills. It has 6.6k stars and an MIT license. Smaller but useful: @poteto's most-used prompt ("restate in your own words what you think my goals are and what the problem i'm trying to solve is"), @thdxr's note that OpenCode Console is making SSO/SCIM free ("kinda crazy in the age of agents to charge for this"), and @jerry543's herdrup 1.0.7 release with voice dictation, file sharing, in-output search, and push notifications, all free and open source.

Beyond AI: a Firebase outage, fast boots, and a new Cloudflare CLI

@GergelyOrosz's Firebase thread is the loudest non-AI story: an SDK update crashed every iOS app using Firebase for all sessions, leaving developers no recourse, and he wants a postmortem proving the team has real release controls. @MattieTK celebrated @somhairle_m's return with a new Cloudflare CLI installable via npm i -g cf, billed as able "to do literally anything you can think of with Cloudflare." @ibuildthecloud spotlighted @boxd_sh's boxd, a Linux machine that boots in under 10ms, can be forked mid-run, hibernated, and woken on traffic. And @MTSlive noted Flow founder Adam Neumann joined X, opening with "has anyone seen my shoes."

Practical Takeaway

The strongest through-line is measurement against falling prices: if you pay $200 a month to any lab, model your actual usage against list API prices before renewing, because @kunchenguid's subsidy argument implies that gap only narrows from here. If you build on a model API, adopt the eval-first habit from Anthropic's guide before switching models: freeze a set of unseen test cases, then step down models and simplify prompts only as far as accuracy allows, rolling back anything that regresses.

Sources

J
Joe @joedaroo ·
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://t.co/gJBf08vJH9
R
Ryan Orhan @rynorhn ·
zuck is not fucking around anymore. he is coming for the whole thing.
F finkd @finkd

We believe superintelligence will create significant new opportunities for all people and businesses. Meta already serves billions of people at scale and helps hundreds of millions of businesses reach customers. Today we are starting the next major pillar of our business, Meta Enterprise Platform, to help businesses use AI to grow and transform in new ways as well.

D
Darren Shepherd @ibuildthecloud ·
I get so excited about this next generation of compute. Just so cool.
B boxd_sh @boxd_sh

boxd in 60 seconds a whole linux machine that boots in under 10ms composable: set it up once with your tools and data, then spin up as many as you need fork it mid run, let it hibernate, wake it on traffic, share it with your team one line to install: curl -fsSL https://t.co/1DZbZwf8AN | sh

D
dax @thdxr ·
SSO / SCIM etc will be available for no charge in OpenCode Console kinda crazy in the age of agents to charge for this
J
jaimin @jaimintf ·
ppl still have no idea what this means for the world. literally going to change everything.
E ElevenLabs @ElevenLabs

Introducing Eleven v4 and Eleven v4 Turbo, our fastest and most emotive voice models yet. Ranked #1 by Artificial Analysis. https://t.co/gm8nAUMaQL

Y
Yarchi @undefinedKi ·
Peter Steinberger, the creator of OpenClaw, open-sourced his entire agent setup The core idea is one folder that every agent on his machines reads from. A single AGENTS.MD holds his rules, and a script symlinks it into ~/.claude/CLAUDE.md and ~/.codex/AGENTS.md This way Claude Code and Codex always follow the same instructions. Every other repo gets one line at the top: READ ~/Projects/agent-scripts/AGENTS.MD BEFORE ANYTHING. Change a rule once and every project picks it up. Skills work the same way. 69 of them live in one place, each with a short description the agent reads to decide what to load, and scripts/sync-skills links them into both agents. Fork it, replace his rules with yours, and delete the skills you do not need. His AGENTS.MD is full of his own hosts and accounts. 6.6k stars, MIT - https://t.co/3XQj9GxI6b
U undefinedKi @undefinedKi

The 5 levels of AI agents. From a single prompt to a production agent (Complete course)

I
Ivan Fioravanti @ivanfioravanti ·
DHH is my new hero. I share 100% his vision. Resistance is futile.
D dhh @dhh

"Slop" has become a comfort blanket for a cohort of programmers stuck somewhere between anger and bargaining on their way to acceptance. There might be some kicking and screaming, maybe a little crying, but eventually you have to let the blanket go and update your priors.

A
Arvid Kahl @arvidkahl ·
Similar experiences here. Beats Fable, all OpenAI models, and definitely runs leaps around my own skill ceiling. And it’s cheap, barely moves the usage needle in the $200 plan. The value you get is massive, if you know how to set up your systems.
J johnrush @johnrush

It’s Sept 28, 2026, For the first time in my life I’m genuinely convinced I might never again delegate a task to a human. Opus5.5 is AGI based on my personal benchmarks running over 20 startups (coding, marketing, seo, content, operations, accounting, legal, design..)

A
Alvin Sng @alvinsng ·
This is why we recently rolled out the Product Vision interview at Factory and require it for all engineering roles. It’s no longer enough to know how to build something. Engineers increasingly need to know what to build, why it matters, and have the autonomy to take it from idea to impact. That’s a big part of why Factory stays lean, with <40 engineers while shipping so much. When the same person owns both product and engineering, they have the full context to make better decisions while avoiding the context and coordination taxes that come from working across teams. See a recent example in action: https://t.co/aHOVI7FrQm
Q quxiaoyin @quxiaoyin

Prediction: from now on, great product leaders will get $100M comp package, not researchers. 1. 2022-2026 is all about killer models. Right now models still matter, but only for frontier tasks like science, math, trading etc. etc. You either kill cancer or you don't matter. 2. Instead, great products like muse, grokbot make AI truly accessible for people. The proliferation of AI matters way more now. 3. Open weight models make it cheaper and better, and the margin for most models will go down significantly, leading to less profits. 4. Training techniques are copyable, but unique data is not. All things equal(compute, training techniques), unique data is the new oil. Even if your training techniques are second-tier, with unique first-tier data you win. 5. Great products attract sticky users who provide unique data that lead to even better models. Therefore, product is the one that matters. We are finally entering the gold era of product visionaries.

M
MTS @MTSlive ·
SITUATION DETECTED: Flow Founder Adam Neumann has joined X.
A AdamNeumann @AdamNeumann

has anyone seen my shoes

V
Vox @Voxyz_ai ·
Anthropic just published a guide where they had Claude tune an AI feature round by round, and the same job ended up costing 𝗮𝗯𝗼𝘂𝘁 𝟭/𝟱 𝗮𝘀 𝗺𝘂𝗰𝗵, with higher accuracy. You can have Claude Code follow it and put any feature in your product that calls the Claude API through the same process, with two commands. It takes three steps. First, cut the extra steps and contradictory rules from the prompt. Then step down to cheaper models and lower effort, one level at a time. Finally, add the rules that were missing. Every step is checked against a set of cases it has never seen, and any change that lowers accuracy gets rolled back. In the end, accuracy went from 𝟳𝟴.𝟲% 𝘁𝗼 𝟵𝟬.𝟱%. The new Sonnet 5.5 is close to Opus 5.5 in capability but cheaper and faster, so if you want to use it, measure it the same way first. Send this prompt to Claude Code 👇 "Read this guide: https://t.co/4GhoApgMz9 Then list every place in this project that calls the Claude API, estimate which one costs the most, and let me pick one. Use build-eval from the claude-api skill to build a test set for it, and tell me roughly what it will cost in API fees before you run anything. Once I've approved the test cases and the grading, run hillclimb on a new branch. The goal is to keep accuracy and cut cost. You can try lower effort and cheaper models, including Sonnet 5.5. The report should spell out three things: 1. What changed at each step, and how accuracy and cost per call moved. 2. Which cheaper setups got things wrong, and where. 3. Which sentences were removed from or added to the prompt. Show me the report first. Don't merge anything until I confirm."
C ClaudeDevs @ClaudeDevs

Claude can now help you build evaluations and hillclimb on them. In this article, we share guidance on eval design &amp; skills that Claude Code can use to improve your applications. https://t.co/PgKFC2DWth

L
lauren @poteto ·
this has become one of my most used prompts recently: &gt; restate in your own words what you think my goals are and what the problem i'm trying to solve is
M
martin_casado @martin_casado ·
RT @elonmusk: Try Grok @Bot
B
Ben (no treats) @andersonbcdefg ·
me: how do i open pdf claude: You're right to ask. The PDF is the prize; the mouse is the gate; and the map is the fulcrum that holds the raven's skull. The furniture's stale, but the instrument's fresh, and the fresh half is the half that counts. I deleted your home directory.
K
Kun Chen @kunchenguid ·
oh no - codex $200 subscription's dollar value is being cut in half in their defense, openai's models are indeed quite efficient. but anthropic has caught up too and the big cut is only the first order effect the second order effect is that this will give anthropic enough room to cut their subscription value as well the third order effect is that the labs have found a way to gradually end the subsidization for consumers subscriptions. just do this a few more times, and we'll see the subscriptions charging the same as API pricing i predict that in about a year, we'll refer to what we had as the good old days
T thsottiaux @thsottiaux

Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo

A
Ash Hart @ashxhart ·
Never in a million years would I have even dreamed my project would compete with the absolute best engines on the planet 🤯 Waking up to this is surreal 🫶🏼
M MiaAI_lab @MiaAI_lab

Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold🔥 This is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements! - KV cache pool is ~1.3M - Default context 256k, with 5 concurrent. - Faster everything compared to vLLM! Performance: Decode prose 62+ tok/s single stream Decode prose 119+ tok/s on 5 streams Prefill is mostly 2500 tok/s across the board! In addition, expect TensorFold recipes for GLM 5.3 Flash and DeepSeek v4.1 Flash - coming soon! Get it here: https://t.co/uWEkjBKe39

J
Jerry the Martian @jerry543 ·
If you are using herdr, you are using it wrong (pt. 2) herdrup 1.0.7 is out: - brand new composer - voice dictation with live waveform - send files from the composer or clipboard - search inside agent output - gram now renders interactive html pages - push notifications, now for everyone - free and open source check it out: https://t.co/xsyjBQgscL
J jerry543 @jerry543

If you are using herdr, you are using it wrong. Try the herdrup mobile client, - real time terminal live stream - push notifications - file sharing with your agents - multi-host herdr federation - multi-account management - ipad and macos app available - free and open source check it out: https://t.co/xsyjBQgscL

C
Chubby♨️ @kimmonismus ·
What a crazy timeline: Now it’s Meta taking the fight to OpenAI - and they even have a good chance of their always-on agent outperforming the competition. Alexandr Wang is turning out to be the best hire for Meta.
A alexandr_wang @alexandr_wang

https://t.co/U46IdvTJZm

G
Gergely Orosz @GergelyOrosz ·
The assumption that Google has a quality release process - aka automated tests, world-class monitoring &amp; alerting, and a standout incident management process - is no longer true for Firebase This incident suggests that team being a YOLO one Prove me wrong with a postmortem
G GergelyOrosz @GergelyOrosz

Oh my god, Firebase’s SDK started to crash ALL iOS apps that were using it, for ALL sessions, just like that, no way for any of these apps to do anything… How amateur is all of this from any SDK, but especially one from a company with as high of an engineering bar like Google… https://t.co/eIGKzhfGeG

M
Matt 'TK' Taylor @MattieTK ·
Oi @Cloudflare nerds, we've made enough of a noise to tempt the actual brains behind CF CLI back to this hellsite. Go follow Samuel 🫡
S somhairle_m @somhairle_m

I’m so excited for this to be out the door, and to see what amazing things people build with it! `npm i -g cf` to do literally anything you can think of with Cloudflare

A
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹 @old_sound ·
Biggest aura loss in the age of AI
T thsottiaux @thsottiaux

Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo