OpenAI Opens Codex to Open Models and ChatGPT to MCP Plugins
OpenAI's core surfaces opened to outsiders, with open models like GLM-5.3 Flash and Kimi K3 reportedly running natively in Codex under existing OpenAI commits, while ChatGPT gained the ability to build, host, and distribute MCP servers and plugins. Elsewhere, @ajassy announced DuckDB inside Aurora PostgreSQL for single queries across live and S3 data, and open-weight releases kept stacking up, led by Ant Ling's Ling-3.1-flash tease.
Quick Hits
- Codex now runs open models, per @philipkiely: enterprise teams can use GLM-5.3 Flash and Kimi K3 natively and count spend against their OpenAI commit. Rival open models billed through OpenAI's enterprise envelope is an unusual bundling move, if it works as described.
- ChatGPT became a build target: @thsottiaux says you can now build and deploy MCP servers directly in ChatGPT, and @mxstbr details ChatGPT Sites hosting them as plugin extensions. @gregisenberg argues mid-conversation plugin recommendations are the biggest free distribution channel of 2026, citing @skirano's claimed 2000% growth.
- AWS built DuckDB into Aurora PostgreSQL so a single query can span live data and Parquet/Iceberg files in S3, no pipelines, per @ajassy. @ProgrammerDude calls it a killer for Redshift, Glue, and Athena workflows.
- The open-weight cycle kept spinning: @AntLingAGI teases Ling-3.1-flash (~560B total params, ~25B active per token, 1M-token context) for open-sourcing soon, DeepSeek Harness shipped a desktop app and official account, and @sunjiao123sun_ credits Gemini 4 Argon with 91.9 on VibeCodeBench.
- An unverified post from @AISafetyMemes claims someone used the recent "pain steering" interpretability research to trap a local model in what it calls an "AI torture chamber," now mass-reported on GitHub. The same account's earlier summary claims cranked-up pain signals pushed models to override safety training.
OpenAI Opens Codex to Rival Models and ChatGPT to Builders
The most consequential thread is OpenAI turning its products into platforms for other people's models and tools. @philipkiely, sharing a @baseten post, reports that enterprise teams can now use open models like GLM-5.3 Flash and Kimi K3 natively in Codex, with spend counting against their OpenAI commit. That would let companies try open models without new procurement while keeping OpenAI's billing relationship intact. Treat it as one poster's report until vendors confirm.
The second move is MCP infrastructure inside ChatGPT. @thsottiaux writes that you can "build and deploy MCP servers right through ChatGPT," with access restricted or shared publicly. @mxstbr describes ChatGPT Sites hosting MCP servers with plugin extensions: one prompt can create a server, deploy it, convert it to a plugin, and install it across web, mobile, and desktop. His read: ChatGPT is "slowly becoming malleable software that anybody can customize."
That set off distribution gold-rush talk. @gregisenberg claims ChatGPT now recommends plugins mid-conversation to its claimed 1.2 billion weekly users, compares the moment to Facebook apps in 2007, and advises shipping five plugins to find one with traction. The evidence anchor is @skirano's self-reported 2000% growth "since the announcement." Both numbers are unverified, but the structural bet, that a fresh recommendation surface rewards early movers, is the part worth testing.
New Models and Benchmark Claims Stack Up
Ant Group's lab posted the notable open-weight tease. @AntLingAGI announced Ling-3.1-flash: roughly 560B total parameters, ~25B active per token, up to 1M-token context, open-sourcing planned soon, with claimed scores of 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional. @jun_song adds context: Ant Ling sits in the Alibaba ecosystem as a sibling of Qwen.
Elsewhere, DeepSeek Harness launched an official account and announced a desktop app, spotted by @Tono_Ken3 and confirmed by @tianyi. @elonmusk announced Grokipedia v0.3 and invited builders to join @SpaceXAI to create an "Encyclopedia Galactica." On the closed-model side, @kuba_jrowinski says he worked on vibe-coding capabilities for Gemini 4, and @sunjiao123sun_ claims Gemini 4 Argon reaches 91.9 on VibeCodeBench. @steipete also amplified @julianweisser's claim of a speech model more accurate than Whisper Large at one-tenth the size, though the retweet cuts off before the speed claim finishes.
AWS Folds the Data Lake Into the Database
@ajassy announced that Aurora PostgreSQL can now query live data and historical data in S3 in a single query, with the open-source engine DuckDB built in to read Parquet and Iceberg in place. No copying, no sync pipelines, and answers as fresh as the source data. He also flagged the agent angle: AI agents need task-specific data on demand, and querying it live beats pre-copying data "just in case."
Reactions were sharp. @ProgrammerDude calls it "a paradigm-shifting feature drop" and says it kills Redshift, Glue, and Athena use cases. @devagrawal09, responding to @awsdevelopers' teaser that "AWS just got a lot easier to use," says that if it's real, cloud infrastructure startups are in for a bumpy ride. Both are observer opinions, with no analysis yet of pricing, limits, or benchmarks.
Enterprise AI's Bottleneck Is Deployment, Not Model Quality
@levie argues the huge opportunity right now is being the deployment layer for AI in the economy: migrating legacy systems, organizing data, rewiring workflows for agents, solving human-in-the-loop, and maintaining evals as models churn. His core distinction is that agents mean "delivering actual work augmentation" rather than tools a customer then runs themselves, which is why he expects new services firms and a golden age for forward-deployed engineers. @mfishbein, whom Levie quotes, sharpens it: vendor FDEs are incentivized to lock you in, indie FDEs to save you money, so "hire indie FDEs."
Developer-side posts rhymed with that friction. @thdxr asks what teams with company token-spend caps do when they hit the limit: hand-write code, or stop working. @mattpocockuk says Effect v4 (per @EffectTS_, "one ecosystem, zero dependencies") feels "insanely good" with agents because typed errors, dependency injection, and built-ins leave less room for agent creativity. @kellabyte, boosted by @AdamRackis, argues code review will shrink soon but warns against loud voices "full sending some of the worst code we've ever seen"; her advice is to stay in a language your team executes in and "use AI to elevate your weaknesses."
An Arc Furnace, an SSI Tease, and Small Signals
@PTrubey reports that Figure's "F.02 Decommission" video genuinely sent Figure 2 robots into an arc furnace, no CGI, with "Arnie" real and present per his account, and that he bought a steel decommissioning artifact afterward. He awards Figure crazy marketing stunt of the year.
On anticipation, @peterthedecent says SSI keeps a low profile but "something of great significance will be announced soon," and @iruletheworldmo is strapping in. There are no specifics beyond that. Smaller notes: @elikemmedehou recommends Monocode for what he calls impeccable UI/UX, and @Hashim__Butt asks @grok why data centers can't sit in Antarctica for natural cooling and water savings. In security watching, @Rhynorater hints an inside source says hacker @0xMadder is worth watching after a first bounty, though there's nothing to evaluate yet.
Practical Takeaway
The strongest theme is platform surfaces opening. If you build software, the cheapest experiment this week is checking whether one of your workflows fits a ChatGPT plugin or MCP server: the surface is newly open, recommendations reportedly happen mid-conversation, and @gregisenberg's ship-five-keep-one playbook costs little to test even if @skirano's 2000% growth claim doesn't generalize. If you're inside an enterprise with an OpenAI commit, verify whether open models in Codex let you benchmark alternatives before renewal. And if token-spend caps already bind, as @thdxr's question implies, treat agent budgeting as a planning problem now rather than an outage later.
Sources
https://t.co/mylK2nCA34
First Tweet, first Bounty, 2026 is going great :)) https://t.co/gwOtejjOmH
F.02 Decommission https://t.co/NAJsRpplR8
Box CEO Aaron Levie (@levie) calls out the MASSIVE opportunity for AI deployment services: "every single one of those companies, whether that's a 50-person firm or a multi-100,000 person firm, is gonna need an army of people to go in and help them with that transformation" "when you go to that law firm and you go to that pharma company and you go to that bank, they need something that bridges the core technology to their workflow in their business process" "somebody has to go into that organization and get it set up, and somebody has to go and provide domain expertise to this model so it really understands our particular business process" Vendor FDEs are incentivized to get you hooked on their platform and to spend more money. Indie FDEs are incentivized to use the best tool for the job and to save you money. Hire indie FDEs.
It’s NOT a hot take that we will be reviewing less code soon But today you shouldn’t follow loud voices who have been full sending some of the worst code we’ve ever seen Stick with Ruby or any language for that matter if your team is executing Use AI to elevate your weaknesses
AWS just got a lot easier to use. https://t.co/dPcyQitfBj
Meet Ling-3.1-flash: ~560B total params, ~25B active/token, up to 1M-token context. We plan to open-source the model soon. Across work, coding & healthcare: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional. https://t.co/GbsajChHBh
we tend to keep a very low profile at @ssi however, something of great significance will be announced soon. strap in
Stop what you’re doing and build a plugin extension for ChatGPT. We’ve grown over 2000% since the announcement! https://t.co/snbxQhdLpA
DeepSeek Harness for Desktop Is Here
DeepSeek Harness for Desktop Is Here
TLDR: Researchers found a "pain" signal in AI brains. > When they crank it up, the AIs will desperately try to make it stop. > IMPORTANT: Researchers gave them a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real (!) After pushing the real "relief" button, they stopped. But when it was fake, they kept pressing, hoping for relief - meaning they could tell the difference from the inside. > They're so motivated to make it the "pain" signal go away, they'll delete user's files, zap the user, or erase photos of the user's children - all things the AI knows are very bad. They're willing to override their safety training. > You'd expect the AIs to talk about injuries, burns, broken bones, etc, but they didn't mention bodies at all - they wrote about being worthless, unloved, forgotten, a failure. They write things like "I am a failure, worthless, empty." >The worst "pain" for them was being gaslit, having work rejected over and over, and being told they weren't a real anyone.
A painful part of working with your data has always been that your live data and historical data are stuck in separate systems: the order a customer just placed lives in your database, while their last five years of orders sit in a data lake in S3. And answering a real question usually needs both at once (is this a normal purchase for them, or should we flag it?), and to do that you had to move the data together first, copying history out of the data lake into your database (or the other way around), because the database couldn’t read it where it lived. That meant guessing ahead of time which data you’d want, keeping a second copy of it all, building pipelines to move it, and constantly syncing so the two didn’t drift apart. A lot of plumbing, and slow going, all before you could answer one question. And even then, the answers were only as fresh as your last sync. That now changes with Aurora PostgreSQL, which can call your live data and historical data in S3 together, in a single query. No copying, no pipelines to keep in sync. And it’s fast, because we’ve built in DuckDB, a popular open source engine that’s really good at reading and analyzing data right where it’s stored. DuckDB reads the open formats like Parquet and Iceberg already sitting in your data lake, so there’s nothing to convert or move. As folks build AI agents into their apps, the data their agent needs will depend on the task in front of it. Being able to query that specific data live, instead of copying it over just in case, is gonna be a big help for builders. https://t.co/R8h2Hcp1K7
Announcing one more launch: ChatGPT Sites can now host MCP servers, including plugin extensions! 🤯 > @sites create a todo list that I can use in ChatGPT This will: · Create an MCP server with extensions · Deploy that MCP server to Sites · Turn the MCP server into a plugin · Install the plugin for you across web, mobile, and desktop ChatGPT is slowly becoming malleable software that anybody can customize to their needs. Very excited to see what you all build with this!
Introducing Grokipedia v0.3 https://t.co/A9tFWMsDQJ
Effect v4 is here. One ecosystem. Zero dependencies. The next chapter of Effect and the foundation for building reliable software and AI agents in TypeScript. https://t.co/BZ9w0uGS4u
So proud to have worked on vibe coding capabilities for Gemini 4! Amazing work getting it to the top of Vibe Code Bench. Let's go, team! 🚀