Jane Street's $200M-per-Megawatt Cerebras Claim Puts OpenAI's Inference Speed in Focus
A widely shared breakdown from @cryptopunk7213 claims Jane Street bought up Cerebras capacity OpenAI was using and pays $200M per megawatt for ultrafast inference, while @sama confirms only a deep partnership on speed. Agent builders also rallied around Pi Durable and OptChat-style memory, and Aleph Alpha open-sourced its 78B-parameter Kolibri model under Apache 2.0.
Quick Hits
- The Cerebras story now has dollar figures, per @cryptopunk7213 (Ejaaz): he claims OpenAI ran out of Cerebras capacity because Jane Street bought the supply, and that Jane Street pays $200M per megawatt against a Cerebras cost of roughly $26M. @sama confirmed only that Cerebras is "a close partner" pushing "the frontiers of speed."
- @Aleph__Alpha open-sourced Kolibri: 78B parameters, 3.46B active, up to 1M tokens of context, Apache 2.0. @charles_maddock claims it beats Qwen, Mistral, and Nematron models at similar active parameter counts.
- Durable execution had a day: @pidotdev posted a Pi Durable deep-dive with @badlogicgames and @mitsuhiko, while @raunakdoesdev and @tokumin paired Pi Durable with @VictorTaelin's OptChat memory recipe. @ishaansehgal countered that the harness was never the hard part.
- @Rames_Jusso flags a writeup on building a production video studio around Opus 5.5, arguing frameworks like @HyperFrames_ matter more than one-shot generation.
- @deanwball shared the "just screwed" scenario he gives senior AI figures: abundant cognitive labor plus the vulnerable world hypothesis.
The Cerebras squeeze, in Ejaaz's math
Altman's statement, that Cerebras is "a close partner" with "a deep engagement pushing on the frontiers of speed," neither confirms nor denies the rumors, and @cryptopunk7213 filled the gap with specifics: OpenAI is using Cerebras chips for ultrafast inference but ran out because Jane Street bought up all the supply; Jane Street is paying Cerebras $200M per megawatt for that speed; and Cerebras' cost is about $26M per megawatt. His conclusion: "cerebras has a good chip + is making bank?"
None of this is verified, and the numbers come from a single thread. But if Ejaaz's figures are even roughly right, the striking detail is the premium: nearly 8x over cost per megawatt for inference speed alone, paid by a trading firm rather than a lab. Worth watching whether any of the three parties says anything on the record.
Durable agents: Pi Durable, OptChat, and a harness disagreement
Pi Durable got the most attention. @pidotdev's Monday Meditations, featuring @badlogicgames and @mitsuhiko, covers why they built a harness around a small task-based workflow engine so long-running, multiplayer agents can suspend and resume anywhere, plus why they passed on Temporal and Effect.ts. @badlogicgames also pointed followers to @lucataco's explainer, calling it "100% accurate."
@raunakdoesdev reports building an agent on @VictorTaelin's OptChat memory idea inside a durable object, one that can modify and redeploy itself and resume durably without issues, wired up with codemode and @RhysSullivan's executor. He says he wouldn't bother with instinct or muse after this. @tokumin compressed the whole stack into "optchat + pi-durable = 🪄✨," quoting @VictorTaelin's "THE RECIPE" post, and @VictorTaelin himself spent the day marveling that "something that dumb is transforming my life."
The counterpoint: @ishaansehgal argues every serious agent team ends up drawing the same architecture tree, claiming Sentry spent about 4 months and 100k lines on theirs and that Stripe, Shopify, Harvey, Ramp, and Sierra built their own. His line: "the harness isn't the hard part. everything around it is," and he's building that surrounding layer at @omnaraai. The @kushbhuwalka post he quotes cuts the other way: "fluffles," a no-restrictions agent on a Mac mini, meant a harness that was painful to build, with an unbounded surface area for things to go wrong, before the team ported to @eve. Read together, the disagreement is less about whether scaffolding matters and more about where the pain concentrates.
A 28-day shipping streak, a 2,500-PR month, and a yolo toggle
@thsottiaux pledged that over the next 28 days, his team will ship one clear improvement relevant to most codex/work users each day, or ship a full reset. It follows his own earlier post that feedback was clear: simplify, and focus only on efficiency, groundbreaking features, or new models. Public daily-shipping commitments are easy to make and hard to keep, so the streak itself is the test.
@thecsguy took notes from @poteto's interview with @mattpocockuk, where @poteto says he landed 2,500 PRs last month and recommends combining the two skill plugins. His takeaway: seemingly small insights from the interview took his agent kitchen "from 10 to 1000." For the reckless, @pvncher flags a yolo setup for computer use, quoting @steipete's one-liner, a defaults-write command that sets ComputerUseAllowForbiddenTargets to YES globally. Unlocking forbidden targets machine-wide is exactly as advisable as it sounds.
When agents pick your programming language
@dhh's two-part argument: he has had agents implement and optimize the Campfire web app in Elixir, Go, and Rust, and Ruby is the slowest, but that mattered less when the payoff was productivity and developer joy. Now the qualifier: if you're no longer reading the code, and frontier agents can't write your language well out of the box, "that language is going to have a hard time in the future." It's one person's experience with one app, but the incentive argument is real, and @_waela's "vibe coding discourse speedrun" meme captures how fast this argument now cycles. Elsewhere in builder notes, @championswimmer says herdr-gpui is the first tool to go from exploration to can't-live-without in under 12 hours, and @enjojoyy reports their personal site took under an hour to build, drew unexpected attention, and they're open to work.
Kolibri's open weights, a loosening Astra quota, and the day's long-arc bets
Per @Aleph__Alpha's announcement: "78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe," with weights available under Apache 2.0 to run on your own hardware. @charles_maddock's framing, a German lab dropping a state-of-the-art model trained from scratch that beats Qwen, Mistral, and Nematron at similar active parameter counts, is his claim rather than a benchmark table, so treat it as a prompt to test rather than a result.
@kunchenguid noticed something odd: Astra handled "SO MANY meaty tasks" and their quota barely moved, unlike at release, and they're asking whether others see the same. No explanation is offered in the thread, so it's one data point.
Two long-arc views rounded out the day. @deanwball recounted the scenario he gives when senior AI figures ask when we're essentially "just screwed": the vulnerable world hypothesis from @dewierwan_'s post, where massive cognitive labor applied to basic science surfaces cheap, ultra-destructive weapons tech, solvable only through global preventative policing like AI nonproliferation. @QuanquanGu pitches the optimistic mirror image at @Geodesiclab: biology as AI's next recursive self-improvement frontier, design in silico, test in the lab, turn experimental feedback into semi-synthetic training data, repeat, so "the wet lab becomes part of the learning algorithm." @FredaDuan calls him one of the few researchers at the intersection of foundation models and AI for science. In small-drama news, @Ayesha_Khan167 tagged @grok with a video and a "what the hell is this," unresolved at time of writing.
Practical Takeaway
The strongest actionable thread is durable, resumable agent scaffolding. If you run long-lived or multiplayer agents, spend an hour with the Pi Durable rationale, especially the Temporal and Effect.ts comparisons, and @VictorTaelin's OptChat recipe before writing your own orchestrator; @raunakdoesdev's combo shows you can prototype a self-redeploying agent in a durable object quickly. But budget for @ishaansehgal's warning: the surface area around the harness, integrations, auth, failure handling, is where projects stall. And if you're choosing a stack today, @dhh's argument is worth a quick test of how well current frontier agents actually write your language before you commit.
Sources
updated my personal website, check it out https://t.co/A0hSzg3Jt0 also, if you're hiring, i'm open to talk!
There is some speculation about our partnership with Cerebras. Cerebras is a close partner, and we have a deep engagement pushing on the frontiers of speed.
I've had agents implement and optimize the Campfire web app in Elixir, Go, and Rust. Yes, Ruby is the slowest. I didn't care about that when the pay-off was huge in terms of productivity and developer joy. But if you're no longer reading the code? https://t.co/MM4yFETaI3 https://t.co/O4ZF7jAuV7
All right, we’re locking in. Only things being worked on are simplifications, more efficiency for more usage, groundbreaking features or new models. Sometimes you have to invest ahead of the curve, but feedback is clear that you all want things to get simpler. On it.
@BenjaminBadejo defaults write -g ComputerUseAllowForbiddenTargets -bool YES
Small bird, fast wings, Kolibri is here. 78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe. Now the weights are yours. Run it on your own hardware, under Apache 2.0. https://t.co/5263xZ9xZN
Biology could be AI’s next recursive self-improvement frontier. Design in silico → test in the lab → turn experimental feedback into semi-synthetic training data → improve the model → repeat. At @Geodesiclab , we’re building toward that flywheel. The wet lab becomes part of the learning algorithm.
Motion Engineering: Build a Video Studio Around Opus 5.5
i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://t.co/RS0Mxgoy8n
before @eve, we had built an agent called fluffles. it ran on a mac mini that sits across me. the goal was for it to be a god agent - an agent so powerful it could operate indistinguishably from a human, with no restrictions whatsoever. we wanted to give it a credit card, phone number, patched browser - the whole 9 yards. we mapped out what primitives it would need, and came up with the architecture shown. we thought it would be straightforward. but building the harness was unbelievably painful. from automations to subagents. there wasn't one specific thing that was hard, it was that the surface area of shit that can go wrong is unbounded. when eve came out, we ported over almost immediately. the scaffolding let us focus on agent behavior instead of the intricacies of a slack integration. we're excited to show you our work, and will be launching this week.
THE RECIPE https://t.co/1oezv7fqg1 Ask your agent to build this and ENJOY FREEDOM 🥳
What is Pi Durable? https://t.co/xwSFGvQYK5
One of my biggest fears in AI safety is that this might be true: > The vulnerable world hypothesis is likely correct. Once enormous amounts of cognitive labor start getting applied to basic science, it will quickly become apparent how many avenues exist to create cheap, ultradestructive weapons technology. Solving this problem without global preventative policing (e.g. AI nonproliferation) is impossible, because hardening civilians against all avenues of attack is too expensive and will take too long.