GPT-6 Sol Scores 29.3 of 105 Hidden Bugs for $9.93 a Run, and the Feed Does the Capability-per-Dollar Math
Paweł Huryn's 105-hidden-bug benchmark puts GPT-6 Sol at 29.3 fixed versus 45 for GPT-6 Astra, but at $9.93 per run versus $33.04, framing a day of cheap-model arithmetic that also brought a $17 Jev-style classifier from Together AI and Epoch AI's 47%-per-quarter cost curve. Agent tooling kept shipping too, with Vercel's Drives storage, Cursor's Rollouts, and local Projects in Claude Code, while Jeffrey Katzenberg's long essay on AI and creativity earned Aaron Levie's endorsement.
Quick Hits
- @PawelHuryn ran GPT-6 Sol against 105 hidden bugs across two real repos and reports 29.3 fixed at max effort versus 45 for GPT-6 Astra, though Sol's run cost $9.93 against Astra's $33.04. His one-line verdict: "GPT-6 Sol is like GPT-5.6 Terra."
- Chart of the day: @EpochAIResearch, shared approvingly by @AymericRoucher, says cost at a fixed performance level has fallen roughly 47% per quarter since 2023, which Epoch calls faster than DNA sequencing, compute, lithium batteries, and pre-1973 electricity.
- Agent infrastructure day: @rauchg introduced Vercel Drives for persistent agent storage (up to 16 TiB per Drive), @cursor_ai launched Rollouts to verify deployments, and @ClaudeDevs added local execution for Projects in Claude Code.
- From Meta Connect, @finkd says Meta Glasses can now serve as FDA-cleared hearing aids for mild to moderate loss at a fraction of typical prices; @DontFearAI jokes that banning them would run into the ADA.
- @levie amplified @jeffreykWNDR's essay on AI and creativity, endorsing its prediction that falling costs mean more films, not fewer, provided consent and compensation are built in.
The capability-per-dollar debate, in one benchmark and two price tags
The sharpest signal is a single unverified benchmark that prices quality tiers against each other. @PawelHuryn's max-effort results on 105 hidden bugs: GPT-6 Astra 45, GPT-5.6 Sol 43.5, Opus 5.5 41.7, Muse Spark 1.3 32.2, GPT-6 Sol 29.3. API-equivalent costs: $33.04, $95.25, $58.53, $18.11, $9.93. He calls the raw score "a huge degradation," then flips to price: Sol delivers roughly two-thirds of Astra's fixes for about a third of the cost, and he quips that GPT-6 Luna maps to GPT-5.6 Asteroid. Effort-level results are promised in his thread. Treat it as one author's harness, not a lab result.
That sits neatly on @EpochAIResearch's chart, which @AymericRoucher recommends following for exactly this framing: at a given performance level, cost down ~47% per quarter since 2023, a decline Epoch compares as 4× faster than DNA sequencing, 6× than compute, 18× than lithium batteries, and 54× than electricity up to 1973.
Two posts show what the cheap end of that curve buys. @neural_avb shares DSPy code that turns raw text into choice-based JEV decision training data, saying gpt-6-luna and deepseek-v4.1-flash are cheap enough to produce 50K examples for roughly $5 to $10. @togethercompute released tev1-4B-experimental, a Jev-like classifier finetuned on Qwen3.5 4B, served at $0.042 per million input tokens with free output, plus the data recipe and a finetuning tutorial; @nutlope's companion post notes it cost $17 to train.
Agents get durable storage, deployment checks, and a shared memory
The agent stack filled three gaps at once. @rauchg frames every successful agent as brain, hands, and files: model plus harness, tools and browsers, memories and repos. A single stateful Mac Mini works, he argues, but cost-efficient cloud agents decompose those parts, and the missing piece was decoupled storage. Hence Drives, which @vercel_dev describes as persistent Sandbox storage in public beta on every plan, mountable up to four per sandbox and up to 16 TiB each. His example is a nightly memory-consolidation job that reads and writes agent files without booting the full machine, and he claims the decomposition also improves security and auditability.
@cursor_ai's Rollouts target the deploy side: the feature writes a monitoring plan, watches changes as they ship, and verifies deployments so regressions get caught before users see them.
@ClaudeDevs added local support to Projects in Claude Code, so the parallel threads that previously ran as cloud sessions can now run on your own machine. On the same account, @JeremyNguyenPhD points paying users to a free $100 or $250 cloud-credit claim, which per @ClaudeDevs requires GitHub connected and expires October 7.
Memory is the third gap: @akshay_pachaar walks through Beacon by @asymptotelabs, an open-sourced layer that normalized 579 sessions across five coding-agent harnesses, with Jev deciding which runs get promoted into reusable skills. The pitch is cross-pollination, a lesson from Claude Code improving the next Codex run, across 20-plus harnesses. Migration friction looks low elsewhere too: @openchamber_dev says its OpenCode 2 move was mostly adopting its own code, with the @opencode team and @thdxr fixing the few issues quickly.
Workflow moats and the rethinking-software pile-on
Several posts argue the enterprise stack is being rebuilt around workflows rather than data stores. @abishekv's article claims the moat is owning the workflow, not being the system of record, and that systems of record will become irrelevant. @alex_prompter retweeted @alex_verem calling some enterprise-AI article the best on the platform, though the post doesn't name the piece. @kentcdodds, quoting @megadevhq's "Towards Autonomous Product Development," says software building itself needs rethinking.
The most concrete artifact is @suraj_sharma14's updated list of 15 systems a forward deployed engineer should build instead of demos, spanning one-click customer environments, SSO/SCIM packs, permission-filtered RAG, MCP gateways over customer ERPs, human-in-the-loop approval gates, per-tenant cost governance, and customer-specific eval suites. It updates his own earlier six-month FDE roadmap for the agent era.
There's also evidence the optimization loop is turning inward: @minimaxir notes Anthropic's writeup on making the Claude frontend faster coincidentally uses the same agentic optimization techniques as his blog post, where prompting agents to optimize code beat state-of-the-art libraries. The mood, compressed by @simeonGriggs: "I used to debug code now I type this and hit enter."
Katzenberg's terms: build with the storytellers, not on top of them
@levie quotes @jeffreykWNDR's essay at length. Jeffrey Katzenberg rehearses his 2023 claim that AI would cut world-class animation time and cost by up to 90% within three years, walks through Sousa's 1906 war on the phonograph and the "canned music" fight as proof that these battles settle on terms rather than technology, and splits reasoning from creating, calling today's AI "statistics, not soul." His ask is credit, consent, and compensation, because the north has the tools and the south has the creative soul. Levie's gloss: every prior shift expanded opportunity, and the need for skill and taste never goes away.
Also on the feed: a gpui rebuild and a Terafab hire
@fabienpenso shows off a heavily improved Herdr desktop interface built in gpui, fully compatible with a running Herdr instance. And @jjkaplowitz congratulates @SemiShinigami, who announced he is joining Terafab (@Tesla @SpaceX) as a Staff Process Engineer for wet etch, long nights and American manufacturing included.
Practical Takeaway
The strongest thread today is capability per dollar, so make it measurable before your next model choice: seed a repo with your own hidden bugs or tickets, run the contenders at fixed effort, and divide the bill by tasks actually solved, exactly the ratio @PawelHuryn's numbers expose ($9.93 for 29.3 bugs versus $33.04 for 45). If a budget model clears your bar, cheap Jev-style classifiers like @togethercompute's $17 recipe and @neural_avb's $5-to-$10 synthetic data flow mean evaluation and training data are now line items rather than projects. And if agents are heading to production for you, the Vercel Drives and Cursor Rollouts betas are worth testing now, since @rauchg's claim that decoupled storage is a security prerequisite is cheap to verify before you need it.
Sources
If I had 6 months to become a Forward Deployed Engineer. I'd do this. Stage 1: Full-Stack and API Foundations Python or TypeScript, FastAPI, async, webhooks, REST/GraphQL, error handling, third-party SDKs. Stage 2: Enterprise Integrations OAuth2, SAML, SSO, SCIM, Salesforce/Slack/HubSpot APIs, webhook verification, retry logic. Stage 3: Multi-Tenancy and Data Isolation Row-level security, tenant identifiers, per-tenant databases vs schemas, data partitioning. Stage 4: Security and Compliance SOC 2, GDPR, HIPAA basics, PII redaction, encryption at rest and in transit, audit logging. Stage 5: AI on Customer Data Secure RAG ingestion, per-document access control, citation tracking, permission filtering. Stage 6: Agents and MCP in the Customer Stack Expose customer systems as MCP tools, human-in-the-loop gates, agent workflows over real SOPs. Stage 7: Deployment Automation Docker, Kubernetes, Helm, CI/CD, infrastructure as code, one-click customer environments. Stage 8: Observability and Customer Dashboards Distributed tracing, AI quality metrics, usage analytics, SLA monitoring, customer health pages. Stage 9: Incident Management Production debugging, rollback strategies, post-mortems, customer communication during outages. Stage 10: Discovery and Requirements Translation Stakeholder interviews, business-to-technical specs, expectation management, scoping. Stage 11: Adoption, Expansion and ROI Proof Usage analysis, adoption tracking, expansion signals, churn reduction, ROI reporting. Stage 12: Portfolio and Case Studies Documented deployments, implementation guides, integration patterns, published success metrics. The role between engineering and revenue. The one that ships AI inside real enterprises. Most people stay stuck watching tutorials. Builders get hired. (Bookmark it)
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Towards Autonomous Product Development
Jev Clearly Explained
Vercel Sandbox now has persistent storage with Drives, in public beta on every plan. ▪︎ Store agent workspaces, data, models, deps ▪︎ Read snapshots across parallel sandboxes ▪︎ Mount up to four Drives per sandbox ▪︎ Up to 16 TiB per Drive https://t.co/5sOGOiREOv
New blog post up: I discovered that you can indeed prompt agents to make your code faster to the point it beats current state-of-the-art libraries. This is not a vaguepost, I include both my prompts and benchmark results. https://t.co/ER6rGcb411
Follow the link below to claim the credit or run /claim-credit in the CLI. You’ll need GitHub connected to start a session. Claim by Oct 7. Terms apply. https://t.co/TzoO6fzJHB
How to train your own Jev for $17
Building Terafab @Tesla @SpaceX - https://t.co/wpL20KIABG > Joining Terafab as Staff Process Engineer, Wet Etch. > High-energy team, extreme ownership > Innovative and Industry veterans > American Manufacturing 🇺🇸 Hoping to stay humble and continuously learn from the team, industry partners and vendors. Looking forward to long weeks, long nights, and hard work to build something new.
Today we're rolling out Projects in Claude Code on desktop and web. A project is one conversation with Claude. It splits the work into threads itself, runs them as parallel cloud sessions, passes context between them, and keeps going when you leave. In beta for select users. https://t.co/j4k7rludhV
The World is Changing: AI For Creativity By Jeffrey Katzenberg A few months ago, I sat in my office in Silicon Valley and watched as a tech founder showed me something extraordinary. On the screen was a fully realized, beautifully lit, well-composed animated scene. It was stunning and it made me feel exactly what I felt in 1986 watching Luxo Jr. That was the first time I watched a computer-animated 3D character take a breath and seem, against all reason, to have life. It left me in awe. Later that day, I received a text from an artist I've known for thirty years, 350 miles to the south, in the city where I spent most of my career. After seeing a similar video, she texted: "Is this the end of us?" My answer was, "Certainly not.” I have spent the better part of the last decade in Silicon Valley, but the heart of my career has been in Hollywood. Being deeply connected to both worlds means I have deep loyalties to each and a responsibility to speak honestly to both. In 2023, I said that these new AI tools would cut the time and cost of producing world-class animation by as much as ninety percent within three years. Some colleagues were alarmed, many were furious. There is growing fear and resistance surrounding AI within the creative community. I deeply understand it, because I've spent countless hours walking through animation studios watching gifted artists bent over their desks, rebuilding a single second of film for the tenth time because the ninth version wasn't quite right. I've sat in screening rooms where four years of people's labor played out in minutes, and I knew the name of every person that had spent countless hours bringing those images to life. The creative process is a calling, there's really no other way to describe it. From the outside some see resistance. From the inside, it is love. People do not fight this hard for things they don't care about. The pushback coming out of Hollywood represents the collective effort of people who are deeply passionate about their craft. Is History Repeating Itself? The history here is more complicated than either side may realize. In 1906, the most famous composer in America, John Philip Sousa, published an essay titled “The Menace of Mechanical Music." He warned that the phonograph would become "a substitute for human skill, intelligence and soul." Sousa's fight was not really about the machine, it was about money. The machines were playing his compositions, and the men who built them weren't paying him a cent. His campaign helped create the Copyright Act of 1909. He did not stop the technology. He changed the terms under which it could use his work. A hundred years ago, sound came to the movies. We remember it now as a miracle, and it was. What we forget is who paid for it. Before sound, tens of thousands of musicians made their living in the orchestra pits of movie houses, scoring every film live, every night, in towns all over the world. When the soundtrack arrived, the work of one composer and one orchestra was recorded for a film that went into thousands of theaters. The union fought back with everything it had, taking out newspaper ads across the country warning against the menace of "canned music," one of them showing a mechanical man tearing the strings out of a harp while an angel wept. They were not fools, and they were not Luddites. They were right. Those pit jobs did not come back. And yet (this is the part we have to be brave enough to admit), sound gave us the movie musical, the modern score, sfx, sound design, audio engineering, and an art form vastly larger than the one it disrupted. And it helped keep Hollywood in the forefront of world entertainment for the rest of the century and into the next. The loss was real. And yet the art form expanded. This is a story that has been told over and over again. To resist technology is to risk irrelevance. Just look at Kodak or Blockbuster. To embrace technology is to open doors of new possibility. Just consider Apple and Netflix. What I Learned From Walt Disney In the mid-1980s, I was tapped to lead Disney's animation division at a moment when the studio was at an inflection point. Animation wasn't just another business unit. It was the soul of the company, a medium revered because of Walt's genius and his passion. But the production system was cumbersome and unforgiving. A single movie was 125,000 individual hand-drawn and painted cels, photographed one frame at a time. Every revision carried a cost measured in months. These degrees of difficulty shaped the kinds of stories we could tell. We found our way forward in an unexpected place: Walt himself. The Disney archives held astonishing recordings of Walt explaining his creative process. His own writings. His notes and storyboards. Work product captured at every stage of his process. This was truly a gift. Listening, reading, sitting with the work itself, we heard him talk about character, about emotion, about how an audience feels when a character truly comes alive. He talked about making bold choices and refining a scene until it genuinely moved people. We didn't hear a word about pencils or paintbrushes. In fact, Walt was famous for being a technologist, forever hunting for state-of-the-art tools, often inventing them himself to achieve the images he saw in his head. But he never defined animation by the tools. He defined it by whether the audience believed the character. His principles were timeless. The tools were not. That realization changed everything. We co-developed the Computer Animation Production System (CAPS) with a young Northern California company called Pixar, replacing hand-painted cels with CGI. In The Little Mermaid, the final scene shimmered with a dimensionality and light that the old process simply couldn't achieve. In Beauty and the Beast, the ballroom sequence moved with a cinematic sweep that placed the audience inside the emotion of the moment. In Aladdin, the Cave of Wonders felt vast and alive, and the Magic Carpet became an intricate, compelling character all its own. In The Lion King, the stampede carried a scale and intensity that raised the emotional stakes beyond anything we'd done before. Technology didn't diminish the craft, it expanded the canvas. It gave artists more room to create. A decade later, the canvas expanded again. When Disney released Pixar's Toy Story, it wasn't simply a technical milestone. It was proof that a fully computer-animated film could carry real emotional weight, that it could make audiences laugh, cry, and believe. At DreamWorks, we made the difficult decision to sunset hand-drawn animation and become a fully computer-animated studio. It was the right thing to do, but it was not without pain. It cost talented people their place in an industry where they had worked their whole lives. Some made the leap to the new tools and did the finest work of their careers. Some never did. Tools are never the point. The instruments change with every generation. What endures is taste and imagination. The magical ability to make an audience feel. One of the greatest storytellers of our generation, George Lucas, succinctly captured the eternal essence of this issue: “It’s not the how, it’s the why.” A Distinction With a Difference I asked one of the leading AI models a question that has been challenging me for months. What is the difference between reasoning and creating? Its answer changed how I think about almost everything happening in this industry. It said . . . Reasoning and creating are two distinct cognitive modes, though they also work together. Reasoning is fundamentally evaluative and analytical. It operates on what already exists: facts, premises, evidence. It moves toward a conclusion that was in a sense already implied by the input. Reasoning is constrained by logic and truth. Its goal is to arrive somewhere correct, not to invent somewhere new. Creating is fundamentally generative. It produces something that didn't exist before. And crucially, there's no single right answer waiting to be found. A blank page has infinite valid responses. Creation involves choices that can't be fully justified by logic alone. Taste, intuition and vision fill the gap where deduction runs out. Reasoning is what Silicon Valley has been perfecting. Creating is what Hollywood has been practicing for more than a century. AI today operates almost entirely on the reasoning side of the line. It can deduce, evaluate, optimize, and pattern-match brilliantly. And while it can create, there is a real distinction to being creative. What it doesn’t yet have is those things that make us human: empathy, devotion, serendipity, the kind of creativity that comes from a person trying to say something only they could say. When the bot generates a piece of art, it is not trying to communicate anything. It is statistics, not soul; it is emulating things that have been done. By contrast, human creativity isn’t about repeating patterns of zeros and ones; it is about doing something new. One day, AI may close this gap. Three years ago, the leaders building AI would have called what they are achieving today, improbable, if not impossible. Impossible is no longer improbable. Today, the line between reasoning and creating is real. Even the leading technologists acknowledge we are not there yet. There is no scientific path to crossing this divide that anyone in the field can articulate today. Understanding that gap is where we will find common ground. A Path Forward In 2016, I closed one chapter in Hollywood with the sale of DreamWorks and opened another in Northern California, co-founding WndrCo. We’ve backed more than 50 founders building the next generation of technology and watched how breakthroughs in Silicon Valley emerge, first as experiments, then as platforms, and finally as infrastructure that reshapes entire industries. It's worth remembering that the last great revolution in animation also came from the north. Pixar was a Northern California company, forged not in the conventions of the Hollywood studio system, but in the technological breakthroughs of Silicon Valley. I've spent years on both sides of this bridge. For sure, I don’t have all the answers (take Quibi, for one!). But, from my past and present vantage points of my long career, here is what I see . . . Brilliant people in Northern California building this technology have made something extraordinary. They have earned the right for the rest of us to be, if not believers, at least optimistic that what comes next will be remarkable. But they have not made an artist. The tools are powerful, but they are not what makes a story matter. That knowledge lives 350 miles to the south, inside people whose life's work has informed the very models you are building. The right path forward includes them by design, with credit, with consent, and with compensation. Build this with the storytellers. Not on top of them. Taste is not something that can be synthesized, it is uniquely human. At the same time, Hollywood needs to accept that AI is not going away. The energy they are spending trying to make it disappear is energy they are not spending deciding the terms on which it will exist. And the terms are everything. The north needs something from it that they cannot build and cannot buy: creativity. The kind that takes a blank page and conjures a single right answer where there was none and has held audiences for a century. Without it, the most powerful reasoning engine ever invented will still be missing the only thing that makes a story worth telling. The artists who learn to wield these new instruments will do things the engineers never dreamed of. They always have. Edison invented the motion picture but made terrible movies. It took Chaplin, Lloyd, Keaton and so many others to make movies emotional. Now, the canvas is about to expand yet again. We should decide now that we intend to paint on it. There are so many valuable lessons in history. This has happened many times before, and it was never settled by the technology. It was settled by the terms. Sousa did not stop the phonograph; he helped write the law that made sure composers got paid. And two years ago, when the writers and the actors walked out, they were fighting for the very things Sousa was fighting for in 1906. Consent, compensation, the basic recognition that human creative work has a price that must be paid. The terms of that fight are still being negotiated, but the principle is older than any of us. The tools-versus-no-tools argument is a trap. First, we must all agree that there should be terms. Then we can have the crucial debate about what fairness requires. What I Learned From Steve Jobs Years ago, Steve Jobs said, "It's in Apple's DNA that technology alone is not enough. It's technology married with the liberal arts, married with the humanities, that yields us the result that makes our hearts sing." He was describing a device. But he could just as easily have been describing this tale of two cities. What I See Coming Soon As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world. In live action, filmmakers like Steven Spielberg, James Cameron and Peter Jackson embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process. Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process. Assuredly, I don’t have all the answers, but I am confident that the creative opportunities will expand yet again. How we come through this is a choice. The north has the new tools. The south has the creative soul. The best future will draw on the best of both worlds.
Meta Glasses can now serve as FDA-cleared hearing aids. I'm pretty proud of this one -- they're medical grade, help with mild to moderate hearing loss, and a fraction of the cost of a normal hearing aid.
The Moat Is the Workflow. Systems of Record Will Become Irrelevant.
What people think the moat is: being the system of record. Owning the data. What the moat actually is: owning the workflow. For far too long, we’ve co...
I finally tested GPT-6 Sol on a real work. 2 repos. 105 hidden bugs. Find and fixed what you can. It looks like a huge degradation. The results: - GPT-6 Astra (max): 45 - GPT-5.6 Sol (max): 43.5 - Opus 5.5 (max): 41.7 - Muse Spark 1.3 (max): 32.2 - GPT-6 Sol (max): 29.3 Until you measure the cost (API-equivalent): - GPT-6 Astra (max): $33.04 - GPT-5.6 Sol (max): $95.25 - Opus 5.5 (max): $58.53 - Muse Spark 1.3 (max): $18.11 - GPT-6 Sol (max): $9.93 More effort levels (xhigh, high, medium, low) dropping in this thread today 🧵