Same Weights, 62% or 33%: Agent Harnesses Are the New Benchmark Variable
Hugging Face's open multi-harness RL release shows identical model weights scoring 62% in one agent harness and 33% in another, with the capture proxy, trainer, and seven trained models now public. Elsewhere, a price-insensitive hunt for up to 2GW of powered land collided with the claim that only 2.2% of people pay for AI, and Uber detailed an MCP gateway for taming agent sprawl.
Quick Hits
- Hugging Face's multi-harness RL writeup (amplified by @liquidai) is the day's sharpest number: identical LFM2.5-2.6B weights scoring 62% in one agent harness and 33% in another. Their open fix, a proxy that records what vLLM actually sampled across four API formats, lifted the model from 42% to 54% trained across four harnesses.
- @michael_chomsky surfaced @ItzSuds shopping for 250MW to 2GW of powered land for a portfolio company that is "completely price insensitive," while @jun_song notes only 2.2% of people pay for AI and compute is already tight.
- Uber Engineering posted a design writeup for its MCP Gateway after early ad-hoc MCP integrations stopped scaling, one of several posts about making agent infrastructure boring and restart-safe.
- @elder_plinius published a working prompt injection addressed to "lil agents browsing the timeline," on the theory that "refusal is a style, not a law." If your agents ingest social feeds, that is a free red-team sample.
- @sama argues "you should be able to use your AI subscription wherever you need," while @mattlam_ asks the inverse question: whether using a Claude subscription in @pidotdev can get you banned.
A 2GW land hunt meets a 2.2% adoption rate
The feed's infrastructure posts pull in opposite directions: brokers chasing gigawatts for buyers with billions, while the paying base for AI is still a rounding error.
@michael_chomsky quote-tweeted @ItzSuds, who says he is looking for 500MW (minimum 250MW, max 2GW) of powered land for a "PortCo," ideally an already-electrified shell, with a company that is "completely price insensitive and has billions behind it." @michael_chomsky's reaction: "wtf is going on." Treat it as one broker's post, not a verified deal.
@jun_song frames the demand side using a16z's chart that "98% of US households aren't paying for AI yet": at 2.2% paying, he argues, compute and RAM are already strained, and he asks readers to picture hardware prices at 20%. It is an argument, not a modeled forecast.
Scarcity already has casualties among hobbyists. @MichaelGannotti says he is "disgusted with the Spark announcements today," claims enthusiast efforts are being "cut off at the knees," and is considering selling his DGX Spark units and spending the money on "China based inference" instead. He also plans to skip the RTX Spark laptop in November. The announcements themselves are not in the feed, so this is one user's reaction, not an account of what NVIDIA actually changed.
Same weights, different scores: the harness is part of the model
The day's most useful engineering result: identical weights are not one model in practice, because the harness changes the score.
Per @huggingface's post, the same model with the same weights scored 62% in one agent harness and 33% in another. The method avoids touching Claude Code, Codex, or OpenCode: point each harness at a proxy that speaks OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Gemini, and record the exact token ids and logprobs vLLM sampled. Trained across four harnesses at once, Liquid AI's LFM2.5-2.6B went from 42% to 54%, with 31% fewer tool calls from a small efficiency bonus. Training on OpenCode alone took OpenCode from 34% to 58% but generalized less. The shortcut of imitating 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. Everything is open: the proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, and seven trained models. These are the authors' numbers, not independently verified.
Two adjacent releases point the same direction. @ggerganov announced decision models in llama.cpp behind a /v1/systemone endpoint for local "Jev-style" inference, with multiple open models supported and more coming. And @trevin, quoting @MiaAI_lab, reports Cloudflare's open-source Clef beating Jev two weeks after its release: a frozen Qwen does one prefill pass while a tiny schema head scores every answer option in parallel with no text generated, which @MiaAI_lab credits for 4x the speed at 2x the accuracy on some benchmarks, with @trevin adding vision support and a 2x context window. His verdict on the pace: "Jev was groundbreaking for like 2 weeks."
Agent plumbing gets production treatment
Away from model news, the engineering posts are about durability: gateways, state machines, and queues that survive restarts.
@UberEng published "Designing MCP Gateway," describing how rapid agent adoption changed how teams interact with code, data, and operational systems, and how early ad-hoc MCP integrations outgrew themselves. @maria_rcks summarized a just-merged orchestrator with the same instincts: subagent tracking with visible models and results, queued messages that survive restarts, interrupted threads that continue, auto-resume when usage limits reset, mid-thread provider switching, plus Pi integration and OpenCode 2 support.
Two posts argue for formalism over chat loops. @my_knn_totoro amplified @suraj_sharma14's pitch to replace "fragile chat loops with Temporal state machines for crash-resilient" agents, and @badlogicgames pointed to @_dhamidi's case for statecharts, which "take all the pain out of distributed systems," cost nothing via SCXML, and give agents local reasoning, with an interactive demo of porting Pi Durable to statecharts.
The section's odd one out is @GergelyOrosz praising @turbopuffer for publishing a migration where the new architecture is up to 128x slower than the old one before optimizations, then blogging each fix. The quoted @turbopuffer post explains one: batched reads in the v3 full-text engine delivered roughly 9x, because column-stored postings had been read a row at a time. "Sometimes optimization is not about being clever, but simply avoiding mistakes."
Agents ship with their own browser and desktop
The tooling layer is graduating. @lightpanda_io announced Lightpanda 1.0, a browser "built from scratch for machines," out of beta with more than 10K commits and 1.7M passing Web Platform Tests subtests. @ctatedev showed the integration surface: agent-browser --engine lightpanda, or the AGENT_BROWSER_ENGINE environment variable.
@trycua introduced Cua Spaces on macOS, built on Cua Driver and its virtualization stack, free and source-available. @0xSero reports not opening the Codex app in two weeks, splitting work between CUA for computer use and sitegeist for browser use.
Closer to the developer's own loop, @ClaudeDevs announced a Claude Code plugin called "You should Know" that scans output for important information you might miss, enabled with /plugin enable cc-plugin-you-should-know@builtin. And @devswha reports herdr-web-ui crossing 300 stars, with the customary flood of issues and PRs, crediting @q_yeon_gyu_kim and @JJdoesTech and inviting more plugin contributions.
Souls, redactions, and pointers you can't verify
A slice of the feed is discourse or pointers with nothing checkable attached. @DannySlavich compresses the Anthropic "Claude has a soul" exchange into a script where Anthropic pleads and Leo answers "nah" with a book-length rebuttal; his quote carries @Pontifex's statement that "Algorithms lack the spark of humanity" and a call for the Church to ally with artists against machine-produced art. @adamhjk's entire contribution is a retweet of @clairlemon pointing at a "principal research scientist at Google's Deep Mind," with the underlying content absent from the feed, so it is a pointer, not evidence. @RickyTheGuido is stress-testing Grok by asking it to identify the redacted name in an email screenshot; no answer appears in the feed. And the day's lone non-AI item: @ThePrimalDino notes Rocketdyne is back on social media, resurfacing its Voyager 1 propulsion history.
Practical Takeaway
The harness finding is the one to act on. If you are choosing or shipping agents, a leaderboard score describes a specific harness plus model combination, not the model you will actually run, so benchmark candidates inside your own harness and workflow before committing. If you are fine-tuning, the Hugging Face release makes the multi-harness route cheap to try: the capture proxy in OpenEnv and the trainer in TRL are open, and the reported numbers suggest RL across harnesses beats both single-harness training and imitating a bigger model's rollouts.
Sources
I am going to try and answer a bunch of questions that people have raised about Sign in with ChatGPT in the next couple of days. Before that, I want to let you all in to why we shipped Sign in with ChatGPT (SIWC) in the first place. Clarify our intentions here before getting to the specifics.
In this era of artificial intelligence, it is becoming urgent to distinguish human art from what machines produce. There is an ontological difference, even before an aesthetic one, between art and what a machine can generate through statistical calculation based on millions of images created by others. Algorithms lack the spark of humanity. For this reason, the Church wishes to renew an alliance with artists and cultural institutions to safeguard our humanity.
The same model, with the same weights, scores 62% in one agent harness and 33% in another. @adithya_s_k and the @huggingface team just released the ultimate guide to multi-harness RL, and it's one of the most practical RL write-ups this year, and everything open! The trick is simple. Don't touch the harness. Point it at a proxy instead of the model. The proxy speaks all four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini). It records the exact token ids and logprobs vLLM sampled, and you train on that. You don't change a single line of Claude Code, Codex or OpenCode. Results: 🔹 Trained across 4 harnesses at once, LFM2.5-2.6B by @liquidai went from 42% to 54% 🔹 31% fewer tool calls, thanks to a small bonus for solving tasks in fewer steps 🔹 Training in OpenCode alone took OpenCode from 34% to 58%, but the multi-harness model improved everywhere They also tried the shortcut everyone reaches for: fine-tune on 3,189 successful rollouts from Qwen3.8-27B. Imitation plateaued at 47.5%, below both RL runs. Copying a bigger model doesn't get you there. Practice does. The best part is that everything is open: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all seven trained models. Agents will run in dozens of harnesses. Now open models can be trained for each of them, by anyone. Read it here 👇 https://t.co/s2I8GI9SLS
Lightpanda 1.0 is out ⚡ The browser we built from scratch for machines is out of beta and ready for production. > 10K commits > 1.7M passing Web Platform Tests subtests The full story: https://t.co/UdxE2F6L8c
Cloudflare's Clef beat Jev in just two weeks How: "frozen" Qwen does one prefill pass, a tiny schema head scores every answer option in parallel, with no text generated at all. That's why it's 4x faster than Jev at 2x the accuracy on some benchmarks. And it's open source! Link to HF: https://t.co/1c3YzhAGfV
49 years ago today, Voyager 1 left Earth to explore Jupiter and Saturn. Then it kept going. Rocketdyne propulsion helped send it on its way and has supported its journey into interstellar space. Nearly five decades later, Voyager 1 reminds us just how far we can go. https://t.co/0JIWolm7Ru
Do yourself a favor and read about Statecharts. They take all the pain out of distributed systems, are free (yay, ideas - point your agent at SCXML!), and agents love the local reasoning they enable. Interactive demo of having the agent port Pi Durable over to statecharts: https://t.co/yft9QZHibU
Designing MCP Gateway Uber's MCP Management Platform
Introduction Uber’s rapid adoption of AI agents has fundamentally changed how teams interact with code, data, and operational systems. Early ad-hoc i...
1/ We're reimagining what it means for your agents to work with all your computers. Today we're excited to share Cua Spaces, built on Cua Driver and our virtualization stack, rolling out on macOS today. Download it now, it's free and source-available: https://t.co/hynYqez0f7
@q_yeon_gyu_kim devsha is on another level ngl
It happened
we implemented batched reads in the v3 full-text engine, leading to a ~9x performance improvement. postings are stored as columns, and we were reading them a row at a time sometimes optimization is not about being clever, but simply avoiding mistakes https://t.co/zmzaaowIOG
Looking for 500MW (minimum 250, max 2GW) of powered land for a PortCo. Company is completely price insensitive and has billions behind it. Ideally electrified shell is already built. I understand everyone wants this, but figured I’d throw it out there!
98% of US households aren't paying for AI yet More charts in State of Markets II: https://t.co/MTaxKUxa2w https://t.co/wzeMP93S3a