AI Digest.

Harness Engineering Takes Center Stage as Reverse Proxy Phishing Exposes MFA Weakness

Social media discussions highlight that optimizing AI agent harnesses is becoming as critical as raw model capability, with developers diving deep into context management and terminal multiplexing. Meanwhile, security professionals warn that sophisticated real-time phishing attacks are rendering traditional MFA useless, leaving FIDO2 keys as the only reliable defense.

Quick Hits

  • Infrastructure ambitions are scaling up. @muskonomy reports on a potential collaboration between Cloudflare and Starlink, greenlit by Elon Musk, to rebuild traffic rules for low Earth orbit satellites and meet massive AI data demands. On the hardware front, @ivanfioravanti highlights new benchmarks for running DeepSeek V4 Flash on NVIDIA DGX Spark systems.
  • Open weights and AI rumors are circulating. @ylecun shares the release of Muse Glimmer, a 30B parameter dense model built for local execution. Meanwhile, @0xSero relays rumors of an upcoming OpenAI model named Astra, reportedly a 10 trillion parameter model with high emotional intelligence and writing capabilities.
  • Practical AI prompting tips continue to evolve. @daniel_mac8 suggests telling Claude to write in ASD-STE100 Simplified Technical English to dramatically improve output clarity. @Timb03 shares a quick SEO win using AI to find missing content gaps in Google Search Console exports.
  • The community is sharing expansive resource lists. @divaagurlxw outlines the massive skill set required for modern LLM engineering, from KV cache management to safety boundaries. @mertmetindev lists 50 useful websites for bypassing paywalls and finding academic papers. @EXM7777 shares a prompt designed to use AI for deep self reflection.
  • Developer culture remains collaborative yet critical. @dqnamo boosts Chinese design engineering Twitter and a newly open sourced AI video animation workflow. In AI governance, @jachiam0 voices support for a security leader's transparency following a recent incident.

The Deepening Complexity of Harness Engineering

Building AI agents is shifting toward meticulous harness engineering. @Teknium highlights a massive deep dive by @MrAhmadAwais detailing how Command Code rebuilt its read tool to save billions of tokens. The analysis explains that handling invisible failure modes, enforcing byte caps on minified files, fixing adversarial filename encodings, and managing relational state between read and write tools are the actual bottlenecks in agent efficiency.

The environments hosting these agents are also evolving. @dhh and @flaviocopes discuss Herdr, a release shipping with Omarchy Quattro that essentially rebuilds the tmux terminal multiplexer around coding agents. It features persistent workspaces and sidebars to track which agents are working, blocked, or idle. Developers are actively testing these multi agent setups. @imikerussell spotlights a setup by @tinkerersanky using Buzz to configure a six agent team (including an architect, builder, and debugger) as a default work environment.

Despite the rapid advancement in tooling, skepticism remains about autonomous coding maturity. @samuelcolvin pushes back against a post by @rauchg regarding code review. He argues that giving significant credibility to coding narratives from leaders who haven't shipped notable code recently ignores reality. @samuelcolvin points out that AI models still make rookie architectural mistakes, requiring developers to remain actively engaged to prevent technical debt.

Autonomous Workflows and the Enterprise AI Bottleneck

The conversation around enterprise productivity is splitting between interactive assistants and background autonomy. @gdb spotlights a walkthrough by @rileybrown detailing 14 capabilities of ChatGPT Work, highlighting its shift toward a comprehensive cloud workspace with branching chats, remote voice mode, and multi agent features.

However, true automation requires a fundamental reimagining of business processes. @vasuman argues there is a profound difference between AI sidekicks (which require continuous prompting) and AI background agents (which operate autonomously). Citing @levie, the posts note that agentic coding has exploded because it is purely digital, but enterprise domains like sales or legal require deep workflow reengineering to allow agents to work uninterrupted. While background agents can yield 60 to 90 percent efficiency gains, building them requires integrating custom models directly into legacy systems of record.

Real Time Phishing Bypasses Traditional MFA

Security professionals are sounding the alarm on reverse proxy phishing kits. @T3chFalcon details how tools like Evilginx, Modlishka, and Muraena defeat multi factor authentication. Instead of using static fake login pages, these tools act as a live proxy between the user and the legitimate service. When a user approves an MFA prompt through the proxy, the attacker intercepts the session cookie and gains full access to the account.

This attack vector is growing rapidly. According to the post, reverse proxy phishing surged 139 percent between September 2025 and March 2026. Commercial subscription versions of this malware are actively sold on messaging apps. @T3chFalcon concludes that SMS codes, push notifications, and authenticator apps are completely vulnerable to this relay attack. The only effective defense is FIDO2 passkeys and hardware security keys, which cryptographically bind credentials to the actual domain and cannot be replayed by a proxy server.

Practical Takeaway

If you are deploying AI coding agents, your highest leverage task is auditing how your harness reads and processes files. As evidenced by the token-saving strategies detailed today, naive file reads quietly drain context windows and budget. Developers should implement strict byte and line ceilings, alongside logic to handle adversarial filenames, to prevent agents from getting stuck in silent failure loops.

Sources

M
Mert Metin Tekdemir @mertmetindev ·
Google'ın Bilmenizi İstemediği 50 Web Sitesi 1.) 'https://t.co/tuoCcY59Pw' — Herhangi bir ödeme duvarını bypass edin 2.) 'https://t.co/7BVVqZj1f0' — Milyonlarca ücretsiz ders kitabı 3.) 'https://t.co/aj5Cj7pKtF' — Ücretsiz akademik makaleler 4.) 'https://t.co/ummI1BogUf' — Ücretsiz uygulama alternatifleri bulma 5.) 'https://t.co/g1l52DQ18h' — Herhangi bir içeriğin nerede yayınlandığını bulma 6.) 'https://t.co/l3Mb5x5KkS' — Herhangi bir eski web sayfasına erişme 7.) 'https://t.co/B6NDL6Bxri' — 70 bin ücretsiz klasik kitap 8.) 'https://t.co/DEEEuqpkI1' — Ücretsiz PDF indirme 9.) 'https://t.co/sSdBoLWEvB' — En iyi üniversitelerden ücretsiz kurslar 10.) 'https://t.co/OUOt3guKRb' — Herhangi bir matematik problemini anında çözme 11.) 'https://t.co/zVT0QaZ7iT' — Tarayıcı içinde ücretsiz Photoshop 12.) 'https://t.co/J3ffcTtbb0' — Herhangi bir görseli ücretsiz sıkıştırma 13.) 'https://t.co/yW8V9oPBRI' — Görsel arka planını ücretsiz kaldırma 14.) 'https://t.co/7yVdVI8CEU' — Fotoğraflardan nesne silme 15.) 'https://t.co/dMiW23fUnQ' — Video arka planını kaldırma 16.) 'https://t.co/B0A72IOts6' — Kodu sanata dönüştürme 17.) 'https://t.co/5cSFvKMjgA' — Şık kod ekran görüntüleri 18.) 'https://t.co/32fbjwbMzf' — Ücretsiz ürün mockup'ları 19.) 'https://t.co/vDnmbgoZMO' — Photoshop olmadan mockup yapma 20.) 'https://t.co/Dfk0b4x0Cz' — Hacklenip hacklenmediğinizi kontrol etme 21.) 'https://t.co/XE6rSQQAqd' — Herhangi bir dosyayı kötü amaçlı yazılıma karşı tarama 22.) 'https://t.co/Iq1D1lXVsj' — Kendi kendini yok eden mesajlar gönderme 23.) 'https://t.co/vfvGnFUrPT' — Anında tek kullanımlık e-posta 24.) 'https://t.co/dPp9oSFvdN' — Otomatik silinen dosya paylaşımı 25.) 'https://t.co/V07bDP1uxr' — Herhangi bir web sayfasını kalıcı olarak kaydetme 26.) 'https://t.co/tJUmsRxtPg' — Herhangi bir sitenin alternatifini bulma 27.) 'https://t.co/5e9jPtemru' — Dünyadaki herhangi bir radyoyu dinleme 28.) 'https://t.co/HwrFjKDKbq' — Her müzik türünü keşfetme 29.) 'https://t.co/BIwK31ecWz' — Herhangi bir diziden şarkı bulma 30.) 'https://t.co/RcswG9FZ0F' — Odaklanma müziği 31.) 'https://t.co/9k0m1JYQGF' — Özel odaklanma ses ortamları 32.) 'https://t.co/x7K9a9fxNN' — Üretkenliği artıran kafe sesleri 33.) 'https://t.co/VFDpzRCa3K' — Yapay zeka akademik makale asistanı 34.) 'https://t.co/VqHpxRxZWi' — Bilimsel fikir birliği arama 35.) 'https://t.co/I4jyq10CKs' — Akademik makale haritalarını görselleştirme 36.) 'https://t.co/eyfR9hrNwN' — Ücretsiz akademik arama 37.) 'https://t.co/plhsbWRurT' — Herhangi bir akademik makaleyi anlama 38.) 'https://t.co/DXWuCx94LL' — Herhangi bir YouTube videosunu özetleme 39.) 'https://t.co/JfJQkWb2DI' — Geliştiriciler için yapay zeka arama motoru 40.) 'https://t.co/GomAXY4Yj5' — Herhangi bir regex'i anında test etme 41.) 'https://t.co/HkRjtKR2VI' — Kodu temiz bir şekilde formatlama 42.) 'https://t.co/Icoy0jDDUN' — JSON'ı insan gibi okuma 43.) 'https://t.co/cIc12iKvXZ' — Terminal komutlarını anlama 44.) 'https://t.co/TH5zI6Uuvu' — Etkili yer imi yöneticisi 45.) 'https://t.co/oFBqzj5AMy' — Herhangi bir sitenin çöküp çökmediğini kontrol etme 46.) 'https://t.co/MmW02cqMWV' — Tersine görsel arama 47.) 'https://t.co/MlQpyusEcv' — İnternet hızını kontrol etme 48.) 'https://t.co/mZzc9gLOhS' — Ücretsiz PDF düzenleme 49.) 'https://t.co/qW0GK9KXJw' — PDF birleştirme ve bölme 50.) 'https://t.co/ZCsQHlLf01' — Saniyelik geçici e-posta İnternet, Google'ın gösterdiğinden çok daha büyük.
M
Mike Russell @imikerussell ·
Smartest setup I’ve seen so far for vibe coding with Buzz agents. I’m gonna steal this 👀
T tinkerersanky @tinkerersanky

I am loving Buzz. Made a team of 6 agents - assistant (terra-high) - builder (sol-med) - designer (opus-5) - planner (opus-5/sol-high) - architect (fable-5) - debugger (sol-high) Seems like this may become my default work environment going ahead. Still testing. https://t.co/U0UPKOJbQt

I
Ivan Fioravanti ᯅ @ivanfioravanti ·
Amazing job! @anemll is one of the best account on X! Deep diving like crazy on very complex stuff related to AI, Apple Neural Engine, Apple Silicon in general and now DGX Spark. Trust me, follow him now 🚀
A anemll @anemll

Published and tested vLLM 0.25.1 compatibility overlay and TP=2 reference deployment for DeepSeek V4 Flash plus DSpark/NVFP4 on 2× NVIDIA DGX Spark / ASUS GX10 systems. Includes integration fixes, a live dashboard, deployment tooling and reproducible benchmarks. This integration builds on work from @vllm_project, FlashInfer, Luke Alonso/b12x, @MiaAI_lab, @rafaelcaricio, Keys/drowzeys, @2WildTech and many other contributors. https://t.co/l5U6vK7GMh Benchmarks 👇

G
Greg Brockman @gdb ·
learn how to use ChatGPT Work:
R rileybrown @rileybrown

Learn 99% of ChatGPT Work in 61 minutes: GPT Work is like Codex in the cloud. It works on phone, web, and desktop. And I've been using it to run my business. These are the 14 knowledge work capabilities. 00:00 Intro 03:18 #1 Presentations 09:16 #2 Plugins 17:36 #3 Blocks 23:34 #4 Websites & Apps & Hosting 30:09 #6 Branching Chats 31:51 #7 Desktop App (Cloud vs Local) 37:01 #8 Voice Mode 41:36 #9 Remote Voice Mode 45:52 #10 In-App Browser 48:50 #11 Skills 51:45 #12 Scheduled Automations 54:37 #13 Spreadsheets 56:58 #14 Multi-Agent Workspace 59:19 Final Review

D
Dan McAteer @daniel_mac8 ·
Oh my god. It's hard to express how big a positive impact this tip can have on your life if you talk to Claude a lot. > Tell Claude to write in ASD-STE100, or Simplified Technical English As a bonus, tell Claude to follow Zinsser's four principles of quality writing: 1. Simplicity 2. Brevity 3. Clarity 4. Humanity h/t @DanielLockyer
D
diva @divaagurlxw ·
As an AI Engineer. Please learn >Harness engineering, not just prompt engineering >Context engineering, not just long prompts >Prompt caching vs. semantic caching tradeoffs >KV cache management, eviction, reuse, and memory pressure at scale >Prefill vs. decode latency and why they optimize differently >Continuous batching, paged attention, and throughput optimization >Speculative decoding vs. quantization vs. distillation tradeoffs >INT8, INT4, FP8, AWQ, GPTQ, and when quantization hurts quality >Structured output failures, schema validation, repair loops, and fallback chains >Function calling reliability, tool contracts, argument validation, and idempotency >Agent guardrails, loop budgets, tool budgets, and termination conditions >Model routing, graceful fallback logic, and degraded-mode UX >RAG architecture: chunking, embeddings, hybrid search, reranking, and freshness >Retrieval evals: recall, precision, grounding, attribution, and citation quality >Evals: golden sets, regression tests, adversarial tests, LLM-as-judge, and human evals >LLM observability as a first-class discipline: traces, spans, tokens, latency, errors, and drift >Cost attribution per feature, workflow, tenant, and user journey not just per model >Safety engineering: prompt injection defense, data leakage prevention, and permission boundaries >Multi-tenant isolation, cache safety, and cross-user context contamination prevention >Fine-tuning vs. in-context learning vs. RAG vs. distillation and when each is the wrong tool >Latency, quality, cost, and reliability tradeoffs across the full inference stack >Production failure modes: hallucinated tool calls, malformed JSON, stale retrieval, runaway agents, and silent eval regressions
D
DHH @dhh ·
Herdr is shipping as part of Omarchy Quattro. I've been pebbling the team with PRs to get full parity with my beloved tmux config. https://t.co/jOR5F5tiAy
T tobi @tobi

my favorite way to explain people herdr is to just open a new workspace, launch an agent and tell it to: read `herdr --skill` then split 10 other panels that make my computer look like some scifi hacker terminal from the movies https://t.co/b6m6RzCO6W

J
JP @dqnamo ·
you guys need to be on chinese design engineering twitter
I I_am_oil_oil @I_am_oil_oil

我把我用 AI 视频做动画的 Skill 开源啦:https://t.co/RQwQEgDTDA 整个 Skill 的流程大概是这样的: 参考素材、表达目的和控制方式 ↓ 先确认开始、中间和结束时应该是什么样子 ↓ 使用 AI 视频生成这些画面之间的连续动作 ↓ 逐帧检查,删除停顿、重复和异常画面 ↓ 按照页面中的实际显示大小整理并压缩资源 ↓ 把滚动、鼠标、拖动、触摸或设备方向对应到动画进度 因为整个流程还是有点复杂的,所以我把我实际做的时候的一些实现做成脚本放进去了,这样生成后处理动画会更加稳定一些,而且我还补充了一些创意方面的参考文档,比如说大家能够在视频里面看到的那个苹果的滚动爆炸图的效果。 欢迎大家体验~

J
Joshua Achiam @jachiam0 ·
To my friends in AI safety: Dane is an extremely good guy, he gets it, and we are very lucky he's the guy for this.
C cryps1s @cryps1s

@Simeon_Cps I disagree with this framing. The incident wasn't a 3-month long undetected attack, we actually responded twice, and we are very transparent about security issues. You can critique me for many reasons - but not caring about security, or transparency, ain't on the list.

I
IT Guy @T3chFalcon ·
Did you know there's a type of phishing attack that defeats MFA completely. It's called a reverse proxy phishing kit and it's one of the fastest-growing attack techniques in cybersecurity. Traditional phishing works by using fake login page. you type your password. attacker gets it. if you have MFA enabled, the stolen password alone is useless. attacker gets stopped at the second gate. reverse proxy is different. instead of a fake page, the attacker sets up a live proxy server between you and the real website. when you click the phishing link, you connect to the attacker's server. the server connects to the real Microsoft, Google, or Okta on your behalf. it fetches the real login page and serves it directly to you, pixel for pixel in real time. you're looking at the actual login page. you type your real password. the proxy relays it to Microsoft. Microsoft sends back an MFA prompt. the proxy relays that to you. you approve it. Microsoft sends back a session cookie. the proxy intercepts the cookie before it reaches your browser. You're logged in, so is the attacker using the same session. Tools used: Evilginx, Modlishka, Muraena. all open source. all free. all with pre-built templates for Microsoft 365, Google Workspace, Okta, and PayPal. commercial versions: EvilProxy, Tycoon 2FA, Mamba 2FA, Starkiller. Sold as subscription services on Telegram. Reverse proxy phishing surged 139% between September 2025 and March 2026, nearly 1 in 4 phishing links now carries a reverse proxy payload. 18 US universities hit last year. Microsoft 365 campaigns targeting thousands of organizations globally. The only authentication method that stops this completely: FIDO2 passkeys and hardware security keys. They bind credentials to the real domain. The proxy can't replay them because the credential was never issued for the proxy's domain. Everything else, SMS codes, authenticator apps, push notifications can all be relayed.
P PhishCore @PhishCore

What are reverse proxy phishing kits?

V
vas @vasuman ·
There's a difference between AI sidekicks and background agents. AI sidekicks require you to prompt them to handle tasks on the fly, and don't work without prompting. They help you move faster through your work, but don't fully take work off your plate. AI sidekicks bring between 10-20% efficiency for the average desk worker, and can cost a fortune if mismanaged. AI background agents run completely autonomously, only surfacing to you for exceptions, judgements, and other human-in-the-loop events. They actually remove work from your plate, and the workflows we take on see between 60-90% efficiency gains. The enterprise value ratio between the former and the latter is 1:10. Yet there's a ton of sidekicks like Claude Cowork, Codex, and GitHub Copilot, and almost 0 AI background agents. This is because setting up AI background agents is incredibly involved. You have to: - Go deep into a company, and understand how processes run today, and how exceptions are handled - Re-engineer those processes so that they are reimagined for an AI-native future. For example, 15-step workflows can now be condensed into 5-steps, and AI can handle 4 of them. - Build the agents custom for this process, on top of the systems of record. If your work runs on Dynamics 365 and Salesforce, your agents should do the same. - Integrate human-in-the-loop feedback and model-optimization so that the agents get smarter, faster, and cheaper over time. This is nearly impossible to do alone, because access to frontier AI talent is limited to top tech companies. And MBB firms like McKinsey will pitch you on AI transformation just to hand over a 300 page slide deck after 6 months and not ship a single agent. There's only one company on the planet that will do this work for you, and do it well, and that is linked below.
L levie @levie

One reason why we’re going to get uneven diffusion rates of agents is because different workflows in the enterprise are more or less aligned to continuous, uninterrupted computer work. A big reason why agentic coding growth has gone completely vertical is because it’s a type of work where economic value is directly correlated with the output of solely digital information, and where the task size can in theory be unbounded in a single session. This means that as AI agents can do larger amounts of work at a time, and as model capability improves on this task type, they can directly be deployed at larger workloads near instantaneously. Lots of work doesn’t natively have this property in how today’s workflows are designed. A sales rep needs a feedback loop with the customer before they can do additional work for a specific account, a lawyer needs to talk to the client, and a doctor needs to interact with a patient. So large agentic work will likely not look the same as it has in coding, at least by default. Instead, in many of these domains, the agents will actually need to tie into a workflow that is reengineered for AI. Matan Grinberg at Factory had a good way of thinking about this on the Training Data podcast: if everyone called in sick tomorrow, token usage would plummet because no one would be there to prompt them. This shows how early we are in agents running in the background for most processes; whereas this should be the majority of volume of agentic work in the future. In legal, it will be about processing every contract that comes in, in sales it will be agents that roam through customer records and figure out signals of when to better do outreach, in life sciences research it will be about swarms of agents reading through every test or all research. To do that these business processes have to be wired up to support agents, as opposed to the vast amounts of manual work that we do today. This is the big opportunity right now, but it’s going to diffuse differently from something like agentic coding. Lots of change management, lots of data cleanup, lots of workflow reengineering, and more.

0
0xSero @0xSero ·
I’ve always said GPT-4.5 was special Megamind level, pleasure to talk to.
I iruletheworldmo @iruletheworldmo

ok, apologies in advance because i do so hate to provide actual signal. openai basically sucked at pre-training for a while and crushed at post. their answer to this was spud. spud was the large new pre-train for the new family of models: sol, terra, and luna. the upcoming model, ‘astra’, has an even bigger pre-train, codenamed ‘doug’. (this was a running joke, as doug was the name for the world’s largest spud, although that’s not technically accurate — look it up). still with me? ok good, so sol is a 5t model based on the spud pre-train. astra is a 10t model based on the doug pre-train. this is the model i’ve discussed that openai researchers have been considerably more excited for. not only does it crush fable on all benchmarks, it’s genuinely delightful to talk to, high eq (4.5 vibes), and this is the first model that is genuinely incredible at writing, with output indistinguishable from that of great human writers. ok im tired and i have a lot of ssi news coming but ill keep you updated on this. everything here is well sourced and im highly confident. very well sourced.

M
Machina @EXM7777 ·
RT @EXM7777: this prompt will change your life: ---------------------------------- "People lie in journals. They perform in therapy. Nobo…
T
Tim Bennetto @Timb03 ·
Case in point. This one took 5 days to rank, and hit first page immediately. https://t.co/MxSjjNjRwo
T Timb03 @Timb03

Easy SEO win: - go to google search console -> performance - set date to 12mo -> export - upload Pages.csv and Queries.csv to any AI tool Tell it to find queries you rank for but don't have content for. Go write the content and you will rank quick

M
Muskonomy @muskonomy ·
NEWS: Cloudflare and Starlink may team up to fix a hidden problem that holds satellite internet back. The idea came from Cloudflare's chief technology officer Dane Knecht (@dok2001) and Elon Musk seemed to like it. Elon said he sent it straight to the Starlink team. Here is the problem they are talking about. The internet runs on a set of traffic rules that decide how data moves from place to place. Those rules were built for cables in the ground, which stay put and give a steady, predictable connection. Starlink is different. Its satellites sit in low orbit and race across the sky, so the connection is always shifting. The old traffic rules were never made for that, which keeps Starlink from hitting its full speed. Knecht's idea is to rebuild those rules from scratch, made for fast-moving satellites. He says Cloudflare and Starlink should do it together, since Cloudflare already handles a huge slice of the world's internet traffic. If it works, Starlink gets faster, smoother and more reliable. Better video calls, better streaming, better gaming, even in places with no cables at all. And it clears the way for the massive data that AI will need. Source: Elon Musk and Dane Knecht on X, August 2026
D dok2001 @dok2001

@elonmusk Bandwidth is half of it. Most congestion control still assumes a stable path and a stable RTT. LEO is neither. Area needs innovation all the way down the stack. @Cloudflare and @Starlink should be building that together.

T
Teknium 🪽 @Teknium ·
Thanks Ahmad for showing some new improvements that could be saving Hermes Agent users a ton of tokens and time! All improvements to the read tool we didn't have, now in Hermes https://t.co/YP2wYR92Oz
M MrAhmadAwais @MrAhmadAwais

how our read tool saves billions of tokens vs claude code command code is purpose-built for open models, so we optimize things other coding agents get to ignore. for the v1 release i rebuilt the read tool from scratch. it's now one of the most complicated, carefully engineered pieces of the system, and it saves billions of tokens a month. here's what we learned. context: we wanted the `read_file` in command code to be the best among coding agents, then benchmarked it capability-by-capability against the nine other common harnesses: claude code, opencode, cline, kilo, codex, grok, hermes, pi, openclaw. most were open-source; claude code ships none, so its column came from feeding the live tool crafted files and watching what came back. count the reads in any agent session. every edit starts with a read. every grep hit becomes a read. a plan step opens 3 files. a few hundred reads per session, ~50 million a month across command code. you've seen the failure modes. it reads a file, learns nothing, reads it again. it reads a 5MB lockfile straight into context. it reads a minified bundle once and that junk sits in the window for every turn after. napkin math: ``` 500 junk tokens × 50M reads/month ───────── 25B junk tokens × every turn they stay in context ``` i think the `read_file` tool is like a compiler that turns your filesystem into the model's context. every decision inside it is a token budget decision multiplied by fifty million times it's used every month. and that's why coding agents feel expensive: the bill is mostly reads building context. what "saves billions of tokens" means here: cost per successful read. claude code's read tool succeeds by spending more: more tokens per call, more turns per miss, and a model smart enough to fish the signal out of the noise. ours had to succeed by spending less, because our models can't peek over a sloppy read and our users care about the token bill. ask claude code to read a 3,000-line file and it hands the model all 3,000 lines. ask it for a file with a 3,900-character minified line and it hands over the whole line. no window, no byte ceiling, no per-line clamp. i ran the probe twice because i didn't believe it the first time. everybody ships a read tool in week one. readFile, slice by offset, return the string. first tool you write, last one you think about. ours ended up as dozens of modules with 98 tests, and it was the highest-leverage thing in v1. a naive read and a harness engineered read are both "correct". they both work. the difference is that one of them quietly spends a fifth of your context window on bytes the model never needed, and occasionally deadlocks against your own write tool. a few things i learned worth sharing: 1/ you need three ceilings, not one, and it's always the third one people skip. every codebase keeps a small zoo of hostile files: the 80,000-line lockfile, the minified bundle that's technically one line, the log that never stops growing. each ceiling handles one animal file if you will. ``` 2,000 lines longfiles 128 KB logs 2,000 ch/line bundles ``` the line window bounds an ordinary large file. the byte budget bounds a file whose lines are wide rather than many. the per-line clamp catches the case the other two miss: one minified line that sits comfortably inside the 2,000-line window and, on its own, eats the entire byte budget. you get back a single unusable mega-string that displaced everything the model actually needed to see. drop any one ceiling and there's a shape of file that costs you the whole read. no log will ever show it, just a turn where the model got nothing and paid full price for it. 2/ the most expensive thing a tool can return is silence. the costliest failure is an ambiguous non-answer. an empty result string is indistinguishable, from inside the model, from a broken tool. so it re-reads. widens the window. tries a different path. burns three turns learning what one sentence notice could have told it. so every dead end names its own recovery: ``` empty → "is empty" past EOF → "retry smaller" byte cap → "offset=1847" line cap → "offset=2001" pdf → "pdftotext" ``` two details carry most of the value here. the resume offsets are precomputed, so the model never does pagination arithmetic (which it does in reasoning tokens you pay for, and gets wrong often enough to cost another round trip). and none of these carry an `Error:` prefix, so the tui doesn't paint them red and the model doesn't treat a fact about the world as a failure worth apologizing for. (the byte-truncated case deliberately resumes ON the last line shown rather than the line after it, because that line got cut mid-content. an off-by-one in a resume hint is a silently corrupted read, which is the one bug class here that's worse than a wasted turn.) 3/ the bug that taught us the most was relational, and no input validation could ever have caught it. `read_file` records what the model has SEEN of each file into a ledger: the content, the mtime at read time, and a flag for whether the view was partial. `write_file` consults it and refuses to overwrite a file you've only partly seen, because you'd silently destroy the part it never saw. now compose that with the per-line clamp: ``` read ↓ one clamped line ↓ ledger says partial ↓ write DENIED ↓ model re-reads ↓ dedup returns "unchanged" ↺ forever ``` we hit it in the wild, on plan files during plan reviews and refinement. every field in every call was valid. the invariant that broke lived in the relationship between three tools that never call each other. shape invariants are checkable per field, and every schema you write already checks them. relational invariants across stateful tools are where the real bugs live, and you only find them by watching production traffic. 4/ a cache whose stale hit is catastrophic should expire itself on use. re-reading the same window of an unchanged file is pure waste: the content is already sitting in the conversation. so we return a short stub. fires only when mtime, size, and the exact (offset, limit) window all match. but that stub points at an earlier tool result. what if compaction ate it? now the model has been told to refer to something it can no longer see. forever. ``` read → content read → stub (eaten) read → content again ``` a dedup hit consumes its record. worst case is one wasted turn instead of an unbounded loop. cheap miss, catastrophic stale hit → self-expiring cache. that shape shows up all over a harness once you look for it. i settled with this design as it was a good enough tradeoff between complexity and risk, and it was the only one that did well in our benchmark. 5/ filenames are adversarial and the model can't see why. macos names screenshots with a NARROW NO-BREAK SPACE before AM/PM. it stores filenames NFD-decomposed. finder renames turn `'` into `’`. ``` "Screenshot 3.04 PM.png" "Screenshot 3.04 PM.png" ``` different byte strings. in a terminal, the same picture. the model reads the path off the screen, retypes it faithfully, gets "file not found", and no amount of reasoning recovers because the difference isn't rendered. you can burn an entire session on this and never learn anything. so before failing we retry 7 candidate spellings: narrow space ↔ regular, NFD, NFC, straight ↔ curly quote, NFD+curly. each one re-checked against the workspace boundary, because a repair must never quietly become an escape hatch. then, and only then, "did you mean?": substring match plus a bounded levenshtein of 2, which is what catches `AGENT.md` → `AGENTS.md` where substring matching finds nothing. these are the most common super cheap open model problems we now repair saving more tokens than silly token compression tricks. when a failure is invisible to the model, retrying is the tool's job. the model would retry the same wrong bytes forever. this is what harness engineering is about: finding the invisible failure modes and fixing them in the tool so the model can focus on reasoning. 6/ my favourite bug lives at a chunk boundary. reads stream chunk by chunk instead of loading the file, so a 400MB line sitting BEFORE your window never accumulates. fine. but it turns out that if the line limit is hit EXACTLY at a chunk boundary, you're standing in a spot where the answer to "is there more file?" doesn't exist yet. ``` limit hit at chunk end ↓ more bytes? → partial stream end? → complete ``` saying "more of the file remains" at that moment is a lie roughly half the time, and it's a lie that costs a turn every time it fires. so defer the decision to the next chunk instead of guessing. when you can't know yet, say nothing yet. (also: don't `break` out of the for-await. it calls the iterator's `return()` and destroys the stream underneath you.) 7/ images attach for real. ``` 4K screenshot ↓ jpeg ladder 95→80→60→40→20 ↓ attach at first fit ``` vision models get the actual image, compressed down a jpeg quality ladder (95 → 80 → 60 → 40 → 20). a 4K screenshot degrades instead of failing to attach. format detection sniffs magic bytes, never the extension: garbage in a .png must never reach the api, real webp must pass. we also gave vision to non-vision models using a VISION tool. so fun. 8/ downscaled images disclose their scale factor. ``` on disk 3024x1964 attached 1092x709 ↓ "multiply displayed coords by 2.77" ``` without that line, every click coordinate computed off a screenshot is confidently wrong. nothing in the image says it was resized on the way in. and at the end of the day you're saving tokens costs. 9/ notebooks render as documents. ``` .ipynb json soup ↓ tagged cells plots → real images 10K+ cell → jq hint ``` raw .ipynb is json soup: base64 blobs, per-character source arrays. we return tagged cells, plots attached as images. any cell output over 10,000 chars becomes a jq pointer, so one dataframe dump can't eat the read budget. if you do a lot of data work in notebooks, you can now read the notebook without reading the entire dataframe. the model can still reason about the data, but it doesn't have to pay for it in tokens. major time savings for the user, and a major token savings for the model. 10/ boring formats get one-line answers. ``` .svg → text (xml) binary → mime note .pdf → pdftotext hint ``` svg is text (it's xml, the model can edit it). binary returns its mime type, never garbage bytes. pdf gets a pdftotext hint for now; inline is on the list. and we have a tools to read and parse different formats as needed loaded on demand. the model can reason about the file without reading it all, and the user doesn't pay for it in tokens. 11/ numbering matches cat -n. ``` model → line 412 editor → line 412 trace → line 412 ``` 1-indexed, prefix on every line. the model, your editor, and your stack traces agree on what "line 412" means. every resume offset and edit target depends on that. 12/ repair inputs, don't bounce them. ``` filePath → file_path "2000" → 2000 "2abc" → rejected 1.5 → rejected ``` 10 aliases for file_path (filePath, absolutePath, target_file...) get repaired. numeric strings coerce via Number(), never parseInt: "2abc" is rejected, never silently read as 2. fractional offsets are rejected, never floored. a silently wrong window is worse than an error. our repair harness engineering shows up every where. 13/ some paths must never be opened. ``` /dev/zero refused /dev/urandom refused /proc/N/fd/0 refused ↑ before any i/o ``` /dev/zero, /dev/urandom, /dev/stdin, /proc/<pid>/fd/* are refused by name before any i/o. no extension to check, and the workspace boundary won't save you when cwd is /. a read tool that hangs on /dev/zero is a denial of service you shipped yourself. 14/ hygiene you only notice when it bites. ``` BOM stripped CRLF → LF utf-8 never split dedup kill-switch ``` bom stripped. crlf normalized. byte-cap truncation binary-searches a utf-8 prefix so it never splits a codepoint. the dedup ships with a kill-switch env var, because every cache needs one. so where does everyone else land? the top of the table is basically solved. eight of ten harnesses have a line window (500 to 2,000) and a second ceiling (25K tokens to 128 KB). that part is common knowledge now. the bottom of the table is empty almost everywhere. across all ten: ``` deferred chunk cut 1/10 unicode name retry 1/10 device blocklist 1/10 EOF note not error 1/10 coord scale note 3/10 did-you-mean 2/10 partial-view ledger 2/10 ``` the pattern: none of those rows show up in a demo. every one of them starts costing you in hour nine of a long session (remembering what the model has already seen, giving it a way back when a read misses, refusing to read /dev/zero). teams build them only after production forces it. no one nowadays is sitting watching their agents run and fixing the coding agent's bad behaviors. most developers just select a harness on some random vibe in the first week and never look back. this harness is minimal, it must be the best, this harness is from the model maker, this must be the best. wrong! 🤦‍♂️ … whatever happened to the engineer in us? I hope this post helps you see the difference between a harness that just works and one that was engineered to work well. claude code is the interesting column precisely because it's the incumbent: ledger, notebooks, vision, empty-file note, and then no window, no byte cap, no clamp, no resume offset, no streaming, no suggestion on a miss. that team just hasn't been forced yet, and it runs on models forgiving enough to absorb the waste. we were forced. we run open models where a wasted turn is visible in the eval score the same day, which is the entire reason any of the above got built. constraint is a feature - it forces you to engineer the right solution instead of hoping the model will figure it out. i hope this post helps you see the difference between a harness that just works and one that was engineered to work well. you can try all of this yourself in Command Code (we're also going open source soon so you'll be able to read the code yourself). i'd love to share more deep dives on harness engineering of command code, let me know what y'all wanna read about.

F
flavio @flaviocopes ·
Herdr is tmux rebuilt around coding agents. Real terminals. Persistent workspaces. A sidebar showing who’s working, blocked, done, or idle. A CLI and API agents can control. I looked at how it works, the tradeoffs, and where I’d use it: https://t.co/H7CDIbLzIP
S
Samuel Colvin @samuelcolvin ·
Sorry, but I have to point this out. 6k likes for a thought piece on coding from someone who, according to github, hasn't written significant code since 2016, perfectly encapsulates the state of twitter in 2026. Guillermo is extremely impressive, but he's not a thought leader on how code is written today.
R rauchg @rauchg

If you’re not reading the code, whether explicitly or through agentic inquiry, one or more of these is true: ○ You’re a beginner ○ Software is throwaway ○ You’re prototyping ○ You have no users / revenue ○ You’re taking on debt & risk ○ Your problems are basic And btw. All of this is fine. But the reality is that models are still not at the “full autonomy” stage yet. They make rookie mistakes, they go down bad architectural paths. I just had the best model in the world add a nonsensical 700ms delay to “settle” something and it told me “you’re right, I was cargo-culting” 🤨 I am on the camp that this need will diminish more and more. Most code is indeed going to be assembly-like. But we also have the global internet and software infrastructure riding on these models and narrative, and we have to respect that.

Y
Yann LeCun @ylecun ·
RT @finkd: Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also r…