Claude Code Bolsters Security While 27B Models Shrink to 8GB
Today's developments highlight a massive shift toward autonomous, multi-agent orchestration and extreme model compression. Developers are getting powerful new tools for local AI deployment and commit-time security, even as the industry grapples with an underlying infrastructure of secretly hijacked smart TVs fueling AI scraping.
Daily Wrap-Up
The AI ecosystem on July 23, 2026, is defined by a fascinating tension between massive leaps in developer productivity and the hidden, messy realities of the AI supply chain. On the surface, the tools we use to build software are becoming remarkably autonomous and deeply integrated into our workflows. Anthropic rolled out a beta for a Claude Security plugin that scans codebases for vulnerabilities directly within the terminal, while leaks suggest OpenAI's Codex is about to get a realtime voice mode where a main assistant orchestrates background worker agents. We are moving rapidly past the era of asking an AI for a snippet of code and into an era where AI agents act as automated sales representatives, security auditors, and autonomous UI designers.
Beneath these productivity gains, the open-source community is achieving seemingly impossible feats of engineering. The release of BTL-3 demonstrates that a 27-billion parameter agentic coding model can be crushed down into an 8.39GB file, proving that extreme sub-2.5-bit quantization is viable without destroying the model's reasoning capabilities. Yet, as local models become incredibly powerful and accessible, the industry's hunger for training data has led to some genuinely dystopian infrastructure practices. The revelation that hundreds of millions of smart TVs have been secretly converted into residential proxy nodes to route third-party AI scraping traffic is a stark reminder that the AI revolution often relies on covertly exploiting consumer hardware.
Navigating this landscape requires developers to balance aggressive adoption of new autonomous tools with a healthy dose of operational security. The most practical takeaway for developers: integrate local commit-time security scans into your CI/CD pipeline immediately, and aggressively evaluate the new wave of sub-3-bit quantized models like BTL-3 Compact to run powerful agentic workflows entirely offline.
Quick Hits
- EdTech platform allows organizations to upload their internal knowledge and instantly transform it into an adaptive learning system for standalone use or LMS integration. (@esandurrani)
- A 16-year-old high schooler builds the wildly popular "Taste Skill" for AI frontend design, accumulating 67,000 GitHub stars and utilizing 813 billion tokens in Codex. (@LLMJunky)
- A comprehensive new guide drops offering a full masterclass in advanced graph engineering for data-intensive applications. (@DeRonin_)
- Physical AI and robotics researchers argue that true embodied intelligence requires experiencing physics, noting that an AI can watch thousands of cycling videos but still not understand the physical feeling of losing balance. (@timi_cypher)
- The intersection of advanced mathematics and AI remains highly relevant as researchers visualize counterexamples to the Jacobian conjecture using high school math concepts to inform AI architectures. (@tarunchitra)
- Buzz Desktop releases version 0.4.23, improving agent reliability when switching between multiple communities and adding safer key backup protocols. (@jack)
Agent Orchestration and Automated Workflows
The trajectory of AI development tools is shifting rapidly from simple code completion to fully autonomous workflow orchestration. We are seeing the emergence of specialized agents that handle distinct phases of the software lifecycle, from generating UI components to autonomously prospecting leads and closing sales. This transition represents a fundamental change in how developers and businesses interact with machine learning, turning static prompts into dynamic, multi-step operational pipelines that require minimal human intervention.
The clearest indicator of this shift is the integration of automated security directly into the coding environment. Anthropic's new Claude Security plugin allows developers to run adversarial red-team scans on their codebases right from the terminal. As developer advocate @Aykutuces noted, this effectively solves a major anxiety in rapid application development: "Solo founder'ın gizli korkusu şuydu: hızlı hızlı ship ediyorsun, üründe security açığı kalıyor." By allowing developers to run hostile, adversarial scans locally without extra API costs, Claude is pushing security testing further left in the development cycle.
Simultaneously, agent architectures are becoming deeply conversational and compartmentalized. Leaks suggest Codex is preparing a realtime voice mode that acts almost like a project manager. "the main assistant holds the conversation and the worker agents handle everything in the background," reported @Ananth7e. This compartmentalization allows humans to direct high-level strategy via voice while invisible worker agents execute the tedious, multi-step tasks required in the background. We are also seeing this orchestration expand beyond pure software engineering into business operations, with @vovudebosh marveling at Claude's ability to autonomously find leads, email them, and book calls. Even UI design is being swallowed by agentic loops, with @trevin highlighting the release of Impeccable 4, a creative engine designed to handle greenfield design by seeding directions from human-approved visual worlds. As @odysseus0z pointed out regarding the push to automate eval engineering, the ultimate goal across the board is the creation of reliable, self-improving agent loops.
Shrinking the Stack: Extreme Quantization and Local AI
Running enterprise-grade AI models locally used to require a massive hardware investment, but novel compression techniques are completely changing the math. The open-source community is relentlessly focused on shrinking massive parameter models into digestible files that can run efficiently on consumer hardware. By aggressively optimizing the KV cache and utilizing advanced quantization, developers are proving that you do not need a massive cloud API budget to run highly capable agentic workflows.
The release of BTL-3 by Bad Theory Labs is a staggering achievement in this space. The model packs 27 billion parameters into a single 8.39GB file, operating at under 2.5 bits per parameter. @ivanfioravanti highlighted the significance of this release, noting that the model "retains 92.2% of the 27B intelligence" while running fully local at 43 tokens per second on an RTX PRO 6000. The creators achieved this by abandoning standard quantization in favor of packed AVQ2 decoder tensors and affine INT4 precision islands. It is a massive technical breakthrough that allows a 27B model to be smaller than a standard 8B model in fp16.
This hyper-compression is breathing intense life into the local agent ecosystem. @kimmonismus pointed out that the first local agent has officially beaten the Hermes benchmark on GAIA Level 1. By running open-weight models like Qwen, Gemma, and Llama via llama.cpp, local agents are utilizing stable-prefix caching to keep long agentic sessions cheap, while TurboQuant shrinks the memory footprint dramatically. Even API providers like Nous Research are feeling the competitive pressure from local inference, offering a 20% discount on all models at the Nous Portal, as @NousResearch announced, to ensure their Hermes agents remain an attractive alternative to fully localized setups.
AI Economics, Policy, and the Talent Shake-Up
As AI capabilities expand, the economic and political realities of the industry are growing increasingly complex. While doomsday predictions about AI-driven unemployment dominated headlines in previous years, the current macroeconomic data tells a completely different story. AI is proving to be highly labor-augmenting rather than replacing, a trend that is forcing economists and tech leaders to rethink how value is actually generated by human-machine teams.
Aaron Levie, CEO of Box, provided an excellent breakdown of a recent report from the Head of Economics at Anthropic. Levie points out that AI has not negatively impacted jobs because it requires human oversight to function properly. "Most jobs can’t be fully automated with AI, only certain tasks in those jobs," Levie explained. By automating specific, tedious tasks, human workers are freed up to multiply their overall output, thereby maintaining or even increasing the demand for skilled labor. This phenomenon, known as the Jevons paradox, is highly visible in software engineering, where AI tools have simply allowed companies of all sizes to initiate software projects that were previously impractical.
However, while domain experts remain in high demand, specialized AI researchers are facing a turbulent job market. @chris_j_paxton mocked the bizarre reality of the current AI labor cycle, exclaiming, "Imagine laying off your LLM pretraining team in 2026." This followed news that Amazon had gutted its AGI foundation MoE pre-training team, signaling that even tech giants are restructuring their core AI talent as pre-training techniques mature and consolidate. Furthermore, geopolitical tensions are actively warping the global AI market. @quxiaoyin amplified concerns regarding US trade restrictions, noting the irony of policies that put Chinese models on the entity list, effectively forcing American companies to purchase more expensive, heavily regulated domestic alternatives. Through all these macroeconomic shifts, the overarching sentiment of the AI community remains one of relentless acceleration, with figures like @gmoneyNFT and @dillon_mulroy reacting to recent industry developments with pure, unadulterated hype and calls of "agi."
Hardware Headaches: Smart TV Botnets and Wearables
The physical infrastructure powering the AI boom is often obscured by flashy software releases, but today brought a stark reminder of how consumer hardware is being covertly weaponized to feed the data-hungry machine. As AI companies desperately seek residential IP addresses to scrape the web without triggering data center blocks, unscrupulous proxy firms have resorted to hijacking everyday household devices, turning the concept of edge computing into a privacy nightmare.
An alarming investigation detailed by @IntCyberDigest revealed that LG Electronics is finally suspending smart TV apps that secretly turn hundreds of millions of screens into residential proxy nodes. Researchers found proxy SDKs embedded in over 42% of apps in LG's webOS store, effectively routing unknown third-party traffic through users' home internet connections. These residential proxy firms pay developers to bundle this hidden software, which is then rented out to entities conducting large-scale scraping operations, often for AI training data. The practice is deeply insidious because the devices, as @IntCyberDigest noted, are "hardware people don't treat as computers and can't easily inspect." Even worse, the problem remains completely unaddressed on Samsung's Tizen store, leaving millions of consumers unknowingly acting as exit nodes for anonymous web traffic that could easily damage their home IP reputation.
While living room hardware is being secretly exploited, wearable tech is openly embracing AI integration. Halliday announced priority access for its G2 display AI glasses, a new generation of eyewear explicitly designed to facilitate smarter conversations and meeting assistance. This juxtaposition perfectly encapsulates the current state of hardware in the AI era. On one hand, we have consumer electronics operating as a shadow infrastructure for data scraping without user consent. On the other, we have highly visible, purpose-built wearables trying to genuinely augment human intelligence. It highlights the desperate need for comprehensive hardware security and transparency as the physical and digital worlds become increasingly entangled.
Sources
so… I was on a vacation a while back just to clear my head & i bumped into an old friend of mine. dude’s lowkey an ai guru, so one topic led to another… from crypto to ai to robots then he said : “around 2022, when Chat gpt blew up, everyone suddenly needed massive amounts of high-quality human data.” “but companies kept running into the same issues:” > fake accounts > bot traffic > duplicate datasets > low-quality labels > ai-generated content training more AI models ( a thread)
This can't be true. No way https://t.co/nKAJa9YKyv
How to master graph engineering (Full Course)
The Claude Security plugin for Claude Code is now available in beta. Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Claude inference you already run. https://t.co/tEM7Tz7f1o
Introducing Impeccable 4 *world builder* However many ways you ask, "be creative!!!" does nothing to an LLM. v4 cracks it, and greenfield work is where it shows. • a creative engine for greenfield and redesign: directions seeded and fused from hundreds of human-approved visual worlds • hyper-optimized for frontier models (GPT 5.6, Fable and class), on a core 58% smaller • far simpler to use: no command to learn, it works out the job itself (blank slate, redesign, added section, scoped refinement) • mobile app design, the #1 request, now in alpha: Apple HIG or Material 3 on top, audit and adapt running as VoiceOver and TalkBack passes • Grok Build and Mistral Vibe join the supported harnesses • plenty more https://t.co/WglrY1uE4B
Today we are Introducing BTL-3. A 27B open-weight agent model built for agentic coding, structural tool use . The complete thing fits in one 8.39GB file under 2.5 bits per parameter smaller than an 8B model in fp16, and retains 92.2% of the 27B itelligence BTL-3 is trained for the loop real agents live in: reason, act, inspect the result, recover, continue. It handles single, sequential, and parallel tool calls and knows when the right move is no tool call at all. HumanEval: 95.12% pass@1 BFCL v4 AST: 88.5% (full 1,240-case set) Multiple tool calls: 95.5% Tool-call abstention: 91.2% 262K context architecture Two editions, both open today. BTL-3 is the maximum-quality checkpoint, for Transformers and vLLM. BTL-3 Compact is the entire model in one standalone 8.39GB GGUF. No base download. No reconstruction. One file, one command, a running agent. Compressing 27B this far normally destroys a model. Standard quantization couldn't do it, so we built the stack ourselves: packed AVQ2 decoder tensors, affine INT4, measured precision islands, packed vocabulary matrices, rank-32 output correction, behavioral repair. 2,416 tensors byte-verified at export. Then we tested whether the agent survived. On a fresh sealed 100-turn tool-contract gate, Compact retained 92.2% of teacher-correct behavior 100% on single, parallel, sequential, and abstention calls. 43 tok/s generation on an RTX PRO 6000. Fully local. Nothing leaves your machine. BTL-3: https://t.co/ddZWWr6i3o Compact: https://t.co/6URHEBGJgG Runtime + source: https://t.co/MjXQR6koKt Apache-2.0 model. MIT runtime.
❗️🚨 An Israeli company has backdoored hundreds of millions of households through countless Smart TV apps, and they're quietly turning Samsung and LG TVs into exit nodes for AI web-scraping. Your TV is relaying strangers' web traffic from your home IP, your bandwidth, your address attached to whatever those scraping jobs touch. Roku, Fire TV and Google TV banned the practice. Samsung and LG didn't. The culprit is Bright Data's proxy SDK, which rides inside Tizen and webOS apps, 200+ on webOS alone. Datacenter IPs get blocked, home IPs don't. Include Security reverse-engineered the SDK and found its relay protocol has no message signing, authentication, or device attestation. Their words: less secure than typical malware command-and-control. To make things worse, they found that in iOS the relay tunnel binds straight to the physical network interface, so it routes around any VPN the user is running. Bright Data's config also ships per-country tiers. Devices in Uzbekistan and Oman are cleared to relay down to 1% battery, with data caps up to 60x the worldwide default. Before the BaCkDoOrEd replies land: technically you agreed. In practice you were enrolled into a global proxy network you were never given the information to refuse. And these exit nodes drag down your IP's reputation, potentially leaving you with blocks from providers.
Towards Automating Eval Engineering
why did i post that so stupid so stpupid so stopid
First local agent to beat Hermes on benchmarks! ✦ runs Qwen, Gemma, Llama via llama.cpp ✦ stable-prefix caching keeps sessions cheap ✦ TurboQuant cuts the KV-cache 6.4× smaller ✦ 37 tasks solved vs Hermes' 31 on GAIA Level 1 Open source on macOS, Windows & Linux 👇 https://t.co/eiDkEeJehN
why we're buzzing
Buzz Desktop v0.4.23 🐝 More of you are running multiple communities — agents now start reliably when you switch between them. Plus safer sign-out (key backup required), cleaner GIF uploads, and UI polish. https://t.co/3Eypwg3swd
Looking for new oppotunities! Most members in our MoE pre-training team got impacted today 😅 Feel free to ask me about what happened inside of Amazon AGI foundation team and I will share my brief thoughts with my tenure period. Also, if you know there is any place that needs people to help/optimize LLM pretraining, please share that with me! I am open to new opportunities in the next few month! #amazon #agi #moe #llm #pretraining
I’ve made a visualization of the counterexample to the jacobian conjecture which involves high school math, three lines and two parabolas. In fact this diagram can be easily be turned into a proof that the construction gives a counterexample. Enjoy! https://t.co/n4k0plZMF9 @__alpoge__
Why hasn’t AI increased unemployment?
feeling codexy after a while. whatever tibo is talking about I'm really exited. also i think we might get a banked reset along with this. https://t.co/0Eo5Rs5VWI