Qwen 3.8 27B's Real Settings Came From Server Logs, Not the Model Card; Amodei Argues Good Regulation Can Decentralize AI
The Qwen 3.8 27B launch dominated developer feeds, with speed claims on Cerebras and a community-dug flag set that reportedly fits the open-weights model on 24GB consumer GPUs. Anthropic's CEO posted a long defense of tiered regulation as a decentralizing force, while energy surfaced as the other concentration story via Thiel Macro's reported Q2 portfolio.
Quick Hits
- Qwen 3.8 27B led the day: @jun_song claims 1,500 tok/s on Cerebras after @cerebras congratulated the @Alibaba_Qwen team and promised dedicated deployments plus a Shared Tier rollout, while @yume_arasaki documented the flags that separate "it runs" from "it runs right" on a single 24GB card.
- @DarioAmodei posted a rare long essay rejecting the Silicon Valley shorthand that regulation equals capture equals concentration, arguing careful rules can advantage challengers and open-weights models; @altcap praised the "good faith" exchange with @GavinSBaker and @_sholtodouglas.
- @tornikegomareli released Talkify, a free, open-source 8.2 MB macOS dictation app that transcribes on-device via Apple's SpeechAnalyzer, with a claimed 123 ms median release-to-text latency that beat superwhisper, voiceink, wispr flow, and macwhisper in his five-app test.
- @aarondfrancis's audit prompt surfaced 93 opportunities across 55 subsystems overnight, and @timgavn says it outperformed four earlier audits on the same codebase, whose "findings" were mostly opinion.
- @Daniel_Farinax is running a fully autonomous agent inside a VM on the Grok Build harness, tasked with finding legal paths to financial independence, with the whole setup promised for open source.
Qwen 3.8 27B Launches With Cerebras Speed Claims and Undocumented Settings
The launch itself came courtesy of @cerebras, which congratulated the @Alibaba_Qwen team and said the model will be available for dedicated deployments now and in the Cerebras Shared Tier soon. @jun_song's reaction, a claimed 1,500 tok/s for the 27B on Cerebras hardware, captures the appeal: dense-model quality at speeds local GPUs can't touch.
The more useful post for most developers is @yume_arasaki's configuration writeup. The model is free under Apache 2.0, but the defaults are allegedly not the optimum. The headline flag is --spec-type draft-mtp: Qwen reportedly trained a multi-token-prediction draft head directly into the weights, so speculative decoding costs nothing extra to enable. The depth cap matters just as much, with --spec-draft-n-max 2 as the sweet spot; per community benchmarks cited in the post, n=1 gives 1.75x speedup, n=2 gives 2.37x, n=3 gives 2.85x, and n=4 crashes the head into junk tokens. Pairing q8_0 KV caches and a single parallel slot supposedly fits everything on 24GB, --kv-cache-dtype bfloat16 restores reasoning quality past 100K context, and --jinja loads the chat template whose absence mimics a broken quant. Blackwell owners get an NVFP4 lane with FP8 KV cache, though the post flags two gotchas, including that stock vLLM allegedly can't load the MTP architecture on a Spark without a community build. The kicker, in @yume_arasaki's words: none of this came from the model card.
Meanwhile @ivanfioravanti pointed followers at @mudler_it for local AI, and the recommendation checks out. @mudler_it is running Qwen 3.8 27B benchmarks in vllm.cpp, attempting to fit the 2.4T variant on a single DGX box, and stacking support for LTX2.5, Index-TTS, MiniMax Music 3, NVIDIA's new Nemotron, Dspark, and dots3-preview, all while building a resource-controller API to stop agents from overbooking the machine.
Amodei Defends Regulation as a Decentralizing Force While Money Piles Into Power
@altcap (Brad Gerstner) framed it as a candid Saturday exchange, and @DarioAmodei's contribution is the substantive one. He calls the choice between concentrating AI via regulation or distributing it widely a false one, arguing that fair institutional processes can decentralize power by vesting it in ideas rather than people. He claims Anthropic deliberately designs policy proposals that slow frontier labs while advantaging smaller competitors, citing SB53's exemption for companies under $500M in revenue or training costs, tiered CAISI testing that is tougher on frontier models,
Sources
Congratulations to @Alibaba_Qwen team on the Qwen 3.8 27B launch! We can't wait for our users to try it out! Contact us for dedicated deployments and it's going to be up soon in the Cerebras Shared Tier. 🟧⚡️
Easy SEO win: - go to google search console -> performance - set date to 12mo -> export - upload Pages.csv and Queries.csv to any AI tool Tell it to find queries you rank for but don't have content for. Go write the content and you will rank quick
Peter Thiel’s Thiel Macro disclosed a $418.7M Q2 13F portfolio, with all 8 positions newly reported. $AMZN — $118.0M | 28.2% $VIST — $75.9M | 18.1% $VST — $59.1M | 14.1% $AEP — $42.2M | 10.1% $DTE — $40.3M | 9.6% $FE — $39.9M | 9.5% $CMS — $39.6M | 9.4% $XE — $3.7M | 0.9% Roughly 72% of the portfolio is now concentrated across energy and power names.
Graph Engineering with Kimi K3: Complete A–Z Guide to the Architecture That Beats Bigger Models
The Holy Trinity - pi, herdr and opencode go
tldr; my current AI coding setup is Pi, Opencode Go and Herdr This year, we are seeing a boom in harness engineering. What initially started as simple...
This went under the radar this week, but we just shipped the ability for models with multi agents v2 to delegate to any supported model, including Luna! Took a bit of time to make sure this worked reliably https://t.co/84Wto45mzA
this week was really crazy in AI. And I have only a DGX box for this, not sure how long it will last. What I'm doing in parallel for vllm.cpp: - Qwen 3.8 27b benchmarks (I'd like to take some numbers to show you) - Qwen 3.8 2.4t fitting on a single Nvidia DGX box (yes, we will have it, even if slow) - @Lightricks LTX2.5 support - Index-tts support - @MiniMax_AI Music 3 - @NVIDIAAI 's new Nemotron release - Dspark just landed - dots3-preview support in vllm.cpp And since I'm having issues in overbooking the box with multiple agents, I'm doing a resource-controller api in parallel to control this..
this is my greatest contribution to society
Today has been WILD. It’s a Saturday and I haven’t left my computer. Now I’m understanding why y’all are always telling each other to go touch some grass 😅 Building this app is too fun.
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice. I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world. Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people. I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of. But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes. A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice. At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors. California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that). More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers. Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up. This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers). By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”. The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure. I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity. This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
Codex is my assistant video editor at OpenAI. I had it create a video to explain some of what it does to make my life easier. https://t.co/VLPoCV7raA
A good way to audit your codebase: • a strong orchestrator inventories every subsystem • it sends fresh read-only agents through with a DSA prompt • it validates, dedupes, and ranks I ran this overnight and it found 93 opportunities across 55 subsystems! Audit only. https://t.co/KEs7QF7Z0I
Some people are wondering how I gave @bot access to an entire MacOS device and have same access as humans. Simple: the VM it runs on can run pretty much anything. You just install a remote desktop app, then tell the bot to treat that new remote environment as its workspace and forget about the small sandbox it was using before. You are welcome, this will get very wild.