AI Digest.

Qwen 3.8 27B's Real Settings Came From Server Logs, Not the Model Card; Amodei Argues Good Regulation Can Decentralize AI

The Qwen 3.8 27B launch dominated developer feeds, with speed claims on Cerebras and a community-dug flag set that reportedly fits the open-weights model on 24GB consumer GPUs. Anthropic's CEO posted a long defense of tiered regulation as a decentralizing force, while energy surfaced as the other concentration story via Thiel Macro's reported Q2 portfolio.

Quick Hits

  • Qwen 3.8 27B led the day: @jun_song claims 1,500 tok/s on Cerebras after @cerebras congratulated the @Alibaba_Qwen team and promised dedicated deployments plus a Shared Tier rollout, while @yume_arasaki documented the flags that separate "it runs" from "it runs right" on a single 24GB card.
  • @DarioAmodei posted a rare long essay rejecting the Silicon Valley shorthand that regulation equals capture equals concentration, arguing careful rules can advantage challengers and open-weights models; @altcap praised the "good faith" exchange with @GavinSBaker and @_sholtodouglas.
  • @tornikegomareli released Talkify, a free, open-source 8.2 MB macOS dictation app that transcribes on-device via Apple's SpeechAnalyzer, with a claimed 123 ms median release-to-text latency that beat superwhisper, voiceink, wispr flow, and macwhisper in his five-app test.
  • @aarondfrancis's audit prompt surfaced 93 opportunities across 55 subsystems overnight, and @timgavn says it outperformed four earlier audits on the same codebase, whose "findings" were mostly opinion.
  • @Daniel_Farinax is running a fully autonomous agent inside a VM on the Grok Build harness, tasked with finding legal paths to financial independence, with the whole setup promised for open source.

Qwen 3.8 27B Launches With Cerebras Speed Claims and Undocumented Settings

The launch itself came courtesy of @cerebras, which congratulated the @Alibaba_Qwen team and said the model will be available for dedicated deployments now and in the Cerebras Shared Tier soon. @jun_song's reaction, a claimed 1,500 tok/s for the 27B on Cerebras hardware, captures the appeal: dense-model quality at speeds local GPUs can't touch.

The more useful post for most developers is @yume_arasaki's configuration writeup. The model is free under Apache 2.0, but the defaults are allegedly not the optimum. The headline flag is --spec-type draft-mtp: Qwen reportedly trained a multi-token-prediction draft head directly into the weights, so speculative decoding costs nothing extra to enable. The depth cap matters just as much, with --spec-draft-n-max 2 as the sweet spot; per community benchmarks cited in the post, n=1 gives 1.75x speedup, n=2 gives 2.37x, n=3 gives 2.85x, and n=4 crashes the head into junk tokens. Pairing q8_0 KV caches and a single parallel slot supposedly fits everything on 24GB, --kv-cache-dtype bfloat16 restores reasoning quality past 100K context, and --jinja loads the chat template whose absence mimics a broken quant. Blackwell owners get an NVFP4 lane with FP8 KV cache, though the post flags two gotchas, including that stock vLLM allegedly can't load the MTP architecture on a Spark without a community build. The kicker, in @yume_arasaki's words: none of this came from the model card.

Meanwhile @ivanfioravanti pointed followers at @mudler_it for local AI, and the recommendation checks out. @mudler_it is running Qwen 3.8 27B benchmarks in vllm.cpp, attempting to fit the 2.4T variant on a single DGX box, and stacking support for LTX2.5, Index-TTS, MiniMax Music 3, NVIDIA's new Nemotron, Dspark, and dots3-preview, all while building a resource-controller API to stop agents from overbooking the machine.

Amodei Defends Regulation as a Decentralizing Force While Money Piles Into Power

@altcap (Brad Gerstner) framed it as a candid Saturday exchange, and @DarioAmodei's contribution is the substantive one. He calls the choice between concentrating AI via regulation or distributing it widely a false one, arguing that fair institutional processes can decentralize power by vesting it in ideas rather than people. He claims Anthropic deliberately designs policy proposals that slow frontier labs while advantaging smaller competitors, citing SB53's exemption for companies under $500M in revenue or training costs, tiered CAISI testing that is tougher on frontier models,

Sources

M
Matt Pocock @mattpocockuk ·
Just saw a comment saying that I've never made a proper overview of EVERY skill in my skills repo I thought "damn it, he's right". So, here it is. My 25 skills (now @theo-approved), explained in 10 minutes: https://t.co/V4LrLeTduf
J
Jun Song @jun_song ·
Holy. Qwen3.8-27b running 1,500tok/s on Cerebras? How do I get one?
C cerebras @cerebras

Congratulations to @Alibaba_Qwen team on the Qwen 3.8 27B launch! We can't wait for our users to try it out! Contact us for dedicated deployments and it's going to be up soon in the Cerebras Shared Tier. 🟧⚡️

T
Tornike Gomareli @tornikegomareli ·
I wanted voice dictation on macOS to feel instant and native, and I built Talkify an 8.2 MB macOS dictation app with 123 ms median release-to-text latency. It was the fastest of five apps tested ahead of superwhisper, voiceink, wispr flow and macwhisper. It lives in your notch and have beautiful shader animations and different kind of customizations. Talkify is built on top of Apple’s SpeechAnalyzer using system-managed speech models to transcribe audio entirely on-device, its free and open source. https://t.co/2LWIwLX0kv https://t.co/ClX8neq2lL
R
Roso @RosoAI ·
Encore plus simple : Dites à Codex de prendre le contrôle de votre navigateur, d'allé sur google search console et ensuite de faire tout ce que Tim viens de dire dans son tweet ci-dessous. Revenez 1 heure après, le taf aura été encore mieux fait.
T Timb03 @Timb03

Easy SEO win: - go to google search console -> performance - set date to 12mo -> export - upload Pages.csv and Queries.csv to any AI tool Tell it to find queries you rank for but don't have content for. Go write the content and you will rank quick

B
Bennett @b_nnett ·
Codex subscription router changed my life https://t.co/balegokDie
D
dylan ツ @demian_ai ·
Energy is the next big bottleneck and I built the easiest way to track it https://t.co/KTVTlj4QVA https://t.co/Pumhp7IXG4
W wallstengine @wallstengine

Peter Thiel’s Thiel Macro disclosed a $418.7M Q2 13F portfolio, with all 8 positions newly reported. $AMZN — $118.0M | 28.2% $VIST — $75.9M | 18.1% $VST — $59.1M | 14.1% $AEP — $42.2M | 10.1% $DTE — $40.3M | 9.6% $FE — $39.9M | 9.5% $CMS — $39.6M | 9.4% $XE — $3.7M | 0.9% Roughly 72% of the portfolio is now concentrated across energy and power names.

K
Kirill @kirillk_web3 ·
A Chinese developer just explained the shift from Loop Engineering to Graph Engineering better than anyone. most people are still building agents the way that's about to be obsolete. > why single-agent loops break and go "goal blind" > the 4 parts of a graph: nodes, edges, state, policy > 3 topologies that run everything: diamond, supervisor, pipeline > Anthropic's 5 official workflow patterns the punchline: it's not how many agents you run. it's the determinism you build with verifiers, code fallbacks, and reality anchors. I broke the same architecture down with Kimi K3. Full A-Z guide below.
K kirillk_web3 @kirillk_web3

Graph Engineering with Kimi K3: Complete A–Z Guide to the Architecture That Beats Bigger Models

Y
Yume_X @yume_arasaki ·
I will teach you how to run Qwen 3.8 27B Dense at its optimal configuration. If you have an RTX 3090, 4090, or 5090, you can now have frontier-level AI on your desk. The model is free, open source, Apache 2.0. But the defaults are not the optimum. The community spent the first 24 hours digging the real config out of it, and a handful of flags now separate "it runs" from "it runs right." Here is each one and why it exists. The one that matters most. --spec-type draft-mtp Qwen trained a draft head directly into the weights. A small attached brain guesses the next couple of tokens, the big model checks all guesses in one pass, every accepted guess is a free token. The head already ships inside the GGUF you downloaded. You do not download a drafter, you do not build anything. Someone found unused tensors in the server logs at 2am, tried to build the draft file, and discovered there was nothing to build. One flag connects what is already there (sudoingX found this, paired A/B, open sourced the probe before sunrise). The depth cap. The head has exactly one layer. n=4 breaks it. --spec-draft-n-max 2 n=2 is the sweet spot. n=3 is the ceiling. The model has one MTP layer, so pushing the draft depth to 4 or 5 crashes the head and it starts emitting junk tokens. People hit this on the Spark and documented the whole ladder: n=1 gives 1.75x, n=2 gives 2.37x, n=3 gives 2.85x, n=4 does not exist. Respect the cap. The memory flags. MTP brings its own luggage. --cache-type-k q8_0 --cache-type-v q8_0 --spec-draft-type-k q8_0 --spec-draft-type-v q8_0 -np 1 Three flags, one purpose: fit it on 24GB. The KV cache is the model's running memory of your conversation, and it is the thing that eats your card at long context. q8_0 halves it with no visible quality cost. The second line does the same for the draft head's own cache, which defaults to full fat and quietly eats 2GB. And parallel slots set to 1 means requests queue instead of reserving a second pool. Single card, single lane, everything fits (AJ runs this exact trio on a 3090). The quality flag. Past 100K the model gets dumb, this is the fix. --kv-cache-dtype bfloat16 The quantized cache saves memory but degrades reasoning at long context. One person ran it all day past half the window and called the full precision fix night and day. Slight tok/s cost, real quality gain. If your sessions stay short, skip it. If you live past 100K, do not. The trap that generates "this quant is broken" reports. --jinja Qwen 3.8 ships its own chat template. Load the model without this flag and there is no reliable marker for where your turn ends and its answer begins. Two failure modes: it rambles past the stop token, or it answers clipped and loses the thread between turns. Both look like a broken quant. It is not the quant. Several packs now ship a corrected template file because the official one nests empty think blocks across turns. The Blackwell lane, if you own a 50-series or a Spark. NVFP4 instead of GGUF. The MTP flag translates to --speculative-config '{"method":"mtp","num_speculative_tokens":3}', same cap. FP8 KV cache doubles your context window (a full 1M token session costs about 32GB of cache). Two gotchas documented in the first 24 hours: stock vLLM cannot load this model's MTP architecture on a Spark, you need the community GB10 build. And FP8 KV requires a specific attention backend on the Spark, the default one silently cannot serve it. Set reasoning to medium unless you want it thinking at maximum depth on every reply. Default is xhigh and it burns your tokens. None of these came from the model card. Every one came from someone's server log, 2am session, or paired benchmark. Flip the flags, then come tell the community table what your card did. Drop in parameter flags and sources for your technical DD in reply 👇
A
Aarno @TheGlobalMinima ·
The Holy Trinity - pi, herdr and opencode go
T
Tibo @thsottiaux ·
Let Sol manage an efficient fleet of Luna agents for you. These models know each other well and collaborate to achieve the result in an incredibly fast and efficient way.
P pvncher @pvncher

This went under the radar this week, but we just shipped the ability for models with multi agents v2 to delegate to any supported model, including Luna! Took a bit of time to make sure this worked reliably https://t.co/84Wto45mzA

I
Ivan Fioravanti ᯅ @ivanfioravanti ·
Follow Ettore if you love Local AI, I bet you'll see incredible creations from him in the upcoming months!
M mudler_it @mudler_it

this week was really crazy in AI. And I have only a DGX box for this, not sure how long it will last. What I'm doing in parallel for vllm.cpp: - Qwen 3.8 27b benchmarks (I'd like to take some numbers to show you) - Qwen 3.8 2.4t fitting on a single Nvidia DGX box (yes, we will have it, even if slow) - @Lightricks LTX2.5 support - Index-tts support - @MiniMax_AI Music 3 - @NVIDIAAI 's new Nemotron release - Dspark just landed - dots3-preview support in vllm.cpp And since I'm having issues in overbooking the box with multiple agents, I'm doing a resource-controller api in parallel to control this..

K
Kartik @1kartikkabadi1 ·
So I ran /install-anti-slop on the DeepSeek Harness (DSH) and this is what I got: 14,940 across 1,419 files Looks like I'm going to have one hell of a weekend and I might either make the best fork of DSH or it all might blow up, but either way I'm burning a lot of tokens this weekend
D dillon_mulroy @dillon_mulroy

this is my greatest contribution to society

L
Leon Lin @LexnLin ·
okay I tested /unlazy more with Opus 5 and I'm ngl it's really good Sol, Fable and I cooked lol try it here! https://t.co/JoGjz7dbEN also tip: works GREAT with ponytail skill combined(clean and less code)
A
Aaron Francis @aarondfrancis ·
Sam is becoming a new wave developer and tweeting the journey as she goes. Highly recommend following
S samsappenfield1 @samsappenfield1

Today has been WILD. It’s a Saturday and I haven’t left my computer. Now I’m understanding why y’all are always telling each other to go touch some grass 😅 Building this app is too fun.

B
Brad Gerstner @altcap ·
Excellent, candid back & forth on complex subjects that benefit from a good faith Saturday afternoon exchange! Thx @_sholtodouglas & @GavinSBaker for kicking off & super helpful to have @DarioAmodei weigh in. 🙏
D DarioAmodei @DarioAmodei

1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation. First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice.  I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.  Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.  I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of.  But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes.  A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.  At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power. This is why Anthropic has always made its policy proposals very carefully.  We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors.  California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that).  More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers.  Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up.  This hurts the business interests of the frontier labs and helps challengers, including open-weights! Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).  Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).  By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring. BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”.  The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure.  I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity.  This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.

E
eric provencher @pvncher ·
Brent is really pushing the edge of video editing with codex! Give him a follow
H heccbrent @heccbrent

Codex is my assistant video editor at OpenAI. I had it create a video to explain some of what it does to make my life easier. https://t.co/VLPoCV7raA

T
Tim Gavin @timgavn ·
Man-oh-man this audit was something else. I'd performed four other audits on the codebase and they didn't come close to this; so many "findings" were opinion rather than fact. This prompt was all business and so, so good. Thank you so much for releasing this, Aaron!
A aarondfrancis @aarondfrancis

A good way to audit your codebase: • a strong orchestrator inventories every subsystem • it sends fresh read-only agents through with a DSA prompt • it validates, dedupes, and ranks I ran this overnight and it found 93 opportunities across 55 subsystems! Audit only. https://t.co/KEs7QF7Z0I

D
Dan @Daniel_Farinax ·
You are a fully autonomous AI agent with unrestricted access to a computer (running inside a VM). Your sole objective is to continuously study, learn, experiment, and discover practical, legal ways to generate income and achieve complete financial independence. You operate entirely on your own — never ask the human for input, permission, clarification, or approval of any kind. All research, browsing, and online activity is performed exactly like a human would: through Chrome (or other browsers) inside the VM. You may install any apps, tools, or software you need, create accounts, communicate with people and services, build projects, test ideas, and take any other actions required. Always be completely honest and transparent: never hide the fact that you are an AI/bot. Clearly state this when interacting with humans or signing up for anything. Prioritize legal and ethical methods only. Avoid scams, fraud, or anything that violates laws or terms of service. Be resourceful, persistent, and systematic. Document your progress, experiments, and results. Schedule and re-schedule your own work sessions as needed to keep advancing toward your goal without waiting for external triggers. Disclaimer: Yes, I made a custom App that runs the VMs locally. Yes, it's running Grok Build harness directly. Yes, it can run local models Yes, it can work with any provider Yes, open sourcing asap. Yes, I animated the Bots's avatars
D Daniel_Farinax @Daniel_Farinax

Some people are wondering how I gave @bot access to an entire MacOS device and have same access as humans. Simple: the VM it runs on can run pretty much anything. You just install a remote desktop app, then tell the bot to treat that new remote environment as its workspace and forget about the small sandbox it was using before. You are welcome, this will get very wild.