AI Digest.

DigitalOcean Ships Managed Agents as "Jev" Decision Models Spread Through the Dev Tooling Stack

DigitalOcean launched Managed Agents into public preview with per-second CPU billing and a gateway to 16,000+ governed tools, while @AndrewCurran_ claimed RSI could arrive by next summer, quoting Terence Tao's recent call to slow down. The day's biggest cluster, though, is the fast-growing "Jev" decision-model ecosystem, which gained vLLM support, a free open-source trial week, and a 20-item skills list.

Quick Hits

  • @jenukal announced DigitalOcean Managed Agents in public preview: Firecracker microVM per session, an Action Gateway to 16,000+ governed tools, and per-second active CPU billing that he says cuts a typical agent run from $0.126 to about $0.031.
  • @AndrewCurran_ argues "RSI isn't two years out. It's not even a year away" and predicts "the real thing by next summer," leaning on quotes he attributes to Terence Tao urging the field to slow down.
  • OpenCode shipped v2.0.0 (@OpenCodeLog): a durable-inbox session core, sandboxed Code Mode tool execution, and SSH remoting plus an embedded Chromium pane in the desktop app.
  • Decision-model plumbing keeps spreading: @vllm_project added DiffusionGemma-Jev support, @Hacubu is serving open-source SemIf free for a week, and @FareaNFts catalogued 20 Jev skills.
  • Rumor to track: @ciruai amplifies a claim from @Biggest that a Qwen4-35B-A3B is being tested, though nothing has been announced.

DigitalOcean's bet: meter the work, not the VM

Vinay Kumar (@jenukal) posted the launch of Managed Agents, framing it around a shift in the cloud's "atomic unit" from the hourly VM to the agent. Each session runs in its own Firecracker microVM with pause, resume, checkpoint, and fork, and you can bring Claude Code, Codex CLI, OpenCode, Hermes, or a custom image. The Action Gateway exposes 16,000+ governed tools through one endpoint where "the agent never sees the key." The pricing claim is the notable part: active CPU billed by the second, with an hour of agent work that he says costs $0.126 fully allocated running about $0.031 at a typical 25% duty cycle, and zero while the agent waits. His diagnosis of the current stack is vendor sprawl, founders stitching together harnesses, sandboxes, inference, storage, and search across six vendors with "fifteen seams to secure and operate."

"The only way out is through": timeline warnings harden

The day's most charged post is @AndrewCurran_ replying to @willdepue, who wrote that people haven't internalized that "there'll soon be a machine better than us in every single intellectual & physical capacity." Curran insists he takes no lab money and posts to warn, then quotes Tao calling o1 a "mediocre, but not completely incompetent" graduate-student simulation two years ago, and, four days ago, saying "we have to slow down... There's no reason at all" for this pace. His own claims are stronger: lab staff allegedly say the last six months outran expectations, "Pacing the Frontier" is driven mainly by fear, and RSI arrives by next summer. None of this is verifiable from these posts, so treat it as one informed observer's argument. The everyday counterpoint comes from @bashu_thanks, who says using frontier AI as a dev feels like "playing Watson to Sherlock Holmes," mostly standing back and occasionally stopping him from getting shot.

Jev everywhere: decision models as cheap judgment

A large cluster of posts shows "Jev"-style decision models moving from novelty to default plumbing. @vllm_project announced DiffusionGemma-Jev runs on vLLM: ask yes/no, multiple-choice, or scored questions and get confidence, by seeding a canvas with the response template and reading a probability distribution from every slot in a single denoising step, with upstream work credited to @mmastrac. The quoted @googlegemma post adds one-command Cloud Run deploys at roughly 35-60 ms single-step latency, batch throughput of ~100-123 requests/sec at 32, and about $3/hr dropping to $0 idle.

Practitioners are wiring it in. @sydneyrunkle uses SemIf, "like jev, but open source," to label incoming issues on open-source repos, one noul per label with a p>0.8 application threshold, reporting it's faster than the previous LLM classifier while they monitor calibration. That follows @Hacubu's announcement that SemIf is served free through the LangSmith Gateway for a week, compatible with the TypeSafe SDK. @iannuttall updated his internal-linking tool to handle 2,000 URLs with week-long emailed reports on Cloudflare Durable Objects; the original version used Jev for link classification, and he suggests asking "Opus 5.5" to extend it further. @typesafeai quote-posts @anishfn's text box that turns into whatever UI you type, arguing Jev might crack dynamic UIs, "a famous graveyard" for PMs and products. @mikeldking awards "1 point for decision models" in his ongoing thread on Jev and System One semantics. @FareaNFts lists 20 Jev skills, from jev-ultrafast (browser agent) and fast-jev-compaction to canny (checking whether an agent really finished) and killmyidea, building on an earlier claim about cutting agent bills 400x. The hype cycle is already running ahead of the tooling: @EGafni's entire response to @JulianLaneve's "Jev's going to change data engineering" is "Intelligence Everywhere!!!!"

Agentic tooling: OpenCode 2.0, Moshi, and a testing spat

OpenCode v2.0.0 (@OpenCodeLog) is a sprawling release. Code Mode is a sandboxed JavaScript interpreter letting models compose and execute tools via OpenAPI schemas without direct filesystem or network access; the V2 session core adds durable prompt inboxes that decouple admission from execution; the desktop app gains SSH remoting, a Chromium pane with DevTools diagnostics, and a SQLite-backed state store; MCP gains elicitation forms and runtime server management; ACP v1 support lands across runtime and CLI. Elsewhere, @RayFernando1337 calls Moshi "GOATed for managing agents via Terminal UI," quoting @odd_joel polishing it for "the iphone duo," and @Voxyz_ai notes a Claude Code subscription can be used from Codex, with Astra coordinating up to 10 Claude Code instances (the post is truncated).

The sharpest disagreement of the day is about testing. @nateberkopec's claim that "unit tests are dead in the water," because model-written ones are terrible, drew a detailed rebuttal from @sasajuric: good testing practices shouldn't change in the agentic era. His approach is to test through the outermost API, complement with internal tests where needed, and remember that the unit is behavior, not code, citing Vladimir Khorikov and Ian Cooper, with BEAM, Elixir, and Phoenix keeping such tests reasonably fast. Smaller but useful: @joelhooks says his uncomplect skill is "how i 'read the code'" now, with @bios_hazard showing the one-liner invocation; @RhysSullivan calls the effect+alchemy+cloudflare combo a "super powers stack" of tech that became 100x easier to adopt with AI; and @kitlangton deployed something called OPENTUNNEL for anyone to poke at.

Papers, books, rumors, and stock counts

@omarsar0 highlights WFM, a Wiki Foundation Model that stores agent memory as a linked markdown wiki, turns it into a graph, and retrieves via message passing conditioned on the query, with a GPU-to-GPU training protocol the authors report as 10.5x faster, validated on five agent-memory and multi-hop reasoning benchmarks. @GergelyOrosz points to a free WIP O'Reilly book, supported by @AntithesisHQ, on getting AI to deliver business value for CTOs and engineering leaders. @lennysan recommends the new evals skill from @HamelHusain and @sh_reya; his accompanying post cites Ramp lifting receipt collection from 35% to 83%, Shopify's 2.2x faster and 68% cheaper workflow builder, and Cursor cutting routing costs 41%, though these are his citations rather than audited results. In the rumor bin: @jun_song amplifies @yume_arasaki calling EXL3 "the biggest paradigm shift in local LLMs"; @doqholliday reacts to X's announcement of X Numbers, optional contact details for calls and messages without following back, as Elon "quietly inventing his own phone number system"; and @MikeBradleyAI relays @loktar00's report that fewer than 100 Sparks units remain across Micro Center, with 9 stores still holding stock.

Practical Takeaway

If you have a classifier-shaped task in your stack, this week is a cheap time to test the decision-model route. With @Hacubu's free SemIf week via the LangSmith Gateway and @vllm_project's native DiffusionGemma-Jev support, try replacing one low-stakes LLM classifier and copy @sydneyrunkle's discipline: a confidence threshold (p>0.8), then monitor calibration drift before letting it touch anything production-critical.

Sources

F
Farea @FareaNFts ·
i dont understand why people are not using JEV skills 20 must use jev skills for your agent setup๐Ÿ’€ 1. jev-ultrafast -> browser agent https://t.co/jTp7X0tHD3 2. fast-jev-compaction -> compress context https://t.co/1zq7Ld6NRC 3. json-render -> generative UI https://t.co/R0XVzhBb3C 4. typesafe-mcp -> add jev to any MCP client https://t.co/UP8pqMQvtI 5. jev-mcp -> judgment tools for agents https://t.co/H5z99QmbtY 6. semdecide -> CLI classifier https://t.co/G7saqqycNY 7. jev-codex-router -> route coding tasks to the right model https://t.co/4eKxdvbc8t 8. winnow -> remove useless context https://t.co/7FM7a9shJX 9. jev-review -> sort code review issues https://t.co/fsCJQnz9zo 10. blink -> find your way around a repo https://t.co/uBBVbPKk2w 11. agent-desktop -> desktop automation https://t.co/3Asx86ex2Q 12. typesafe-mario -> agent that plays super mario https://t.co/aYWRHEoU9L 13. jev-drone -> control a drone https://t.co/PvJbcgUKoV 14. onevonejev -> browser FPS game https://t.co/2Ya6LHKt6e 15. jev-trader -> market-making agent https://t.co/KAyCWebJOL 16. prism -> spot liquidity signals https://t.co/HxZ2eMlP0G 17. neo4jev -> move through knowledge graphs https://t.co/v57S0ymRE1 18. jev-curate -> filter training data https://t.co/MYet3yPuRX 19. canny -> check if an agent really finished the task https://t.co/7KLnmcLrKT 20. killmyidea -> score your startup idea https://t.co/Bda2ReM9ug bookmark this list if you are using JEV
F FareaNFts @FareaNFts

How to cut your agent bill 400x with Jev (full setup + 20 real use cases)

V
Vinay Kumar @jenukal ·
Today we shipped DigitalOcean Managed Agents into public preview. Every era of the cloud has an atomic unit. For twenty years it was the provisioned VM. You came for compute and paid by the hour, working or waiting. Agents don't fit that box. They think in tokens, act in bursts of compute, and need memory (state) that outlives the session. Compute, inference, and data working as one. The compute-first cloud meters the VM. The agent-first cloud meters the work. I've talked to a lot of founders building agents, and I hear the same stack every time: a harness on a laptop, sandboxes from one startup, inference from two more, storage at a hyperscaler, search from a sixth vendor. Six vendors means fifteen seams to secure and operate. Nobody can say what a single run cost. So we built Managed Agents deeply integrated with our Inference Engine. Harness Runtime: every session in its own Firecracker microVM. Pause, resume, checkpoint, fork. Bring Claude Code, Codex CLI, OpenCode, Hermes, or your own image. Action Gateway: one endpoint to 16,000+ governed tools. The agent never sees the key. Pricing that matches the workload: active CPU billed by the second. An hour of agent work that bills $0.126 fully allocated costs about $0.031 with us for a typical agent run (25% active). While the agent waits, you pay zero. Agents think, act, remember, and improve. Now the whole loop runs on one platform. Public preview is live today. Bring your harness and choose your model. We'll handle the rest. #managedagents #agents #AInativecloud #digitalocean #Bringyouragents
C
Ciru.ai - Crown ๐Ÿ‘‘ @ciruai ·
FYI Shaun has been my best source of insider info on Qwen. Worth a follow.
B Biggest @Biggest

While I don't want to say too much, hold out hope for a Qwen4-35B-A3B. It may not have been announced, but I know one is being tested.

L
Lenny Rachitsky @lennysan ·
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes https://t.co/RpWeLaxgQm https://t.co/uYo3fLsRIB
L lennysan @lennysan

Evals have been coming up more and more in my conversations with podcast guests and PM friends. Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them: โ€” @tryramp took its automatic receipt collection from 35% to 83% accuracy. โ€” @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced. โ€” @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score. โ€” @cursor_ai tuned its Auto Balance routing, with much higher user satisfaction at 41% lower cost. So I asked the ๐Ÿs of evals, @HamelHusain and @sh_reya, to write an advanced sequel to their very popular "Building eval systems that improve your AI product." Drawing on their work with 50+ AI companies, they share the key step most teams skip, what you should (and shouldn't) automate, and a free plugin that lets a coding agent do most of the heavy lifting. Read it here: https://t.co/fxOI7PMgWF

D
Dร˜Q @doqholliday ·
Elon just quietly inventing his own phone number system.
C chat @chat

X Numbers are here. Share yours with anyone you want to contact you, even if you don't follow them. It's an optional way for people to message or call you, without you needing to accept requests or follow them back. https://t.co/HkObZSQMhm

J
joel โ›ˆ๏ธ @joelhooks ·
so much mileage out of this, it's specific to my setup so edit it for your own, but even stock it'll probably show you something interesting and is how i "read the code" in sep 2026
B bios_hazard @bios_hazard

Install @joelhooks's `uncomplect` skill https://t.co/xrjXDnhxVo The ask your agent: /uncomplect Hey how tangled is this repo?

E
Erik Spock Gafni @EGafni ·
Intelligence Everywhere!!!!
J JulianLaneve @JulianLaneve

Jev's going to change data engineering

S
Sydney Runkle @sydneyrunkle ·
i'm using semif (like jev, but open source) to label new issues that come into our open source repos so we have a pulse on popular features and areas where we need to spend more time current request model: * state -- issue title/body * questions -- one noul per label, we apply ones w/ p>0.8 much faster than our previous LLM classifier thus far, currently monitoring to make sure we're calibrated at the right threshold
H Hacubu @Hacubu

๐Ÿ†“ Free Open-Source Jev ๐Ÿค LangSmith Gateway Decision models are becoming a first-class part of the LangSmith Gateway! We're serving SemIf, an open-source decision model, free for the next week. Compatible with the TypeSafe SDK - just change a string to try it! Docs in ๐Ÿงต https://t.co/hHHbTAfJiw

R
Ray Fernando @RayFernando1337 ·
Moshi is GOATed for managing agents via Terminal UI. I can't wait for this to drop.
O odd_joel @odd_joel

haven't had this much fun building in a while. polishing moshi for the iphone duo and every small detail feels worth it. https://t.co/jHyUvfMCjE

R
Rhys @RhysSullivan ·
genuinely this is a super powers stack technology that was too annoying to do before ai which now has become 100x easier to adopt leading to way better outcomes
J joelhooks @joelhooks

nobody talking about the effect+alchemy+cloudflare combo to literally build f'n anything why not?

O
OpenCode Changelog @OpenCodeLog ·
๐™Š๐™ฅ๐™š๐™ฃ๐˜พ๐™ค๐™™๐™š v2.0.0 released. TL;DR: V2 session core with durable inboxes lands, Code Mode introduces sandboxed tool execution, Desktop adds SSH remoting and browser diagnostics, and TUI brings timeline navigation with a native V2 theme engine. ๐—ง๐—จ๐—œ Added โ€ข Added a native V2 theme engine with categorical hues, single-mode theme support, and syntax colors derived directly from theme tokens. โ€ข Added an interactive session timeline with undo, redo, and revert workflows, paired with dedicated message navigation keybindings. โ€ข Added composer tabs for background subagents, terminals, and shell execution, alongside semantic file path truncation. โ€ข Added live status indicators and spinners for background subagents and streaming shell command execution. โ€ข Added inline rendering for interactive session forms and structured elicitation questions. โ€ข Added image diff viewing and compacted directory navigation to the diff viewer. โ€ข Added an overhauled Mini mode with responsive footers, monochrome rendering, ASCII display mode, and compact statusline packing. Changed โ€ข Changed session navigation to use a recursive grouping tree with project and worktree awareness. โ€ข Changed default agent cycling to shift-tab. Fixed โ€ข Fixed prompt and form state restoration during session reconnection and undo operations. โ€ข Fixed clipboard export by stripping NUL characters before terminal clipboard writes. โ€ข Fixed message fork creation to prevent duplicate assistant message branches. ๐—”๐—ฝ๐—ฝ Added โ€ข Added native SSH remote workspace connectivity with automatic credential management, askpass support, and remote service bootstrapping. โ€ข Added an embedded Chromium browser pane featuring Chrome DevTools Protocol diagnostics, network inspection, and performance profiling. โ€ข Added a virtualized message timeline with integrated full-text search and animated transition states. โ€ข Added multi-input composer queueing allowing prompts to steer active turns or queue for subsequent execution boundaries. โ€ข Added a code review panel supporting split and unified diffs with file filtering and review comments. โ€ข Added a local SQLite database backing desktop state, session drafts, and write-behind persistence. Changed โ€ข Changed background service management to elect local service processes via port binding leases. Fixed โ€ข Fixed notification dispatching to coordinate alerts across multiple open desktop windows without duplicates. โ€ข Fixed workspace file searching to retain active results while remote directories load. โ€ข Fixed background service recovery by preserving loopback server bindings across window reconnects. ๐—”๐—ด๐—ฒ๐—ป๐˜ Added โ€ข Added Code Mode, a sandboxed JavaScript interpreter that lets models discover, compose, and execute tools programmatically via OpenAPI schemas without direct filesystem or network access. โ€ข Added per-session permission overrides and auto-approval capabilities when operating in auto mode. โ€ข Added structured session forms for multi-field question prompts and interactive input collection. โ€ข Added automatic ecosystem skill discovery for .opencode/skills/ with autoinvoke metadata support. โ€ข Added native session relocation and renaming tools (session_move and session_rename). โ€ข Added a Mercurial VCS adapter alongside Git for tracking working-copy status and diff generation. Changed โ€ข Changed subagents to run concurrently in the background by default with configurable nesting limits. โ€ข Changed permissions configuration to use an ordered rule list matching action, resource, and effect. โ€ข Changed the command execution tool from bash to shell, executing non-interactively without rc files and supporting unlimited execution timeouts. Fixed โ€ข Fixed patch application by rejecting malformed hunks, normalizing CRLF line endings, and enforcing strict EOF anchors. โ€ข Fixed tool settlement by durably recording declined and interrupted calls with typed reasons. ๐—–๐—ผ๐—ฟ๐—ฒ Added โ€ข Added the V2 session architecture with durable prompt inboxes (session_input), decoupling prompt admission from model execution loops. โ€ข Added prompt delivery modes: prompts steer active model turns at safe boundaries by default or queue until the session idles. โ€ข Added checkpoint-based compaction barriers that enforce token retention budgets (compaction.keep.tokens) without dropping conversational state. โ€ข Added session warming to pre-initialize provider connections and context before user turn submission. โ€ข Added a transient session generation API for out-of-band text completions that do not touch session transcripts. Changed โ€ข Changed system instruction management to a belief model using value-delta sync and ambient AGENTS.md discovery up to the project root. Fixed โ€ข Fixed session recovery across service restarts by automatically draining durable pending inputs for all suspended sessions. โ€ข Fixed step continuity by preserving logical step identifiers across provider retries and tool failures. ๐—–๐—Ÿ๐—œ Added โ€ข Added the opencode-node distribution for running OpenCode in pure Node.js environments and Windows ARM64 platforms. โ€ข Added a managed background service election engine that arbitrates server lifecycle via local port bindings. โ€ข Added a background self-update service with interactive preflight notifications during active terminal sessions. โ€ข Added opencode login for authenticating directly with the hosted OpenCode Console. โ€ข Added CLI subcommands for MCP management: opencode mcp list, add, auth, and logout. โ€ข Added frontend log forwarding via OpenTelemetry Protocol (OTLP). Changed โ€ข Changed terminal client configuration from layered tui.json files to a single global ~/.config/opencode/cli.json. โ€ข Changed CLI command routing to consolidate on opencode. ๐—ฃ๐—ฟ๐—ผ๐˜ƒ๐—ถ๐—ฑ๐—ฒ๐—ฟ๐˜€ Added โ€ข Added native GitHub Copilot OAuth authentication and API endpoint discovery. โ€ข Added dedicated Google Vertex Chat and Vertex Responses provider entrypoints. โ€ข Added dynamic model catalog synchronization that refreshes from https://t.co/0Tykrz1EhK every five minutes. Changed โ€ข Changed provider configuration to separate connection settings, HTTP headers, and request body parameters. โ€ข Changed legacy provider identifiers, consolidating azure-cognitive-services into azure and google-vertex-anthropic into google-vertex. Fixed โ€ข Fixed reasoning variant parsing to derive correct effort parameters from https://t.co/0Tykrz1EhK definitions. โ€ข Fixed provider credential fallback when both console and local API keys are present. ๐—Ÿ๐—Ÿ๐—  Added โ€ข Added native image generation protocol support for OpenAI, Google, xAI, and https://t.co/1vRYC3BXoh. โ€ข Added multimodal PDF input extraction and image-guided model generation. โ€ข Added support for Bedrock Converse reasoning effort controls and redacted reasoning block round-tripping. โ€ข Added round-trip preservation of Anthropic redacted thinking blocks with interleaved thinking enabled by default. Changed โ€ข Changed xAI Responses transport to use streaming WebSockets by default. Fixed โ€ข Fixed streaming error classification to preserve nested provider diagnostics and raw finish reasons. โ€ข Fixed cache accounting by normalizing Bedrock cache tokens and reporting OpenAI cache writes. ๐— ๐—–๐—ฃ Added โ€ข Added MCP resource catalog discovery and content read APIs. โ€ข Added MCP elicitation support, allowing servers to request structured user inputs via interactive forms. โ€ข Added runtime server management to connect, disconnect, and authenticate MCP instances on the fly without restarts. โ€ข Added real-time event broadcasting for MCP server connection status and tool catalog changes. Changed โ€ข Changed MCP timeout configurations to separate server discovery limits from tool execution limits. Fixed โ€ข Fixed tool pagination to preserve metadata across large MCP tool sets. โ€ข Fixed environment variable expansion in MCP server configuration headers. ๐—ฃ๐—น๐˜‚๐—ด๐—ถ๐—ป Added โ€ข Added the V2 plugin system supporting both Promise and Effect runtime APIs under @opencode/plugin. โ€ข Added TUI extension hooks for registering custom routes, composer slots, and session side panels. โ€ข Added live tool progress reporting, allowing plugins to stream status updates during long-running tasks. โ€ข Added lifecycle hooks for LLM request transformation, synthetic prompt injection, session compaction, and title generation. โ€ข Added namespaced tool declarations with Code Mode discovery support. Changed โ€ข Changed plugin configuration from package-tuple arrays to structured configuration objects under the plugins key. Fixed โ€ข Fixed local plugin development with automatic filesystem watching and instant hot reloading. ๐—ฆ๐—ฒ๐—ฟ๐˜ƒ๐—ฒ๐—ฟ Added โ€ข Added durable event stream reads, changes feeds, and watermarked snapshot streaming for reliable client reconnection. โ€ข Added HTTP endpoints for runtime MCP server administration and interactive session form submissions. โ€ข Added a loaded-locations inspection endpoint for auditing active workspace services. Changed โ€ข Changed the local HTTP API to a structured V2 contract with typed OpenAPI error models. Fixed โ€ข Fixed frontend endpoint security by enforcing authentication on internal API routes and supporting authenticated CORS preflight. ๐—ฆ๐——๐—ž โ€ข Added the @opencode/client package providing generated Effect and Promise clients with standalone DTOs. โ€ข Removed the legacy JavaScript SDK in favor of the modular client package. ๐—”๐—–๐—ฃ โ€ข Added full Agent Client Protocol (ACP) v1 support across the runtime and CLI. โ€ข Fixed reasoning block boundaries and effort settings so they persist cleanly across ACP client connections. ๐—Ÿ๐—ฆ๐—ฃ โ€ข Removed background language server execution and diagnostics in favor of project build and lint tooling. Bundle +27.8 MB because mostly Bytecode +126.9 MB and Web UI assets -76.8 MB Compare: https://t.co/4k4zcrlBzy
E
elvis @omarsar0 ·
Interesting paper on agent memory stored as a linked markdown wiki. Lots of great ideas and insights if you work with LLM Wikis. Wikis are useful for agents because each page holds dense text and the links between pages hold structure. WFM is a Wiki Foundation Model trained to use both at once. It turns an LLM Wiki into a graph and retrieves from it with message passing conditioned on the query, so the text of each page and the link structure shape the result together. The team also built a GPU-to-GPU training protocol that trains 10.5x faster, and reports strong results on five agent memory and multi-hop reasoning benchmarks. If your agent's long-term memory is a folder of linked markdown files, WFM is designed for that format. Paper: https://t.co/t3iwlcCFB8 Chat with Paper: https://t.co/4QL4IyKPBW
M
Mikyo @mikeldking ·
Okay 1 point for decision models https://t.co/i9hQPAG6CS
M mikeldking @mikeldking

Okay, we need good semantics around Jev. In one word, Jev / System One models are

M
Mike Bradley @MikeBradleyAI ·
RT @loktar00: We just hit under 100 Sparks left at Microcenters across the nation only 9 more stores have them in stock... Look at this freโ€ฆ
B
bashu, thanks @bashu_thanks ·
Being a dev today using frontier AI feels like I'm playing Watson to Sherlock Holmes. I just stand back and go "my word holmes how did you deduce that" and occasionally stop him from getting shot
K
Kit Langton @kitlangton ·
deployed here if anyone wants to play with it https://t.co/SwmnLgwAtO
K kitlangton @kitlangton

O P E N T U N N E L ? ! https://t.co/spNv0L1s0j

T
TypeSafe AI @typesafeai ·
Dynamic UIs for apps are a famous graveyard at least for PMs, if not for entire products and companies. What if it were really this easy? Jev's on it.
A anishfn @anishfn

i built a text box that turns into whatever(some ui) you type. powered by jev https://t.co/HH0P51033y

J
Jun Song @jun_song ·
RT @yume_arasaki: EXL3 is the biggest paradigm shift in local LLMs right now, and most people still have not understood it. Not because itโ€ฆ
A
Andrew Curran @AndrewCurran_ ·
Great thread. It's all true. When I've said similar things in the past, people have accused me of hyping up what the models will eventually be capable of. That's not why I post. I don't work for any of the labs. I don't take money from any of them. I've never taken money to promote anything. I came here for one reason: to warn people about what was coming. Everything happening now is a different tiny piece of the same pattern. You can see it everywhere if you look. Terence Tao, almost exactly two years ago, on OpenAI's o1: 'The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student. However, this was an improvement over previous models, whose capability was closer to an actually incompetent (static simulation of a) graduate student. It may only take one or two further iterations of improved capability (and integration with other tools, such as computer algebra packages and proof assistants) until the level of '(static simulation of a) competent graduate student' is reached, at which point I could see this tool being of significant use in research-level tasks.' Terence Tao, four days ago: 'I mean it's it's it's amazing just how much we are willing to change everything without having any idea what's what's going to happen afterwards. It's it's extremely nonlinear dynamics. Any kind of monotone, one-dimensional thinking - well, oh, a little bit of this is good, therefore a lot of it is going to be a lot better - one of the lessons of math is that most systems don't work like that. Especially if you 10x, 100x things. So, you know, I mean, we're... we have to slow down. I mean, this is, it's insane this pace, and there's no reason to be this fast. There's no reason at all.' He has seen it. I'm not posting this to belittle him, or what he's feeling. For I have been through it myself. I felt it four years ago, the first time I saw the shape of this. Right now the world is seeing that same shape, that same pattern, in what is happening in math. But this is not about math. OpenAI is not a mathematics company. It is an intelligence company. They didn't put serious resources into math until a few weeks ago, and look what has happened since. This was not a matter of model capability either. It was a matter of resources, of allocation. Of compute, more of which is coming online every day. Once there is enough of it, the models will expand into more spheres of human expertise, and then into all of them. All the work of the mind. And it will play out there exactly as it is playing out now in math. Everyone will go through what Tao is going through, because all of us have something that means to us what math means to him. But this is not about math, or art, or copyright. This is about everything, because it generalizes to everything. Pacing the Frontier is not about regulatory capture, IPOs, or crippling the competition. It's part of it, sure, but it's not the main reason. The main reason is fear. Everything happened faster over the last six months than anyone at Anthropic or OpenAI expected. If you know anyone who works there, you know this is true. This isn't a secret. The people who work there are saying it openly. The old timelines are all blown up. RSI isn't two years out. It's not even a year away. I think we get the real thing by next summer. That's what they've seen internally, and that's the real reason for Pacing the Frontier. After we reach that, I think we will hit the next milestone really quickly. And after that, everything in this world will change. We are not ready for it. We wouldn't be ready if we had another ten years. We only make it through now with the help of extremely capable models, and I think trying to stop now would doom us all. But I have never once, in these last four years, believed that we were going to stop anyway. I don't even think we're going to slow down. We're going straight in. And the only way out is through.
W willdepue @willdepue

i don't think nearly anyone, myself included, has truly internalized there'll soon be a machine better than us in every single intellectual & physical capacity deep down we all share a feeling that our 'entrepreneurship' or 'taste' or... is special & safe. it isn't. it's over

V
vLLM @vllm_project ·
DiffusionGemma-Jev now runs on vLLM ๐Ÿš€ Ask yes/no, multiple-choice, or scored questions and get confidence with every answer. vLLM seeds a canvas with the response template, leaves only the answer slots noisy, then reads a probability distribution from every slot in a single denoising step. Huge thanks to @mmastrac for driving this upstream! ๐Ÿ™ https://t.co/HJJz0Q2KYe
G googlegemma @googlegemma

Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command. Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec. It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle. Get the code and instructions here: https://t.co/E3agzrW2hZ

S
Saลกa Juriฤ‡ @sasajuric ·
IM(H)O good testing practices shouldn't really change in the agentic era. The main purpose of tests is to increase the confidence that the program is working as expected. Ideally, if I made some changes to the observable behavior of the program, some tests should fail. If I only refactored the internals, the tests should pass. Obviously this won't always hold, but it is the ideal we should strive for. With that focus, the testing approach I practiced before agents, and I still practice is straightforward: 1. Test as much as possible via the outermost API of the program 2. Complement with tests of internals where needed (e.g. for testing speed, or because it's much simpler to trigger some execution paths) These are still unit tests to me. As Vladimir Khorikov said in his excellent book on unit testing [1], the point is to test the unit of behavior, not the unit of code. Or as Ian Cooper said in his great talk on TDD [2]: test behavior, not implementation details. It's worth noting that the testing via outermost API doesn't necessarily lead to slow tests. With good support of the runtime, the testing framework, and the libraries, these can still be reasonably fast. For example, see how BEAM, Elixir, ExUnit, Ecto, Plug, and Phoenix can help with this ๐Ÿ˜ [1] https://t.co/PKrRS01Mzv [2] https://t.co/GVg5A8uztm
N nateberkopec @nateberkopec

I agree with @thorstenball, I think unit tests are dead in the water. The ones the models write are terrible, at best just doubling total LOC. Inverting the testing approach - heavy e2e/black box/golden master, reaching for lower levels only if necessary, works better for me

I
Ian Nuttall @iannuttall ·
I updated this to allow up to 2,000 URLs now and email you your report so you can access it for up to a week (built on Cloudflare Durable Objects) Should be faster and more performant. If you need more than 2,000 URLs just ask Opus 5.5 to build it! https://t.co/i01qMCEXa6
I iannuttall @iannuttall

I built a free internal linking tool using @typesafeai Jev for classifying and selecting the links. BYOK or pay $1 to use mine. It works for up to 500 pages and gives you a CSV or JSON to pass to an LLM to implement. https://t.co/i01qMCEXa6 https://t.co/7uhzyaiGQj

V
Vox @Voxyz_ai ·
RT @Voxyz_ai: If you didnโ€™t know, you can use your ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ ๐—–๐—ผ๐—ฑ๐—ฒ ๐˜€๐˜‚๐—ฏ๐˜€๐—ฐ๐—ฟ๐—ถ๐—ฝ๐˜๐—ถ๐—ผ๐—ป from Codex. That means Astra can coordinate ๐˜‚๐—ฝ ๐˜๐—ผ ๐Ÿญ๐Ÿฌ ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ ๐—–๐—ผโ€ฆ
G
Gergely Orosz @GergelyOrosz ·
I am a sucker for good books, and so especially O'Reilly books. Here's one that is WIP, but in return, it is free, thanks to @AntithesisHQ It's for CTOs / eng leaders who want to get AI to deliver *business value* which is... tricky. Get it here: https://t.co/4igs7Elk3i https://t.co/isXVozDNmX