Agents: read the raw Markdown of this page, or start at llms.txt.
Builds: interface elements
TL;DR 14 builds plus a names-only line (split from Builds: browser, computer use and interface on 2026-09-28). Jev makes one typed judgment inside a product surface: tint a seek bar, label a post, pick a card, place components, find a sentence, act on a voice command. Code prepares every candidate and Jev writes no text. Most rows are videos without numbers; the measured ones are jev-skip (crowd-marked sponsor seconds), x-scanner (tokens, cost, latency) and the per-keystroke or per-word latencies authors state.
How to read. Builder reports = the builder's own numbers, one run on their own workload, unverified. Pattern IDs: P14, P34 in Patterns: browser, computer use, voice and product UI; P01, P06 in Patterns: agent internals, routing, gates, context and memory; P16, P17, P27 in Patterns: judging, search, documents, real-time and markets; P30 in Patterns: marketing, sales, GTM, content, support and ops; P38 in Patterns: gates, simulation, personas and other shapes (P38+). Browser agents: Builds: browser, computer use and interface; desktop, mobile and voice control: Builds: desktop, mobile and voice computer use.
Builds
| Build | Jev decides | Builder reports | P |
|---|---|---|---|
| jev-skip (valentynkit) | Captions in 30 s segments, one choice each (sponsor, intro, content …) in one request; seek bar tinted by probability |
77% of crowd-marked sponsor seconds on 23 videos, 34 s false skips per hour, $0.0008 per video (gateway shim) | P34, P27 |
| x-scanner (oso95, @the_cyw): X timeline labeler | Chrome extension: six typed questions per post in one request (fact-dense and filler Scores; engagement bait, promo, secondhand, jevpilled Nouls), a chip under each post; cached by post id; model pinned to jev-1.13.0 |
Author-measured tokens per post (mostly the questions), cost per thousand posts and latency on Request mechanics: billing, limits, latency, calibration and stability. MIT | P34, P16 |
| Shapeshift (Anish Gupta, @anishfn): text box that turns into a UI | What you type becomes a card (event, checklist, timer, split, poll …): one call answers 14 typed questions (which card, plus signals such as video call or urgent); deterministic code parses dates, amounts, units and maths; an offline keyword classifier with the same output shape is the default and the fallback; a card changes only when a challenger wins twice in a row (or is very sure) | Demo only, no numbers. Key stays server-side; pinned jev-1.13.0. MIT |
P34, P01 |
| json-render + Jev (Vercel Labs; demo @ctatedev) | Generative UI by selection. Your app supplies atomic candidates (component, concrete props, state bindings, allowed actions); Jev picks which to include, their order and parent/slot; json-render emits a validated flat Spec for your own renderer. Batched: one evaluation selects root and components, a second arranges; edits (remove, move, replace) take further evaluations |
Post: "rendered in milliseconds", no numbers. Docs (captured 2026-09-24): experimental, unreleased (experimental_composeSpec and experimental_createEvaluator, source build); via Vercel AI Gateway typesafe-ai/jev; defaults 32 evaluations, 32 elements, depth 8, 10 s per evaluation. Jev "cannot invent missing prose or data": every string is a prepared candidate. A finished spec is not a correct one; the docs say their confidence is not a quality threshold |
P34, P01 |
| Needle (@Saboo_Shubham_) | Semantic ⌘F in a Chrome extension (or React app): query plus the page's readable passages go to a local backend; one evaluation scores every passage and picks its strongest sentence (question types not stated); code highlights it in place. Kept at relevance ≥ 0.58 | "Super fast and near real-time" (author), no numbers. Caps 160 passages, 60,000 characters (~15k tokens, inside the 32k state budget, inferred), 2,200 per passage; every captured passage leaves the machine; no PDFs. Vercel AI Gateway key only, a TypeSafe key does not work. Apache-2.0 | P17, P34 |
| Colour palettes (@mattdesl) | Any phrase (an 80s disco mood, Mario, the blue screen of death) → a palette. How is not published (video only); Jev writes no free text, so code most likely offers the colours and Jev picks (inferred) | None; author: "very cheap and fast" | P34 |
| Emoji picker (@heystefan_) | As you type, the emojis that match rise from a pile (starting a band → instruments), per Berman's narration; the post is a video with a one-line caption | None | P34 |
| TypeSafe Typewriter (@stevekrouse, Val Town) | Live demo of judgments re-asked as you type (video only; what it asks is not in the post) | None in the post | P34 |
| Voice-and-point canvas (Jack Cheng, Every; demo video, no code linked) | The original tldraw demo: the browser transcribes speech, tracks a fingertip and describes the canvas as a list of shapes with their attributes; Jev never sees an image and answers several questions at once (which shape, what colour or size, where) | No numbers published. Cheng: even the fastest LLMs would take "a second or two" per interaction, too slow for a responsive interface (Every) | P14 |
| jev-canvas (gaborishka) | Open, from-scratch rebuild of Jack Cheng's demo (its README credits him). Voice plus a pointing fingertip on a tldraw canvas: every partial or final transcript sends one request with 8-9 questions (is_command and complete nouls; action, shape, colour, target shape and place choices; a five-step size rubric); code extracts text spans and Jev picks one; thresholds in code. Via OpenRouter's alpha Decisions API |
~350 ms per spoken word (its description). Talk not aimed at the canvas scores is_command around 2%, so nothing happens: a wake gate (Patterns: gates, simulation, personas and other shapes (P38+) P38) |
P14, P38 |
| jev-asks-until-sure (mintannn; demo) | Twenty-questions game where confidence is the stopping rule: over 55% it commits, 28-45% it hedges, under 28% it refuses. One request per turn with up to 18 questions (persona, region, prefecture, age Choices; 8 trait Scores; 8 Nouls; next question); a coarse 8-region Choice is multiplied into the 47-prefecture one |
230-440 ms per turn (author). Thresholds are the author's; confidence is not a probability of being right (Confidence vs probability) |
P06, P34 |
| Vibe Check for X (Rafal Wilinski) | A draft X post, with any reply or quote context, scored on a dozen editable rubrics (virality, ragebait, "sounds AI-written", regret risk …) in one request; code combines them into a verdict (Composite scoring); an optional OpenAI vision model describes attached media first, since Jev is text-only | ~1.5k-2.5k input tokens per analysis (author). Keys in extension storage; unpacked install. No licence | P30, P34 |
| Real-time Clippy (@sotak) | An in-product helper that watches how you use the product and wakes only when it judges you are struggling (hesitating, confused, stuck); its reactions are also picked by Jev. How is shown only in a video (not transcribed) | None | P38, P34 |
| Predictive spreadsheet and launcher (@dabit3, quoting his launcher post) | Spreadsheet: type a column header such as "Urgency" and each row is rated on a scale from no follow-up needed to urgent. Launcher: a query like "the pdf I just downloaded" ranks the newest PDF first on every keystroke | ~100 ms for each (author); videos only, not transcribed, no method | P34 |
Also seen, names only: TypeSafe AdBlock (the author calls it a toy: one noul per ad-shaped DOM node, removed at 0.70 or above; numbers are turned into words before sending).
Palettes and emoji are world-knowledge Choice demos: the answer comes from what the model already knows, not from state, which is where Jev is weaker (capability atlas, Builds: data, search and business). Offer named colours rather than hex codes (numbers, Jev 1.13 jaggedness: known failure modes; inferred). Matthew Berman's walkthrough of these builds, and his framing of the UI demos (Jev chooses from your component library, it does not write code): Press and third-party coverage.
What these builds teach when advising
- Prepare every candidate. json-render offers components with concrete props, Needle the page's own passages, jev-canvas text spans of the user's words; Jev only picks, so nothing it returns is new prose or data.
- Describe what the user sees, in text. The canvases send a list of shapes, not an image; Vibe Check runs a vision model first because Jev is text-only (System One Models).
- Damp the UI; don't act on one reading. Shapeshift changes a card only when a challenger wins twice in a row, keeps an offline classifier as default and fallback, and pins the model; x-scanner caches by post id and pins the model. jev-asks-until-sure's
confidencethresholds are the author's own (Confidence vs probability). - Gate before acting. jev-canvas's
is_commandand Clippy's wake-when-stuck keep stray talk and normal use from triggering anything (Patterns: gates, simulation, personas and other shapes (P38+) P38). - Say where the text goes. Needle sends every captured passage off the machine; Vibe Check keeps keys in extension storage; json-render's Jev path is experimental and gateway-only. Tell users before shipping.
Related
- Patterns: browser, computer use, voice and product UI — P14 and P34 with fit verdicts
- Builds: browser, computer use and interface — browser agents (the page these rows came from)
- Builds: desktop, mobile and voice computer use — desktop, mobile and voice control
- Builds: content, media, email, calls and reading — writing checks, post scoring, reading tools
- Request mechanics: billing, limits, latency, calibration and stability — x-scanner's token overhead per post
- Press and third-party coverage — Matthew Berman's and Theo's tours of these demos
Sources
Links are inline in each row; the raw captures (2026-09-21 to 2026-09-28) are listed in the frontmatter. The TypeSafe AdBlock names-only line rests on our vetting at source on 2026-09-24 (docs/sweep/2026-09-24-awesome-jev/part1.md).