Agents: read the raw Markdown of this page, or start at llms.txt.
Rival decision models: Clef, Liquid d1, GLiDE, Drex and other non-Jev System One models
TL;DR None of these is Jev: each is another vendor's or author's own decision model, most taking a Jev-shaped request (
stateplus typednoul/choice/scorequestions). Every number below is the vendor's own run,unverified; four vendors claim the top of the Decision Index from their own runs. Pick one for a gap Jev has (images: Clef, pplx-decider; very long input: pplx-decider; local or open weights: Clef, Strands Decider, pplx-decider; harder reasoning at a latency cost: GLiDE), then re-measure on your own labels: Jev-tuned thresholds do not transfer. Independent comparisons live on Head-to-head: Jev vs small, fast and local models; open replicas of Jev on Open replicas and Jev-compatible servers.
Scope (split from Open replicas and Jev-compatible servers 2026-10-04): models launched as alternatives to Jev under their own name, hosted or open. Not here: replicas and servers built to copy Jev's interface over open models (Open replicas and Jev-compatible servers, Open replicas: servers and logit readers over other models); routes that serve the real Jev (Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev); resellers and look-alike sites (Warnings: not-Jev services, key safety, look-alikes and install names). Numbers are each vendor's unless marked; "Jev" in their tables is their reading of Jev, often copied from a published record rather than run live.
Hosted rivals from vendors
Vendor claims, unverified. Prices are per million input tokens; Jev's list price is $0.042, output free (Models, aliases, pricing, rate limits, context).
| Model (vendor) | Access | Their numbers | Caveats |
|---|---|---|---|
| Clef and Clef-flash (Cloudflare, 2026-10-01; blog, post) | Workers AI @cf/cloudflare/clef, model clef or clef-flash, Jev-shaped body plus an images extension (up to 4); 1 to 64 questions; 65,536-token context; vision; $0.24 (Clef) and $0.09 (Clef-flash) per M input. Weights on Hugging Face, Apache-2.0: Clef 27B (frozen Qwen3.8-27B), Clef-flash 9B (Qwen3.5-9B), each with a joint schema head and rank-256 LoRA, trained with a Brier term and its own "RLCD". Cloudflare says it does not read, store or train on requests; fine-tuning offered through its engineers |
Cloudflare's own Decision Index 0.2.1 run ("currently the leader"): median latency 209.3 ms (Clef) and 38.8 ms (Clef-flash) vs Jev 524.1 ms, p95 238.6 / 122.4 vs 536.0 ms; ahead of Jev on BANKING77 (94.2 vs 79.7 macro-F1) and CLINC150+OOS (97.4 vs 89.3); Jev ahead on GPQA Diamond (78.3 vs 48.0), MMLU-Pro (82.7 vs 65.9), BBH (92.9 vs 73.7) and When2Call (81.0 vs 72.4). TypeSafe's four workflow evals: Clef ahead on three (invoice 64.7 vs 61.8), Jev ahead on agent-trace observability (71.6 vs 68.5) | Cloudflare Workers AI also serves the real Jev (typesafe/jev, Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev): the model ID decides which you get. Clef is 5.7× Jev's list price per input token, Clef-flash 2.1× (our arithmetic). The blog's "Jev's 32k" context contradicts docs (64k per request; 32k is state plus the longest question). "Fully API-compatible": unverified; the Workers AI schema requires model to be clef or clef-flash, and long text state is truncated. Builders' own runs (some slower and pricier than Jev): Head-to-head: Jev vs small, fast and local models |
| GLiDE (Fastino Labs, 2026-10-01; blog) | POST https://api.fastino.ai/v1/systemone, model fastino/glide, Fastino key; "conforms to the System One model schema"; 40k-token context, oversized requests rejected, thinking tokens share the window. Price: the post points to Fastino's docs (not captured) |
Fastino's own full run with the official Decision Index 0.2.1 scorer: 64.81 vs Jev's published 57.91; ahead in all five areas and 31 of 38 benchmarks (Knowledge and Reasoning 62.9 vs 51.4; Tools and Automation 83.5 vs 75.1; CLadder 88.7% vs 72.6%; CRUXEval 92.6% vs 73.0%) | A "thinking" decision model: it reasons when its first answer is uncertain, so latency varies with difficulty; not System One speed on hard items (builders' runs: Head-to-head: Jev vs small, fast and local models). Not on the public board (0.2.1 submissions paused, per Fastino). Fastino's earlier open-weight GLiNER2.5-Decide (launch post linked from the GLiDE post, not captured) appears in one builder's local bake-off (Open replicas and Jev-compatible servers, trust notes) |
| pplx-decider-v1-27b (Perplexity Decisions API; docs, card) | POST https://api.perplexity.ai/v1/decisions (not /v1/systemone), Perplexity key; state as text, JSON or images (base64 data URLs); choice 1 to 255 options, score up to 10 levels, a noul with neither instructions nor criteria returns 400; input under 262,144 tokens; 10 requests/second per organisation; $0.04 per M input, output free. Weights on Hugging Face (fine-tuned Qwen3.8-27B, ~49 GiB; licence not shown in the captured card) |
Card, 11 benchmarks, its side run through Perplexity's API: overall 85.71% vs Jev 84.51% (base Qwen3.8-27B 74.76%); Jev ahead on 6 of 11 (BBH 94.27% vs 82.80%, WinoGrande 90.70% vs 83.30%, JevBench public hard 73.27% vs 70.30%), pplx-decider ahead on RAGTruth (88.80% vs 77.27%) and FinancialPhraseBank (84.18% vs 76.98%). Docs (2026-09-30 tests): under 2 s for a few hundred input tokens, 23 s near the input limit | A different path, so the TypeSafe SDKs are not a drop-in (inferred: they post to /v1/systemone). Rate limit 10/s vs Jev's 80/s. Builders' cost and accuracy runs: Head-to-head: Jev vs small, fast and local models |
| Instinct (ZooWork / Serendipity One Inc.; @ninghu, blog, repo) | Hosted API at instinct.zoowork.ai, free preview (billing not enabled as of 2026-09-29), listed $0.03 (instinct, 27B) and $0.01 (instinct-dual-4b, instinct-tuned-4b); Apache-2.0 runtime |
@ninghu reports 87.45% vs Jev's 86.58%, 5.5× faster, up to 76% cheaper. ZooWork's blog: 202 of 231 JevBench public items vs Jev's 200 (Jev's side re-aggregated from the board's v1.3.0 per-item records, not run live; latencies from different clients and times), p50 120.5 vs 665.0 ms; "76% cheaper" is list price per input token. ZooWork's reading of JevBench v1.5.4 (1,624 decisions, 720 sealed): Jev leads on quality (Intelligence 72.00 vs 62.75 for the 27B and 43.12 for Dual 4B) | "All weights open": two of the three run unchanged Qwen weights (Qwen3.8-27B; Qwen3.5-4B in two option orders) with a logit readout; only Tuned 4B (HF srpone/instinct-tuned-4b) is ZooWork's own checkpoint (repo). ZooWork itself says the 2-item gap shows no lead. Board rows: Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites |
| Canonopy Labs (site) | POST /v1/build with the JSON you send Jev (plus rules in plain words) builds a per-decision model under 1 MB (text models share a 34 MB reader); then the same JSON to api.canonopylabs.com with its model name; answers name the rules applied. 5-day trial, then $20 a month flat |
Doom: 45 kills to 39 and 2 deaths to 5 over 13 games against recorded live-Jev games at the same decision rate (45 to 21 against Jev's aggressive style); Snake: 42.0 vs 40.0 food over 20 games; 3,080 real bank messages, nine next steps, five hard rules: 85.5% on day one, 91.7% after about 400 reviewed cases vs Jev 86.6%, about 95% with the bank's history | One trained model per decision site, built from your example and rules, not a general model; the text gains need your reviewed cases. Jev's game side is recorded play, not a live rerun |
| Model (vendor) | Access | Their numbers | Caveats |
| --- | --- | --- | --- |
| d1 (Liquid AI, 2026-09-29; launch, digest by @som_dutt_) | Liquid API d1:free at https://api.liquid.ai/decisions/v1/systemone; Liquid's docs call it with TypeSafe's own Python and JS SDKs via base_url (Noul, Choice, Score; confidence on Choice and Score; output_tokens 0). Vercel AI Gateway liquid/d1, OpenRouter "soon" (per @som_dutt_) |
Decision Index 0.2.1, Liquid's own reproduction: 58.9 vs Jev 1.13 57.9; Arts +7.8, Language +5.6, Retrieval +5.3, Tools −1.0, Knowledge −8.0 (@som_dutt_ reading Liquid's chart). Multilingual, injection and long-input wins claimed with no numbers | API only; paid pricing not published (per @som_dutt_); one integrator reports HTTP 422 on a Choice with fewer than 2 options (single report). An official TypeSafe SDK pointed at another vendor's base_url sends that vendor your requests and key (key safety: Warnings: not-Jev services, key safety, look-alikes and install names) |
| Drex 1.5 (Nace.AI) | Nace-managed, your cloud or on-premises; Nace tunes it on your labels and ships you the weights | Decision Index 0.2.1 rank 1 of 71: 58.28 vs Jev 1.13.0 57.91 (the board treats gaps under 0.25 as ties), official kit on all 38 tests, 2026-09-28, ahead on 21. Its own JevBench run puts Jev ahead: 231 public items, Jev 87.0% (201) vs Drex 1.5 86.2% (199). Board games vs Jev 122-87 with 47 draws (256 matches, Kaggle Game Arena harness, no lookahead). Under 10B parameters, 128K context | Jev leads on knowledge in Nace's own table (GPQA Diamond 78.6 vs 45.4, MMLU-Pro 82.7 vs 58.7, HLE 20.4 vs 11.8); Drex leads chord recognition (POP909-CL 73.1 vs 16.6). Vendor page |
Decision Index claims. d1, Drex 1.5, Clef and GLiDE each claim the top of Decision Index 0.2.1 from the vendor's own run; we did not check the board, and GLiDE says it is not on it. Board: multimodalart/jev-decision-index ("Which is which" on Open replicas and Jev-compatible servers).
Open models from vendors
| Model (vendor) | What | Their numbers | Caveats |
|---|---|---|---|
| Strands Decider 2B (AWS Strands Labs; card) | LoRA on Qwen3.5-2B-Base plus a small readout head; pip install strands-decider, a CLI and an HTTP /v1/systemone server; full recipe, data inventory and evaluations published (training/recipe.sh; about 11 h to retrain on one RTX 3090) |
JevBench public 167/231, Brier 0.348, ECE 0.050; internal held-out short tasks 0.641 (n = 6,000) | Its server binds to 127.0.0.1 with no authentication. Its own limits: questions are read less than documents (a reworded question often gets the same answer); long multi-step documents are the weak spot; score and noul transfer poorly to unfamiliar rubrics; one calibration temperature per primitive. Trained partly on a frozen Qwen3.5-4B teacher's output distributions |
| Clef, Clef-flash, pplx-decider-v1-27b | open weights of the hosted models above | rows above | — |
Small open alternatives (individual authors)
Closer to the replicas on Open replicas and Jev-compatible servers; listed here because their authors position them as alternatives to Jev rather than copies.
| Project (author) | What | ★ · push · licence | Their numbers and caveats |
|---|---|---|---|
| JevAlt (mertkayacs): Deem-4B (English), Karar-4B (Turkish), Wähler-4B (German) | 4B models from internlm/Intern-Decision-4B (Qwen3.5-4B); jevalt serve on /v1/systemone, GGUF on a CPU in about 3 GB RAM; extras reasoning, abstain (an "unknown" option) and coverage |
4 · 2026-10-04 · Apache-2.0 | Author's live tests 2026-10-04, 130 known-answer requests per language: Wähler-4B 122, Deem-4B 122, Karar-4B 113; Wähler-4B matched only 14 of 20 long-policy cases. Its Jev rows come from TypeSafe's jaggedness notes and a third-party audit, not a live Jev run. It warns dates are shaky (Wähler-4B miscounted a return window across August and September, reasoning on) |
jevos (feder-cr; repo jev) |
Local yes/no decisions (choice and score "early") in one C++ binary on OpenVINO int8, CPU only; 8,192-token context; accepts any jev-* model name and answers as jevos-v4 |
1203 · 2026-10-04 · MIT (repo created 2024-08-15, before Jev) | Author, jevos-v2 on an Intel laptop: 2,000 yes/no questions from unseen policies 0.810 vs Jev 0.927 vs Laya 0.489; latency 26 vs Jev 344 ms (short), 112 vs 345 ms (long). Model-name key safety: Warnings: not-Jev services, key safety, look-alikes and install names (c) |
| GenClass (MeharPro) | 32M-parameter ModernBERT/Ettin encoder with typed heads and per-option attention isolation (option order cannot change the answer); runs in the browser (WebGPU/WASM) behind a voice-control Chrome extension; /v1/systemone server; calls itself "the version of Jev that runs in your browser" |
2 · 2026-10-04 · Apache-2.0 | Repo table: mid-sentence command recognition 90.4% vs Jev 66.4%; general held-out questions 80.5% vs Jev 94.7%; on-screen element 82.4% vs 86.6%; option-order flips 0 vs 10 to 13% (Jev's flips from a third party). Its harness carries 549 published Jev numbers, so which Jev figures were run live is unclear: unverified. Not affiliated with TypeSafe (README) |
Proposals and platform APIs
- Chrome "Decisions API" (Intent to Prototype, Mike Wasserman, blink-dev, 2026-10-01; explainer): a proposed on-device
window.DecisionModelthat scores a structured set of questions and options in one forward pass, with calibrated confidence, "tens or low hundreds of milliseconds" on laptop CPUs and GPUs. The explainer, by the Google Chrome Built-in AI Team, says it has not been approved to ship; it names Jev, Laya, Kev and Open-Jev as examples of the model class and names no model of its own. Rick Byers' reply asks whether it would rely on closed-source models and about a public eval suite. No numbers. - OpenAI Decisions API (limited preview, announced 2026-09-29): facts on When not to use Jev: rules, embeddings, trained classifiers, small and frontier LLMs; comparisons against GPT-6 Luna on Head-to-head: Jev against other models and methods.
- Ollama 0.35 serves Bespoke's nimble and Together AI's tev1 on a local
/v1/systemone: a local server, so its row is on Open replicas: servers and logit readers over other models.
Name collision. "Decisions API" names at least five different things: OpenAI's (GPT-6 Luna), Perplexity's (/v1/decisions, pplx-decider), Chrome's proposal, OpenRouter's alpha endpoint serving the real Jev (Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev), and the /v1/decisions paths of B.AI (a route) and TokenBazaar (Warnings: not-Jev services, key safety, look-alikes and install names). Check which one a post means.
No methodology, unverified: "100% accuracy, independently verified" for ConceptNet (conceptnet.co.uk), a one-founder rival, set against a "68%" Jev figure it does not source (@tonylab_net; the Medium write-up is members-only, fetched 2026-10-01, not captured; moved from Open replicas and Jev-compatible servers 2026-10-04).
Advising with rivals
- Vendor numbers are self-run. Every table above is the vendor's or author's own; several read Jev from a published record instead of calling it. Treat a gap of a few points as a tie, as ZooWork itself does.
- Thresholds do not transfer. Jev's
confidenceformulas are published (Confidence vs probability); none of the rival docs captured here says itsconfidenceuses the same formula. Re-fit every threshold on your own labels (Testing and evaluating a Jev workflow). - Key safety. Pointing an official TypeSafe SDK at another vendor's base URL sends that vendor your requests and the key in
TYPESAFE_API_KEY(Warnings: not-Jev services, key safety, look-alikes and install names (c), TYPESAFE_* environment variables across SDKs). - Where a rival fills a gap Jev has: images (Clef, pplx-decider; Jev is text only), inputs past Jev's 64k (pplx-decider, 262,144), open weights to run locally (Clef, Clef-flash, pplx-decider, Strands Decider, Instinct Tuned 4B), harder reasoning where latency matters less (GLiDE). Where Jev's own docs and the boards still favour Jev: knowledge and reasoning benchmarks in most vendors' own tables (GPQA, MMLU-Pro, BBH).
Related
- Head-to-head: Jev vs small, fast and local models — builders' own comparisons of Jev with Clef, Perplexity and small LLMs (independent of the vendors)
- Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites — JevBench by version, where some rivals are ranked
- Open replicas and Jev-compatible servers — open replicas of Jev's interface; Open replicas: servers and logit readers over other models — servers over other models, incl. Ollama 0.35
- Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev — routes to the real Jev, incl. Cloudflare Workers AI; Warnings: not-Jev services, key safety, look-alikes and install names — resellers, look-alikes, key safety
- When not to use Jev: rules, embeddings, trained classifiers, small and frontier LLMs — the OpenAI Decisions API and when not to use Jev at all
- Jev 1.13 jaggedness: known failure modes — Jev's documented failure modes, against which rival claims are worth reading
Sources
Links inline; raw captures in frontmatter (vendor pages, model cards and READMEs captured 2026-10-04 under raw/community/ and raw/x-repos/; the d1 and Drex rows and the ConceptNet note captured 2026-09-29 to 2026-10-01, moved from Open replicas and Jev-compatible servers unchanged). Not captured: Perplexity's and Fastino's launch posts on X, the Strands Decider code repository, OpenRouter's pplx-decider page, GLiNER2.5-Decide and press coverage.