$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Rival decision models: Clef, Liquid d1, GLiDE, Drex and other non-Jev System One models

[ community tier ][ updated 2026-10-04 ][ confidence medium ][ jev-1.13.0 ]#community · rivals · decision-models · alternatives · vendors

TL;DR None of these is Jev: each is another vendor's or author's own decision model, most taking a Jev-shaped request (state plus typed noul / choice / score questions). Every number below is the vendor's own run, unverified; four vendors claim the top of the Decision Index from their own runs. Pick one for a gap Jev has (images: Clef, pplx-decider; very long input: pplx-decider; local or open weights: Clef, Strands Decider, pplx-decider; harder reasoning at a latency cost: GLiDE), then re-measure on your own labels: Jev-tuned thresholds do not transfer. Independent comparisons live on Head-to-head: Jev vs small, fast and local models; open replicas of Jev on Open replicas and Jev-compatible servers.

Scope (split from Open replicas and Jev-compatible servers 2026-10-04): models launched as alternatives to Jev under their own name, hosted or open. Not here: replicas and servers built to copy Jev's interface over open models (Open replicas and Jev-compatible servers, Open replicas: servers and logit readers over other models); routes that serve the real Jev (Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev); resellers and look-alike sites (Warnings: not-Jev services, key safety, look-alikes and install names). Numbers are each vendor's unless marked; "Jev" in their tables is their reading of Jev, often copied from a published record rather than run live.

Hosted rivals from vendors

Vendor claims, unverified. Prices are per million input tokens; Jev's list price is $0.042, output free (Models, aliases, pricing, rate limits, context).

Model (vendor) Access Their numbers Caveats
Clef and Clef-flash (Cloudflare, 2026-10-01; blog, post) Workers AI @cf/cloudflare/clef, model clef or clef-flash, Jev-shaped body plus an images extension (up to 4); 1 to 64 questions; 65,536-token context; vision; $0.24 (Clef) and $0.09 (Clef-flash) per M input. Weights on Hugging Face, Apache-2.0: Clef 27B (frozen Qwen3.8-27B), Clef-flash 9B (Qwen3.5-9B), each with a joint schema head and rank-256 LoRA, trained with a Brier term and its own "RLCD". Cloudflare says it does not read, store or train on requests; fine-tuning offered through its engineers Cloudflare's own Decision Index 0.2.1 run ("currently the leader"): median latency 209.3 ms (Clef) and 38.8 ms (Clef-flash) vs Jev 524.1 ms, p95 238.6 / 122.4 vs 536.0 ms; ahead of Jev on BANKING77 (94.2 vs 79.7 macro-F1) and CLINC150+OOS (97.4 vs 89.3); Jev ahead on GPQA Diamond (78.3 vs 48.0), MMLU-Pro (82.7 vs 65.9), BBH (92.9 vs 73.7) and When2Call (81.0 vs 72.4). TypeSafe's four workflow evals: Clef ahead on three (invoice 64.7 vs 61.8), Jev ahead on agent-trace observability (71.6 vs 68.5) Cloudflare Workers AI also serves the real Jev (typesafe/jev, Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev): the model ID decides which you get. Clef is 5.7× Jev's list price per input token, Clef-flash 2.1× (our arithmetic). The blog's "Jev's 32k" context contradicts docs (64k per request; 32k is state plus the longest question). "Fully API-compatible": unverified; the Workers AI schema requires model to be clef or clef-flash, and long text state is truncated. Builders' own runs (some slower and pricier than Jev): Head-to-head: Jev vs small, fast and local models
GLiDE (Fastino Labs, 2026-10-01; blog) POST https://api.fastino.ai/v1/systemone, model fastino/glide, Fastino key; "conforms to the System One model schema"; 40k-token context, oversized requests rejected, thinking tokens share the window. Price: the post points to Fastino's docs (not captured) Fastino's own full run with the official Decision Index 0.2.1 scorer: 64.81 vs Jev's published 57.91; ahead in all five areas and 31 of 38 benchmarks (Knowledge and Reasoning 62.9 vs 51.4; Tools and Automation 83.5 vs 75.1; CLadder 88.7% vs 72.6%; CRUXEval 92.6% vs 73.0%) A "thinking" decision model: it reasons when its first answer is uncertain, so latency varies with difficulty; not System One speed on hard items (builders' runs: Head-to-head: Jev vs small, fast and local models). Not on the public board (0.2.1 submissions paused, per Fastino). Fastino's earlier open-weight GLiNER2.5-Decide (launch post linked from the GLiDE post, not captured) appears in one builder's local bake-off (Open replicas and Jev-compatible servers, trust notes)
pplx-decider-v1-27b (Perplexity Decisions API; docs, card) POST https://api.perplexity.ai/v1/decisions (not /v1/systemone), Perplexity key; state as text, JSON or images (base64 data URLs); choice 1 to 255 options, score up to 10 levels, a noul with neither instructions nor criteria returns 400; input under 262,144 tokens; 10 requests/second per organisation; $0.04 per M input, output free. Weights on Hugging Face (fine-tuned Qwen3.8-27B, ~49 GiB; licence not shown in the captured card) Card, 11 benchmarks, its side run through Perplexity's API: overall 85.71% vs Jev 84.51% (base Qwen3.8-27B 74.76%); Jev ahead on 6 of 11 (BBH 94.27% vs 82.80%, WinoGrande 90.70% vs 83.30%, JevBench public hard 73.27% vs 70.30%), pplx-decider ahead on RAGTruth (88.80% vs 77.27%) and FinancialPhraseBank (84.18% vs 76.98%). Docs (2026-09-30 tests): under 2 s for a few hundred input tokens, 23 s near the input limit A different path, so the TypeSafe SDKs are not a drop-in (inferred: they post to /v1/systemone). Rate limit 10/s vs Jev's 80/s. Builders' cost and accuracy runs: Head-to-head: Jev vs small, fast and local models
Instinct (ZooWork / Serendipity One Inc.; @ninghu, blog, repo) Hosted API at instinct.zoowork.ai, free preview (billing not enabled as of 2026-09-29), listed $0.03 (instinct, 27B) and $0.01 (instinct-dual-4b, instinct-tuned-4b); Apache-2.0 runtime @ninghu reports 87.45% vs Jev's 86.58%, 5.5× faster, up to 76% cheaper. ZooWork's blog: 202 of 231 JevBench public items vs Jev's 200 (Jev's side re-aggregated from the board's v1.3.0 per-item records, not run live; latencies from different clients and times), p50 120.5 vs 665.0 ms; "76% cheaper" is list price per input token. ZooWork's reading of JevBench v1.5.4 (1,624 decisions, 720 sealed): Jev leads on quality (Intelligence 72.00 vs 62.75 for the 27B and 43.12 for Dual 4B) "All weights open": two of the three run unchanged Qwen weights (Qwen3.8-27B; Qwen3.5-4B in two option orders) with a logit readout; only Tuned 4B (HF srpone/instinct-tuned-4b) is ZooWork's own checkpoint (repo). ZooWork itself says the 2-item gap shows no lead. Board rows: Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites
Canonopy Labs (site) POST /v1/build with the JSON you send Jev (plus rules in plain words) builds a per-decision model under 1 MB (text models share a 34 MB reader); then the same JSON to api.canonopylabs.com with its model name; answers name the rules applied. 5-day trial, then $20 a month flat Doom: 45 kills to 39 and 2 deaths to 5 over 13 games against recorded live-Jev games at the same decision rate (45 to 21 against Jev's aggressive style); Snake: 42.0 vs 40.0 food over 20 games; 3,080 real bank messages, nine next steps, five hard rules: 85.5% on day one, 91.7% after about 400 reviewed cases vs Jev 86.6%, about 95% with the bank's history One trained model per decision site, built from your example and rules, not a general model; the text gains need your reviewed cases. Jev's game side is recorded play, not a live rerun
Model (vendor) Access Their numbers Caveats
--- --- --- ---
d1 (Liquid AI, 2026-09-29; launch, digest by @som_dutt_) Liquid API d1:free at https://api.liquid.ai/decisions/v1/systemone; Liquid's docs call it with TypeSafe's own Python and JS SDKs via base_url (Noul, Choice, Score; confidence on Choice and Score; output_tokens 0). Vercel AI Gateway liquid/d1, OpenRouter "soon" (per @som_dutt_) Decision Index 0.2.1, Liquid's own reproduction: 58.9 vs Jev 1.13 57.9; Arts +7.8, Language +5.6, Retrieval +5.3, Tools −1.0, Knowledge −8.0 (@som_dutt_ reading Liquid's chart). Multilingual, injection and long-input wins claimed with no numbers API only; paid pricing not published (per @som_dutt_); one integrator reports HTTP 422 on a Choice with fewer than 2 options (single report). An official TypeSafe SDK pointed at another vendor's base_url sends that vendor your requests and key (key safety: Warnings: not-Jev services, key safety, look-alikes and install names)
Drex 1.5 (Nace.AI) Nace-managed, your cloud or on-premises; Nace tunes it on your labels and ships you the weights Decision Index 0.2.1 rank 1 of 71: 58.28 vs Jev 1.13.0 57.91 (the board treats gaps under 0.25 as ties), official kit on all 38 tests, 2026-09-28, ahead on 21. Its own JevBench run puts Jev ahead: 231 public items, Jev 87.0% (201) vs Drex 1.5 86.2% (199). Board games vs Jev 122-87 with 47 draws (256 matches, Kaggle Game Arena harness, no lookahead). Under 10B parameters, 128K context Jev leads on knowledge in Nace's own table (GPQA Diamond 78.6 vs 45.4, MMLU-Pro 82.7 vs 58.7, HLE 20.4 vs 11.8); Drex leads chord recognition (POP909-CL 73.1 vs 16.6). Vendor page

Decision Index claims. d1, Drex 1.5, Clef and GLiDE each claim the top of Decision Index 0.2.1 from the vendor's own run; we did not check the board, and GLiDE says it is not on it. Board: multimodalart/jev-decision-index ("Which is which" on Open replicas and Jev-compatible servers).

Open models from vendors

Model (vendor) What Their numbers Caveats
Strands Decider 2B (AWS Strands Labs; card) LoRA on Qwen3.5-2B-Base plus a small readout head; pip install strands-decider, a CLI and an HTTP /v1/systemone server; full recipe, data inventory and evaluations published (training/recipe.sh; about 11 h to retrain on one RTX 3090) JevBench public 167/231, Brier 0.348, ECE 0.050; internal held-out short tasks 0.641 (n = 6,000) Its server binds to 127.0.0.1 with no authentication. Its own limits: questions are read less than documents (a reworded question often gets the same answer); long multi-step documents are the weak spot; score and noul transfer poorly to unfamiliar rubrics; one calibration temperature per primitive. Trained partly on a frozen Qwen3.5-4B teacher's output distributions
Clef, Clef-flash, pplx-decider-v1-27b open weights of the hosted models above rows above —

Small open alternatives (individual authors)

Closer to the replicas on Open replicas and Jev-compatible servers; listed here because their authors position them as alternatives to Jev rather than copies.

Project (author) What ★ · push · licence Their numbers and caveats
JevAlt (mertkayacs): Deem-4B (English), Karar-4B (Turkish), Wähler-4B (German) 4B models from internlm/Intern-Decision-4B (Qwen3.5-4B); jevalt serve on /v1/systemone, GGUF on a CPU in about 3 GB RAM; extras reasoning, abstain (an "unknown" option) and coverage 4 · 2026-10-04 · Apache-2.0 Author's live tests 2026-10-04, 130 known-answer requests per language: Wähler-4B 122, Deem-4B 122, Karar-4B 113; Wähler-4B matched only 14 of 20 long-policy cases. Its Jev rows come from TypeSafe's jaggedness notes and a third-party audit, not a live Jev run. It warns dates are shaky (Wähler-4B miscounted a return window across August and September, reasoning on)
jevos (feder-cr; repo jev) Local yes/no decisions (choice and score "early") in one C++ binary on OpenVINO int8, CPU only; 8,192-token context; accepts any jev-* model name and answers as jevos-v4 1203 · 2026-10-04 · MIT (repo created 2024-08-15, before Jev) Author, jevos-v2 on an Intel laptop: 2,000 yes/no questions from unseen policies 0.810 vs Jev 0.927 vs Laya 0.489; latency 26 vs Jev 344 ms (short), 112 vs 345 ms (long). Model-name key safety: Warnings: not-Jev services, key safety, look-alikes and install names (c)
GenClass (MeharPro) 32M-parameter ModernBERT/Ettin encoder with typed heads and per-option attention isolation (option order cannot change the answer); runs in the browser (WebGPU/WASM) behind a voice-control Chrome extension; /v1/systemone server; calls itself "the version of Jev that runs in your browser" 2 · 2026-10-04 · Apache-2.0 Repo table: mid-sentence command recognition 90.4% vs Jev 66.4%; general held-out questions 80.5% vs Jev 94.7%; on-screen element 82.4% vs 86.6%; option-order flips 0 vs 10 to 13% (Jev's flips from a third party). Its harness carries 549 published Jev numbers, so which Jev figures were run live is unclear: unverified. Not affiliated with TypeSafe (README)

Proposals and platform APIs

Name collision. "Decisions API" names at least five different things: OpenAI's (GPT-6 Luna), Perplexity's (/v1/decisions, pplx-decider), Chrome's proposal, OpenRouter's alpha endpoint serving the real Jev (Platforms and gateways: Zapier, Cloudflare, Netlify, Vercel AI Gateway, OpenRouter, Pydantic AI Gateway, Opper, Fly.io, OpenCode Zen and other hosted routes to Jev), and the /v1/decisions paths of B.AI (a route) and TokenBazaar (Warnings: not-Jev services, key safety, look-alikes and install names). Check which one a post means.

No methodology, unverified: "100% accuracy, independently verified" for ConceptNet (conceptnet.co.uk), a one-founder rival, set against a "68%" Jev figure it does not source (@tonylab_net; the Medium write-up is members-only, fetched 2026-10-01, not captured; moved from Open replicas and Jev-compatible servers 2026-10-04).

Advising with rivals

Sources

Links inline; raw captures in frontmatter (vendor pages, model cards and READMEs captured 2026-10-04 under raw/community/ and raw/x-repos/; the d1 and Drex rows and the ConceptNet note captured 2026-09-29 to 2026-10-01, moved from Open replicas and Jev-compatible servers unchanged). Not captured: Perplexity's and Fastino's launch posts on X, the Strands Decider code repository, OpenRouter's pplx-decider page, GLiNER2.5-Decide and press coverage.