Agents: read the raw Markdown of this page, or start at llms.txt.
Open replicas and Jev-compatible servers
TL;DR None of these is TypeSafe or runs Jev. They copy the wire shape of
/v1/systemone, not Jev's calibration, and most computeconfidencetheir own way, so thresholds tuned on Jev do not transfer. On the one independent board with a sealed set, JevBench v1.4.2 (2026-09-24, 93 systems), Jev 1.13.0 is #2 at 63.29 behind decider-4b v2 at 64.13, and leads on reasoning and calibration; the board weighs speed and cost equally with those. Pick a replica only for offline, private or cheaper runs, and re-measure on your own labels (Measurements, access routes and open replicas). Read the Warnings first. Split from Community SDKs and libraries on 2026-09-25 (nothing dropped).
Numbers are each author's unless marked. ★ · push as captured 2026-09-23 to 25. Pattern IDs: Decision patterns from the community (with fit verdicts). "Our check" = our 2026-09-25 source and registry vetting (docs/sweep record), not a raw capture.
Quick picks by need
From the rows below; none is Jev.
- Highest on JevBench v1.4.2: decider-4b v2 (#1, 64.13 vs Jev 63.29), but Jev leads reasoning and calibration there, decider leads speed and cost, its 2026-09-24 models failed their own pre-registered release rules, and "nothing distilled" is
unverified. Next: JevK5 (#3, 62.04; self-disclosed MMLU-Pro test-item replay in v0.2), reflex 4B (#7, 53.99; MIT, ships a frozen model and publishes its negatives). - Small or edge: Laya (32.8 ms per question on a T4; Node port @receptron/laya), but its base checkpoints fall below the majority baseline on typed decisions and the English one stays confident on Khmer and Hindi it fails; openJev-verdict-2.0 (149.6M, WebGPU); Von (395M; its own table 72.0% vs Jev 96.6% on 49 tasks); OpenThai-SystemOne (0.8B; 74.3 vs Jev 76.0 on Nimble's suite). Any logprob endpoint: NotJev (Qwen3-8B Q4 p50 23 ms).
- Avoid: SiliconLabAI/OpenJev (any web page you visit while it runs can take its keys), solar-mini4-jev when you need
confidence(hard labels), ollaya and deepopen (look-alike, renamed copy), and servers that silently answer tojev-latest(check the responsemodel). With razorback16/openjev, never put a TypeSafe key beside the Codiv base URL. Details: Warnings. - Before any swap: Jev-tuned thresholds do not transfer; re-measure on your own labels and test option order, irrelevant state, missing evidence and general judgement ("What the replicas' own negatives say").
Which is which
| Name | Means | Not to be confused with |
|---|---|---|
| OpenJev (four) | TheoLeeCJ/openjev, now SemIf (Qwen3.5-4B logit reader; JevBench v1.4.2 #11, 47.69; row on Repos: data, documents, judging, real-time, markets, business apps, replicas; since the 2026-09-19 capture it added CPU via llama.cpp and Torch, and a 27B EXL3 bridge, our check; MLX is in the capture) · razorback16/openjev (DiffusionGemma server, hosted on Codiv) · IamBusy/OpenJev-Vision (image events) · Zefan Cai's Open-Jev (a dataset plus 2B and 9B models; v1.4.2 #76 and #75) | also zhangcy122/OpenJev (sold as OpenJevPro) and SiliconLabAI/OpenJev (Warnings). PyPI openjev is balys's, none of these (our check) |
| jevify (three) | fidecastro/jevify: a server, owns PyPI jevify |
altryne/jevify (Tools and integrations hub: MCP servers, plugins and connectors), ryana/jevify, a prompt (Repos: coding agents, orchestration, memory, browser and computer use, integrations) |
| reflex | kshetrajna12/reflex, whose server command is reflex-serve (the "Reflex" local server some tools name) |
lateos-ai/reflex (a Rust/CUDA engine without /v1/systemone); PyPI reflex (a web framework) |
| djev (three) | Davipar/djev-dev with api.djev.dev (Maisa; the one JevBench ranks, v1.4.2 #8) | mmastrac/djev; taeold/djev-run, built on mmastrac's |
| AnyJev / JevAny | nokia-applied-research/AnyJev | weitianxin/JevAny |
| OmniJev | OmniJev/PlayJev (games) | tinnel123666888/OmniJev (multimodal); Iron-LYK/OmniJev returned 404 (our check) |
| Laya | NandhaKishorM/laya = HF convaiinnovations/laya |
HF convaiinnovations/laya-typed-decisions is a separate checkpoint fine-tuned on that benchmark's train split |
confidence is not one thing. Jev's field is its own (Confidence vs probability). Replicas use normalised entropy 1 − H(p)/ln K (notjev, razorback16/openjev, JevEmbed; localjev and hearim say entropy-based), the top probability (Nemotron_Jev), a "peak" formula (jevify, OmniJev, ollaya, JevAny, our check) or nothing real (solar-mini4-jev answers one letter, so every answer is certain). Same field name, different number.
Replicas that answer to Jev's model names. localjev, razorback16/openjev, jeff and arbiter accept jev-latest (localjev and openjev also jev-preview). An SDK whose base URL points at one gets local answers with no error; only the response model field tells. fastjev rejects jev-latest.
Warnings
Plain names on purpose: no links. Flagged, not adjudicated; our reading of the MCA (Legal: MCA, DPA, privacy, data retention), not legal advice.
- (a) Public Jev on the operator's key (possible MCA §2.3(a): offering the Services as a standalone service). typesafe.pro: a keyless front door with free and paid token tiers that pools several operator master keys (
TYPESAFE_MASTER_API_TOKEN_N, 1,200 requests/minute and 250,000 tokens/second each, per its docs); spreading load over several keys may also meet §2.3(j), Usage Limits (inferred). classifier.dev: free with no key (20,000 requests/day per IP), paid Pro tier; says its fast tier is Jev; its /about page names the operator as Michael Ryaboy, an independent developer; its terms mention partner keys, and whether it holds a TypeSafe agreement is unknown. safer-with-jev: a public demo API with no caller auth that runs anyone's text through Jev on the owner's key (1,000 Jev calls/day deployment-wide). Warnings, never routes. - (b) Distillation (MCA §2.3(b) forbids training a model to imitate the Services' output). SargeDev jev-distill-corpus-v3: its own card says its largest stream, 498,010 of 740,957 rows, has labels distilled from Jev 1.13 via OpenRouter. JevEmbed-Data includes about 38k rows of it (our check of the dataset). laya-jev-GraphRAG pitches paired Laya/Jev answers as training data for Laya (our check). Bespoke Nimble ships Jev-labelling scripts and says saved Jev probabilities are "available for future soft-target distillation"; the released training set looks clean (our check), and its "did not distill from Jev" is
unverified. stuntd's paid-upstream mode learns from TypeSafe's answers and then takes the decision over. - (c) Key safety. SiliconLabAI/OpenJev attaches the server's TypeSafe, decider or OpenAI key to a URL taken from the request body, with CORS open to all origins: any web page you visit while it runs can take those keys (our check; not listed). razorback16/openjev's instructions put a Codiv key in
TYPESAFE_API_KEYbesideTYPESAFE_BASE_URL=https://api.codiv.ai; a builder who changes only the base URL sends a real TypeSafe key to Codiv (inferred). Replicas that answer tojev-latest(above) swap the model silently. - (d) Install names that don't exist or belong to someone else (our registry check): PyPI
open-alternative-jevanddeepopendo not exist; PyPIllm2jevis tic-top's, not Yinsongxu's; PyPIopenjevis balys's; npmfastjevis an unrelated squat (use PyPIfastjev); PyPIreflex,glance,snapare unrelated. On tools pages: npmjevwireis 404, npmpi-jev-guardandpi-jev-routerbelong to other publishers (Tools and integrations hub: MCP servers, plugins and connectors).
Models with their own weights
| Tool (owner) | What | Install | ★ · push | Their numbers and caveats |
|---|---|---|---|---|
| Kev (Jared Palmer) | 0.8B/4B/9B/27B on Qwen bases, plus the original Kev-0.5B (Qwen2.5-0.5B) kept for reference; serves /v1/systemone (+ /v1/systemone/permute, 1–64 option orders) |
HF weights, uv sync |
6875 · 2026-09-25 · Apache-2.0 | README of 2026-09-25: Kev-27B 0.848 vs Jev 0.857 on new sources; TypeSafe's 102 public eval questions, agreement Kev-9B 0.809 vs Jev 0.891; MMLU Kev-9B 0.74 vs Jev 0.90, MMLU-Pro 0.52 vs 0.84; trained on states ≤ 384 tokens, serves 8,192 per state and per question, accuracy drops on long documents. JevBench v1.4.2: kev 4B #28 (36.14), 8B #49, 0.6B #53 (these three marked research preview), 0.5B #59 (18.88). Version-dependent: quote the model and date. "No Jev outputs were used for training": unverified |
| decider (Mapika) | Qwen3.5 fine-tunes (2B, 4B, 35B-A3B), local Qwen3.5-27B teacher | pip install decider-ai |
394 · 2026-09-25 · Apache-2.0 | README: public hard tier Jev 0.730 vs decider-35b-a3b 0.676; publishes that its 2026-09-24 models failed their pre-registered release rules. JevBench v1.4.2: decider-4b v2 #1 at 64.13 vs Jev 63.29; Jev leads Intelligence (53.1 vs 49.4) and calibration, decider leads speed and cost (board's note); decider-35b-a3b #19, decider-2b #39. "Nothing was distilled from Jev": unverified |
| reflex (Kshetrajna Raghavan) | frozen Qwen3.5-4B (default) or Qwen3.8-27B read by option letter, two option orders averaged; no weights of its own; reflex-serve |
git checkout stable && uv run reflex-serve --stable |
147 · 2026-09-23 · MIT | By version: JevBench v1.2 run by the benchmark author (534 decisions, 2026-09-21): reflex 4B (a LoRA configuration, not today's stable) #5 at 71.7 vs Jev #1 75.4; reflex-27b hard tier 75.9% vs Jev 74.1%, top calibration axis 86.2, but #22 on cost and speed. v1.4.2: reflex 4B #7 (53.99), reflex-27b #64. Self-run public items: hard 0.685 / ECE 0.081 vs Jev 0.730 / 0.031. Publishes that every fine-tune and GEPA prompt tuning won in-distribution and lost general judgement, so it ships the frozen model |
| Bespoke Nimble (Bespoke Labs) | Nimble-9B, LoRA on Qwen3.5-9B; data from GPT-5.6 | HF bespokelabs/Bespoke-Nimble-9B |
1795 · 2026-09-24 · no repo licence (weights Apache-2.0, our check) | Independent suite, Jev run live on human labels (13 public subsets, 3,880 records, contributed by Edgar Dyck): Jev 1.13.0 76.0% macro vs Nimble 74.8%; Jev leads Noul 84.6 vs 80.2 and Choice 82.9 vs 81.6, Nimble leads Score 54.6 vs 50.1; Jev 35.0% on SummEval-relevance; Jev lower ECE on 11 of 13; MASSIVE en → de Jev 87.4 → 86.9, Nimble 86.9 → 83.4 (Head-to-head benchmarks: Jev on public datasets and suites). JevBench v1.4.2 #60. Distillation note: Warnings (b) |
| JevK5 (Alibi Serikbay) | Qwen3.5-4B + LoRA reading option letters (SemIf method, credited); also 9B, 2B, Lite, GGUF | HF alibiserikbay/JevK5 |
107 · 2026-09-25 · Apache-2.0 | Author on JevBench public 231: hard 0.784 vs base 0.613; the 9B is worse than the 4B. JevBench v1.4.2 #3, 62.04 vs Jev 63.29; sealed 308 items Jev 36.7% vs JevK5 33.1%. Self-disclosed: v0.2 replayed MMLU-Pro test items. Teacher includes GPT-6 Luna (an OpenAI-terms question). "No Jev outputs": unverified |
| JevAny (Tianxin Wei) | 27B pointer-head model with an RLCR calibration reward; includes Kev code (NOTICE) | HF tianxinwei/JevAny-27B-SFT / -RLCR |
16 · 2026-09-25 · Apache-2.0 | Author: transfer accuracy SFT 82.41% vs Jev 85.37%; MMLU-Pro 73.0% vs Jev 84.0%. No per-item Jev file, so live vs copied is unknown. Own negatives: RLCR no gain; label-free self-training negative. "No Jev outputs": unverified |
| open-spark-jev (Abhishek Rai) | LoRA of Qwen3.5-4B for DGX Spark (spark-s1), 11,792 rows from its own open-model data factory; /v1/evaluate with a /v1/systemone alias |
HF abhishek085/spark-s1-4b-v6 |
18 · 2026-09-22 · Apache-2.0 | Where Jev leads (author, Jev figures from third-party per-row records, not run live): 60 tool calls 0.900 vs Jev 0.917 (Jev p50 421.6 ms with network, local 74.9 ms); Kev transfer 0.756 vs 0.857; vulnerable code at chance (0.502). RLCD made results worse twice; not shipped. Its JevBench 83.1 is a self-computed proxy; v1.4.2 official #15, 44.62. Its "20–200x" is the CEO's launch-post range (post), not the blog's 40x–200x |
| OpenThai-SystemOne (iApp Technology) | Qwen3.5-0.8B with a 256-way slot head, Thai continued pretraining, synthetic data from Qwen3.6-35B-A3B | pip install openthai-systemone; hosted API (our check) |
60 · 2026-09-21 · Apache-2.0 | On Nimble's 13-subset suite: 74.3 vs Jev 76.0 (Jev's copied from Bespoke). Weak spots it states: SummEval-relevance 21.7, PubMedQA 64, English ECE ~0.15; v0.3 trained on train splits of 5 bench sets; option-order fix cuts flips 18.9% → 3.3% on its Thai set |
| AgentJev (malevrigns) | Qwen3-0.6B with a candidate head; own POST /api/evaluate, not /v1/systemone (no SDK route) |
HF aimeigaoshou/agent-jev (a different account) |
303 · 2026-09-23 · Apache-2.0 | typed-decisions test: 79.25% vs a local Laya run 77.00%; Jev's 72.7% is quoted from the typed-decisions card, and its own README says the rows are not comparable (specialists trained on the benchmark). Its Claude Code hook fails open |
| JevForge (zwliJay) | next-click scoring on website-disjoint Mind2Web splits; /v1/systemone without auth |
HF AndeyTait/JevForge-0.8B |
53 · 2026-09-23 · bare MIT line | Author: Choice top-1 test/OOD 0.579 / 0.637 vs Jev 0.543 / 0.610; Jev wins Noul (0.910, Brier 0.092 vs 0.826 / 0.128) and score MAE. Its Jev benchmark caps at 200 records per split and drops failed calls (our check), so not like-for-like. Label source unnamed. P12 |
| open-jev-typed-decision-engine (intikhab49) | ModernBERT-base 150M, labels in the input; trained on typed-decisions train (1,016 cases) | clone | 44 · 2026-09-21 · Apache-2.0 | Test: fine-tune 0.624; ensemble 0.6965 / ECE 0.0565, only for known question sets; unseen questions ≈ 0.624. Jev 0.727 quoted. Its "single human annotator 0.659" misreads the card, whose gold is a model teacher |
| CLM (Contrastive-LM) | frozen Qwen3-8B embedder + 20M two-tower head; serves /v1/systemone; calls Jev only as a baseline |
pip install contrastive-lm |
1187 · 2026-09-24 · Apache-2.0 | T-Rex, 5 seeds × 60 s, safety shield on for both: both survive 5/5, but Jev agreed with the physics planner 0.987 vs CLM 0.658, so CLM's survival leans on the shield; p50 149.8 vs 16.5 ms. Its best-of-N verifier claims against Jev (DeepSWE, Terminal-Bench) have no Jev file in the repo: unverified. JevBench v1.4.2 #78 (8.57) |
| solar-mini4-jev (Sung Kim) | Upstage Solar chat model answers one option letter per question (BYOK Upstage key) | clone; serves /v1/systemone |
43 · 2026-09-24 · no licence | Hard labels: no usable confidence, so no confidence routing. Author, 886 public items: Jev 75.6 vs solar-pro4 76.9, solar-mini4 71.4; Jev tied only on classifier-benchmark v1 (97.4). Prompt tuned on public Jev failure cases (overfit risk) |
| Nemotron_Jev (Alex Steiner) | masked-token scoring on NVIDIA's Nemotron diffusion LMs via modified vLLM | clone (no main branch, our check) |
14 · 2026-09-25 · MIT repo, base model under NVIDIA's Nemotron licence | Author: 174/231 on JevBench public items with answer reordering. No auth, TLS or rate limit (author warns). Probabilities uncalibrated. LoRA labels from an OpenAI model (OpenAI terms, not MCA) |
| OmniJev (tinnel123666888) | decisions over images, video, screens, robot scenes, spectrograms; no HTTP server | HF tinnel123/OmniJev |
70 · 2026-09-25 · NOASSERTION | Author, held out: LIBERO-10 0.807, Mind2Web 0.733, POPE 0.902; A800 294 ms per question. Institutional affiliation self-stated only. Jev is text-only (Models, aliases, pricing, rate limits, context) |
| PlayJev (OmniJev) | Qwen3.5-0.8B-Base playing ten browser games from raw pixels | HF | 36 · 2026-09-24 · Apache-2.0 | 43 ms per move on an H200; 0.57 of its game teachers (trained with DAgger) |
| JEVfire (kikoncuo) | Jev-inspired parallel decisions on vLLM (CUDA) and WebGPU | clone | 67 · 2026-09-18 | 28 decisions in 497 ms, 10.3× faster than generating the same constrained JSON on the same 27B model; Mario 71 ms/action in a browser (Qwen3.5 0.8B, M4 Max); credits the RLCD checkpoint |
| Von (wfzyx) | 395M encoder | pip install von-sdk / npm install von-sdk (plain von is unrelated) |
526 · 2026-09-23 · Apache-2.0 | own table: Jev 96.6% vs Von 72.0% on 49 tasks. JevBench v1.4.2 #46 |
| openJev-verdict-2.0 (Heman10x) | 149.6M ModernBERT, confidence head, WebGPU | Hugging Face | 273 · 2026-09-20 · NOASSERTION | claims 77.10% vs Jev 72.70%; Jev's figure is from the typed-decisions card, where the dataset authors ran Jev live on 2026-09-18, and gold is agreement with a ~4B teacher (ceiling 0.735). JevBench v1.4.2: Verdict 1.4 #58, Verdict #62 |
| visual-jev | image decisions, Qwen3-VL | uv | 2 · 2026-09-20 · no licence | Jev is text-only. Not hr98w/jev-visual (a teaching repo) |
| JevTuner | Brier-loss training recipe | clone | 12 · 2026-09-21 · no licence | not TypeSafe's method |
Servers and readers over other models
| Tool (owner) | What | Install | ★ · push | Their numbers and caveats |
|---|---|---|---|---|
| jeff (Logan Markewich) | /v1/systemone over GLiFormer (knowledgator, 400M); accepts jev-latest; 529 on a full queue, 429 with retry-after-ms |
clone | 246 · 2026-09-20 · MIT | Jev run live, 1,600 labelled items: AG News 75.5% vs Jev 90.5%; p50 151 vs 129 ms. JevBench v1.2.2: jeff 66.9 (#9 of 18) vs Jev 75.3 (#2); hard tier 37.7% vs 74.1%. v1.4.2: #40, 30.58. Self-run hard items: Jev lost temporal/numeric 4/15 vs 6/15 (matches Jev 1.13 jaggedness: known failure modes), led ambiguous 6/7 vs 0/7 |
| localjev (GitHub Next) | Bun bridge: any OpenAI-compatible local model writes probabilities as JSON (self-reported, not logits); accepts jev-latest, jev-preview; 529 above 64 queued |
clone | 775 · 2026-09-18 · MIT | M5 Max bake-off, 120 items: short-input macro Qwen3.6-35B-A3B 76.7%. 2,048 unrelated words lowered every model's macro accuracy (Qwen 76.7% → 69.2%; Gemma E4B SST-5 50.0% → 12.5%). Evidence that prompted-probability replicas degrade under irrelevant state |
| OpenJev (razorback16) | DiffusionGemma 26B-A4B via patched vLLM or MLX, read as a diffusion canvas; also serves Laya and Verdict; extensions incl. images; accepts jev-latest |
clone; hosted free on Codiv (100M input tokens) | 421 · 2026-09-24 · Apache-2.0 | Author, one GPU: p50 27 ms (1 question), 31 ms (3); 57.4 req/s at concurrency 64 (p95 1,109 ms). JevBench v1.4.2 #27 (36.85; thinking mode #68). Codiv key routing: Warnings (c) |
| NotJev (Nathanael Braun) | letters the options, asks any logprob endpoint for one token (26 options max); abstains below theta |
npm i notjev |
21 · 2026-09-22 · Apache-2.0 | Author: Qwen3-8B Q4 p50 23 ms; 27B NVFP4 on vLLM 101 ms; own judge 0.947 agreement on n = 1,224. Its "hosted Jev 419 ms" has no source in the repo. In serve mode one key is both upstream bearer and caller guard |
| jevify (Felipe Infante de Castro) | any LLM as a Jev-like endpoint: hashed YAML recipe, state sent once, logprob ladder; endpoint, embedding, NLI or rerank backends | uv tool install jevify (PyPI) |
43 · 2026-09-23 · MIT | Unmodified typesafe-sdk works against it (author). policy-hard-52: DeepSeek-V4-Flash 47/52, Gemma 4 E4B 45/52. No Jev comparison |
| fastjev (chengyongru) | SDK-first SemIf fork: Torch, vLLM, MLX, llama.cpp GGUF, EXL3; 2–16 options; rejects jev-latest |
pip install 'fastjev[torch]' (npm fastjev is unrelated) |
21 · 2026-09-24 · MIT | RTX 5090: vLLM batched 36.26 vs Torch 18.52 decisions/s. Jev side copied (SemIf's 0.845 vs Jev 0.883 on 102 rows; count once). Headline says "open source implementation" while its results say Jev is not reproduced |
| OpenDecision (Deepan Wadhwa) | zero-shot NLI (ModernBERT-large) behind the Jev shape, plus retrieval | pip install OpenDecision (Python ≥ 3.13) |
57 · 2026-09-21 · Apache-2.0 | Treat scores as uncalibrated (README). JevBench v1.4.2 #56, 21.64. P24 |
| open-alternative-jev (Iker Moel Tacher) | in-process library: all questions about one state in one pass, option-letter logits; not SDK-compatible | clone (the PyPI name in its README does not exist) | 54 · 2026-09-25 · Apache-2.0 | Stock Qwen3.6-27B 8-bit on typed-decisions: 73.7% / ECE 0.020 vs Jev 72.7% / 0.144 (Jev copied; gold is teacher agreement, ceiling 0.735). Reversing options moved 4B yes/no accuracy 13.5 pts; packing hurts below ~4B. v1.4.2 #36. P01 |
| choosekit (Felix Koba) | TS: option-label logprobs from llama.cpp, Ollama or OpenRouter, with images; choosekit-mcp |
npm i choosekit |
23 · 2026-09-24 · Apache-2.0 | Jev run live (OpenRouter): SuperGPQA 1,000 items, Jev 53.6% at $0.0244 per 1,000 decisions, p50 325 ms; Kimi K3 59.3% at ~25× the cost. Graduate knowledge questions are a poor fit (Jev 1.13 jaggedness: known failure modes). SemIf authored144: tie at 139/144 with local Qwen3.8 27B Q4 |
| OpenJevPro (Corel Zhang) | letter-logprob scoring on Ollama or OpenAI-compatible models; "guard" wrapper around Jev | pip install openjevpro |
30 · 2026-09-25 · PolyForm Noncommercial + paid licence | Author, 60 Banking77 cases (benchmark_results.json): Jev 98.3% = an Ollama model 98.3% (gpt-oss:20b-cloud per its benchmark script, our check; the results file says only "ollama"), mean 732 vs 1,937 ms; separately, 36 queries: Jev 97.2% (35/36), p50 740 ms. Its "Jev ~25–45 ms TTFT" contradicts its own 732–750 ms; "60%+ cost cut" and "99.99% SLA" are unmeasured. Tiny n |
| rizzo-flow (Simone Rizzo) | option-letter logits from one shared prefill; adds numeric and an "insufficient evidence" option |
clone | 459 · 2026-09-25 · Apache-2.0 | Author (RTX 5060 Ti): SemIf authored144 0.812 vs SemIf 0.819; p50 49 ms. 6 of 36 confident wrong answers when evidence is missing. "No Jev labels": unverified |
| djev-run (Daniel Lee) | Cloud Run deployment of mmastrac/djev on DiffusionGemma | ghcr.io image | 551 · 2026-09-24 · no licence | Author: $3.19/h on one RTX PRO 6000, scales to zero; cold start ~47.5 s; 117 ms median. Its "JevBench 73.4" is self-run, not an official row. Answer shape differs from the docs (1-indexed Score probabilities, no type field; our check) |
| SNAP (Emanuele Menon) | Rust over vendored llama.cpp; adds numeric, allow_abstain, and a confidence on Noul (Jev has none) |
build | 21 · 2026-09-25 · no licence file | Author: Qwen3.8-4B agrees 73.2% with frontier-consensus labels over 373 TypeSafe public decisions; option-order reversal flips a third of MiniCPM's choices; 51 ms vs Ollama 149 ms. Binary snap shadows Ubuntu's |
| Glance (Yohei Nakajima) | logits from a frozen Qwen3-VL-4B for image decisions; /v1/decide |
pip install glance-vlm |
23 · 2026-09-24 · Apache-2.0 | Pre-registered fresh photos: yes/no 0.939 (541) vs Gemini 3.1 Flash-Lite 0.961; exact rating 0.669 vs 0.763. vs Laya Vision on an M5: 0.886 vs 0.689 (page). P34 |
| JevEmbed (HIT Shenzhen) | any embedding model → Choice/Score/Noul by cosine and temperature; normalised-entropy confidence | clone | 26 · 2026-09-25 · Apache-2.0 | Author, JevBench public 231: best Qwen3-Embedding-8B 58.44%; that report was deleted from the default branch on 2026-09-25 (captured from commit 9cb60bb). Dataset licences mixed, incl. CC-BY-NC; distillation note: Warnings (b). P17 |
| AnyJev (Nokia Applied Research) | any HF LLM, debiased, optional head on 100–300 labels | pip install "anyjev[hf]" |
267 · 2026-09-23 · Apache-2.0 | order flips 0.230 → 0.073 (Qwen3-8B) |
| Simple Jev (Eugene Cheah) | logits server, free demo API | pip from repo | 491 · 2026-09-21 · Apache-2.0 | own contract. typed-decisions card: 0.716 (Jev 0.727). JevBench v1.4.2 Qwen3.8-27B #32 |
| jev-rs (Yijun Yu) | Rust over llama-server; serve, MCP, eval | cargo install jev-rs |
5 · 2026-09-22 · Apache-2.0 | SDKs reach it via TYPESAFE_BASE_URL |
| hearim (ziozzang) | Go gateway over Ollama, vLLM, SGLang | go build |
5 · 2026-09-22 · NOASSERTION | entropy-based confidence |
| Recipe (@theanandprasad) | softmax over yes/no logits | — | post 2026-09-22 | 21/22 on Llama 3.3 70B; raw logits flip with option order (AnyJev) |
What the replicas' own negatives say: option order moves answers (open-alternative-jev, SNAP, AnyJev), irrelevant state hurts prompted replicas (localjev), missing evidence produces confident errors (rizzo-flow), and fine-tuning tends to trade general judgement for in-distribution wins (reflex, open-spark-jev). Test all four before swapping one in (Testing and evaluating a Jev workflow).
Laya and its servers
Laya (NandhaKishorM/laya, Nandakishor M; pip install laya; 18593★ · 2026-09-23 · Apache-2.0): three encoder checkpoints plus a router, laya-serve; 32.8 ms per question on a T4. Its Jev figures are quoted, not measured; Jev leads Banking77 0.870 vs 0.425. Independent caveats: zero-shot on typed decisions the base checkpoints score 0.362 (English) and 0.352 (multilingual), below the 0.461 majority baseline (Laya's model card, quoted by arbiter; the 0.766 belongs to the checkpoint fine-tuned on that benchmark); the English checkpoint scores 0.000 on Khmer at 0.952 confidence and 0.100 on 20-way Hindi intent (chance 0.050) while staying confident (arbiter); lev's encoders reach 61–67% on SemIf's authored144; Laya Vision 0.689 vs Glance 0.886. JevBench v1.4.2 #41, 30.25.
| Tool (owner) | What | ★ · push | Notes |
|---|---|---|---|
| @receptron/laya | Node/TypeScript ONNX Runtime port, npm i @receptron/laya |
466 · 2026-09-21 · MIT | matches Python to four decimal places; ~140 ms per 3 questions on Apple-silicon CPU; ~1.7 GB download; options within 192 tokens, state cut at 512 |
| arbiter (Khaled Bakeer) | serving layer: router, micro-batching, CUDA graphs; accepts jev-latest; MCP server; Claude Code hook that fails open |
33 · 2026-09-24 · MIT | GB10 20.9 ms per question; the Hindi collapse above; six games vs baselines: beats random on only 3 |
| lev (jlt-commons) | Clojure: Laya encoder first, escalates below a gate to Qwen3.5-4B | 19 · 2026-09-24 · Apache-2.0 | encoders 61–67% vs a MiniCPM5-2B thinker 95.1% on authored144; gate at 0.5 gives 92.4% at 270 ms per case. Its "24/64 encoder vs 60/64 hosted Jev" has no file. P06 |
Names only (C, not captured; our check): laya-mps (Apple MPS; its own /v1/decisions, not a drop-in), sys1 (Rust server, entropy confidence), laya-server (1Panel; own keys and admin UI), ollaya (an Ollama look-alike in name, CLI and port, installed by curl | sh that adds a systemd service), deepopen (mostly a renamed copy of Laya presenting Laya's numbers as its own; PyPI deepopen does not exist), laya-jev-GraphRAG (its Jev client reads the wrong response fields and hides errors; distillation: Warnings (b)).
Names only
Not captured (C; our check, 2026-09-25): OpenJev-Vision (IamBusy; image events, not general VQA), JEV-CPU (a CPU port of SemIf, superseded by SemIf's own CPU backends), dev-0.4b (no licence file), Dohnuts (weights CC BY-NC-SA; states Jev leads), JevBERT / typed-decision-bert (NLI shell), jev_local (Japanese; weights CC-BY-SA), LLM2Jev (Yinsongxu; PyPI llm2jev is another project's), OpenSourceJev (sabeel111; JevBench v1.4.2 #21, 40.87), Valen (image and video; Jev is text-only), jevper (a typesafe-sdk-style client over OpenAI/Anthropic models; its "SDKs retry 409" contradicts docs, HTTP status codes, rate limits, retry semantics), Lateos reflex (see Which is which). Earlier names only: LitJev (v1.4.2 #57), jevmlx, jev-visual. Elsewhere: SemIf, NanoJev, jevlike, openjev-sglang, mini-jev → Repos: data, documents, judging, real-time, markets, business apps, replicas.
Reading the numbers
- JevBench changes by version. v1.2 (reflex's record, 2026-09-21): Jev #1 at 75.4. v1.2.2 (jeff's record): classifier.dev fast tier #1 at 84.8, Jev #2 at 75.3. v1.4.2 (2026-09-24): 534 frozen plus 308 sealed decisions; Jev #2 at 63.29, sealed accuracy 36.7% (chance 0.293). Jev's row is flagged API: TypeSafe's endpoint received the sealed item text, without answers, as every API row did. classifier.dev (fast tier = Jev) is listed, not ranked. Always quote the version.
- typed-decisions (LocalLLaMA): the dataset authors ran Jev live on 2026-09-18 (
jev-latest→jev-1.13.0, 400 cases, 2,000 decisions): 0.727, p50 710 ms. Gold is the mean of three samples from a ~4B-class teacher, so the score is agreement with that teacher (ceiling 0.735), not correctness. Replicas quoting Jev 0.727 copy this row.
Related
- Community SDKs and libraries — community SDKs and libraries that call the real Jev
- Measurements, access routes and open replicas — replica trust notes; Head-to-head benchmarks: Jev on public datasets and suites — JevBench and public datasets
- Platforms and gateways: Zapier, LangChain, Spring AI, Cloudflare, Netlify, Vercel, OpenRouter, Fly.io, Pydantic AI and other hosted routes to Jev — hosted routes to the real Jev, and the same warnings
- Legal: MCA, DPA, privacy, data retention — MCA §2.3; TYPESAFE_* environment variables across SDKs —
TYPESAFE_BASE_URL
Sources
Links inline; raw captures (2026-09-23 to 25) in frontmatter. "Our check" items come from the 2026-09-25 vetting record (docs/sweep/2026-09-25-mrjev), not raw captures.