$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites

[ community tier ][ updated 2026-09-25 ][ confidence medium ][ jev-1.13.0 ]#head-to-head · benchmarks · eval-boards · replicas · community

TL;DR On independent boards Jev is at or near the top on accuracy and calibration and loses rank on speed or cost to small open models: JevBench v1.4.2 puts it #2 (63.29) behind decider-4b v2 (64.13), with 86.6% on public items but 36.7% on 308 sealed ones. Always quote a JevBench number with its version; scoring changed at v1.4. Typed-decisions scores are agreement with an unnamed teacher, not accuracy. Split from Head-to-head benchmarks: Jev on public datasets and suites on 2026-09-25.

Numbers as each operator published them, on jev-1.13.0. None of these operators is TypeSafe; TypeSafe publishes no benchmark table (Blog: Lies, Damned Lies, and Benchmarks) and its own board is Workflow evals: how TypeSafe measures Jev. Vendors of rival models are labelled.

Independent eval boards

Board (operator) Task Numbers as published Jev placed Caveat
Jevals.com noul board (anonymous operator; method) PubMedQA, 300 items × 5 repeats, human labels Decision Score (0 = guessing base rates) Jev 69.0 vs Gemini 3.8 Flash 73.0; accuracy 91.3% vs 92.5%; $0.029 vs $0.8 per 1,000; p95 653 ms vs 5.5 s. At 95% accuracy Jev can act alone on 86% vs 79% Tied first: paired difference −2.5 to +10.4 includes 0 Operator anonymous; the site showed a job-board ad (our 2026-09-24 sweep note, not in the text capture); no frontier LLM listed (budget $25). LLM probabilities are verbalized, not logprobs
Jevals.com choice board Banking77, 77 intents, 300 × 5 Jev 67.8 vs Gemini 3.8 Flash 74.1 vs GLM-5.3 66.8; accuracy 79.7% vs 84.6%; $0.043 vs $1.4 per 1,000; 41.5% of Jev answers at confidence 1.00, ECE 9.8 vs 3.2 points; 10.3% of Jev picks change when options are reordered LOST to Gemini (significant); tied 2nd with GLM-5.3 Options shuffled identically for every system; 2.7% of Jev picks change on an identical rerun
Jevals.com score board HelpSteer2 helpfulness, 5 levels Jev 9.2, GLM-5.3 7.8, Gemini 4.6: none beats guessing the label base rates; Jev accuracy 41.3%, ECE 19.7 Nobody wins No model reaches 95% accuracy at any gate; keep a person in the loop
Bespoke Nimble public suite (bespokelabsai), 13 human-labelled public subsets, 3,880 records, Jev run live Nimble-9B vs Jev 1.13.0, identical records and wording Macro 76.0% vs 74.8%. Noul 84.6 vs 80.2, Choice 82.9 vs 81.6, Score 50.1 vs 54.6. Significant gaps favour Jev on civil_comments (81.0 vs 70.3), PAWS, MASSIVE-de, VitaminC; Nimble on SummEval-relevance, where Jev scores 35.0%. Jev ECE lower on 11 of 13; HelpSteer2 near floor (Jev 34.1%). English → German: Jev 87.4 → 86.9, Nimble 86.9 → 83.4 Jev overall; LOST Score tasks Vendor of the rival model; measures agreement with human annotation. Its repo also ships Jev-labelling scripts (distillation flag on the replica pages); "did not distill from Jev" unverified
Typed-decisions card (LocalLLaMA): 4 workflows, 400 test cases, 2,000 decisions, Jev run live 2026-09-18 (jev-latest → jev-1.13.0) Zero-shot general models vs fitted specialists Jev 0.727 (teacher ceiling 0.735, factor ceiling 0.704, majority 0.520, prior 0.470); ModernBERT-base fitted 0.646, MiniLM-L6 0.587. Jev p50 710 ms, $0.016 in all, zero errors. The card lists meraGPT Decider 1 (proprietary) at 0.768 Near the ceiling Gold is the mean of three samples from an unnamed ~4B teacher: agreement, not accuracy; the card says a score far above 0.75 learns the teacher's quirks. Replicas quote this 0.727. Distributions: Failure reports: confidence misread and calibration
JevBench v1.3.0, 534 frozen decisions (fstandhartinger) 48 ranked systems Jev 1.13.0 #1 at 74.4; SemIf 73.1; djev 73.0 Composite Mixes accuracy, calibration, speed, cost; hard items written by Opus 5 and GPT-5.6 Sol; self-hosted latency ×2 + 0.15 s by assumption. Superseded by v1.4.2 (next row)
JevBench v1.4.2 (artifact 2026-09-24, changelog 2026-09-25): the 534 frozen plus 308 sealed decisions 93 systems, 89 ranked decider-4b v2 (Mapika) 64.13, Jev 1.13.0 63.29, JevK5 v0.2.0 62.04, Cygnet 61.76, Hopper 59.43. Jev leads Intelligence (53.1 vs 49.4) and Calibration (76.3 vs 75.0); decider leads Speed and Cost. Jev public 86.6% vs sealed 36.7% (gap 49.9 pts; chance 29.3%). Under v1.3.0 scoring Jev would be 74.40, rank 3 on this board #2 New formula: harmonic mean, Speed and Cost gates below 50, penalty for a public-to-sealed gap over 25 pts. Jev's row is flagged API: TypeSafe's endpoint saw the sealed item text (no answers); offline rows did not. classifier.dev's fast tier (runs Jev) is listed unranked at 70.82. decider-4b v2's private training rows could not be audited
Earlier JevBench revisions, as recorded by entrants v1.2 (scored 2026-09-21, 36 ranked; reflex's record); v1.2.2 (18 ranked; jeff's record) v1.2: Jev 75.4 #1, SemIf 74.7, djev 74.3, reflex 4B 71.7 #5, reflex-27b 64.2 #22 (hard tier 75.9% vs Jev 74.1%; Calibration 86.2 vs 82.7). v1.2.2: classifier.dev 84.8 #1, Jev 75.3 #2, jeff 66.9 #9 #1 or #2 Every JevBench number depends on the version: quote it. classifier.dev became an unranked honorable mention in v1.2.4

Reading JevBench: label every figure with its version (v1.2, v1.2.2, v1.3.0, v1.4.2 above); the sealed set is aggregate-only and built to be hard, so a small gap there is not a proven ranking (author). Reading Jevals: the boards list Jev first on both noul and choice because they order by "wins" (top-2 finishes across score, accuracy, calibration gap, price and speed), not by Decision Score. The capture says PubMedQA is a statistical tie with Gemini 3.8 Flash (Jev lower on the point estimate) and Banking77 a real loss, so "behind Gemini Flash on both" overstates PubMedQA. Its 67.8 vs 74.1 on Banking77 is a Decision Score and only coincides with TypeSafe's workflow-eval 67.8% vs 74.1% (Workflow evals: how TypeSafe measures Jev); different measures. Jevals and ASSAY-001 both report Jev probabilities rounded to 2 decimals (sums of 0.99): unverified, not in HTTP API: POST /v1/systemone and GET /v1/models. Jevals is not openlayer-ai/jevals.

Open replicas measured against live Jev

Rows where the replica author ran Jev live; replicas that only copy Jev's numbers are on Open replicas and Jev-compatible servers.

Dataset (builder) Compared Numbers as published Won Caveat
1,600 labelled items over 8 datasets, plus JevBench public hard items (jeff, logan-markewich; not Alurith/jeff) GLiFormer 400M server vs Jev AG News 75.5% vs Jev 90.5%; p50 151 ms on a laptop vs Jev 129 ms; about $2.6 vs $15.6 per 1M single-question requests. Hard tier (111 public items): Jev led ambiguous 6/7 vs 0/7, long_policy 12/19 vs 2/19, but lost temporal_numeric 4/15 vs 6/15 Jev on accuracy; LOST date and number items Matches Jev's documented weakness (Jev 1.13 jaggedness: known failure modes; Jev 8/30 on the full 30 in v1.4.2). The replica is cheaper but weaker on reasoning (author)
JevBench public items, self-run during development (reflex, kshetrajna12) Frozen Qwen3.5-4B option-letter readout vs Jev Hard 0.685 / ECE 0.081 vs Jev 0.730 / 0.031. Official rows by version: above Jev Items were consulted during development. Negative result it publishes: every fine-tune and GEPA prompt optimisation won in-distribution and lost general judgement, so it ships the frozen model
60 tool-call cases; Kev transfer dev set (open-spark-jev, abhishek085) Qwen3.5-4B LoRA vs Jev figures from third-party per-row records (no live calls) 0.900 vs Jev 0.917 (Jev p50 421.6 ms with network vs 74.9 ms local); Kev dev 0.756 vs 0.857. An earlier 1.7B trailed Jev by 18-22 pts; vulnerable code at chance (0.502) Jev Its training method (RLCD) made results worse twice and was never shipped. Its JevBench 83.1 is a self-computed proxy on public items, not an official score
SemIf authored144 (choosekit, NotXf1le) Local Qwen3.8 27B Q4 on an RTX 4090 vs Jev Tie at 139/144; same pick on 136/144; p50 local 239 ms vs Jev 368 ms Tie One 144-case set (author)

How to read these

Sources

Links inline; raw captures listed in the frontmatter (private repo).