Agents: read the raw Markdown of this page, or start at llms.txt.
Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites
TL;DR On independent boards Jev is at or near the top on accuracy and calibration and loses rank on speed or cost to small open models: JevBench v1.4.2 puts it #2 (63.29) behind decider-4b v2 (64.13), with 86.6% on public items but 36.7% on 308 sealed ones. Always quote a JevBench number with its version; scoring changed at v1.4. Typed-decisions scores are agreement with an unnamed teacher, not accuracy. Split from Head-to-head benchmarks: Jev on public datasets and suites on 2026-09-25.
Numbers as each operator published them, on jev-1.13.0. None of these operators is TypeSafe; TypeSafe publishes no benchmark table (Blog: Lies, Damned Lies, and Benchmarks) and its own board is Workflow evals: how TypeSafe measures Jev. Vendors of rival models are labelled.
Independent eval boards
| Board (operator) | Task | Numbers as published | Jev placed | Caveat |
|---|---|---|---|---|
| Jevals.com noul board (anonymous operator; method) | PubMedQA, 300 items × 5 repeats, human labels | Decision Score (0 = guessing base rates) Jev 69.0 vs Gemini 3.8 Flash 73.0; accuracy 91.3% vs 92.5%; $0.029 vs $0.8 per 1,000; p95 653 ms vs 5.5 s. At 95% accuracy Jev can act alone on 86% vs 79% | Tied first: paired difference −2.5 to +10.4 includes 0 | Operator anonymous; the site showed a job-board ad (our 2026-09-24 sweep note, not in the text capture); no frontier LLM listed (budget $25). LLM probabilities are verbalized, not logprobs |
| Jevals.com choice board | Banking77, 77 intents, 300 × 5 | Jev 67.8 vs Gemini 3.8 Flash 74.1 vs GLM-5.3 66.8; accuracy 79.7% vs 84.6%; $0.043 vs $1.4 per 1,000; 41.5% of Jev answers at confidence 1.00, ECE 9.8 vs 3.2 points; 10.3% of Jev picks change when options are reordered | LOST to Gemini (significant); tied 2nd with GLM-5.3 | Options shuffled identically for every system; 2.7% of Jev picks change on an identical rerun |
| Jevals.com score board | HelpSteer2 helpfulness, 5 levels | Jev 9.2, GLM-5.3 7.8, Gemini 4.6: none beats guessing the label base rates; Jev accuracy 41.3%, ECE 19.7 | Nobody wins | No model reaches 95% accuracy at any gate; keep a person in the loop |
| Bespoke Nimble public suite (bespokelabsai), 13 human-labelled public subsets, 3,880 records, Jev run live | Nimble-9B vs Jev 1.13.0, identical records and wording | Macro 76.0% vs 74.8%. Noul 84.6 vs 80.2, Choice 82.9 vs 81.6, Score 50.1 vs 54.6. Significant gaps favour Jev on civil_comments (81.0 vs 70.3), PAWS, MASSIVE-de, VitaminC; Nimble on SummEval-relevance, where Jev scores 35.0%. Jev ECE lower on 11 of 13; HelpSteer2 near floor (Jev 34.1%). English → German: Jev 87.4 → 86.9, Nimble 86.9 → 83.4 | Jev overall; LOST Score tasks | Vendor of the rival model; measures agreement with human annotation. Its repo also ships Jev-labelling scripts (distillation flag on the replica pages); "did not distill from Jev" unverified |
Typed-decisions card (LocalLLaMA): 4 workflows, 400 test cases, 2,000 decisions, Jev run live 2026-09-18 (jev-latest → jev-1.13.0) |
Zero-shot general models vs fitted specialists | Jev 0.727 (teacher ceiling 0.735, factor ceiling 0.704, majority 0.520, prior 0.470); ModernBERT-base fitted 0.646, MiniLM-L6 0.587. Jev p50 710 ms, $0.016 in all, zero errors. The card lists meraGPT Decider 1 (proprietary) at 0.768 | Near the ceiling | Gold is the mean of three samples from an unnamed ~4B teacher: agreement, not accuracy; the card says a score far above 0.75 learns the teacher's quirks. Replicas quote this 0.727. Distributions: Failure reports: confidence misread and calibration |
| JevBench v1.3.0, 534 frozen decisions (fstandhartinger) | 48 ranked systems | Jev 1.13.0 #1 at 74.4; SemIf 73.1; djev 73.0 | Composite | Mixes accuracy, calibration, speed, cost; hard items written by Opus 5 and GPT-5.6 Sol; self-hosted latency ×2 + 0.15 s by assumption. Superseded by v1.4.2 (next row) |
| JevBench v1.4.2 (artifact 2026-09-24, changelog 2026-09-25): the 534 frozen plus 308 sealed decisions | 93 systems, 89 ranked | decider-4b v2 (Mapika) 64.13, Jev 1.13.0 63.29, JevK5 v0.2.0 62.04, Cygnet 61.76, Hopper 59.43. Jev leads Intelligence (53.1 vs 49.4) and Calibration (76.3 vs 75.0); decider leads Speed and Cost. Jev public 86.6% vs sealed 36.7% (gap 49.9 pts; chance 29.3%). Under v1.3.0 scoring Jev would be 74.40, rank 3 on this board | #2 | New formula: harmonic mean, Speed and Cost gates below 50, penalty for a public-to-sealed gap over 25 pts. Jev's row is flagged API: TypeSafe's endpoint saw the sealed item text (no answers); offline rows did not. classifier.dev's fast tier (runs Jev) is listed unranked at 70.82. decider-4b v2's private training rows could not be audited |
| Earlier JevBench revisions, as recorded by entrants | v1.2 (scored 2026-09-21, 36 ranked; reflex's record); v1.2.2 (18 ranked; jeff's record) | v1.2: Jev 75.4 #1, SemIf 74.7, djev 74.3, reflex 4B 71.7 #5, reflex-27b 64.2 #22 (hard tier 75.9% vs Jev 74.1%; Calibration 86.2 vs 82.7). v1.2.2: classifier.dev 84.8 #1, Jev 75.3 #2, jeff 66.9 #9 | #1 or #2 | Every JevBench number depends on the version: quote it. classifier.dev became an unranked honorable mention in v1.2.4 |
Reading JevBench: label every figure with its version (v1.2, v1.2.2, v1.3.0, v1.4.2 above); the sealed set is aggregate-only and built to be hard, so a small gap there is not a proven ranking (author). Reading Jevals: the boards list Jev first on both noul and choice because they order by "wins" (top-2 finishes across score, accuracy, calibration gap, price and speed), not by Decision Score. The capture says PubMedQA is a statistical tie with Gemini 3.8 Flash (Jev lower on the point estimate) and Banking77 a real loss, so "behind Gemini Flash on both" overstates PubMedQA. Its 67.8 vs 74.1 on Banking77 is a Decision Score and only coincides with TypeSafe's workflow-eval 67.8% vs 74.1% (Workflow evals: how TypeSafe measures Jev); different measures. Jevals and ASSAY-001 both report Jev probabilities rounded to 2 decimals (sums of 0.99): unverified, not in HTTP API: POST /v1/systemone and GET /v1/models. Jevals is not openlayer-ai/jevals.
Open replicas measured against live Jev
Rows where the replica author ran Jev live; replicas that only copy Jev's numbers are on Open replicas and Jev-compatible servers.
| Dataset (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| 1,600 labelled items over 8 datasets, plus JevBench public hard items (jeff, logan-markewich; not Alurith/jeff) | GLiFormer 400M server vs Jev | AG News 75.5% vs Jev 90.5%; p50 151 ms on a laptop vs Jev 129 ms; about $2.6 vs $15.6 per 1M single-question requests. Hard tier (111 public items): Jev led ambiguous 6/7 vs 0/7, long_policy 12/19 vs 2/19, but lost temporal_numeric 4/15 vs 6/15 | Jev on accuracy; LOST date and number items | Matches Jev's documented weakness (Jev 1.13 jaggedness: known failure modes; Jev 8/30 on the full 30 in v1.4.2). The replica is cheaper but weaker on reasoning (author) |
| JevBench public items, self-run during development (reflex, kshetrajna12) | Frozen Qwen3.5-4B option-letter readout vs Jev | Hard 0.685 / ECE 0.081 vs Jev 0.730 / 0.031. Official rows by version: above | Jev | Items were consulted during development. Negative result it publishes: every fine-tune and GEPA prompt optimisation won in-distribution and lost general judgement, so it ships the frozen model |
| 60 tool-call cases; Kev transfer dev set (open-spark-jev, abhishek085) | Qwen3.5-4B LoRA vs Jev figures from third-party per-row records (no live calls) | 0.900 vs Jev 0.917 (Jev p50 421.6 ms with network vs 74.9 ms local); Kev dev 0.756 vs 0.857. An earlier 1.7B trailed Jev by 18-22 pts; vulnerable code at chance (0.502) | Jev | Its training method (RLCD) made results worse twice and was never shipped. Its JevBench 83.1 is a self-computed proxy on public items, not an official score |
| SemIf authored144 (choosekit, NotXf1le) | Local Qwen3.8 27B Q4 on an RTX 4090 vs Jev | Tie at 139/144; same pick on 136/144; p50 local 239 ms vs Jev 368 ms | Tie | One 144-case set (author) |
How to read these
- Versions. JevBench v1.2 to v1.3.0 used a geometric mean of four axes on 534 frozen decisions; v1.4 added 308 sealed decisions, a harmonic mean, Speed and Cost gates and a generalisation penalty. Jev's score fell from 74.4 (v1.3.0, #1) to 63.29 (v1.4.2, #2) with the same
jev-1.13.0measurements plus the sealed run: the change is the formula and the sealed set. Under v1.3.0 scoring it would rank 3rd among v1.4.2's entrants. - Sealed vs public. Public-to-sealed gaps in v1.4.2 ranged from 0.3 to 57.5 points across rows (our count from the capture); Jev's was 49.9, decider-4b v2's 48.8. The public half can be trained on or selected against (the author says so).
- Exposure. API rows (Jev, classifier.dev and some others) sent sealed item text to the operator's endpoint; the flag reports exposure, not misuse.
- Teacher-labelled sets (typed-decisions) cap at teacher self-agreement; human-labelled suites (Nimble, Jevals, PubMedQA) measure agreement with annotators.
- Per-task public-set results: Head-to-head benchmarks: Jev on public datasets and suites; calibration: Failure reports: confidence misread and calibration.
Related
- Head-to-head benchmarks: Jev on public datasets and suites, Head-to-head: Jev against other models and methods, Head-to-head: Jev inside agents, routers and tool gates, Cost ledger: published cost per Jev decision
- Open replicas and Jev-compatible servers: the replicas themselves; Workflow evals: how TypeSafe measures Jev: TypeSafe's own board
Sources
Links inline; raw captures listed in the frontmatter (private repo).