---
title: "Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites"
type: community
source_tier: community
tags: [head-to-head, benchmarks, eval-boards, replicas, community]
created: 2026-09-25
updated: 2026-09-25
confidence: medium
sources:
  - raw/community/jevals-com-home.md
  - raw/community/jevals-com-methodology.md
  - raw/community/jevals-com-noul.md
  - raw/community/jevals-com-choice.md
  - raw/community/jevals-com-score.md
  - raw/community/jevals-com-notes-2026-09-18.md
  - docs/sweep/2026-09-24-awesome-jev/part3.md
  - raw/x-repos/fstandhartinger__jevbench.md
  - raw/x-repos/fstandhartinger__jevbench__readme-v1-4-2.md
  - raw/x-repos/fstandhartinger__jevbench__results-v1-4-2.md
  - raw/x-repos/fstandhartinger__jevbench__docs-method-v1-4.md
  - raw/x-repos/fstandhartinger__jevbench__docs-release-v1-4-0.md
  - raw/x-repos/fstandhartinger__jevbench__docs-release-v1-4-1.md
  - raw/x-repos/fstandhartinger__jevbench__docs-release-v1-4-2.md
  - raw/x-repos/fstandhartinger__jevbench__changelog.md
  - raw/x-repos/kshetrajna12__reflex.md
  - raw/x-repos/kshetrajna12__reflex__docs-results-jevbench-official.md
  - raw/x-repos/kshetrajna12__reflex__docs-results-readme.md
  - raw/x-repos/logan-markewich__jeff.md
  - raw/x-repos/logan-markewich__jeff__bench-results.md
  - raw/x-repos/bespokelabsai__nimble.md
  - raw/x-repos/bespokelabsai__nimble__docs-public-benchmarks.md
  - raw/community/huggingface-co-datasets-localllama-typed-decisions.md
  - raw/x-repos/abhishek085__open-spark-jev.md
  - raw/x-repos/abhishek085__open-spark-jev__docs-benchmarks.md
  - raw/x-repos/abhishek085__open-spark-jev__model-card.md
  - raw/x-repos/NotXf1le__choosekit.md
jev_version: "jev-1.13.0"
summary: "Independent boards and suites that score Jev: JevBench by version (v1.2 to v1.4.2), Jevals.com, the typed-decisions card, Bespoke Nimble's human-labelled suite, and replicas run against live Jev."
---

# Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites

> **TL;DR** On independent boards Jev is at or near the top on accuracy and calibration and loses rank on speed or cost to small open models: JevBench v1.4.2 puts it #2 (63.29) behind decider-4b v2 (64.13), with 86.6% on public items but 36.7% on 308 sealed ones. Always quote a JevBench number with its version; scoring changed at v1.4. Typed-decisions scores are agreement with an unnamed teacher, not accuracy. Split from [[ideas/head-to-head-benchmarks]] on 2026-09-25.

Numbers as each operator published them, on `jev-1.13.0`. None of these operators is TypeSafe; TypeSafe publishes no benchmark table ([[entities/blog-antibenchmaxxing]]) and its own board is [[concepts/workflow-evals]]. Vendors of rival models are labelled.

## Independent eval boards

| Board (operator) | Task | Numbers as published | Jev placed | Caveat |
|---|---|---|---|---|
| [Jevals.com](https://jevals.com/) noul board (anonymous operator; [method](https://jevals.com/methodology/)) | PubMedQA, 300 items × 5 repeats, human labels | Decision Score (0 = guessing base rates) Jev 69.0 vs Gemini 3.8 Flash 73.0; accuracy 91.3% vs 92.5%; $0.029 vs $0.8 per 1,000; p95 653 ms vs 5.5 s. At 95% accuracy Jev can act alone on 86% vs 79% | **Tied** first: paired difference −2.5 to +10.4 includes 0 | Operator anonymous; the site showed a job-board ad (our 2026-09-24 sweep note, not in the text capture); no frontier LLM listed (budget $25). LLM probabilities are verbalized, not logprobs |
| Jevals.com [choice board](https://jevals.com/choice/) | Banking77, 77 intents, 300 × 5 | Jev 67.8 vs Gemini 3.8 Flash 74.1 vs GLM-5.3 66.8; accuracy 79.7% vs 84.6%; $0.043 vs $1.4 per 1,000; 41.5% of Jev answers at confidence 1.00, ECE 9.8 vs 3.2 points; 10.3% of Jev picks change when options are reordered | **LOST** to Gemini (significant); tied 2nd with GLM-5.3 | Options shuffled identically for every system; 2.7% of Jev picks change on an identical rerun |
| Jevals.com [score board](https://jevals.com/score/) | HelpSteer2 helpfulness, 5 levels | Jev 9.2, GLM-5.3 7.8, Gemini 4.6: none beats guessing the label base rates; Jev accuracy 41.3%, ECE 19.7 | Nobody wins | No model reaches 95% accuracy at any gate; keep a person in the loop |
| [Bespoke Nimble public suite](https://github.com/bespokelabsai/nimble) (bespokelabsai), 13 human-labelled public subsets, 3,880 records, Jev run live | Nimble-9B vs Jev 1.13.0, identical records and wording | Macro **76.0% vs 74.8%**. Noul 84.6 vs 80.2, Choice 82.9 vs 81.6, Score 50.1 vs **54.6**. Significant gaps favour Jev on civil_comments (81.0 vs 70.3), PAWS, MASSIVE-de, VitaminC; Nimble on SummEval-relevance, where Jev scores **35.0%**. Jev ECE lower on 11 of 13; HelpSteer2 near floor (Jev 34.1%). English → German: Jev 87.4 → 86.9, Nimble 86.9 → 83.4 | Jev overall; **LOST** Score tasks | **Vendor** of the rival model; measures agreement with human annotation. Its repo also ships Jev-labelling scripts (distillation flag on the replica pages); "did not distill from Jev" `unverified` |
| [Typed-decisions card](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) (LocalLLaMA): 4 workflows, 400 test cases, 2,000 decisions, Jev run live 2026-09-18 (`jev-latest` → `jev-1.13.0`) | Zero-shot general models vs fitted specialists | Jev **0.727** (teacher ceiling 0.735, factor ceiling 0.704, majority 0.520, prior 0.470); ModernBERT-base fitted 0.646, MiniLM-L6 0.587. Jev p50 710 ms, $0.016 in all, zero errors. The card lists meraGPT Decider 1 (proprietary) at 0.768 | Near the ceiling | Gold is the mean of three samples from an unnamed ~4B teacher: agreement, not accuracy; the card says a score far above 0.75 learns the teacher's quirks. Replicas quote this 0.727. Distributions: [[ideas/failure-reports-confidence]] |
| JevBench **v1.3.0**, 534 frozen decisions ([fstandhartinger](https://github.com/fstandhartinger/jevbench)) | 48 ranked systems | Jev 1.13.0 #1 at 74.4; SemIf 73.1; djev 73.0 | Composite | Mixes accuracy, calibration, speed, cost; hard items written by Opus 5 and GPT-5.6 Sol; self-hosted latency ×2 + 0.15 s by assumption. Superseded by v1.4.2 (next row) |
| JevBench **v1.4.2** (artifact 2026-09-24, changelog 2026-09-25): the 534 frozen plus 308 sealed decisions | 93 systems, 89 ranked | decider-4b v2 (Mapika) 64.13, **Jev 1.13.0 63.29**, JevK5 v0.2.0 62.04, Cygnet 61.76, Hopper 59.43. Jev leads Intelligence (53.1 vs 49.4) and Calibration (76.3 vs 75.0); decider leads Speed and Cost. Jev public 86.6% vs sealed **36.7%** (gap 49.9 pts; chance 29.3%). Under v1.3.0 scoring Jev would be 74.40, rank 3 on this board | #2 | New formula: harmonic mean, Speed and Cost gates below 50, penalty for a public-to-sealed gap over 25 pts. Jev's row is flagged **API**: TypeSafe's endpoint saw the sealed item text (no answers); offline rows did not. classifier.dev's fast tier (runs Jev) is listed unranked at 70.82. decider-4b v2's private training rows could not be audited |
| Earlier JevBench revisions, as recorded by entrants | v1.2 (scored 2026-09-21, 36 ranked; [reflex's record](https://github.com/kshetrajna12/reflex)); v1.2.2 (18 ranked; [jeff's record](https://github.com/logan-markewich/jeff)) | v1.2: Jev 75.4 #1, SemIf 74.7, djev 74.3, reflex 4B 71.7 #5, reflex-27b 64.2 #22 (hard tier 75.9% vs Jev 74.1%; Calibration 86.2 vs 82.7). v1.2.2: classifier.dev 84.8 #1, Jev 75.3 #2, jeff 66.9 #9 | #1 or #2 | Every JevBench number depends on the version: quote it. classifier.dev became an unranked honorable mention in v1.2.4 |

Reading JevBench: label every figure with its version (v1.2, v1.2.2, v1.3.0, v1.4.2 above); the sealed set is aggregate-only and built to be hard, so a small gap there is not a proven ranking (author). Reading Jevals: the boards list Jev **first on both noul and choice** because they order by "wins" (top-2 finishes across score, accuracy, calibration gap, price and speed), not by Decision Score. The capture says PubMedQA is a statistical **tie** with Gemini 3.8 Flash (Jev lower on the point estimate) and Banking77 a real **loss**, so "behind Gemini Flash on both" overstates PubMedQA. Its 67.8 vs 74.1 on Banking77 is a Decision Score and only coincides with TypeSafe's workflow-eval 67.8% vs 74.1% ([[concepts/workflow-evals]]); different measures. Jevals and ASSAY-001 both report Jev probabilities rounded to 2 decimals (sums of 0.99): `unverified`, not in [[reference/http-api]]. Jevals is not openlayer-ai/jevals.

## Open replicas measured against live Jev

Rows where the replica author ran Jev live; replicas that only copy Jev's numbers are on [[ideas/open-replicas]].

| Dataset (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| 1,600 labelled items over 8 datasets, plus JevBench public hard items ([jeff](https://github.com/logan-markewich/jeff), logan-markewich; not Alurith/jeff) | GLiFormer 400M server vs Jev | AG News 75.5% vs Jev **90.5%**; p50 151 ms on a laptop vs Jev 129 ms; about $2.6 vs $15.6 per 1M single-question requests. Hard tier (111 public items): Jev led ambiguous 6/7 vs 0/7, long_policy 12/19 vs 2/19, but **lost temporal_numeric 4/15 vs 6/15** | Jev on accuracy; **LOST** date and number items | Matches Jev's documented weakness ([[concepts/jaggedness-jev-1-13]]; Jev 8/30 on the full 30 in v1.4.2). The replica is cheaper but weaker on reasoning (author) |
| JevBench public items, self-run during development ([reflex](https://github.com/kshetrajna12/reflex), kshetrajna12) | Frozen Qwen3.5-4B option-letter readout vs Jev | Hard 0.685 / ECE 0.081 vs Jev 0.730 / 0.031. Official rows by version: above | Jev | Items were consulted during development. Negative result it publishes: every fine-tune and GEPA prompt optimisation won in-distribution and lost general judgement, so it ships the frozen model |
| 60 tool-call cases; Kev transfer dev set ([open-spark-jev](https://github.com/abhishek085/open-spark-jev), abhishek085) | Qwen3.5-4B LoRA vs Jev figures from third-party per-row records (no live calls) | 0.900 vs Jev **0.917** (Jev p50 421.6 ms with network vs 74.9 ms local); Kev dev 0.756 vs **0.857**. An earlier 1.7B trailed Jev by 18-22 pts; vulnerable code at chance (0.502) | Jev | Its training method (RLCD) made results worse twice and was never shipped. Its JevBench 83.1 is a self-computed proxy on public items, not an official score |
| SemIf authored144 ([choosekit](https://github.com/NotXf1le/choosekit), NotXf1le) | Local Qwen3.8 27B Q4 on an RTX 4090 vs Jev | Tie at 139/144; same pick on 136/144; p50 local 239 ms vs Jev 368 ms | Tie | One 144-case set (author) |

## How to read these

- **Versions.** JevBench v1.2 to v1.3.0 used a geometric mean of four axes on 534 frozen decisions; v1.4 added 308 sealed decisions, a harmonic mean, Speed and Cost gates and a generalisation penalty. Jev's score fell from 74.4 (v1.3.0, #1) to 63.29 (v1.4.2, #2) with the same `jev-1.13.0` measurements plus the sealed run: the change is the formula and the sealed set. Under v1.3.0 scoring it would rank 3rd among v1.4.2's entrants.
- **Sealed vs public.** Public-to-sealed gaps in v1.4.2 ranged from 0.3 to 57.5 points across rows (our count from the capture); Jev's was 49.9, decider-4b v2's 48.8. The public half can be trained on or selected against (the author says so).
- **Exposure.** API rows (Jev, classifier.dev and some others) sent sealed item text to the operator's endpoint; the flag reports exposure, not misuse.
- **Teacher-labelled sets** (typed-decisions) cap at teacher self-agreement; human-labelled suites (Nimble, Jevals, PubMedQA) measure agreement with annotators.
- Per-task public-set results: [[ideas/head-to-head-benchmarks]]; calibration: [[ideas/failure-reports-confidence]].

## Related

- [[ideas/head-to-head-benchmarks]], [[ideas/head-to-head]], [[ideas/head-to-head-agents]], [[ideas/cost-ledger]]
- [[ideas/open-replicas]]: the replicas themselves; [[concepts/workflow-evals]]: TypeSafe's own board

## Sources

Links inline; raw captures listed in the frontmatter (private repo).
