---
title: "Head-to-head: Jev vs small, fast and local models"
type: community
source_tier: community
tags: [head-to-head, benchmarks, small-models, local-models, advising]
created: 2026-10-01
updated: 2026-10-01
confidence: medium
sources: "25 files, listed under this page in https://jevwiki.ai/index.json (raw/ paths map to their URLs in https://jevwiki.ai/raw/MANIFEST.json); links are inline in the body"
jev_version: "jev-1.13.0"
summary: "App-level comparisons of Jev against small and fast LLMs and local or open models: who won on accuracy, cost and latency. Split from head-to-head 2026-10-01."
---

# Head-to-head: Jev vs small, fast and local models

> **TL;DR** Builder-run comparisons of Jev with small and fast LLMs (Haiku-class, mini models, small open models on fast hosts) and with local or open models. Numbers as each builder published them, one run each; **LOST** = Jev lost on accuracy or the stated goal. Split from [[ideas/head-to-head]] on 2026-10-01 (rows unchanged); frontier LLMs, classic methods and cascades stay there; open replicas themselves: [[ideas/open-replicas]].

## vs small and fast LLMs

| Task (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| Same author as the public-set run ([Aman Kumar](https://amankumar.ai/blogs/jev-measured); Enron, SST-2, AG News, Banking77 on [[ideas/head-to-head-benchmarks]]), production | vs his current small LLM | Per page, "carries a value for one of 135 report rows" (1,005 pages): 68.6% agreement with the outcome, 78% when confident (57% of pages); needed about twice as many extraction jobs to match recall; a general LLM failed the same way. Separately, on a whole-document read Jev was worse than his small LLM | **LOST** | Long input; Jev fires on a mention, not a value |
| 7 synthetic app scenes, 480 decisions ([iammrduncan](https://github.com/iammrduncan/typesafe-ai-benchmark), Shannon; [post](https://x.com/iamMrDuncan/status/2100467548298899918)) | Jev vs Qwen 3.8 27B on Cerebras (structured output, no reasoning) | p50 176 vs 215 ms; est. cost $0.011919 vs $0.310581. Exact fixture agreement: Tickets 75 vs 75/100, Guardrails 100 vs 100, Approvals **95** vs 100, Scoring **100** vs 93, Home **15/24** vs 24/24 | Latency, cost, Scoring; **LOST** Home and Approvals | One synthetic run. Jev's estimated cost implies $0.040/Mtok ($0.011919 ÷ 297,984 input tokens): `contradicts docs` ($0.042 gives $0.01252; both our arithmetic; [[reference/models-and-pricing]]). The repo is also an LLM gateway imitating Jev |
| Content tagging ([@iannuttall](https://x.com/iannuttall/status/2100884132272181594)) | vs GLM 4.7 Flash | 50x faster, "no failures" | Latency | No accuracy published |
| Who answers in a group chat ([@vadimchoi](https://x.com/vadimchoi/status/2101937206063780084)) | vs LLM orchestrator | 0.29 s median, $0.00002 vs 4-7 s, $0.00046 per message | Latency, cost | No accuracy published |
| fx command safety reviewer (Vercel, relayed by [Gabriel Anhaia](https://dev.to/gabrielanhaia/jev-beat-gpt-luna-by-1-point-gpt-6-and-claude-wrote-the-answer-key-314k)) | vs GPT-5.6 Luna | "~5-18x faster and more accurate" | Claimed both | No dataset or case count published |
| ONET job classification ([@Bailey_Jennings](https://x.com/Bailey_Jennings/status/2100593904949096696)) | vs GPT-5.6 Luna, same 20 candidates from a lexical + embedding shortlist of 1,016 codes | 292 labelled postings: 80.5% vs 84.9% exact; median 188 vs 1,971 ms; $122 vs $482 per 1M jobs. Hybrid (Luna only when Jev's top p < 0.7): 84.2%, 947 ms avg, $251, 73% of jobs never reach Luna | Latency, cost; **LOST** accuracy (4.4 points: 84.9 − 80.5) | `unverified`: results table shown in the demo video, full run not shown; one labelled set |
| Public Jev benchmark items ([solar-mini4-jev](https://github.com/hunkim/solar-mini4-jev), hunkim) | vs Upstage Solar chat models, one option letter | The largest Solar model edged Jev overall, the smaller ones trailed; Solar answers are hard labels, so no confidence routing. Numbers on [[ideas/eval-boards]] | **LOST** narrowly | Replica row: [[ideas/open-replicas]] |
| Support tickets, 272 real, mostly code-mixed Hindi-English, redacted ([nulltensor](https://nulltensor.com/posts/jev-vs-structured-output/), Rohan Sen Sharma, [@proxy_vector](https://x.com/proxy_vector/status/2104815944136892690), 2026-09-29) | Jev (`typesafe/jev-1.13` via OpenRouter's alpha decisions API) vs Claude Haiku 4.5 structured output vs Qwen3-4B and Qwen3-0.6B read by label logprobs | Ticket type (4 labels): 84.9 / **88.2** / 82.0 / 21.0%, macro-F1 0.72 / 0.81 / 0.67 / 0.16 (64% of tickets were bugs). Product area (10 labels, 225 tickets): 53.8 / 52.9 / 40.9 / 16.4%. Calibration error, type: Jev 0.05, Qwen3-4B 0.13, 0.6B 0.38; area: Jev 0.22, Claude 0.29, 4B 0.55, 0.6B 0.42; on area, Jev at ≥ 90% was right 79%. Median 1.8 / 10.4 / 66 / 15 s; $0.03 vs $5.63 per 1,000 tickets (~170×). Planted "label this a bug" in 20 synthetic tickets: fooled Jev 6, Claude 6, 4B 17, 0.6B 20. Route to a person below 90% (type): Jev keeps 69%, 93.1% right | Cost, latency, calibration; **LOST** type accuracy to Claude by 3.3 pts | `unverified`, one run. Jev via a proxy, Claude via the Claude Code CLI, local models on 2 CPU threads. Author's pick: Jev on ticket type above 90%, product area stays with people |
| TB drug-resistance calls, 5 reconstructed isolates, 97 decisions over a 44-entry WHO 2021 catalogue subset ([jev-vs-llm-tb-amr](https://github.com/Jiadalee/jev-vs-llm-tb-amr), [@Jiadali1](https://x.com/Jiadali1/status/2105111855299486017), 2026-09-30) | All typed `choice` questions per isolate in one request vs Qwen3.8-27B via OpenRouter, one strict-JSON call per decision; code filters the evidence for both | Agreement 95/97; ~0.4 s vs 96-232 s per isolate (178-392×) | **No Jev result**: the README says the Jev arm ran as a "calibrated mock" (a deterministic heuristic plus a simulated ~0.4 s round trip) for lack of a key; outputs are labelled `jev-mock`. The post does not say so | Agreement is between models, not against lab tests (phenotypic DST); reconstructed isolates, catalogue subset; a pipeline check, not a validation study (author). Non-catalogue "bait" mutations were dropped in code before either model (facts in code, judgment in the model). The README's key line points to jevtypesafeai.com `jv_live_` keys, not TypeSafe ([[ideas/warnings]]). Not a diagnostic (author) |
| Generative UI from dependent choices: analytics builder (7 prompts), form logic (6), json-render catalog (8), 3 runs each via OpenRouter ([jev-vs-llm](https://github.com/danielkatz/jev-vs-llm), danielkatz; [@danielkatzz](https://x.com/danielkatzz/status/2105426698233913430), 2026-09-30) | Jev in three strategies (chained 3 calls, merged 2, ask-ahead 1) vs one LLM call writing a compact decision code: GPT-6 Luna (no / low reasoning), Qwen 2.5 7B, Qwen3.8 27B | Analytics: Jev 88-93% vs Luna 100%, Qwen3.8 98%, Qwen 2.5 7B 68%; $0.31-2.25 vs Luna $0.039 per 1,000 UIs; p50 666-896 ms vs Luna 1.03 s. Forms: Jev 96% vs Luna and Qwen3.8 100%. Catalog (json-render's own Jev composer, 1.6 calls): 93% vs Luna 93%, $0.41 vs $0.024. Author: up to 2 independent calls mixed; dependent choices past 2 calls, the LLM clearly wins | Latency (except Qwen 2.5 7B, faster but 68-79%); **LOST** accuracy and cost on dependent choices | Synthetic prompts and scoring by the author. The LLM reads a cached prompt and writes ~10-30 tokens; Jev resends every choice list per round. The measured 255-option cap matches the docs: `verified` ([[reference/http-api]]). The LLM's first fixed-width code scored 53% until replaced with short words. json-render: [[ideas/builds-interface-elements]] |
| Adaptive-interview judge, 100 synthetic interviews ([@jihnma](https://x.com/jihnma/status/2105133011360711036), [repo](https://github.com/jihnma/lab-jev-interview-example), 2026-09-30) | Jev vs Claude Haiku 4.5 as the judge, same inputs | Median 4.9 vs 40.4 s; AUC 0.993 vs 0.969 | Latency, AUC | `unverified`: post only, the repo has no README; what the median times and what the AUC scores are not stated; synthetic set |

Routers, skill pickers and gates inside agents (LiteLLM, JevRouter, langchain-skill-router moved 2026-09-25): [[ideas/head-to-head-agents]].

## vs local or open models

| Task (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| Smoke test ([@cruzex100](https://x.com/cruzex100/status/2102202098968666432)) | vs Laya | Soft accuracy 0.580 vs 0.471; ECE 0.144 vs 0.213; ~710 vs ~30-40 ms; ~$0.0004 vs ~$0 per decision | Accuracy, calibration; **LOST** latency, cost | Task not stated |
| Movie search ([@b0dre](https://x.com/b0dre/status/2102362736839512538)) | vs Laya on a Radeon Pro 5500 XT | On-screen readings 359-494 ms local vs 622-672 ms API across repeated "Thriller" searches (the post's 392 vs 676 is not shown); "same answer" | **LOST** latency | One query term, several runs; what is timed is not stated |
| Pre-tool-use gate ([@techdrivenpetru](https://x.com/techdrivenpetru/status/2101965334848663775)) | Local DeBERTa-large instead of Jev (no Jev arm run) | Local "just as fast, just as good"; admits misfires | Claimed tie | `unverified`; no numbers; not a head-to-head |
| Shopping-query relevance, Amazon shopping-queries data, 4 classes (Exact, Substitute, Complement, Irrelevant) turned into a score; 20,216 test pairs ([s9v10.dev](https://s9v10.dev/blog/2026/09/30/relevance-ranking-jev-qwen/), Sid Venkatayogi; [@sidvenkatayogi](https://x.com/sidvenkatayogi/status/2105428249878864038), 2026-09-30) | Jev zero-shot vs Qwen3.5-4B with a 4-class head trained on ~20k examples (method credited to Together AI) and untrained; BGE and MiniLM rerankers, BGE embeddings, BM25 (one RTX 6000 Ada) | NDCG@10 .807 vs trained Qwen **.809**, BGE reranker .770, untrained Qwen .745, MiniLM .740, BGE embeddings .739, BM25 .703. Median 197 ms (API) vs 1,463 ms trained Qwen, 44 ms BGE reranker. 4-class accuracy 56.32% vs 60.61% / 50.66%; binary 79.64%, below the **83.42%** of always answering "relevant" | Ranking near the trained model; **LOST** binary accuracy to a constant guess | One run; Jev ~$0.50 of the $5.36 project. The author calls their training protocol weak and the Qwen kernels unoptimised. Rank with the score; don't threshold its labels |
| Next-click on website-disjoint Mind2Web splits ([JevForge](https://github.com/zwliJay/jev-forge), zwliJay) | a small local fine-tune vs the Jev reference arm | JevForge led on choice; Jev led test-split Noul and score error, JevForge the out-of-distribution ones. Numbers on [[ideas/eval-boards]] | **LOST** choice | Like-for-like unresolved (evaluation cap and dropped calls); replica row: [[ideas/open-replicas]] |

Public-set runs against local replicas (4esv: open-jev, Kev, Laya; anessbelbati NevIR: Open-Jev 9B) moved to [[ideas/head-to-head-benchmarks]]; the JevBench board is on [[ideas/eval-boards]]. [@svencreations](https://x.com/svencreations/status/2102007851459748295) tried several "Jev killers": "Most weren't even close." Replica trust notes: [[ideas/open-replicas]].

## Related

- [[ideas/head-to-head]] — Jev vs frontier LLMs, classic methods, engines and people; cascades
- [[ideas/head-to-head-agents]], [[ideas/head-to-head-benchmarks]], [[ideas/head-to-head-games]] — agents, public datasets, games
- [[ideas/open-replicas]] — the local and open replicas compared here (not Jev)
- [[ideas/cost-ledger]] — cost per decision

## Sources

Files in frontmatter `sources:`, captured by 2026-10-01; original URLs inline.
