Agents: read the raw Markdown of this page, or start at llms.txt.
Head-to-head: Jev vs small, fast and local models
TL;DR Builder-run comparisons of Jev with small and fast LLMs (Haiku-class, mini models, small open models on fast hosts) and with local or open models. Numbers as each builder published them, one run each; LOST = Jev lost on accuracy or the stated goal. Split from Head-to-head: Jev against other models and methods on 2026-10-01 (rows unchanged); frontier LLMs, classic methods and cascades stay there; open replicas themselves: Open replicas and Jev-compatible servers.
vs small and fast LLMs
| Task (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| Same author as the public-set run (Aman Kumar; Enron, SST-2, AG News, Banking77 on Head-to-head benchmarks: Jev on public datasets and suites), production | vs his current small LLM | Per page, "carries a value for one of 135 report rows" (1,005 pages): 68.6% agreement with the outcome, 78% when confident (57% of pages); needed about twice as many extraction jobs to match recall; a general LLM failed the same way. Separately, on a whole-document read Jev was worse than his small LLM | LOST | Long input; Jev fires on a mention, not a value |
| 7 synthetic app scenes, 480 decisions (iammrduncan, Shannon; post) | Jev vs Qwen 3.8 27B on Cerebras (structured output, no reasoning) | p50 176 vs 215 ms; est. cost $0.011919 vs $0.310581. Exact fixture agreement: Tickets 75 vs 75/100, Guardrails 100 vs 100, Approvals 95 vs 100, Scoring 100 vs 93, Home 15/24 vs 24/24 | Latency, cost, Scoring; LOST Home and Approvals | One synthetic run. Jev's estimated cost implies $0.040/Mtok ($0.011919 ÷ 297,984 input tokens): contradicts docs ($0.042 gives $0.01252; both our arithmetic; Models, aliases, pricing, rate limits, context). The repo is also an LLM gateway imitating Jev |
| Content tagging (@iannuttall) | vs GLM 4.7 Flash | 50x faster, "no failures" | Latency | No accuracy published |
| Who answers in a group chat (@vadimchoi) | vs LLM orchestrator | 0.29 s median, $0.00002 vs 4-7 s, $0.00046 per message | Latency, cost | No accuracy published |
| fx command safety reviewer (Vercel, relayed by Gabriel Anhaia) | vs GPT-5.6 Luna | "~5-18x faster and more accurate" | Claimed both | No dataset or case count published |
| ONET job classification (@Bailey_Jennings) | vs GPT-5.6 Luna, same 20 candidates from a lexical + embedding shortlist of 1,016 codes | 292 labelled postings: 80.5% vs 84.9% exact; median 188 vs 1,971 ms; $122 vs $482 per 1M jobs. Hybrid (Luna only when Jev's top p < 0.7): 84.2%, 947 ms avg, $251, 73% of jobs never reach Luna | Latency, cost; LOST accuracy (4.4 points: 84.9 − 80.5) | unverified: results table shown in the demo video, full run not shown; one labelled set |
| Public Jev benchmark items (solar-mini4-jev, hunkim) | vs Upstage Solar chat models, one option letter | The largest Solar model edged Jev overall, the smaller ones trailed; Solar answers are hard labels, so no confidence routing. Numbers on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites | LOST narrowly | Replica row: Open replicas and Jev-compatible servers |
| Support tickets, 272 real, mostly code-mixed Hindi-English, redacted (nulltensor, Rohan Sen Sharma, @proxy_vector, 2026-09-29) | Jev (typesafe/jev-1.13 via OpenRouter's alpha decisions API) vs Claude Haiku 4.5 structured output vs Qwen3-4B and Qwen3-0.6B read by label logprobs |
Ticket type (4 labels): 84.9 / 88.2 / 82.0 / 21.0%, macro-F1 0.72 / 0.81 / 0.67 / 0.16 (64% of tickets were bugs). Product area (10 labels, 225 tickets): 53.8 / 52.9 / 40.9 / 16.4%. Calibration error, type: Jev 0.05, Qwen3-4B 0.13, 0.6B 0.38; area: Jev 0.22, Claude 0.29, 4B 0.55, 0.6B 0.42; on area, Jev at ≥ 90% was right 79%. Median 1.8 / 10.4 / 66 / 15 s; $0.03 vs $5.63 per 1,000 tickets (~170×). Planted "label this a bug" in 20 synthetic tickets: fooled Jev 6, Claude 6, 4B 17, 0.6B 20. Route to a person below 90% (type): Jev keeps 69%, 93.1% right | Cost, latency, calibration; LOST type accuracy to Claude by 3.3 pts | unverified, one run. Jev via a proxy, Claude via the Claude Code CLI, local models on 2 CPU threads. Author's pick: Jev on ticket type above 90%, product area stays with people |
| TB drug-resistance calls, 5 reconstructed isolates, 97 decisions over a 44-entry WHO 2021 catalogue subset (jev-vs-llm-tb-amr, @Jiadali1, 2026-09-30) | All typed choice questions per isolate in one request vs Qwen3.8-27B via OpenRouter, one strict-JSON call per decision; code filters the evidence for both |
Agreement 95/97; ~0.4 s vs 96-232 s per isolate (178-392×) | No Jev result: the README says the Jev arm ran as a "calibrated mock" (a deterministic heuristic plus a simulated ~0.4 s round trip) for lack of a key; outputs are labelled jev-mock. The post does not say so |
Agreement is between models, not against lab tests (phenotypic DST); reconstructed isolates, catalogue subset; a pipeline check, not a validation study (author). Non-catalogue "bait" mutations were dropped in code before either model (facts in code, judgment in the model). The README's key line points to jevtypesafeai.com jv_live_ keys, not TypeSafe (Warnings: not-Jev services, key safety, look-alikes and install names). Not a diagnostic (author) |
| Generative UI from dependent choices: analytics builder (7 prompts), form logic (6), json-render catalog (8), 3 runs each via OpenRouter (jev-vs-llm, danielkatz; @danielkatzz, 2026-09-30) | Jev in three strategies (chained 3 calls, merged 2, ask-ahead 1) vs one LLM call writing a compact decision code: GPT-6 Luna (no / low reasoning), Qwen 2.5 7B, Qwen3.8 27B | Analytics: Jev 88-93% vs Luna 100%, Qwen3.8 98%, Qwen 2.5 7B 68%; $0.31-2.25 vs Luna $0.039 per 1,000 UIs; p50 666-896 ms vs Luna 1.03 s. Forms: Jev 96% vs Luna and Qwen3.8 100%. Catalog (json-render's own Jev composer, 1.6 calls): 93% vs Luna 93%, $0.41 vs $0.024. Author: up to 2 independent calls mixed; dependent choices past 2 calls, the LLM clearly wins | Latency (except Qwen 2.5 7B, faster but 68-79%); LOST accuracy and cost on dependent choices | Synthetic prompts and scoring by the author. The LLM reads a cached prompt and writes ~10-30 tokens; Jev resends every choice list per round. The measured 255-option cap matches the docs: verified (HTTP API: POST /v1/systemone and GET /v1/models). The LLM's first fixed-width code scored 53% until replaced with short words. json-render: Builds: interface elements |
| Adaptive-interview judge, 100 synthetic interviews (@jihnma, repo, 2026-09-30) | Jev vs Claude Haiku 4.5 as the judge, same inputs | Median 4.9 vs 40.4 s; AUC 0.993 vs 0.969 | Latency, AUC | unverified: post only, the repo has no README; what the median times and what the AUC scores are not stated; synthetic set |
Routers, skill pickers and gates inside agents (LiteLLM, JevRouter, langchain-skill-router moved 2026-09-25): Head-to-head: Jev inside agents, routers and tool gates.
vs local or open models
| Task (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| Smoke test (@cruzex100) | vs Laya | Soft accuracy 0.580 vs 0.471; ECE 0.144 vs 0.213; ~710 vs ~30-40 ms; ~$0.0004 vs ~$0 per decision | Accuracy, calibration; LOST latency, cost | Task not stated |
| Movie search (@b0dre) | vs Laya on a Radeon Pro 5500 XT | On-screen readings 359-494 ms local vs 622-672 ms API across repeated "Thriller" searches (the post's 392 vs 676 is not shown); "same answer" | LOST latency | One query term, several runs; what is timed is not stated |
| Pre-tool-use gate (@techdrivenpetru) | Local DeBERTa-large instead of Jev (no Jev arm run) | Local "just as fast, just as good"; admits misfires | Claimed tie | unverified; no numbers; not a head-to-head |
| Shopping-query relevance, Amazon shopping-queries data, 4 classes (Exact, Substitute, Complement, Irrelevant) turned into a score; 20,216 test pairs (s9v10.dev, Sid Venkatayogi; @sidvenkatayogi, 2026-09-30) | Jev zero-shot vs Qwen3.5-4B with a 4-class head trained on ~20k examples (method credited to Together AI) and untrained; BGE and MiniLM rerankers, BGE embeddings, BM25 (one RTX 6000 Ada) | NDCG@10 .807 vs trained Qwen .809, BGE reranker .770, untrained Qwen .745, MiniLM .740, BGE embeddings .739, BM25 .703. Median 197 ms (API) vs 1,463 ms trained Qwen, 44 ms BGE reranker. 4-class accuracy 56.32% vs 60.61% / 50.66%; binary 79.64%, below the 83.42% of always answering "relevant" | Ranking near the trained model; LOST binary accuracy to a constant guess | One run; Jev ~$0.50 of the $5.36 project. The author calls their training protocol weak and the Qwen kernels unoptimised. Rank with the score; don't threshold its labels |
| Next-click on website-disjoint Mind2Web splits (JevForge, zwliJay) | a small local fine-tune vs the Jev reference arm | JevForge led on choice; Jev led test-split Noul and score error, JevForge the out-of-distribution ones. Numbers on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites | LOST choice | Like-for-like unresolved (evaluation cap and dropped calls); replica row: Open replicas and Jev-compatible servers |
Public-set runs against local replicas (4esv: open-jev, Kev, Laya; anessbelbati NevIR: Open-Jev 9B) moved to Head-to-head benchmarks: Jev on public datasets and suites; the JevBench board is on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites. @svencreations tried several "Jev killers": "Most weren't even close." Replica trust notes: Open replicas and Jev-compatible servers.
Related
- Head-to-head: Jev against other models and methods — Jev vs frontier LLMs, classic methods, engines and people; cascades
- Head-to-head: Jev inside agents, routers and tool gates, Head-to-head benchmarks: Jev on public datasets and suites, Head-to-head: games, simulators and real-time control — agents, public datasets, games
- Open replicas and Jev-compatible servers — the local and open replicas compared here (not Jev)
- Cost ledger: published cost per Jev decision — cost per decision
Sources
Files in frontmatter sources:, captured by 2026-10-01; original URLs inline.