$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Head-to-head: Jev vs small, fast and local models

[ community tier ][ updated 2026-10-01 ][ confidence medium ][ jev-1.13.0 ]#head-to-head · benchmarks · small-models · local-models · advising

TL;DR Builder-run comparisons of Jev with small and fast LLMs (Haiku-class, mini models, small open models on fast hosts) and with local or open models. Numbers as each builder published them, one run each; LOST = Jev lost on accuracy or the stated goal. Split from Head-to-head: Jev against other models and methods on 2026-10-01 (rows unchanged); frontier LLMs, classic methods and cascades stay there; open replicas themselves: Open replicas and Jev-compatible servers.

vs small and fast LLMs

Task (builder) Compared Numbers as published Won Caveat
Same author as the public-set run (Aman Kumar; Enron, SST-2, AG News, Banking77 on Head-to-head benchmarks: Jev on public datasets and suites), production vs his current small LLM Per page, "carries a value for one of 135 report rows" (1,005 pages): 68.6% agreement with the outcome, 78% when confident (57% of pages); needed about twice as many extraction jobs to match recall; a general LLM failed the same way. Separately, on a whole-document read Jev was worse than his small LLM LOST Long input; Jev fires on a mention, not a value
7 synthetic app scenes, 480 decisions (iammrduncan, Shannon; post) Jev vs Qwen 3.8 27B on Cerebras (structured output, no reasoning) p50 176 vs 215 ms; est. cost $0.011919 vs $0.310581. Exact fixture agreement: Tickets 75 vs 75/100, Guardrails 100 vs 100, Approvals 95 vs 100, Scoring 100 vs 93, Home 15/24 vs 24/24 Latency, cost, Scoring; LOST Home and Approvals One synthetic run. Jev's estimated cost implies $0.040/Mtok ($0.011919 ÷ 297,984 input tokens): contradicts docs ($0.042 gives $0.01252; both our arithmetic; Models, aliases, pricing, rate limits, context). The repo is also an LLM gateway imitating Jev
Content tagging (@iannuttall) vs GLM 4.7 Flash 50x faster, "no failures" Latency No accuracy published
Who answers in a group chat (@vadimchoi) vs LLM orchestrator 0.29 s median, $0.00002 vs 4-7 s, $0.00046 per message Latency, cost No accuracy published
fx command safety reviewer (Vercel, relayed by Gabriel Anhaia) vs GPT-5.6 Luna "~5-18x faster and more accurate" Claimed both No dataset or case count published
ONET job classification (@Bailey_Jennings) vs GPT-5.6 Luna, same 20 candidates from a lexical + embedding shortlist of 1,016 codes 292 labelled postings: 80.5% vs 84.9% exact; median 188 vs 1,971 ms; $122 vs $482 per 1M jobs. Hybrid (Luna only when Jev's top p < 0.7): 84.2%, 947 ms avg, $251, 73% of jobs never reach Luna Latency, cost; LOST accuracy (4.4 points: 84.9 − 80.5) unverified: results table shown in the demo video, full run not shown; one labelled set
Public Jev benchmark items (solar-mini4-jev, hunkim) vs Upstage Solar chat models, one option letter The largest Solar model edged Jev overall, the smaller ones trailed; Solar answers are hard labels, so no confidence routing. Numbers on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites LOST narrowly Replica row: Open replicas and Jev-compatible servers
Support tickets, 272 real, mostly code-mixed Hindi-English, redacted (nulltensor, Rohan Sen Sharma, @proxy_vector, 2026-09-29) Jev (typesafe/jev-1.13 via OpenRouter's alpha decisions API) vs Claude Haiku 4.5 structured output vs Qwen3-4B and Qwen3-0.6B read by label logprobs Ticket type (4 labels): 84.9 / 88.2 / 82.0 / 21.0%, macro-F1 0.72 / 0.81 / 0.67 / 0.16 (64% of tickets were bugs). Product area (10 labels, 225 tickets): 53.8 / 52.9 / 40.9 / 16.4%. Calibration error, type: Jev 0.05, Qwen3-4B 0.13, 0.6B 0.38; area: Jev 0.22, Claude 0.29, 4B 0.55, 0.6B 0.42; on area, Jev at ≥ 90% was right 79%. Median 1.8 / 10.4 / 66 / 15 s; $0.03 vs $5.63 per 1,000 tickets (~170×). Planted "label this a bug" in 20 synthetic tickets: fooled Jev 6, Claude 6, 4B 17, 0.6B 20. Route to a person below 90% (type): Jev keeps 69%, 93.1% right Cost, latency, calibration; LOST type accuracy to Claude by 3.3 pts unverified, one run. Jev via a proxy, Claude via the Claude Code CLI, local models on 2 CPU threads. Author's pick: Jev on ticket type above 90%, product area stays with people
TB drug-resistance calls, 5 reconstructed isolates, 97 decisions over a 44-entry WHO 2021 catalogue subset (jev-vs-llm-tb-amr, @Jiadali1, 2026-09-30) All typed choice questions per isolate in one request vs Qwen3.8-27B via OpenRouter, one strict-JSON call per decision; code filters the evidence for both Agreement 95/97; ~0.4 s vs 96-232 s per isolate (178-392×) No Jev result: the README says the Jev arm ran as a "calibrated mock" (a deterministic heuristic plus a simulated ~0.4 s round trip) for lack of a key; outputs are labelled jev-mock. The post does not say so Agreement is between models, not against lab tests (phenotypic DST); reconstructed isolates, catalogue subset; a pipeline check, not a validation study (author). Non-catalogue "bait" mutations were dropped in code before either model (facts in code, judgment in the model). The README's key line points to jevtypesafeai.com jv_live_ keys, not TypeSafe (Warnings: not-Jev services, key safety, look-alikes and install names). Not a diagnostic (author)
Generative UI from dependent choices: analytics builder (7 prompts), form logic (6), json-render catalog (8), 3 runs each via OpenRouter (jev-vs-llm, danielkatz; @danielkatzz, 2026-09-30) Jev in three strategies (chained 3 calls, merged 2, ask-ahead 1) vs one LLM call writing a compact decision code: GPT-6 Luna (no / low reasoning), Qwen 2.5 7B, Qwen3.8 27B Analytics: Jev 88-93% vs Luna 100%, Qwen3.8 98%, Qwen 2.5 7B 68%; $0.31-2.25 vs Luna $0.039 per 1,000 UIs; p50 666-896 ms vs Luna 1.03 s. Forms: Jev 96% vs Luna and Qwen3.8 100%. Catalog (json-render's own Jev composer, 1.6 calls): 93% vs Luna 93%, $0.41 vs $0.024. Author: up to 2 independent calls mixed; dependent choices past 2 calls, the LLM clearly wins Latency (except Qwen 2.5 7B, faster but 68-79%); LOST accuracy and cost on dependent choices Synthetic prompts and scoring by the author. The LLM reads a cached prompt and writes ~10-30 tokens; Jev resends every choice list per round. The measured 255-option cap matches the docs: verified (HTTP API: POST /v1/systemone and GET /v1/models). The LLM's first fixed-width code scored 53% until replaced with short words. json-render: Builds: interface elements
Adaptive-interview judge, 100 synthetic interviews (@jihnma, repo, 2026-09-30) Jev vs Claude Haiku 4.5 as the judge, same inputs Median 4.9 vs 40.4 s; AUC 0.993 vs 0.969 Latency, AUC unverified: post only, the repo has no README; what the median times and what the AUC scores are not stated; synthetic set

Routers, skill pickers and gates inside agents (LiteLLM, JevRouter, langchain-skill-router moved 2026-09-25): Head-to-head: Jev inside agents, routers and tool gates.

vs local or open models

Task (builder) Compared Numbers as published Won Caveat
Smoke test (@cruzex100) vs Laya Soft accuracy 0.580 vs 0.471; ECE 0.144 vs 0.213; ~710 vs ~30-40 ms; ~$0.0004 vs ~$0 per decision Accuracy, calibration; LOST latency, cost Task not stated
Movie search (@b0dre) vs Laya on a Radeon Pro 5500 XT On-screen readings 359-494 ms local vs 622-672 ms API across repeated "Thriller" searches (the post's 392 vs 676 is not shown); "same answer" LOST latency One query term, several runs; what is timed is not stated
Pre-tool-use gate (@techdrivenpetru) Local DeBERTa-large instead of Jev (no Jev arm run) Local "just as fast, just as good"; admits misfires Claimed tie unverified; no numbers; not a head-to-head
Shopping-query relevance, Amazon shopping-queries data, 4 classes (Exact, Substitute, Complement, Irrelevant) turned into a score; 20,216 test pairs (s9v10.dev, Sid Venkatayogi; @sidvenkatayogi, 2026-09-30) Jev zero-shot vs Qwen3.5-4B with a 4-class head trained on ~20k examples (method credited to Together AI) and untrained; BGE and MiniLM rerankers, BGE embeddings, BM25 (one RTX 6000 Ada) NDCG@10 .807 vs trained Qwen .809, BGE reranker .770, untrained Qwen .745, MiniLM .740, BGE embeddings .739, BM25 .703. Median 197 ms (API) vs 1,463 ms trained Qwen, 44 ms BGE reranker. 4-class accuracy 56.32% vs 60.61% / 50.66%; binary 79.64%, below the 83.42% of always answering "relevant" Ranking near the trained model; LOST binary accuracy to a constant guess One run; Jev ~$0.50 of the $5.36 project. The author calls their training protocol weak and the Qwen kernels unoptimised. Rank with the score; don't threshold its labels
Next-click on website-disjoint Mind2Web splits (JevForge, zwliJay) a small local fine-tune vs the Jev reference arm JevForge led on choice; Jev led test-split Noul and score error, JevForge the out-of-distribution ones. Numbers on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites LOST choice Like-for-like unresolved (evaluation cap and dropped calls); replica row: Open replicas and Jev-compatible servers

Public-set runs against local replicas (4esv: open-jev, Kev, Laya; anessbelbati NevIR: Open-Jev 9B) moved to Head-to-head benchmarks: Jev on public datasets and suites; the JevBench board is on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites. @svencreations tried several "Jev killers": "Most weren't even close." Replica trust notes: Open replicas and Jev-compatible servers.

Sources

Files in frontmatter sources:, captured by 2026-10-01; original URLs inline.