$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Head-to-head benchmarks: retrieval, reranking, screening and judging

[ community tier ][ updated 2026-10-04 ][ confidence medium ][ jev-1.13.0 ]#head-to-head · benchmarks · retrieval · reranking · judging

TL;DR Builder- and paper-run scores for Jev as a reranker, retriever filter or screener on public sets, and as a judge on LLM-as-a-judge sets. Numbers as published, one run each; vendor rows marked. Split from Head-to-head benchmarks: Jev on public datasets and suites on 2026-10-04 (rows unchanged); classification, structured data, robots and reasoning suites stay there.

Judging (LLM-as-a-judge sets)

Dataset (builder) Compared Numbers as published Won Caveat
RewardBench (1,500 pairs), JudgeBench (350), HaluEval (3,000), 150 final answers; 1,610 held-out pairs for routing; blinded human adjudication; a pre-registered live test (arXiv 2609.26550 v3, "JEV-as-a-Judge", Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman, Carnegie Mellon; v1 2026-09-22) Jev 1.13 (one Choice per request) vs 16 judges: GPT-4.1 mini to GPT-6 Astra, Claude Sonnet 5, Gemini 3 Flash and 3.1 Pro, Qwen, GPT-OSS 120B, PairRM, Skywork V2 RewardBench 92.5 (= GPT-6), JudgeBench 78.6 vs 93.1, HaluEval 87.3 vs 88.4, final answers 94.0 vs 96.7; 8th of 15 LLM judges pooled; $0.044 vs $12.182 per 1,000 judgments, median 0.15 vs 1.89 s. Adjudication sided with GPT-6 (JudgeBench 57 to 1; reasoning, coding, math 27-0, 9-0, 7-0). Within ~2 pts where the verdict can be read off the text, 7-28 pts behind where it must be derived (math, code, logic); a more elaborately written wrong answer cost it 8.8 pts (RM-Bench). Reference-free prose: Jev and two GPTs near chance yet confident (Jev error AUROC 0.498). Frozen cascade (τ 0.9, both orders averaged, threshold set on 96 pilot pairs): 68.5% accepted, 93.4% vs GPT-6's 92.5% at 41.4% of its fee; live replay of 510 pairs −0.59 pts at 57.2%, median 0.27 vs 2.10 s. Prospective test on two new hard workloads (PPE correctness, an unseen JudgeBench split; τ 0.95 / 0.99): matched GPT-6 exactly (88.2%, 94.7%), escalating 68% / 89%, at 75.5% of its fee Cost, latency; cascade ≥ GPT-6; LOST alone on derived verdicts Preprint; only the frozen policies, the adjudication outcome and the live test were pre-specified; one annotator; fees from price rules. Reversing a pair flipped 3.7% (RewardBench) and 11.1% (JudgeBench) of verdicts, 92 of 95 flips below 0.9 confidence; 4 of 48 paraphrased rubrics changed a decision, 0 of 96 repeats. Recipe: ~100 local labels and a lower-confidence-bound rule (a pooled threshold lost > 2 pts in 40-50% of draws). The promoting post (@28qortex) calls the live test "production workloads" and says Jev absorbs "70%+" of traffic: the paper used benchmark workloads, and acceptance was 68.5% held out, 11-32% prospective. P15: Patterns: judging, search, documents, real-time and markets
MLflow QA answers, the authors' own sets: Part 1, 30 human-labelled answers (15 right, 15 wrong); Part 2, 72 near-miss answers (12 questions × 1 right + 2 subtly wrong, English and Japanese), two runs (Part 1, Yuki Watanabe, 2026-09-22; Part 2, Takaaki Yayoi and Yuki Watanabe, 2026-09-25; @MLflow, relayed by @kral; MLflow (Databricks) is adding Jev to its judges) One Noul "correct?" at p ≥ 0.5 (jev-1.13.0; typesafe/jev-1.13 in Part 2) vs GPT-5.6 Terra and Luna, Claude Sonnet 4.6 and Opus 4.8, DeepSeek-V4.1-Flash (reasoning off, 128 output tokens) in Part 1; GPT-OSS-120B in Part 2 Part 1: agreement 30/30 for Jev, Terra and Luna; Sonnet 27/30, Opus 28/30, DeepSeek 29/30; median 369 ms vs 910-1,966 ms; $0.0247 vs $0.0624-3.775 per 1,000. Part 2: 64/72 vs 72/72 in both runs (32/36 per language); median 0.20 vs 1.44 s; ~$0.020 per 1,000 (GPT-OSS averaged 224 output tokens). Every miss was a false accept of an almost-right answer (reversed MERGE roles, a missing USE CATALOG permission, the wrong purge command); every correct answer, paraphrases included, scored ≥ 0.91, false accepts ~0.5-0.8. Routing 0.2-0.8 to GPT-OSS: 12 and 13 of 72 sent, 72/72 after routing Part 1: tie at the lowest cost and latency; LOST Part 2 Small own sets; the routing band was chosen after seeing the data, and one false accept scored 0.82 on a repeat (authors). A Noul returns only a yes probability, so it cannot abstain: verified (Noul (yes/no) questions); the authors suggest a Choice with an unsure option. Jev gives no rationale. Cascade summary: Head-to-head: Jev against other models and methods

Retrieval, search and screening

Dataset (builder) Compared Numbers as published Won Caveat
Next-item reranking, Amazon Reviews 2023: Movies and TV (954 users), Video Games (1,000), Books (626); K = 20 to 200 hard SASRec candidates (arXiv 2609.40241, Hanjia Lyu, Yinglong Xia, 2026-09-30; @_reachsumit) Jev Choice over the K items, 10 recent items as state (jev-latest on 2026-09-30) vs SASRec, DCNv2, Qwen2.5 7B Instruct pointwise and listwise (local A800) MRR, K = 20: Jev 0.238-0.346 vs pointwise Qwen 0.232-0.282, listwise 0.199-0.245; K = 200: 0.080-0.140 vs 0.061-0.074 and 0.033-0.039 (SASRec 0.093 beat Jev on Movies at K = 200). Median latency at K = 200: Jev 627-967 ms (hosted) vs pointwise Qwen 6.1-14.9 s, listwise 0.4-1.5 s, SASRec ~1.2 ms, DCNv2 ~0.55 ms Quality, gentler scaling with K; LOST latency to recommender models Preprint; one small LLM; latency not hardware-matched (API vs local GPU, authors say so); candidate order randomised. K = 200 fits the 255-option cap (HTTP API: POST /v1/systemone and GET /v1/models). P17: Patterns: judging, search, documents, real-time and markets
Amazon ESCI (Shopping Queries), 250 US-English test queries, 4,754 judged pairs (the text later says 4,716), every method reordering the same judged candidates; questions, model and policy frozen on 50 training queries (Elastic Search Labs, Dustin Coates, 2026-09-30; @elastic; vendor: Elastic sells the retrieval side) Jev 1.13.0 per query/product pair (a 4-way relationship Choice plus 4 Nouls, or one ordered Score), a short Python policy turning the probabilities into a rank, vs Elasticsearch BM25 and BM25 + semantic RRF hybrid nDCG@10: best Jev policy (expected utility over the Choice distribution) 0.9565 vs hybrid 0.9351 (+0.0214, 95% CI +0.0125 to +0.0307), BM25 0.9234, random 0.8595; direct Score 0.9557, composite 0.9555, P(exact) 0.9533. Exact MRR best with P(exact): 0.9616 vs hybrid 0.9201 Jev over its first stage One run, one dataset skewed to Exact (2,960 of 4,716 pairs, so random scores 0.86); a Jina reranker was set up as a comparator but has no row in the results table; the text calls the composite best overall, the table ranks expected utility first. The post's "Latency: 70–500ms" is TypeSafe's launch figure quoted in the blog, not a measurement (verified as TypeSafe's claim, raw/site/blog-introducing-system-one.txt). $0.042 per M input, output free: verified (Models, aliases, pricing, rate limits, context). Same dataset as a score ranker: s9v10 on Head-to-head: Jev vs small, fast and local models
Evidence judging for agent memory: 599 LoCoMo and LongMemEval-S questions, 14,359 retrieved records, 1,004 gold (Mnemon, arXiv 2609.36059, Guangren Wang; preview in mnemon-dev/mnemon; @grivn_eth, 2026-09-30) Jev yes/no "does the reply need this record" vs DeepSeek (V4.1-Flash, inferred from the setup) and gpt-4.1-mini asked the same proposition Gold-evidence AUC 0.942 vs 0.900 vs 0.853 (LongMemEval-S alone 0.939 / 0.872 / 0.798); a 24-record Jev call answering two propositions per record 0.34 s vs 1.05-3.74 s for one (3-11×). Whole system, gpt-4.1-mini answering: LoCoMo 91.7% (first of 15 under OmniMemEval's protocol), LongMemEval-S 83.8%; reasoning model answering 92.2% / 94.4%. Cost per question ×1.11 from BEAM-100K to BEAM-10M (80× the records); 8th of 13 on HaluMem, 10th of 12 on BEAM-100K. Jev-Mem under the same protocol: 84.4% vs 91.7% on LoCoMo, spending 9.5× Mnemon's write-time cost per history Evidence ranking, speed Preprint by the builder; developed on LoCoMo and LongMemEval-S, none held out; one run per setting (repeats differed up to 2.6 pts); no Jev substitute tested in the full system; 10-15 s per question end to end. Jev-Mem: Builds: coding agents, harnesses and orchestration
Rerank BM25 top 30, 8 datasets, 1,617 queries (anessbelbati) vs Cohere Rerank 4 Pro, zerank-2, DeepSeek nDCG@10 0.692 / 0.691 / 0.682 / 0.682; BM25 0.486. 422 vs 844 ms; $0.45 vs $2.51 per 1,000 queries Tie; latency. Per-query weighting: LOST (0.738 vs 0.756) Reversing passage order changed the top pick of Jev's Choice setup on 24.7% of queries
Turkish XQuAD, 1,044 questions over 240 passages, the same BM25 + BGE-M3 top 20 for every reranker (jev-rag-benchmark, erendikmenn, via OpenRouter) Batched Jev Nouls vs Cohere Rerank 3.5 Recall@5 99.425% both; nDCG@10 98.097 vs 98.626; p50 532.0 vs 466.1 ms; $0.410866 vs $1.044; no rerank 97.605% Cost; LOST nDCG and latency Pointwise Jev gained nothing at +56.5% cost (200 questions); a full-corpus hierarchical Jev retriever reached 76% Recall@5 (100 questions), worse than BM25; community run, unverified. Candidate order changed rankings (Spearman 0.262 over 5 permutations). Author: Jev 1.13 is strongest in English
Reranking, same 8 English datasets plus NevIR negation pairs (anessbelbati) vs Open-Jev 9B, Laya nDCG on the 8 datasets: Jev yes/no per pair beats Open-Jev 9B by 7.0 points; Laya ≈ BM25. NevIR negation pairs (paired accuracy, never in the averages): Open-Jev 9B 77% vs Jev 71%, Laya 36% nDCG: Jev. LOST negation to 9B Open-Jev 9B: 19.1 s per query
BEIR NFCorpus (323) and SciFact (300) (llama-index-jev, WiktorB2004) MiniLM top-10, Jev Score rerank to top 5, vs MiniLM and rank-bm25 nDCG@5 NFCorpus 0.340 → 0.396 (+0.056, CI 0.042-0.072); SciFact 0.629 → 0.715; BGE-small 0.375 → 0.415. ~$0.0003 per query (OpenRouter usage.cost) Beats its first stage Not a BEIR leaderboard run (no Pyserini, no Cohere column); rerank fails open. BEIR SciFact, NFCorpus, FiQA reranking vs Cohere, Voyage and LLM listwise (hev reranker): Tools: CLIs, linters, semantic grep and dev tools
SWE-bench Lite file localisation, frozen 202-issue test split (siftr, Bentlybro; method) Jev outlines search vs BM25 vs grep Right file in top 5: 82% vs 52% vs 22%; pick test file among ~550: 81% vs BM25 38%; read kept 92% of edited lines while cutting 59%. ~2 s, 1-2¢ per search on a 4,000-file repo All; LOST on scikit-learn (71% vs 79%) Log filter lost to grep on LogHub's BGL sample (both kept 100% of alert lines; a grep for fatal, error or fail at 23% of the log vs 46%), so filter ships experimental. Snippets go to OpenRouter
Title and abstract screening, 4,527 records from two systematic reviews of bipolar disorder treatments (light therapy, adjunctive pharmacotherapy); reviewers' retain decisions as reference (medRxiv preprint, Kentaro Matsui and Yoshikazu Takaesu, posted 2026-10-01; @matsuikentaro1; code and predictions) Jev Noul "retain?" at a 50% cutoff fixed before the final runs, plus lower cutoffs; Jev Choice include/exclude vs GPT-5 mini and GPT-6 Astra (default reasoning) on the same prompts Noul at 50%: sensitivity / specificity 99.6 / 98.9% (light therapy), 93.2 / 93.5% (pharmacotherapy). Lower cutoffs kept every reference-positive record and cut manual screening by 96.9% and 61.3%. Sensitivity at 50% with explicit options: Jev 92.3 / 90.0%, GPT-5 mini 87.2 / 90.0%, GPT-6 Astra 92.3 / 80.8%. One pass over all 4,527: $0.20 (Jev Choice) vs $2.59 vs $33.45 Cost; sensitivity tie or better Abstract capture only (full text not captured); the lossless cutoffs were selected on the same data (optimistic, like CLEF TAR below); one clinical area. Pattern: Patterns: judging, search, documents, real-time and markets P16
CLEF TAR 2019, 12 Cochrane reviews (Paper Radar, Eliot5566) Jev criteria Nouls vs reviewers' title/abstract decisions 4 held-out reviews: 19,447 records, 96.9% recall, 78.0% pooled work saved, $0.60. 8 development reviews: 38-48% pooled (38.9% with v2 criteria, 47.7% with v1; 0.1% to 82% per review) Held-out: yes; dev: weak Thresholds fitted on the same judgments (optimistic); pre-registration predicted 30-70% and was wrong; unverified. The author: performance is dominated by the review, not the tool. Its "waitlist" line is stale. Failure reports: where Jev broke, lost, or was the wrong tool
Web research inside coding agents (webctl, dorkitude) webctl arms vs the agent's native search Own question set, not a public suite: judged-quality and cost A/B on Head-to-head: Jev inside agents, routers and tool gates
Loghub HDFS and BGL samples (Jev Logs, reachjalil) Jev line triage vs block labels Default routing kept 0.84% of a 30%-anomalous HDFS sample (5 of 750 labelled "anomalies", all successful block verifications); BGL alerts caught 100% by the local FATAL rule, not Jev LOST Block-level labels vs line-level judgment; community run, unverified. Synthetic pager win: Head-to-head: Jev against other models and methods

Sources

Files in frontmatter sources:, captured by 2026-10-04; original URLs inline.