---
title: "Head-to-head benchmarks: retrieval, reranking, screening and judging"
type: community
source_tier: community
tags: [head-to-head, benchmarks, retrieval, reranking, judging]
created: 2026-10-04
updated: 2026-10-04
confidence: medium
sources: "28 files, listed under this page in https://jevwiki.ai/index.json (raw/ paths map to their URLs in https://jevwiki.ai/raw/MANIFEST.json); links are inline in the body"
jev_version: "jev-1.13.0"
summary: "Jev on public retrieval, reranking and screening sets (BEIR, ESCI, CLEF TAR, recommendation) and on LLM-as-a-judge sets (MLflow, CMU). Split from head-to-head-benchmarks 2026-10-04."
---

# Head-to-head benchmarks: retrieval, reranking, screening and judging

> **TL;DR** Builder- and paper-run scores for Jev as a reranker, retriever filter or screener on public sets, and as a judge on LLM-as-a-judge sets. Numbers as published, one run each; vendor rows marked. Split from [[ideas/head-to-head-benchmarks]] on 2026-10-04 (rows unchanged); classification, structured data, robots and reasoning suites stay there.

## Judging (LLM-as-a-judge sets)

| Dataset (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| RewardBench (1,500 pairs), JudgeBench (350), HaluEval (3,000), 150 final answers; 1,610 held-out pairs for routing; blinded human adjudication; a pre-registered live test (arXiv [2609.26550](https://arxiv.org/abs/2609.26550) v3, "JEV-as-a-Judge", Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman, Carnegie Mellon; v1 2026-09-22) | Jev 1.13 (one Choice per request) vs 16 judges: GPT-4.1 mini to GPT-6 Astra, Claude Sonnet 5, Gemini 3 Flash and 3.1 Pro, Qwen, GPT-OSS 120B, PairRM, Skywork V2 | RewardBench 92.5 (= GPT-6), JudgeBench **78.6 vs 93.1**, HaluEval 87.3 vs 88.4, final answers 94.0 vs 96.7; 8th of 15 LLM judges pooled; $0.044 vs $12.182 per 1,000 judgments, median 0.15 vs 1.89 s. Adjudication sided with GPT-6 (JudgeBench 57 to 1; reasoning, coding, math 27-0, 9-0, 7-0). Within ~2 pts where the verdict can be read off the text, 7-28 pts behind where it must be derived (math, code, logic); a more elaborately written wrong answer cost it 8.8 pts (RM-Bench). Reference-free prose: Jev and two GPTs near chance yet confident (Jev error AUROC 0.498). Frozen cascade (τ 0.9, both orders averaged, threshold set on 96 pilot pairs): 68.5% accepted, 93.4% vs GPT-6's 92.5% at 41.4% of its fee; live replay of 510 pairs −0.59 pts at 57.2%, median 0.27 vs 2.10 s. Prospective test on two new hard workloads (PPE correctness, an unseen JudgeBench split; τ 0.95 / 0.99): matched GPT-6 exactly (88.2%, 94.7%), escalating 68% / 89%, at 75.5% of its fee | Cost, latency; cascade ≥ GPT-6; **LOST** alone on derived verdicts | Preprint; only the frozen policies, the adjudication outcome and the live test were pre-specified; one annotator; fees from price rules. Reversing a pair flipped 3.7% (RewardBench) and 11.1% (JudgeBench) of verdicts, 92 of 95 flips below 0.9 confidence; 4 of 48 paraphrased rubrics changed a decision, 0 of 96 repeats. Recipe: ~100 local labels and a lower-confidence-bound rule (a pooled threshold lost > 2 pts in 40-50% of draws). The promoting post ([@28qortex](https://x.com/28qortex/status/2105402111852969990)) calls the live test "production workloads" and says Jev absorbs "70%+" of traffic: the paper used benchmark workloads, and acceptance was 68.5% held out, 11-32% prospective. P15: [[ideas/patterns-data]] |
| MLflow QA answers, the authors' own sets: Part 1, 30 human-labelled answers (15 right, 15 wrong); Part 2, 72 near-miss answers (12 questions × 1 right + 2 subtly wrong, English and Japanese), two runs ([Part 1](https://mlflow.org/blog/jev-llm-judge/), Yuki Watanabe, 2026-09-22; [Part 2](https://mlflow.org/blog/jev-llm-judge-part-2/), Takaaki Yayoi and Yuki Watanabe, 2026-09-25; [@MLflow](https://x.com/MLflow/status/2106038472779710951), relayed by [@kral](https://x.com/kral/status/2106298445132636397); MLflow (Databricks) is adding Jev to its judges) | One Noul "correct?" at p ≥ 0.5 (`jev-1.13.0`; `typesafe/jev-1.13` in Part 2) vs GPT-5.6 Terra and Luna, Claude Sonnet 4.6 and Opus 4.8, DeepSeek-V4.1-Flash (reasoning off, 128 output tokens) in Part 1; GPT-OSS-120B in Part 2 | Part 1: agreement 30/30 for Jev, Terra and Luna; Sonnet 27/30, Opus 28/30, DeepSeek 29/30; median 369 ms vs 910-1,966 ms; $0.0247 vs $0.0624-3.775 per 1,000. Part 2: **64/72** vs 72/72 in both runs (32/36 per language); median 0.20 vs 1.44 s; ~$0.020 per 1,000 (GPT-OSS averaged 224 output tokens). Every miss was a false accept of an almost-right answer (reversed MERGE roles, a missing `USE CATALOG` permission, the wrong purge command); every correct answer, paraphrases included, scored ≥ 0.91, false accepts ~0.5-0.8. Routing 0.2-0.8 to GPT-OSS: 12 and 13 of 72 sent, 72/72 after routing | Part 1: tie at the lowest cost and latency; **LOST** Part 2 | Small own sets; the routing band was chosen after seeing the data, and one false accept scored 0.82 on a repeat (authors). A Noul returns only a yes probability, so it cannot abstain: `verified` ([[concepts/noul]]); the authors suggest a Choice with an `unsure` option. Jev gives no rationale. Cascade summary: [[ideas/head-to-head]] |

## Retrieval, search and screening

| Dataset (builder) | Compared | Numbers as published | Won | Caveat |
|---|---|---|---|---|
| Next-item reranking, Amazon Reviews 2023: Movies and TV (954 users), Video Games (1,000), Books (626); K = 20 to 200 hard SASRec candidates (arXiv [2609.40241](https://arxiv.org/abs/2609.40241), Hanjia Lyu, Yinglong Xia, 2026-09-30; [@_reachsumit](https://x.com/_reachsumit/status/2105516211434049854)) | Jev Choice over the K items, 10 recent items as `state` (`jev-latest` on 2026-09-30) vs SASRec, DCNv2, Qwen2.5 7B Instruct pointwise and listwise (local A800) | MRR, K = 20: Jev 0.238-0.346 vs pointwise Qwen 0.232-0.282, listwise 0.199-0.245; K = 200: 0.080-0.140 vs 0.061-0.074 and 0.033-0.039 (SASRec 0.093 beat Jev on Movies at K = 200). Median latency at K = 200: Jev 627-967 ms (hosted) vs pointwise Qwen 6.1-14.9 s, listwise 0.4-1.5 s, SASRec ~1.2 ms, DCNv2 ~0.55 ms | Quality, gentler scaling with K; **LOST** latency to recommender models | Preprint; one small LLM; latency not hardware-matched (API vs local GPU, authors say so); candidate order randomised. K = 200 fits the 255-option cap ([[reference/http-api]]). P17: [[ideas/patterns-data]] |
| Amazon ESCI (Shopping Queries), 250 US-English test queries, 4,754 judged pairs (the text later says 4,716), every method reordering the same judged candidates; questions, model and policy frozen on 50 training queries ([Elastic Search Labs](https://www.elastic.co/search-labs/blog/ecommerce-search-reranking-llm-alternative-jev), Dustin Coates, 2026-09-30; [@elastic](https://x.com/elastic/status/2105693854695395548); **vendor**: Elastic sells the retrieval side) | Jev 1.13.0 per query/product pair (a 4-way relationship Choice plus 4 Nouls, or one ordered Score), a short Python policy turning the probabilities into a rank, vs Elasticsearch BM25 and BM25 + semantic RRF hybrid | nDCG@10: best Jev policy (expected utility over the Choice distribution) **0.9565** vs hybrid 0.9351 (+0.0214, 95% CI +0.0125 to +0.0307), BM25 0.9234, random 0.8595; direct Score 0.9557, composite 0.9555, P(exact) 0.9533. Exact MRR best with P(exact): 0.9616 vs hybrid 0.9201 | Jev over its first stage | One run, one dataset skewed to Exact (2,960 of 4,716 pairs, so random scores 0.86); a Jina reranker was set up as a comparator but has no row in the results table; the text calls the composite best overall, the table ranks expected utility first. The post's "Latency: 70–500ms" is TypeSafe's launch figure quoted in the blog, not a measurement (`verified` as TypeSafe's claim, raw/site/blog-introducing-system-one.txt). $0.042 per M input, output free: `verified` ([[reference/models-and-pricing]]). Same dataset as a score ranker: s9v10 on [[ideas/head-to-head-small-models]] |
| Evidence judging for agent memory: 599 LoCoMo and LongMemEval-S questions, 14,359 retrieved records, 1,004 gold (Mnemon, arXiv [2609.36059](https://arxiv.org/abs/2609.36059), Guangren Wang; preview in [mnemon-dev/mnemon](https://github.com/mnemon-dev/mnemon/tree/master/experimental/memory-agent); [@grivn_eth](https://x.com/grivn_eth/status/2105414167318733181), 2026-09-30) | Jev yes/no "does the reply need this record" vs DeepSeek (V4.1-Flash, inferred from the setup) and gpt-4.1-mini asked the same proposition | Gold-evidence AUC 0.942 vs 0.900 vs 0.853 (LongMemEval-S alone 0.939 / 0.872 / 0.798); a 24-record Jev call answering two propositions per record 0.34 s vs 1.05-3.74 s for one (3-11×). Whole system, gpt-4.1-mini answering: LoCoMo 91.7% (first of 15 under OmniMemEval's protocol), LongMemEval-S 83.8%; reasoning model answering 92.2% / 94.4%. Cost per question ×1.11 from BEAM-100K to BEAM-10M (80× the records); 8th of 13 on HaluMem, 10th of 12 on BEAM-100K. Jev-Mem under the same protocol: 84.4% vs 91.7% on LoCoMo, spending 9.5× Mnemon's write-time cost per history | Evidence ranking, speed | Preprint by the builder; developed on LoCoMo and LongMemEval-S, none held out; one run per setting (repeats differed up to 2.6 pts); no Jev substitute tested in the full system; 10-15 s per question end to end. Jev-Mem: [[ideas/builds-agents]] |
| Rerank BM25 top 30, 8 datasets, 1,617 queries ([anessbelbati](https://github.com/anessbelbati/jev-rerank-bench)) | vs Cohere Rerank 4 Pro, zerank-2, DeepSeek | nDCG@10 0.692 / 0.691 / 0.682 / 0.682; BM25 0.486. 422 vs 844 ms; $0.45 vs $2.51 per 1,000 queries | Tie; latency. Per-query weighting: **LOST** (0.738 vs 0.756) | Reversing passage order changed the top pick of Jev's Choice setup on 24.7% of queries |
| Turkish XQuAD, 1,044 questions over 240 passages, the same BM25 + BGE-M3 top 20 for every reranker ([jev-rag-benchmark](https://github.com/erendikmenn/jev-rag-benchmark), erendikmenn, via OpenRouter) | Batched Jev Nouls vs Cohere Rerank 3.5 | Recall@5 99.425% both; nDCG@10 98.097 vs **98.626**; p50 532.0 vs **466.1 ms**; $0.410866 vs $1.044; no rerank 97.605% | Cost; **LOST** nDCG and latency | Pointwise Jev gained nothing at +56.5% cost (200 questions); a full-corpus hierarchical Jev retriever reached 76% Recall@5 (100 questions), worse than BM25; community run, `unverified`. Candidate order changed rankings (Spearman 0.262 over 5 permutations). Author: Jev 1.13 is strongest in English |
| Reranking, same 8 English datasets plus NevIR negation pairs ([anessbelbati](https://github.com/anessbelbati/jev-rerank-bench)) | vs Open-Jev 9B, Laya | nDCG on the 8 datasets: Jev yes/no per pair beats Open-Jev 9B by 7.0 points; Laya ≈ BM25. NevIR negation pairs (paired accuracy, never in the averages): Open-Jev 9B **77%** vs Jev 71%, Laya 36% | nDCG: Jev. **LOST** negation to 9B | Open-Jev 9B: 19.1 s per query |
| BEIR NFCorpus (323) and SciFact (300) ([llama-index-jev](https://github.com/WiktorB2004/llama-index-jev), WiktorB2004) | MiniLM top-10, Jev Score rerank to top 5, vs MiniLM and `rank-bm25` | nDCG@5 NFCorpus 0.340 → **0.396** (+0.056, CI 0.042-0.072); SciFact 0.629 → **0.715**; BGE-small 0.375 → 0.415. ~$0.0003 per query (OpenRouter `usage.cost`) | Beats its first stage | Not a BEIR leaderboard run (no Pyserini, no Cohere column); rerank fails open. BEIR SciFact, NFCorpus, FiQA reranking vs Cohere, Voyage and LLM listwise (hev reranker): [[ideas/tools-code]] |
| SWE-bench Lite file localisation, frozen 202-issue test split ([siftr](https://github.com/Bentlybro/siftr), Bentlybro; [method](https://github.com/Bentlybro/siftr/blob/4984c23dd338596a2ec62d836f742d98e7bb4722/BENCHMARKS.md)) | Jev outlines search vs BM25 vs grep | Right file in top 5: 82% vs 52% vs 22%; `pick` test file among ~550: 81% vs BM25 38%; `read` kept 92% of edited lines while cutting 59%. ~2 s, 1-2¢ per search on a 4,000-file repo | All; **LOST** on scikit-learn (71% vs 79%) | Log `filter` lost to grep on LogHub's BGL sample (both kept 100% of alert lines; a grep for fatal, error or fail at 23% of the log vs 46%), so `filter` ships experimental. Snippets go to OpenRouter |
| Title and abstract screening, 4,527 records from two systematic reviews of bipolar disorder treatments (light therapy, adjunctive pharmacotherapy); reviewers' retain decisions as reference ([medRxiv preprint](https://www.medrxiv.org/content/10.64898/2026.09.25.26364021v1), Kentaro Matsui and Yoshikazu Takaesu, posted 2026-10-01; [@matsuikentaro1](https://x.com/matsuikentaro1/status/2105852120650158189); [code and predictions](https://github.com/matsuikentaro1/jev-title-abstract-screening)) | Jev Noul "retain?" at a 50% cutoff fixed before the final runs, plus lower cutoffs; Jev Choice include/exclude vs GPT-5 mini and GPT-6 Astra (default reasoning) on the same prompts | Noul at 50%: sensitivity / specificity 99.6 / 98.9% (light therapy), 93.2 / 93.5% (pharmacotherapy). Lower cutoffs kept every reference-positive record and cut manual screening by 96.9% and 61.3%. Sensitivity at 50% with explicit options: Jev 92.3 / 90.0%, GPT-5 mini 87.2 / 90.0%, GPT-6 Astra 92.3 / 80.8%. One pass over all 4,527: $0.20 (Jev Choice) vs $2.59 vs $33.45 | Cost; sensitivity tie or better | Abstract capture only (full text not captured); the lossless cutoffs were selected on the same data (optimistic, like CLEF TAR below); one clinical area. Pattern: [[ideas/patterns-data]] P16 |
| CLEF TAR 2019, 12 Cochrane reviews ([Paper Radar](https://github.com/Eliot5566/JEV-Paper-Radar), Eliot5566) | Jev criteria Nouls vs reviewers' title/abstract decisions | 4 held-out reviews: 19,447 records, 96.9% recall, 78.0% pooled work saved, $0.60. 8 development reviews: **38-48%** pooled (38.9% with v2 criteria, 47.7% with v1; 0.1% to 82% per review) | Held-out: yes; dev: weak | Thresholds fitted on the same judgments (optimistic); pre-registration predicted 30-70% and was wrong; `unverified`. The author: performance is dominated by the review, not the tool. Its "waitlist" line is stale. [[ideas/failure-reports]] |
| Web research inside coding agents ([webctl](https://github.com/dorkitude/webctl), dorkitude) | webctl arms vs the agent's native search | Own question set, not a public suite: judged-quality and cost A/B on [[ideas/head-to-head-agents]] | | |
| Loghub HDFS and BGL samples ([Jev Logs](https://github.com/reachjalil/jevlogs), reachjalil) | Jev line triage vs block labels | Default routing kept **0.84%** of a 30%-anomalous HDFS sample (5 of 750 labelled "anomalies", all successful block verifications); BGL alerts caught 100% by the local FATAL rule, not Jev | **LOST** | Block-level labels vs line-level judgment; community run, `unverified`. Synthetic pager win: [[ideas/head-to-head]] |

## Related

- [[ideas/head-to-head-benchmarks]] — classification, structured data, robots and reasoning suites
- [[ideas/patterns-data]] — P15 judging, P17 rerank
- [[cookbooks/rerank]] — the official reranking cookbook

## Sources

Files in frontmatter `sources:`, captured by 2026-10-04; original URLs inline.
