$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Cost ledger: published cost per Jev decision

[ community tier ][ updated 2026-09-25 ][ confidence medium ][ jev-1.13.0 ]#cost · pricing · head-to-head · community

TL;DR Across the builders' own runs a Jev decision costs roughly $0.000007 to $0.0004, from about 3x (GPT-5.6 Luna) to several hundred times less than the LLM or agent it was compared with. Most totals reconcile with input tokens × $0.042 per million, output free. Price your own workload as calls × state tokens × $0.042/M. Accuracy is on Head-to-head: Jev against other models and methods, Head-to-head: Jev inside agents, routers and tool gates and Head-to-head benchmarks: Jev on public datasets and suites. Split from head-to-head on 2026-09-25.

Community tier: the figures are the builders' own, and ÷ means our arithmetic: their total divided by their volume, or their per-1,000 figure divided by 1,000. Where a total does not reconcile with the price, the row says so. TypeSafe's docs win every conflict (Models, aliases, pricing, rate limits, context).

Published cost per decision

Per-decision figures are the source's, or ÷ where it publishes both numbers. Price is verified: $0.042 per million input tokens, output free (Models, aliases, pricing, rate limits, context). Independently, jev-fanout-bench (blowxian) found usage.cost equal to input tokens × $0.042/M on all 2,976 requests, with 122,844 output tokens not charged: verified.

Workload Volume Total Per decision Source
Paper topics 1,000 papers, 1,392,141 tokens $0.0585 (reconciles at list price) $0.000058; Opus judge $0.00894 Sorokin
Policy gate, projected 10,000 $2.27 vs Sonnet $129.74 ÷ $0.000227 vs $0.0130 Greenberg
Pre-registered tests ~9,750 calls ~$0.38 ÷ ~$0.000039 Robin
JevBench 534 decisions, 950 tokens average $0.0399 per 1,000 ÷ $0.0000399 fstandhartinger
Chess 80 games ≈ $0.12 $0.0015 per game, ~$0.00002 per move Saplin
Relevance labels 9,831 pairs $0.069 vs Haiku $2.37 ÷ $0.0000070 vs $0.00024 zhuyansen
Rerank 328 queries $0.066 ~$0.0002 zhuyansen
Classical-ML benchmark 38,922 attempts ~$4.19, an estimate, not an invoice ÷ ~$0.00011 QuicqDev
Tetris game 150 pieces, ~530k tokens "about two cents" ÷ ~$0.00013 per move Dinh
Email triage / page gate / admin gate per 1,000 15¢ vs $120 agent; 11¢ vs $1.50; 6¢ vs $3.50 as stated Kumar
ONET job classification per 1M jobs $122 vs Luna $482 ÷ $0.000122 (≈ 2.9k input tokens per job at list price, inferred) Bailey Jennings
Banking77 full test 3,080, 2,605 mean input tokens $0.34 billed via OpenRouter (reconciles) vs Opus 5 $7.44 cached ÷ $0.00011 vs $0.00242 (their $0.11 vs $2.42 per 1,000 requests) OpenRouter (vendor)
Phishing, two passes 3.66M input + 0.90M output tokens $0.15 on the TypeSafe dashboard = input only (verified: output free) $0.038 vs Haiku $0.462 per 1,000 emails anisselbd
Kepler dispositions 8,054 objects, 8,025,931 input tokens ~$0.34 estimate (reconciles) ÷ ~$0.000042 ipaulsmith
LexGLUE test 23,607 examples $4.02 vs Luna $16.45 ÷ ~$0.00017 chepyle
Who&When Pro 6,257 traces ~$1.28 ÷ ~$0.00020 TokenTrim
Smoking extraction 1,000 notes $0.18 vs chat-latest $24.54 (both the builder's estimates) ÷ $0.00018 vclic
Three tasks, dimensions vs one question 34.1M input tokens $1.43 (reconciles; 846,765 output tokens unpriced) one question $0.019-0.026, 12 dimensions $0.042 per 1,000 rows agentjournal
Auto Router tier 240 calls, 183,492 input tokens $0.0077 vs Haiku $0.1985 (reconciles) ÷ $0.000032 vs $0.00083 LiteLLM (vendor)
App scenes 480 decisions, 297,984 input tokens $0.011919, the builder's estimate: implies $0.040/Mtok (our arithmetic), contradicts docs; $0.01252 at $0.042 (our arithmetic) ÷ ~$0.000025 iammrduncan
Tool-call gate injection test 300 calls, 449 mean input tokens $0.0057 (reconciles) $0.0000189 eugeniughelbur
Headline severity per 1,000, 618 input tokens each ~$0.026 (reconciles) ÷ ~$0.000026 World Monitor
Calibration audit 8,576 responses "about fifty cents" ÷ ~$0.00006 ASSAY-001

The JevBench row is the v1.3.0 README's figure over the frozen v1.2 items; v1.4.2 keeps that basis (below).

Added 2026-09-25

Workload Volume Total Per decision Source
Typed-decisions test split, run live 2026-09-18 400 cases, 2,000 decisions $0.016 at list price ÷ $0.000008 typed-decisions card
Banking77 / Web of Science (145 classes) 500 each $0.0507 vs DeepSeek V4-Pro $0.2207; $0.1006 vs $0.7355 ÷ ~$0.00010; ~$0.00020 Janus
Rerank top 20, Turkish XQuAD 1,044 questions $0.410866 vs Cohere Rerank 3.5 $1.044 ÷ ~$0.00039 per question jev-rag-benchmark
Meta-World robot arm, list-price estimate 6 episodes $0.018308 vs GPT-6 Astra $3.605290 ÷ ~$0.0031 per episode embodied-jev
LIBERO drawer, GPT-6 plans and Jev acts vs GPT-6 alone 1 task each $0.39400 vs $2.15002 (81.7% less) as stated embodied-jev
Model-tier routing 120 routes $0.0050 $0.0418 per 1,000 tiershift
Screening tool calls and results 1,947 requests, 1.47M input tokens $0.062 (reconciles) ÷ ~$0.000032 agent-chaperone
Coding-agent hooks, one developer 740 live calls, ~1.53M input tokens (tokens counted over the whole log) ~$0.064 (reconciles) ÷ ~$0.000086 jevwire
SMS spam as a described pattern 5,574 messages $0.07 ÷ ~$0.000013 jgrep
Readability sort, pairwise Nouls 300 excerpts, 1,500 calls $0.046 ÷ ~$0.00003 per call jsort
Prose lint, one Noul per rule per paragraph 166 paragraphs $0.0215 $0.0129 per 100 paragraphs vs Haiku 4.5 $0.43, GPT-5.6 Sol $1.64, Sonnet 5 $1.34, Opus 5 $3.08 Sniff Test
SuperGPQA multiple choice 1,000 items $0.0244 per 1,000 decisions vs Kimi K3 $0.6243 choosekit
XSTest labels, 32 items packed per request vs one per request 30,000 judgements $0.430 vs $0.729 (41% less) ÷ ~$0.000014 vs ~$0.000024 jev-ultralightspeed
Solidity file audit one file, 14 checks (a full scan asks 359 Nouls per file) 0.015¢ vs GPT-5.6 high 3.9¢ as stated jevscan-evm
Training-row screening via OpenRouter 81 requests $0.0023 ÷ ~$0.000028 jev-dataops
Approval assay, published-rate estimate 44 calls per model Jev $0.00229, GPT-5.6 Luna $0.00621, GPT-5.4 mini $0.02330 as stated; billing not checked Bear Huddleston
JevBench v1.4.2, Jev row 220 hard items, 1,467 mean input tokens $0.0399 per 1,000 on the v1.2 basis; $0.0616 per 1,000 hard decisions ÷ $0.0000399; $0.0000616 fstandhartinger

How to read the ledger

Sorokin post; Greenberg; Robin; JevBench; Saplin; zhuyansen; QuicqDev; Dinh; Kumar; Bailey Jennings; OpenRouter; anisselbd; ipaulsmith; chepyle; TokenTrim; vclic; agentjournal; LiteLLM; iammrduncan; eugeniughelbur; World Monitor; ASSAY-001; typed-decisions card; Janus (FirasSX914); jev-rag-benchmark (erendikmenn); embodied-jev (FBddcz); tiershift (iamvatsalpatel); agent-chaperone; jevwire (Brainwires); jgrep and jsort (keltokhy); Sniff Test (DanRWilloughby); choosekit (NotXf1le); jev-ultralightspeed (collapseindex); jevscan-evm (devtooligan); jev-dataops (RenaGao); Bear Huddleston's study of anpicasso/hermes-jev-approvals.

Sources

Raw captures listed in the frontmatter (private repo); original URLs above.