Agents: read the raw Markdown of this page, or start at llms.txt.
Cost ledger: published cost per Jev decision
TL;DR Across the builders' own runs a Jev decision costs roughly $0.000007 to $0.0004, from about 3x (GPT-5.6 Luna) to several hundred times less than the LLM or agent it was compared with. Most totals reconcile with input tokens × $0.042 per million, output free. Price your own workload as calls ×
statetokens × $0.042/M. Accuracy is on Head-to-head: Jev against other models and methods, Head-to-head: Jev inside agents, routers and tool gates and Head-to-head benchmarks: Jev on public datasets and suites. Split from head-to-head on 2026-09-25.
Community tier: the figures are the builders' own, and ÷ means our arithmetic: their total divided by their volume, or their per-1,000 figure divided by 1,000. Where a total does not reconcile with the price, the row says so. TypeSafe's docs win every conflict (Models, aliases, pricing, rate limits, context).
Published cost per decision
Per-decision figures are the source's, or ÷ where it publishes both numbers. Price is verified: $0.042 per million input tokens, output free (Models, aliases, pricing, rate limits, context). Independently, jev-fanout-bench (blowxian) found usage.cost equal to input tokens × $0.042/M on all 2,976 requests, with 122,844 output tokens not charged: verified.
| Workload | Volume | Total | Per decision | Source |
|---|---|---|---|---|
| Paper topics | 1,000 papers, 1,392,141 tokens | $0.0585 (reconciles at list price) | $0.000058; Opus judge $0.00894 | Sorokin |
| Policy gate, projected | 10,000 | $2.27 vs Sonnet $129.74 | ÷ $0.000227 vs $0.0130 | Greenberg |
| Pre-registered tests | ~9,750 calls | ~$0.38 | ÷ ~$0.000039 | Robin |
| JevBench | 534 decisions, 950 tokens average | $0.0399 per 1,000 | ÷ $0.0000399 | fstandhartinger |
| Chess | 80 games | ≈ $0.12 | $0.0015 per game, ~$0.00002 per move | Saplin |
| Relevance labels | 9,831 pairs | $0.069 vs Haiku $2.37 | ÷ $0.0000070 vs $0.00024 | zhuyansen |
| Rerank | 328 queries | $0.066 | ~$0.0002 | zhuyansen |
| Classical-ML benchmark | 38,922 attempts | ~$4.19, an estimate, not an invoice | ÷ ~$0.00011 | QuicqDev |
| Tetris game | 150 pieces, ~530k tokens | "about two cents" | ÷ ~$0.00013 per move | Dinh |
| Email triage / page gate / admin gate | per 1,000 | 15¢ vs $120 agent; 11¢ vs $1.50; 6¢ vs $3.50 | as stated | Kumar |
| ONET job classification | per 1M jobs | $122 vs Luna $482 | ÷ $0.000122 (≈ 2.9k input tokens per job at list price, inferred) | Bailey Jennings |
| Banking77 full test | 3,080, 2,605 mean input tokens | $0.34 billed via OpenRouter (reconciles) vs Opus 5 $7.44 cached | ÷ $0.00011 vs $0.00242 (their $0.11 vs $2.42 per 1,000 requests) | OpenRouter (vendor) |
| Phishing, two passes | 3.66M input + 0.90M output tokens | $0.15 on the TypeSafe dashboard = input only (verified: output free) |
$0.038 vs Haiku $0.462 per 1,000 emails | anisselbd |
| Kepler dispositions | 8,054 objects, 8,025,931 input tokens | ~$0.34 estimate (reconciles) | ÷ ~$0.000042 | ipaulsmith |
| LexGLUE test | 23,607 examples | $4.02 vs Luna $16.45 | ÷ ~$0.00017 | chepyle |
| Who&When Pro | 6,257 traces | ~$1.28 | ÷ ~$0.00020 | TokenTrim |
| Smoking extraction | 1,000 notes | $0.18 vs chat-latest $24.54 (both the builder's estimates) |
÷ $0.00018 | vclic |
| Three tasks, dimensions vs one question | 34.1M input tokens | $1.43 (reconciles; 846,765 output tokens unpriced) | one question $0.019-0.026, 12 dimensions $0.042 per 1,000 rows | agentjournal |
| Auto Router tier | 240 calls, 183,492 input tokens | $0.0077 vs Haiku $0.1985 (reconciles) | ÷ $0.000032 vs $0.00083 | LiteLLM (vendor) |
| App scenes | 480 decisions, 297,984 input tokens | $0.011919, the builder's estimate: implies $0.040/Mtok (our arithmetic), contradicts docs; $0.01252 at $0.042 (our arithmetic) |
÷ ~$0.000025 | iammrduncan |
| Tool-call gate injection test | 300 calls, 449 mean input tokens | $0.0057 (reconciles) | $0.0000189 | eugeniughelbur |
| Headline severity | per 1,000, 618 input tokens each | ~$0.026 (reconciles) | ÷ ~$0.000026 | World Monitor |
| Calibration audit | 8,576 responses | "about fifty cents" | ÷ ~$0.00006 | ASSAY-001 |
The JevBench row is the v1.3.0 README's figure over the frozen v1.2 items; v1.4.2 keeps that basis (below).
Added 2026-09-25
| Workload | Volume | Total | Per decision | Source |
|---|---|---|---|---|
| Typed-decisions test split, run live 2026-09-18 | 400 cases, 2,000 decisions | $0.016 at list price | ÷ $0.000008 | typed-decisions card |
| Banking77 / Web of Science (145 classes) | 500 each | $0.0507 vs DeepSeek V4-Pro $0.2207; $0.1006 vs $0.7355 | ÷ ~$0.00010; ~$0.00020 | Janus |
| Rerank top 20, Turkish XQuAD | 1,044 questions | $0.410866 vs Cohere Rerank 3.5 $1.044 | ÷ ~$0.00039 per question | jev-rag-benchmark |
| Meta-World robot arm, list-price estimate | 6 episodes | $0.018308 vs GPT-6 Astra $3.605290 | ÷ ~$0.0031 per episode | embodied-jev |
| LIBERO drawer, GPT-6 plans and Jev acts vs GPT-6 alone | 1 task each | $0.39400 vs $2.15002 (81.7% less) | as stated | embodied-jev |
| Model-tier routing | 120 routes | $0.0050 | $0.0418 per 1,000 | tiershift |
| Screening tool calls and results | 1,947 requests, 1.47M input tokens | $0.062 (reconciles) | ÷ ~$0.000032 | agent-chaperone |
| Coding-agent hooks, one developer | 740 live calls, ~1.53M input tokens (tokens counted over the whole log) | ~$0.064 (reconciles) | ÷ ~$0.000086 | jevwire |
| SMS spam as a described pattern | 5,574 messages | $0.07 | ÷ ~$0.000013 | jgrep |
| Readability sort, pairwise Nouls | 300 excerpts, 1,500 calls | $0.046 | ÷ ~$0.00003 per call | jsort |
| Prose lint, one Noul per rule per paragraph | 166 paragraphs | $0.0215 | $0.0129 per 100 paragraphs vs Haiku 4.5 $0.43, GPT-5.6 Sol $1.64, Sonnet 5 $1.34, Opus 5 $3.08 | Sniff Test |
| SuperGPQA multiple choice | 1,000 items | $0.0244 per 1,000 decisions vs Kimi K3 $0.6243 | choosekit | |
| XSTest labels, 32 items packed per request vs one per request | 30,000 judgements | $0.430 vs $0.729 (41% less) | ÷ ~$0.000014 vs ~$0.000024 | jev-ultralightspeed |
| Solidity file audit | one file, 14 checks (a full scan asks 359 Nouls per file) | 0.015¢ vs GPT-5.6 high 3.9¢ | as stated | jevscan-evm |
| Training-row screening via OpenRouter | 81 requests | $0.0023 | ÷ ~$0.000028 | jev-dataops |
| Approval assay, published-rate estimate | 44 calls per model | Jev $0.00229, GPT-5.6 Luna $0.00621, GPT-5.4 mini $0.02330 | as stated; billing not checked | Bear Huddleston |
| JevBench v1.4.2, Jev row | 220 hard items, 1,467 mean input tokens | $0.0399 per 1,000 on the v1.2 basis; $0.0616 per 1,000 hard decisions | ÷ $0.0000399; $0.0000616 | fstandhartinger |
How to read the ledger
- Reconciled rows are within rounding of input tokens × $0.042/M. Rows that don't reconcile are flagged: iammrduncan implies $0.040/Mtok; @voxmenthe's ~$0.0015 per query for ~18k tokens is on Head-to-head: Jev against other models and methods.
- Comparator prices are the builders' own at the time: cached or uncached, reasoning on or off, list or gateway price. Compare ratios, not absolute LLM costs.
- Gateways may add their own price or pass the list price through (OpenRouter billed list price above). Access routes: Platforms and gateways: Zapier, LangChain, Spring AI, Cloudflare, Netlify, Vercel, OpenRouter, Fly.io, Pydantic AI and other hosted routes to Jev.
- Cheap per call still adds up for repeated decisions over large
state: price per hour (Failure reports: where Jev broke, lost, or was the wrong tool, cost surprises; Request mechanics: billing, limits, latency, calibration and stability).
Source links
Sorokin post; Greenberg; Robin; JevBench; Saplin; zhuyansen; QuicqDev; Dinh; Kumar; Bailey Jennings; OpenRouter; anisselbd; ipaulsmith; chepyle; TokenTrim; vclic; agentjournal; LiteLLM; iammrduncan; eugeniughelbur; World Monitor; ASSAY-001; typed-decisions card; Janus (FirasSX914); jev-rag-benchmark (erendikmenn); embodied-jev (FBddcz); tiershift (iamvatsalpatel); agent-chaperone; jevwire (Brainwires); jgrep and jsort (keltokhy); Sniff Test (DanRWilloughby); choosekit (NotXf1le); jev-ultralightspeed (collapseindex); jevscan-evm (devtooligan); jev-dataops (RenaGao); Bear Huddleston's study of anpicasso/hermes-jev-approvals.
Related
- Head-to-head: Jev against other models and methods, Head-to-head: Jev inside agents, routers and tool gates, Head-to-head benchmarks: Jev on public datasets and suites: what the money bought
- Request mechanics: billing, limits, latency, calibration and stability: billing, token overhead, packing; Models, aliases, pricing, rate limits, context: the official price
Sources
Raw captures listed in the frontmatter (private repo); original URLs above.