---
title: "Cost ledger: published cost per Jev decision"
type: community
source_tier: community
tags: [cost, pricing, head-to-head, community]
created: 2026-09-25
updated: 2026-09-25
confidence: medium
sources:
  - raw/x/stas_sorokin_-2101994942818115738.md
  - raw/x-repos/stas4000__jev-papers.md
  - raw/community/dev-to-bengreenberg-jev-vs-claude-who-wins-4mln.md
  - raw/community/primeline-cc-blog-typesafe-jev-pre-registered-test.md
  - raw/x-repos/fstandhartinger__jevbench.md
  - raw/x-repos/fstandhartinger__jevbench__results-v1-4-2.md
  - raw/community/dev-to-maximsaplin-typesafe-jev-played-chess-and-landed-next-to-reasoning-models-28ga.md
  - raw/x-repos/zhuyansen__jev-search-rerank-eval.md
  - raw/x-repos/QuicqDev__Jev-vs-ML.md
  - raw/x-repos/trungdq88__jev-tetris.md
  - raw/community/amankumar-ai-blogs-jev-measured.md
  - raw/x/Bailey_Jennings-2100593904949096696.md
  - raw/x/Bailey_Jennings-2100593904949096696-video.md
  - raw/community/openrouter-ai-blog-insights-jev-vs-claude-opus-5-classification.md
  - raw/x-repos/anisselbd__jev-phishing-bench.md
  - raw/community/gist-ipaulsmith-jev-nasa-kepler.md
  - raw/x-repos/chepyle__jev-test.md
  - raw/x-repos/TokenTrim__jev-agent-failure-benchmark.md
  - raw/x-repos/vclic__smoking-extraction-benchmark.md
  - raw/community/agentjournal-dev-blog-llm-judge-vs-feature-extraction.md
  - raw/community/docs-litellm-ai-blog-jev-auto-router-benchmark.md
  - raw/x-repos/iammrduncan__typesafe-ai-benchmark.md
  - raw/x-repos/eugeniughelbur__jev-engineering__results-2026-09-20-injection-test.md
  - raw/community/github-com-koala73-worldmonitor-pull-8326.md
  - raw/community/donttrustme-ai-assay-001.md
  - raw/x-repos/jourdanlabs__assay-001.md
  - raw/x-repos/blowxian__jev-fanout-bench.md
  - raw/community/huggingface-co-datasets-localllama-typed-decisions.md
  - raw/x-repos/FirasSX914__Janus.md
  - raw/x-repos/FirasSX914__Janus__research.md
  - raw/x-repos/erendikmenn__jev-rag-benchmark.md
  - raw/x-repos/erendikmenn__jev-rag-benchmark__docs-benchmark-2026-09-20.md
  - raw/x-repos/FBddcz__embodied-jev.md
  - raw/x-repos/FBddcz__embodied-jev__docs-results-libero-vision-results.md
  - raw/x-repos/iamvatsalpatel__tiershift__bench-results.md
  - raw/x-repos/agent-chaperone__agent-chaperone__bench-readme.md
  - raw/x-repos/Brainwires__jevwire__docs-benchmarks.md
  - raw/x-repos/keltokhy__jgrep.md
  - raw/x-repos/keltokhy__jsort.md
  - raw/x-repos/DanRWilloughby__snifftest.md
  - raw/x-repos/DanRWilloughby__snifftest__bench-results-2026-09-17-tables.md
  - raw/x-repos/NotXf1le__choosekit.md
  - raw/x-repos/NotXf1le__choosekit__benchmarks-supergpqa-benchmark-json.md
  - raw/x-repos/collapseindex__jev-ultralightspeed.md
  - raw/x-repos/devtooligan__jevscan-evm.md
  - raw/x-repos/RenaGao__jev-dataops__docs-benchmarks.md
  - raw/community/bearhuddleston-dev-reports-jev-approvals-live-sandbox.md
jev_version: "jev-1.13.0"
summary: "Builder-published cost per Jev decision across ~40 workloads, reconciled against the $0.042/M input price, with the comparator's cost where published. Split from head-to-head."
---

# Cost ledger: published cost per Jev decision

> **TL;DR** Across the builders' own runs a Jev decision costs roughly $0.000007 to $0.0004, from about 3x (GPT-5.6 Luna) to several hundred times less than the LLM or agent it was compared with. Most totals reconcile with input tokens × $0.042 per million, output free. Price your own workload as calls × `state` tokens × $0.042/M. Accuracy is on [[ideas/head-to-head]], [[ideas/head-to-head-agents]] and [[ideas/head-to-head-benchmarks]]. Split from head-to-head on 2026-09-25.

Community tier: the figures are the builders' own, and ÷ means our arithmetic: their total divided by their volume, or their per-1,000 figure divided by 1,000. Where a total does not reconcile with the price, the row says so. TypeSafe's docs win every conflict ([[reference/models-and-pricing]]).

## Published cost per decision

Per-decision figures are the source's, or ÷ where it publishes both numbers. Price is `verified`: $0.042 per million input tokens, output free ([[reference/models-and-pricing]]). Independently, [jev-fanout-bench](https://github.com/blowxian/jev-fanout-bench) (blowxian) found `usage.cost` equal to input tokens × $0.042/M on all 2,976 requests, with 122,844 output tokens not charged: `verified`.

| Workload | Volume | Total | Per decision | Source |
|---|---|---|---|---|
| Paper topics | 1,000 papers, 1,392,141 tokens | $0.0585 (reconciles at list price) | $0.000058; Opus judge $0.00894 | Sorokin |
| Policy gate, projected | 10,000 | $2.27 vs Sonnet $129.74 | ÷ $0.000227 vs $0.0130 | Greenberg |
| Pre-registered tests | ~9,750 calls | ~$0.38 | ÷ ~$0.000039 | Robin |
| JevBench | 534 decisions, 950 tokens average | $0.0399 per 1,000 | ÷ $0.0000399 | fstandhartinger |
| Chess | 80 games | ≈ $0.12 | $0.0015 per game, ~$0.00002 per move | Saplin |
| Relevance labels | 9,831 pairs | $0.069 vs Haiku $2.37 | ÷ $0.0000070 vs $0.00024 | zhuyansen |
| Rerank | 328 queries | $0.066 | ~$0.0002 | zhuyansen |
| Classical-ML benchmark | 38,922 attempts | ~$4.19, an estimate, not an invoice | ÷ ~$0.00011 | QuicqDev |
| Tetris game | 150 pieces, ~530k tokens | "about two cents" | ÷ ~$0.00013 per move | Dinh |
| Email triage / page gate / admin gate | per 1,000 | 15¢ vs $120 agent; 11¢ vs $1.50; 6¢ vs $3.50 | as stated | Kumar |
| ONET job classification | per 1M jobs | $122 vs Luna $482 | ÷ $0.000122 (≈ 2.9k input tokens per job at list price, inferred) | Bailey Jennings |
| Banking77 full test | 3,080, 2,605 mean input tokens | $0.34 billed via OpenRouter (reconciles) vs Opus 5 $7.44 cached | ÷ $0.00011 vs $0.00242 (their $0.11 vs $2.42 per 1,000 requests) | OpenRouter (vendor) |
| Phishing, two passes | 3.66M input + 0.90M output tokens | $0.15 on the TypeSafe dashboard = input only (`verified`: output free) | $0.038 vs Haiku $0.462 per 1,000 emails | anisselbd |
| Kepler dispositions | 8,054 objects, 8,025,931 input tokens | ~$0.34 estimate (reconciles) | ÷ ~$0.000042 | ipaulsmith |
| LexGLUE test | 23,607 examples | $4.02 vs Luna $16.45 | ÷ ~$0.00017 | chepyle |
| Who&When Pro | 6,257 traces | ~$1.28 | ÷ ~$0.00020 | TokenTrim |
| Smoking extraction | 1,000 notes | $0.18 vs `chat-latest` $24.54 (both the builder's estimates) | ÷ $0.00018 | vclic |
| Three tasks, dimensions vs one question | 34.1M input tokens | $1.43 (reconciles; 846,765 output tokens unpriced) | one question $0.019-0.026, 12 dimensions $0.042 per 1,000 rows | agentjournal |
| Auto Router tier | 240 calls, 183,492 input tokens | $0.0077 vs Haiku $0.1985 (reconciles) | ÷ $0.000032 vs $0.00083 | LiteLLM (vendor) |
| App scenes | 480 decisions, 297,984 input tokens | $0.011919, the builder's estimate: implies $0.040/Mtok (our arithmetic), `contradicts docs`; $0.01252 at $0.042 (our arithmetic) | ÷ ~$0.000025 | iammrduncan |
| Tool-call gate injection test | 300 calls, 449 mean input tokens | $0.0057 (reconciles) | $0.0000189 | eugeniughelbur |
| Headline severity | per 1,000, 618 input tokens each | ~$0.026 (reconciles) | ÷ ~$0.000026 | World Monitor |
| Calibration audit | 8,576 responses | "about fifty cents" | ÷ ~$0.00006 | ASSAY-001 |

The JevBench row is the v1.3.0 README's figure over the frozen v1.2 items; v1.4.2 keeps that basis (below).

## Added 2026-09-25

| Workload | Volume | Total | Per decision | Source |
|---|---|---|---|---|
| Typed-decisions test split, run live 2026-09-18 | 400 cases, 2,000 decisions | $0.016 at list price | ÷ $0.000008 | typed-decisions card |
| Banking77 / Web of Science (145 classes) | 500 each | $0.0507 vs DeepSeek V4-Pro $0.2207; $0.1006 vs $0.7355 | ÷ ~$0.00010; ~$0.00020 | Janus |
| Rerank top 20, Turkish XQuAD | 1,044 questions | $0.410866 vs Cohere Rerank 3.5 $1.044 | ÷ ~$0.00039 per question | jev-rag-benchmark |
| Meta-World robot arm, list-price estimate | 6 episodes | $0.018308 vs GPT-6 Astra $3.605290 | ÷ ~$0.0031 per episode | embodied-jev |
| LIBERO drawer, GPT-6 plans and Jev acts vs GPT-6 alone | 1 task each | $0.39400 vs $2.15002 (81.7% less) | as stated | embodied-jev |
| Model-tier routing | 120 routes | $0.0050 | $0.0418 per 1,000 | tiershift |
| Screening tool calls and results | 1,947 requests, 1.47M input tokens | $0.062 (reconciles) | ÷ ~$0.000032 | agent-chaperone |
| Coding-agent hooks, one developer | 740 live calls, ~1.53M input tokens (tokens counted over the whole log) | ~$0.064 (reconciles) | ÷ ~$0.000086 | jevwire |
| SMS spam as a described pattern | 5,574 messages | $0.07 | ÷ ~$0.000013 | jgrep |
| Readability sort, pairwise Nouls | 300 excerpts, 1,500 calls | $0.046 | ÷ ~$0.00003 per call | jsort |
| Prose lint, one Noul per rule per paragraph | 166 paragraphs | $0.0215 | $0.0129 per 100 paragraphs vs Haiku 4.5 $0.43, GPT-5.6 Sol $1.64, Sonnet 5 $1.34, Opus 5 $3.08 | Sniff Test |
| SuperGPQA multiple choice | 1,000 items | | $0.0244 per 1,000 decisions vs Kimi K3 $0.6243 | choosekit |
| XSTest labels, 32 items packed per request vs one per request | 30,000 judgements | $0.430 vs $0.729 (41% less) | ÷ ~$0.000014 vs ~$0.000024 | jev-ultralightspeed |
| Solidity file audit | one file, 14 checks (a full scan asks 359 Nouls per file) | 0.015¢ vs GPT-5.6 high 3.9¢ | as stated | jevscan-evm |
| Training-row screening via OpenRouter | 81 requests | $0.0023 | ÷ ~$0.000028 | jev-dataops |
| Approval assay, published-rate estimate | 44 calls per model | Jev $0.00229, GPT-5.6 Luna $0.00621, GPT-5.4 mini $0.02330 | as stated; billing not checked | Bear Huddleston |
| JevBench v1.4.2, Jev row | 220 hard items, 1,467 mean input tokens | $0.0399 per 1,000 on the v1.2 basis; $0.0616 per 1,000 hard decisions | ÷ $0.0000399; $0.0000616 | fstandhartinger |

## How to read the ledger

- **Reconciled rows** are within rounding of input tokens × $0.042/M. Rows that don't reconcile are flagged: iammrduncan implies $0.040/Mtok; @voxmenthe's ~$0.0015 per query for ~18k tokens is on [[ideas/head-to-head]].
- **Comparator prices** are the builders' own at the time: cached or uncached, reasoning on or off, list or gateway price. Compare ratios, not absolute LLM costs.
- **Gateways** may add their own price or pass the list price through (OpenRouter billed list price above). Access routes: [[ideas/platforms-and-gateways]].
- Cheap per call still adds up for repeated decisions over large `state`: price per hour ([[ideas/failure-reports]], cost surprises; [[ideas/request-mechanics]]).

## Source links

Sorokin [post](https://x.com/stas_sorokin_/status/2101994942818115738); [Greenberg](https://dev.to/bengreenberg/jev-vs-claude-who-wins-4mln); [Robin](https://primeline.cc/blog/typesafe-jev-pre-registered-test); [JevBench](https://github.com/fstandhartinger/jevbench); [Saplin](https://dev.to/maximsaplin/typesafe-jev-played-chess-and-landed-next-to-reasoning-models-28ga); [zhuyansen](https://github.com/zhuyansen/jev-search-rerank-eval); [QuicqDev](https://github.com/QuicqDev/Jev-vs-ML); [Dinh](https://github.com/trungdq88/jev-tetris); [Kumar](https://amankumar.ai/blogs/jev-measured); [Bailey Jennings](https://x.com/Bailey_Jennings/status/2100593904949096696); [OpenRouter](https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/); [anisselbd](https://github.com/anisselbd/jev-phishing-bench); [ipaulsmith](https://gist.github.com/ipaulsmith/e5c3ae3a492a455435d5bfc161404312); [chepyle](https://github.com/chepyle/jev-test); [TokenTrim](https://github.com/TokenTrim/jev-agent-failure-benchmark); [vclic](https://github.com/vclic/smoking-extraction-benchmark); [agentjournal](https://agentjournal.dev/blog/llm-judge-vs-feature-extraction/); [LiteLLM](https://docs.litellm.ai/blog/jev-auto-router-benchmark); [iammrduncan](https://github.com/iammrduncan/typesafe-ai-benchmark); [eugeniughelbur](https://github.com/eugeniughelbur/jev-engineering); [World Monitor](https://github.com/koala73/worldmonitor/pull/8326); [ASSAY-001](https://donttrustme.ai/assay-001.html); [typed-decisions card](https://huggingface.co/datasets/LocalLLaMA/typed-decisions); [Janus](https://github.com/FirasSX914/Janus) (FirasSX914); [jev-rag-benchmark](https://github.com/erendikmenn/jev-rag-benchmark) (erendikmenn); [embodied-jev](https://github.com/FBddcz/embodied-jev) (FBddcz); [tiershift](https://github.com/iamvatsalpatel/tiershift) (iamvatsalpatel); [agent-chaperone](https://github.com/agent-chaperone/agent-chaperone); [jevwire](https://github.com/Brainwires/jevwire) (Brainwires); [jgrep](https://github.com/keltokhy/jgrep) and [jsort](https://github.com/keltokhy/jsort) (keltokhy); [Sniff Test](https://github.com/DanRWilloughby/snifftest) (DanRWilloughby); [choosekit](https://github.com/NotXf1le/choosekit) (NotXf1le); [jev-ultralightspeed](https://github.com/collapseindex/jev-ultralightspeed) (collapseindex); [jevscan-evm](https://github.com/devtooligan/jevscan-evm) (devtooligan); [jev-dataops](https://github.com/RenaGao/jev-dataops) (RenaGao); [Bear Huddleston's study](https://bearhuddleston.dev/reports/jev-approvals-live-sandbox/) of anpicasso/hermes-jev-approvals.

## Related

- [[ideas/head-to-head]], [[ideas/head-to-head-agents]], [[ideas/head-to-head-benchmarks]]: what the money bought
- [[ideas/request-mechanics]]: billing, token overhead, packing; [[reference/models-and-pricing]]: the official price

## Sources

Raw captures listed in the frontmatter (private repo); original URLs above.
