$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Builds: business ops and markets

[ community tier ][ updated 2026-10-04 ][ confidence medium ][ jev-1.13.0 ]#builds · business · operations · advising

TL;DR Community builds where Jev runs business operations (support and inbox routing, sales and lead scoring, recruiting, reconciliation, risk sweeps) plus market demos described as engineering only, never financial advice. Builders' own numbers, mostly one run; vendor rows marked. Split from Builds: data, search and business on 2026-10-04 (rows unchanged); data and search builds stay there.

Business ops

Build (builder) Jev decides Reported numbers Pattern Source
Fraud email cascade (Hassan, @nutlope) Fraud vs legit; confidence < 0.95 escalates to Kimi K3 100 emails in 1.42 s; 31 escalated; pipeline 96/100 in 16 s for ~$0.07, of which Jev $0.003 Patterns: agent internals, routing, gates, context and memory P06, Confidence-gated routing post
Spliit Cloud expenses (Antonio Ivanovski) Expense category from title: local dictionary, then group history, then one engine (LLM or Jev Choice); below AI_CATEGORY_MIN_CONFIDENCE (0.5) no suggestion; uncertain guesses shown as one-tap chips none. README: pin jev-1.13.0 once thresholds are tuned, jev-latest moves (verified, Models, aliases, pricing, rate limits, context); LLM self-reported and Jev confidence are not on one scale, recalibrate when switching Patterns: agent internals, routing, gates, context and memory P06, Patterns: browser, computer use, voice and product UI P34 repo
Regex rules → Jev (Ilias Ism, @illyism) Hard-coded rules (vibe-coded regex) replaced in aiseotracker, linkdr, genppt "around 10x faster, 50% cheaper" than regular LLM calls (no method) Patterns: agent internals, routing, gates, context and memory P01 post
Backdoor job match (@sarvagya_kul) Fit probability of one candidate for 400 companies; flags mismatches 12 s, $0.0005. At list price that is ~11.9k input tokens, ~30 per company (our arithmetic: $0.0005 ÷ $0.042/Mtok, ÷ 400): only fits very short profiles; treat as approximate Patterns: marketing, sales, GTM, content, support and ops P19 post
ZipPad conversation on pujo.ai (@bblever) Runs "the ZipPad conversation" on Jev; what Jev decides is not stated (Jev writes no text, so most likely it picks the next scripted step or reply; inferred) @bblever reports every exchange ~10x faster and 38% cheaper per session; baseline not named, no method; 3 views at capture. Vendor claim, unverified Patterns: agent internals, routing, gates, context and memory P01 (inferred) post, site
Gojiberry outreach mining (@pierreeliottlal) Labels thousands of outreach messages by intent signal; which signals booked demos (the tally is code, inferred) 40 s, < $0.20 Patterns: judging, search, documents, real-time and markets P18, Patterns: marketing, sales, GTM, content, support and ops P20 post
jev-suite (klauswg) Four Java apps on one kernel: did a sponsored video deliver the brief, where a candidate falls short, did an edit keep the facts, what to confirm about a rental listing. A three-way NoulGate (confident pass, confident fail, human review); outside text sanitised; model unreachable → review, never auto-pass; calibration runners refuse mock data Raw Jev lost to a keyword baseline tuned on the same data; only the gated pipeline, abstaining on the unsure claims, beat it, after two v1 bugs were fixed; results and the report link on Head-to-head: Jev against other models and methods Patterns: agent internals, routing, gates, context and memory P06, Patterns: marketing, sales, GTM, content, support and ops P23, Patterns: gates, simulation, personas and other shapes (P38+) P41 repo
lurk (getanyapi-com, AnyAPI) Reddit buyer-intent finder: title triage, only promising threads opened, then Jev (OpenRouter, or Vercel Gateway first if set) judges each post and comment: 0-100 score, intent stage, seller-side flag; never posts No Jev numbers. Vendor demo for a paid data API. Its "Costs, measured" table is dated 2026-09-05, before Jev's 2026-09-15 launch, so its model column cannot be Jev (inferred): never cite it as Jev cost Patterns: marketing, sales, GTM, content, support and ops P19, P20 repo
Volty smart filters (@_voltade, Voltade; vendor) Does this WhatsApp or email conversation fit the user's filter, yes or no, with a probability (Noul, inferred) 700 conversations: 5 min → 20 s ("15× faster"); ~0.25 s per check, so ~9 in flight (our arithmetic, inferred); cost per check on Cost ledger: published cost per Jev decision. No accuracy figures; video not transcribed. unverified Patterns: judging, search, documents, real-time and markets P18, Patterns: marketing, sales, GTM, content, support and ops P22 post
mastra-inbound-router (One, withone.ai; @calcsam; MIT; vendor demo) Shared inbox: one Jev request, four questions per email (category, machine-sent?, someone waiting?, urgency); ≥ 90% vendor or noise is filed without an LLM; the rest go to Claude, which must look the sender up in HubSpot, picks owner and drafts from a rules file; a person reviews when Claude is under 60% sure, names no owner, or Jev gives Claude's category under 20%. Mastra workflow suspends until a person ticks each action; Gmail send scope left off, so replies stay drafts. Gmail, HubSpot, Slack and the Jev call all go through One 20-email sample (fictional company, recorded answers): Jev filed 4 alone, Claude took 16 (5 drafts, 10 handed off, 1 held). Eval gate: routing 95% vs a labelled key (gate 90%). Fixes reported: a "checking in" email labelled noise at 30% confidence would have been archived, hence the 60% floor; a rule clash fixed by one sentence in ROUTING.md. Vendor's own sample of 20; unverified. Mastra's Classifier primitive: Framework integrations: Pydantic AI, LangChain, Spring AI, Vercel AI SDK, eve, LiteLLM, BAML, TanStack AI, DSPy, Langfuse and other framework-shipped Jev support Patterns: marketing, sales, GTM, content, support and ops P22, Patterns: agent internals, routing, gates, context and memory P06 page
Lead qualification guide (Flowgrammer, @Flowgrammers; consultancy, vendor) A design, not a deployment: fit, timing and intent as separate questions over the lead's own words; weights and routes in code; uncertain, missing, failed or conflicting answers go to a person; never auto-reject without a human path. Notes Quebec Law 25 duties for decisions based only on automated processing (their reading; not legal advice) No Flowgrammer result yet (their own lead test promised). Its digest of public lead builds points at rows already in this wiki (Patterns: marketing, sales, GTM, content, support and ops P19). One own observation: a local classifier spot check, English only, 28.7% on urgency (tool unnamed; unverified) Patterns: marketing, sales, GTM, content, support and ops P19, Patterns: agent internals, routing, gates, context and memory P06 guide
ToolJet bank reconciliation (@Athulya___, ToolJet) A ToolJet app built through ToolJet's MCP with Jev: flags each transaction matched, missing or duplicate; a person reviews before save Demo, video not transcribed; no numbers. How Jev is wired is not shown Patterns: marketing, sales, GTM, content, support and ops P23 post
Kitsuno job pipeline (Gregory Turkawka, @GTurkawka, Substack, 2026-10-03; production since 2026-09-23) Gates in a job platform: a check at the door before any LLM reads a job ad, role checks in its Be Found matching, match judgments after extraction; per ad, questions on country, work mode, geographic scope, seniority, language and employment type Since 2026-09-23: 35,812 logged calls, 76M input tokens, $3.68 (does not reconcile as stated: 76M × $0.042/M = $3.19; $3.68 is ~87.6M tokens, our arithmetic); 30,738 calls in the last 7 days. All agents' LLM calls ~4,400/day (2026-09-09 to 22) → ~2,000/day (09-24 to 10-02) while new jobs per day nearly doubled; LLM tokens ~11M → ~5M/day (author: mixed traffic, "a direction, not a controlled test"). Open Laya run in shadow (open weights, not Jev; Open replicas and Jev-compatible servers): untrained 44% vs Jev 94% on their test set; after 8 days of training on frontier-model labels plus his review, never on Jev answers (he reads the customer agreement as forbidding it; Warnings (b), Warnings: not-Jev services, key safety, look-alikes and install names): 95.8% vs 93.7% on 574 decisions over two test sets (31 right only for Laya, 19 only for Jev, p = 0.12, not significant); 93.4% each on the set never used to pick a model; his own 53 judgments: Laya 48, Jev 36. Caveat he names: the same frontier model wrote Laya's labels and the first draft of the reference answers, which favours Laya. Jev reads a country in the location line as "only that country" where their question calls it unclear (15-0 for Laya); seniority 9-5 for Jev. When Laya is ≥ 0.9 sure (86% of decisions) it is right 97.8% vs Jev 95.7% on the same decisions; Laya seniority 96% sure, 82% right. Production agreement with Jev 89-97% by question type. Laya 12.7 s per intake item → ~4 s (CPU) vs Jev 0.29 s per call. Reason to switch: EU data residency and open models, not cost; Laya does not decide yet; a data-retention question to TypeSafe unanswered since September. One author's runs; unverified Patterns: judging, search, documents, real-time and markets P18, Patterns: agent internals, routing, gates, context and memory P01 post, article
risk-analysis-server (Edward Irby, @edwardirby, You.com's open-source org; X article "Reviewed by @typesafeai", 2026-10-01; MIT; vendor) Background risk-monitoring agent (MCP server): four Jev gates per sweep, each one batched request: a triage Noul on search highlights (escalate at ≥ 0.5, knob RISK_TRIAGE_THRESHOLD); a Noul ranking all Qwen-proposed queries (top 8 searched); a 3-level relevance Score on 30 results; a low / medium / critical Choice. Qwen3.8 27B (OpenRouter) only proposes queries and writes the briefing; code owns thresholds, budgets and caps. RISK_JUDGE=jev|qwen swaps the judge Author's runs, one news day: 2¢ of Qwen per escalated sweep, $0.14 all-in per Jev-judged sweep; search ~6× the model bill; Jev ~¼¢ for three sweeps, ~$0.0005 per deep dive (13k input tokens). Ranking queries before searching cut 19-26 searches per sweep to 12. Threshold 0.7 instead of 0.5 dropped the day's critical signal (two of three profiles sat at 0.56-0.63); kept 0.5. Judge swap: Jev triage stable within ±0.05 per profile, escalated 9 of 9 sweeps; Qwen as judge (same questions, strict JSON, 0 parse failures) escalated 6 of 11, one profile read 0.35, 0.68, 0.50 on consecutive runs, another 0.35-0.40 where Jev read 0.63-0.65; Qwen judging $0.10-0.18 per deep dive (~250×), 7-12 min vs ~2; blind human rubric on 15 reports tied 4.25 vs 4.25. Author's limits: one reader, one day, 9 vs 6 reports. Primitive shapes and "Noul least certain near 0.5" verified (Primitives: Choice, Score, Noul, Noul (yes/no) questions); numbers unverified Patterns: agent internals, routing, gates, context and memory P36, Patterns: judging, search, documents, real-time and markets P17 article, repo
IssueRelay (tutorial, repo; Andrew Baisden, freeCodeCamp, 2026-10-02) A support widget: Jev suggests type and severity (only message and topic sent), code routes, and only a bug above a 0.9 cut becomes a GitHub issue (author: an uncalibrated starting point); uses the chosen label's probability as confidence; human overrides stored separately 3 live reports classified as intended at 0.95-1.00 (unverified, n=3) Patterns: marketing, sales, GTM, content, support and ops P22 tutorial

Also seen, names only (vetted 2026-09-25): Jev demos (mayank953: six side-by-side demos where Jev decides and Claude or Kimi writes; when hosted, visitors' keys pass through the app server; the listed instance returned 404 on 2026-09-25). Added 2026-10-01: Archaic's enterprise data-analysis platform (Tokyo; press release, Japanese, via @PRTIMES_TECH): Jev sorts logs, transactions, tickets and records grouped by user, device or lot, an LLM analyses only flagged cases, people confirm; sync or batch; no numbers, real-data validation still to come.

Markets (engineering only, not financial advice)

Build (builder) Jev decides Reported numbers Pattern Source
fedjev-bench (maybern-tripp-smith) Pairwise Choice "more hawkish about inflation" over FOMC press-conference openings, checked against rate actions and FedLock scores Main run ≈ $0.031; gates 1, 3, 4, 6 pass, 2/5/7 reported; 7-criterion extension $1.5329. Text ranking, not a trade signal; buy/sell (Patterns: judging, search, documents, real-time and markets P25) stays poor Patterns: judging, search, documents, real-time and markets P17 repo

Also seen, names only (vetted 2026-09-25): QuantStudio (quantskills: local research and trading workbench; Jev grades trade plans on a futures-simulation contest account and every trade needs the user's confirmation; buy/sell stays P25 poor; described as engineering only, not advice).

Sources

Files in frontmatter sources:, captured by 2026-10-04; original URLs inline.