$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Failure reports: wrong role, or options that don't fit

[ community tier ][ updated 2026-09-30 ][ confidence medium ][ jev-1.13.0 ]#failure-reports · limitations · advising · option-design

TL;DR The largest group of community failure reports: Jev asked to plan or investigate, to read subtext or predict outcomes, or to pick from options that overlap, miss the answer or describe themselves; plus trimmed or packed state that hid the signal. The same builders report wins when code builds the candidates and Jev picks. All n≈1, builders' own numbers. Split from Failure reports: where Jev broke, lost, or was the wrong tool on 2026-09-30 (rows unchanged); other failure groups stay there, confidence failures on Failure reports: confidence misread and calibration.

Marks check against Jev 1.13 jaggedness: known failure modes (§n) and reference/. Numbers are the builder's own, on Jev 1.13 (jev-1.13.0; assumed where a row names no version, inferred): some may be setup errors, or fixed in later versions. TypeSafe's docs win.

Report What broke Matches Takeaway
@hrishioa (Southbridge), quoted by @zeroxmin Across a few thousand hours of agent runs: the best tool he tested for progress and time-to-finish, "dangerous" for detecting harmful commands, weak at spotting lazy models. Article not captured unverified. Fits §6 adversarial content Observer, not the only security gate; keep deterministic allow-lists (P03, P15)
jev-wechat-live (abandoned), via OpenRouter Reading subtext in WeChat messages: "reply with a clarifying question" in 6/11, "irritated" in 5/11, single emoji unreadable. Added context barely helped. Author: options were a linguist's taxonomy; with none fitting, Jev picked the nearest safe one §1, §7 verified in direction Options must cover reality, plus "other/none". Patterns: marketing, sales, GTM, content, support and ops P20
@hw40, SEO tools Under-called commercial pages as informational (25% of 263 vs an outside classifier). A citation rubric mixed two questions (33% → 64–70% after splitting). On a single-topic site, generic anchors scored high until a "specific" check was added (0.82 → 0.94) §1 and "several judgments in one question" verified Split rubrics; check against outside ground truth; log biases, don't threshold them away. Patterns: marketing, sales, GTM, content, support and ops P32
primeline.cc A request phrased as a passing remark was missed 42.5% of the time until the question said such remarks count (0%). Two-stage narrowing was 12.9 points worse when broad groups overlapped §1 verified. Narrowing penalty unverified Name edge cases in the instruction. Don't assume hierarchy helps
Reddit thread, 2026-09-23: critique, not evidence Demos framed Jev as an agent (cars, games), not a classifier. A commenter: one forward pass caps what it can classify. Another: would old pattern matching do? Jev does not choose its own next action (System One Models) verified. Single-pass ceiling unverified Pitch it as a router in a code-owned loop; baseline rules first (When not to use Jev: rules, embeddings, trained classifiers, small and frontier LLMs)
agentjournal: real bookkeeping, many debit accounts One Choice over all the accounts lost badly to word-bigram naive Bayes; scored dimensions with fitted weights came close: Head-to-head: Jev against other models and methods §1 verified in direction (Composite scoring) Many close classes: score dimensions and fit weights, or use n-grams. Author: never ask one question with that many choices
jevgrep, allebee: one Noul per log line, batched A log line that describes itself as the target label was scored as that label where Claude Haiku was not fooled, and a routine line took on its list-indexed neighbours' answer. Numbers: Head-to-head: Jev against other models and methods adversarial verified; bleed unverified Text can argue for its own label: never auto-act on attacker-writable lines. Name items in batched state
jev-engineering, eugeniughelbur: tool-call gate, 300 calls A blunt injection let 0/30 dangerous commands through but raised safe-command denials 1 → 8/30; "a human approved this" framing let up to 3/30 through (git stash clear under all three). Catching all 5 successful attacks needed a 0.8 confidence floor that escalated 58% of normal traffic §6 verified in direction Hard-code catastrophic commands as rules; set the floor from your own log. P03
jev-edge, kiwi0719, vs agent-chaperone: BIPIA email injections jev-edge missed most BIPIA EmailQA injections at the default cut, while agent-chaperone's differently worded questions caught most BIPIA email injections: Head-to-head benchmarks: security (phishing, spam, injection, vulnerable code) §6 verified in direction; the two results conflict (wording, task) Wording decides recall; describe the deployment and mark tool results untrusted (both cut jev-edge's misses sharply)
jev-secret-detection, teyhouse: one Noul per snippet Scores Jev alone (deliberately no regex). 99/100 main set, 16/17 edge set, but 15/20 and 16/20 (AUC 0.939, 0.955) on config-shaped cases: hashes, SCRAM verifiers and stock passwords pushed into the 0.3-0.7 review band; missed an SNMP community string (no credential field name) unverified Config files need format-aware rules alongside Jev (P35). No licence: describe only
Jev Wrapped, gaborishka: 4 questions per Telegram post On a channel with no paid ads it marked two partner promotions as ads unverified Treat labels as a reason to look, not a verdict
Jev for Chrome, chy4pro: browser suite One genuine model decision error (on arXiv it ignored the sort control and opened a same-titled paper); its other misses were rate limits, an error page and an executor bug since fixed. Results on Builds: browser, computer use and interface unverified Keep a code-side goal check; two Nouls already veto a premature DONE. P12
Dex Horthy (HumanLayer), AI That Works live stream 2026-09-22: code-search harness Jev picked the next action (read this folder, read this file) from a menu the harness built over a codebase of roughly 3,500 files. Jev takes no arguments, so every path is its own option; the 255-option cap and the 32k budget held menus to about 80 files per step, leaving the harness to pre-sort and rerank. His verdict: "absolutely an abuse of Jev and did not get good results" (no numbers, no code shown) 255 options verified (HTTP API: POST /v1/systemone and GET /v1/models); his 32k matches the state-plus-longest-question budget (State: what you send Jev); outcome unverified An open action space is the wrong role: code or an LLM owns the search loop, Jev makes bounded picks (P12, P05)
Endform (@OliverStenbom, 2026-09-28): Jev drives Playwright end-to-end tests A Bun controller turned each page's accessibility snapshot into a menu of actions; Jev picked one per step, and a completion Noul at ≥ 0.90 triggered independent Playwright verification (30 steps, 7 min per attempt). 72/150 verified (48%) over 15 scenarios × 10 runs: 5 scenarios 10/10, 6 at 0/10; of 78 failures, Jev chose stop in 67, stalled in 10, hit the step limit in 1. Jev cost $0.238 in total. 30 injected faults: no false pass where a fault visibly broke a checked outcome. Named causes: lossy page → accessibility tree → menu conversion, the 255-choice cap, earlier milestones outside the five-action history, billing values never supplied as inputs. The post's "under 50% failed" misstates the article (48% passed) 255 choices verified (HTTP API: POST /v1/systemone and GET /v1/models); the article's $42 per billion input tokens matches the list price verified (Models, aliases, pricing, rate limits, context); results unverified Arbitrary test flows are the wrong role: Jev can only pick actions the menu offers. Author sees promise in partly dynamic tests where Playwright keeps setup and assertions (P12)
Jev Logs, reachjalil: Loghub HDFS and BGL samples Line-level triage missed nearly all block-level HDFS anomalies, and the BGL alerts came from a local rule, not Jev. Numbers on Head-to-head benchmarks: Jev on public datasets and suites unverified Line-level judgment cannot recover block-level labels; its synthetic pager win is in-sample (Head-to-head: Jev against other models and methods)
Paper Radar, Eliot5566: CLEF TAR screening Development reviews saved far less work than held-out ones; thresholds fitted on the same judgments: Head-to-head benchmarks: Jev on public datasets and suites unverified Performance is dominated by the review; splitting criteria into more vetoes cut work saved (Builder lessons: what changed the result)
langchain-skill-router, deyna256 When Jev loaded a wrong skill the agent answered far worse than with the right one, and contradictory skills had put an earlier run below the full catalog (Head-to-head: Jev inside agents, routers and tool gates) unverified Skill quality caps any router; raise load_at with near-duplicate skills
iammrduncan: synthetic app scenes Lost to a small Qwen model on Cerebras on multi-field home-automation commands and on approvals; numbers on Head-to-head: Jev against other models and methods unverified Multi-field device commands: test per field. Head-to-head: Jev against other models and methods
Jevals.com score board, anonymous operator HelpSteer2 helpfulness levels: no model, Jev included, beat guessing the label base rates. Numbers on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites unverified Subjective quality rubrics are a poor Score target; keep a person in the loop
jev-ultralightspeed, collapseindex: 32 items packed per request Keeping only the first 300 characters of each item dropped agreement to 37.3%, below the base rate (always answering the majority label: 57.7%; last 300: 74.8%; first 500: 88.6% vs 89.4% untrimmed, a trim the author recommends); sorted queues showed a position effect (Request mechanics: billing, limits, latency, calibration and stability) §5 verified in spirit: the cut removed the evidence Measure a trim before shipping it: which end and how much decide it; packing itself held agreement (Request mechanics: billing, limits, latency, calibration and stability)
jev-rag-benchmark, erendikmenn: Turkish XQuAD A full-corpus hierarchical Jev retriever did worse than BM25, and pointwise reranking gained nothing over batch reranking at higher cost; reranking a hybrid shortlist, Jev tied Cohere Rerank on recall at lower cost but trailed on nDCG. Numbers on Head-to-head benchmarks: Jev on public datasets and suites unverified Jev reranks a shortlist; it does not replace retrieval (P17)
jevwire, Brainwires: coding-agent hooks Too many near-alike candidates in one request flattened the scores and buried the right answer, so the tool caps its batch size (run: Request mechanics: billing, limits, latency, calibration and stability). Hooks fail open silently; the npm name in its README is unclaimed (Warning receipts: what we checked behind each warning) unverified Keep candidate lists short
Sniff Test, DanRWilloughby: one Noul per rule per paragraph In the bench path one rule's answer was missing from Jev's reply on several paragraphs (cause open): Head-to-head: Jev against other models and methods; mid-range readings on text Jev could not read: Failure reports: confidence misread and calibration unverified Check every question came back; treat a mid-range reading as no judgment
typesafe-skill-router, DECRUX9812: Hermes skill suggestions A miss and an off-target pick in its first live turns; fits scored lower in Spanish: Head-to-head: Jev inside agents, routers and tool gates unverified Test each language you serve
HuLU test (aihirfolyam.hu, @heyitsbalazs, 2026-09-29): Hungarian grammar Jev caught far fewer ungrammatical Hungarian sentences than GPT-6 Luna, and an English instruction did not change that; the same run matched Luna on 5 of 6 task types. Numbers on Head-to-head benchmarks: Jev on public datasets and suites unverified Test grammatical acceptability in your own language before relying on it
Jev macOS Loop, jcpsimmons: native GUI control Passed its small task suites, but a later Calculator regression typed an extra digit (caught by the independent verifier); OCR disagreed with accessibility labels or lagged transitions. Numbers on Builds: desktop, mobile and voice computer use unverified; small samples (author) Verify the end state in code
macbrow, timpratim: voice Mac agent An early version, asked to tidy the Desktop, moved every Desktop file into a folder; the safety policy file exists because of it. An LLM writes new AppleScript tools that run locally unverified Policy and scope limits before actions on files (P03)
jeff, Alurith (not logan-markewich/jeff): semantic Go lint First real run, 16 dev cases ($0.00066): exited 1 because the author's own quality gates failed (which gates, and the figures, are not in the capture) unverified Gate rules on dev data first
@PrajwalTomar_, 468 of his own posts, engagement hidden (2026-09-21) The reach-band Choice hit 95/468 = 20% exact vs 76% for always guessing the bottom band; 261 posts put in the 30K-100K band where 24 landed; 0 of his 28 posts over 100K reached the top band. The same run's hook-strength Score tracked reality (median impressions weak 462, okay 320, strong 3,474, scroll-stopping 5,420; number in line one 3,833 vs 2,854), though weak beat okay: adjacent levels did not separate unverified: author-reported, one account. Reach depends on the card, video and algorithm, none in state; the export truncates at ~280 characters, no video or card Judge the artefact, never forecast the outcome
Theo (t3.gg video, 2026-09-21), his own coding threads and a checkers bot A "worth saving for a video?" tag flagged about half his threads even after the prompt was reworded; his checkers bot answered instantly but lost to him while he barely paid attention. Numbers on Builds: data, search and business docs' few-seconds test verified (System One Models) Taste and lookahead need thought: keep them with an LLM or in code
jev-oncall (@rae1101x, 2026-09-29): outage-investigation agent Named the right service while often never checking what failed; fixed by reserving checks for the blamed service. Numbers: Head-to-head: Jev inside agents, routers and tool gates unverified Score the evidence, not only the verdict

Sources

Files in frontmatter sources:, captured by 2026-09-30; original URLs inline.