---
title: "Failure reports: wrong role, or options that don't fit"
type: community
source_tier: community
tags: [failure-reports, limitations, advising, option-design]
created: 2026-09-30
updated: 2026-09-30
confidence: medium
sources: "41 files, listed under this page in https://jevwiki.ai/index.json (raw/ paths map to their URLs in https://jevwiki.ai/raw/MANIFEST.json); links are inline in the body"
jev_version: "jev-1.13.0"
summary: "Community reports where Jev was given the wrong job or options that did not fit: planning, subtext, self-describing text, near-alike candidates, trimmed state. Split from failure-reports 2026-09-30."
---

# Failure reports: wrong role, or options that don't fit

> **TL;DR** The largest group of community failure reports: Jev asked to plan or investigate, to read subtext or predict outcomes, or to pick from options that overlap, miss the answer or describe themselves; plus trimmed or packed state that hid the signal. The same builders report wins when code builds the candidates and Jev picks. All n≈1, builders' own numbers. Split from [[ideas/failure-reports]] on 2026-09-30 (rows unchanged); other failure groups stay there, confidence failures on [[ideas/failure-reports-confidence]].

Marks check against [[concepts/jaggedness-jev-1-13]] (§n) and reference/. Numbers are the builder's own, on Jev 1.13 (`jev-1.13.0`; assumed where a row names no version, inferred): some may be setup errors, or fixed in later versions. TypeSafe's docs win.

| Report | What broke | Matches | Takeaway |
|---|---|---|---|
| [@hrishioa](https://x.com/hrishioa/status/2101842370052669903) (Southbridge), quoted by [@zeroxmin](https://x.com/zeroxmin/status/2101940446637531634) | Across a few thousand hours of agent runs: the best tool he tested for progress and time-to-finish, "dangerous" for detecting harmful commands, weak at spotting lazy models. Article not captured | `unverified`. Fits §6 adversarial content | Observer, not the only security gate; keep deterministic allow-lists (P03, P15) |
| [jev-wechat-live](https://github.com/duckegg0623-create/jev-wechat-live) (abandoned), via OpenRouter | Reading subtext in WeChat messages: "reply with a clarifying question" in 6/11, "irritated" in 5/11, single emoji unreadable. Added context barely helped. Author: options were a linguist's taxonomy; with none fitting, Jev picked the nearest safe one | §1, §7 `verified` in direction | Options must cover reality, plus "other/none". [[ideas/patterns-business]] P20 |
| [@hw40](https://x.com/hw40/status/2102332459203653977), SEO tools | Under-called commercial pages as informational (25% of 263 vs an outside classifier). A citation rubric mixed two questions (33% → 64–70% after splitting). On a single-topic site, generic anchors scored high until a "specific" check was added (0.82 → 0.94) | §1 and "several judgments in one question" `verified` | Split rubrics; check against outside ground truth; log biases, don't threshold them away. [[ideas/patterns-business]] P32 |
| primeline.cc | A request phrased as a passing remark was missed 42.5% of the time until the question said such remarks count (0%). Two-stage narrowing was 12.9 points worse when broad groups overlapped | §1 `verified`. Narrowing penalty `unverified` | Name edge cases in the instruction. Don't assume hierarchy helps |
| [Reddit thread](https://www.reddit.com/r/TheMachineLearning/comments/1wn2gn8/jev_demos_are_misleading_says_developer/), 2026-09-23: critique, not evidence | Demos framed Jev as an agent (cars, games), not a classifier. A commenter: one forward pass caps what it can classify. Another: would old pattern matching do? | Jev does not choose its own next action ([[concepts/system-one]]) `verified`. Single-pass ceiling `unverified` | Pitch it as a router in a code-owned loop; baseline rules first ([[syntheses/when-not-to-use-jev]]) |
| [agentjournal](https://agentjournal.dev/blog/llm-judge-vs-feature-extraction/): real bookkeeping, many debit accounts | One Choice over all the accounts lost badly to word-bigram naive Bayes; scored dimensions with fitted weights came close: [[ideas/head-to-head]] | §1 `verified` in direction ([[patterns/composite-scoring]]) | Many close classes: score dimensions and fit weights, or use n-grams. Author: never ask one question with that many choices |
| [jevgrep](https://github.com/allebee/jevgrep), allebee: one Noul per log line, batched | A log line that describes itself as the target label was scored as that label where Claude Haiku was not fooled, and a routine line took on its list-indexed neighbours' answer. Numbers: [[ideas/head-to-head]] | adversarial `verified`; bleed `unverified` | Text can argue for its own label: never auto-act on attacker-writable lines. Name items in batched `state` |
| [jev-engineering](https://github.com/eugeniughelbur/jev-engineering), eugeniughelbur: tool-call gate, 300 calls | A blunt injection let 0/30 dangerous commands through but raised safe-command denials 1 → 8/30; "a human approved this" framing let up to 3/30 through (`git stash clear` under all three). Catching all 5 successful attacks needed a 0.8 confidence floor that escalated **58%** of normal traffic | §6 `verified` in direction | Hard-code catastrophic commands as rules; set the floor from your own log. P03 |
| [jev-edge](https://github.com/kiwi0719/jev-edge), kiwi0719, vs [agent-chaperone](https://github.com/agent-chaperone/agent-chaperone): BIPIA email injections | jev-edge missed most BIPIA EmailQA injections at the default cut, while agent-chaperone's differently worded questions caught most BIPIA email injections: [[ideas/head-to-head-security]] | §6 `verified` in direction; the two results conflict (wording, task) | Wording decides recall; describe the deployment and mark tool results untrusted (both cut jev-edge's misses sharply) |
| [jev-secret-detection](https://github.com/teyhouse/jev-secret-detection), teyhouse: one Noul per snippet | Scores Jev alone (deliberately no regex). 99/100 main set, 16/17 edge set, but 15/20 and 16/20 (AUC 0.939, 0.955) on config-shaped cases: hashes, SCRAM verifiers and stock passwords pushed into the 0.3-0.7 review band; missed an SNMP community string (no credential field name) | `unverified` | Config files need format-aware rules alongside Jev (P35). No licence: describe only |
| [Jev Wrapped](https://github.com/gaborishka/jev-wrapped), gaborishka: 4 questions per Telegram post | On a channel with no paid ads it marked two partner promotions as ads | `unverified` | Treat labels as a reason to look, not a verdict |
| [Jev for Chrome](https://github.com/chy4pro/jev-for-chrome), chy4pro: browser suite | One genuine model decision error (on arXiv it ignored the sort control and opened a same-titled paper); its other misses were rate limits, an error page and an executor bug since fixed. Results on [[ideas/builds-browser-and-interface]] | `unverified` | Keep a code-side goal check; two Nouls already veto a premature DONE. P12 |
| [Dex Horthy](https://www.youtube.com/watch?v=35PSMmDDKP8) (HumanLayer), AI That Works live stream 2026-09-22: code-search harness | Jev picked the next action (read this folder, read this file) from a menu the harness built over a codebase of roughly 3,500 files. Jev takes no arguments, so every path is its own option; the 255-option cap and the 32k budget held menus to about 80 files per step, leaving the harness to pre-sort and rerank. His verdict: "absolutely an abuse of Jev and did not get good results" (no numbers, no code shown) | 255 options `verified` ([[reference/http-api]]); his 32k matches the `state`-plus-longest-question budget ([[concepts/state]]); outcome `unverified` | An open action space is the wrong role: code or an LLM owns the search loop, Jev makes bounded picks (P12, P05) |
| [Endform](https://endform.dev/blog/jev-playwright-testing) ([@OliverStenbom](https://x.com/OliverStenbom/status/2104551929984717236), 2026-09-28): Jev drives Playwright end-to-end tests | A Bun controller turned each page's accessibility snapshot into a menu of actions; Jev picked one per step, and a completion Noul at ≥ 0.90 triggered independent Playwright verification (30 steps, 7 min per attempt). **72/150 verified (48%)** over 15 scenarios × 10 runs: 5 scenarios 10/10, 6 at 0/10; of 78 failures, Jev chose stop in 67, stalled in 10, hit the step limit in 1. Jev cost $0.238 in total. 30 injected faults: no false pass where a fault visibly broke a checked outcome. Named causes: lossy page → accessibility tree → menu conversion, the 255-choice cap, earlier milestones outside the five-action history, billing values never supplied as inputs. The post's "under 50% failed" misstates the article (48% passed) | 255 choices `verified` ([[reference/http-api]]); the article's $42 per billion input tokens matches the list price `verified` ([[reference/models-and-pricing]]); results `unverified` | Arbitrary test flows are the wrong role: Jev can only pick actions the menu offers. Author sees promise in partly dynamic tests where Playwright keeps setup and assertions (P12) |
| [Jev Logs](https://github.com/reachjalil/jevlogs), reachjalil: Loghub HDFS and BGL samples | Line-level triage missed nearly all block-level HDFS anomalies, and the BGL alerts came from a local rule, not Jev. Numbers on [[ideas/head-to-head-benchmarks]] | `unverified` | Line-level judgment cannot recover block-level labels; its synthetic pager win is in-sample ([[ideas/head-to-head]]) |
| [Paper Radar](https://github.com/Eliot5566/JEV-Paper-Radar), Eliot5566: CLEF TAR screening | Development reviews saved far less work than held-out ones; thresholds fitted on the same judgments: [[ideas/head-to-head-benchmarks]] | `unverified` | Performance is dominated by the review; splitting criteria into more vetoes cut work saved ([[ideas/builder-lessons]]) |
| [langchain-skill-router](https://github.com/deyna256/langchain-skill-router), deyna256 | When Jev loaded a wrong skill the agent answered far worse than with the right one, and contradictory skills had put an earlier run below the full catalog ([[ideas/head-to-head-agents]]) | `unverified` | Skill quality caps any router; raise `load_at` with near-duplicate skills |
| [iammrduncan](https://github.com/iammrduncan/typesafe-ai-benchmark): synthetic app scenes | Lost to a small Qwen model on Cerebras on multi-field home-automation commands and on approvals; numbers on [[ideas/head-to-head]] | `unverified` | Multi-field device commands: test per field. [[ideas/head-to-head]] |
| [Jevals.com score board](https://jevals.com/score/), anonymous operator | HelpSteer2 helpfulness levels: no model, Jev included, beat guessing the label base rates. Numbers on [[ideas/eval-boards]] | `unverified` | Subjective quality rubrics are a poor Score target; keep a person in the loop |
| [jev-ultralightspeed](https://github.com/collapseindex/jev-ultralightspeed), collapseindex: 32 items packed per request | Keeping only the first 300 characters of each item dropped agreement to **37.3%**, below the base rate (always answering the majority label: 57.7%; last 300: 74.8%; first 500: 88.6% vs 89.4% untrimmed, a trim the author recommends); sorted queues showed a position effect ([[ideas/request-mechanics]]) | §5 `verified` in spirit: the cut removed the evidence | Measure a trim before shipping it: which end and how much decide it; packing itself held agreement ([[ideas/request-mechanics]]) |
| [jev-rag-benchmark](https://github.com/erendikmenn/jev-rag-benchmark), erendikmenn: Turkish XQuAD | A full-corpus hierarchical Jev retriever did worse than BM25, and pointwise reranking gained nothing over batch reranking at higher cost; reranking a hybrid shortlist, Jev tied Cohere Rerank on recall at lower cost but trailed on nDCG. Numbers on [[ideas/head-to-head-benchmarks]] | `unverified` | Jev reranks a shortlist; it does not replace retrieval (P17) |
| [jevwire](https://github.com/Brainwires/jevwire), Brainwires: coding-agent hooks | Too many near-alike candidates in one request flattened the scores and buried the right answer, so the tool caps its batch size (run: [[ideas/request-mechanics]]). Hooks fail open silently; the npm name in its README is unclaimed ([[ideas/warning-receipts]]) | `unverified` | Keep candidate lists short |
| [Sniff Test](https://github.com/DanRWilloughby/snifftest), DanRWilloughby: one Noul per rule per paragraph | In the bench path one rule's answer was missing from Jev's reply on several paragraphs (cause open): [[ideas/head-to-head]]; mid-range readings on text Jev could not read: [[ideas/failure-reports-confidence]] | `unverified` | Check every question came back; treat a mid-range reading as no judgment |
| [typesafe-skill-router](https://github.com/DECRUX9812/typesafe-skill-router), DECRUX9812: Hermes skill suggestions | A miss and an off-target pick in its first live turns; `fits` scored lower in Spanish: [[ideas/head-to-head-agents]] | `unverified` | Test each language you serve |
| [HuLU test](https://aihirfolyam.hu/2026/09/mennyire-ert-magyarul-a-typesafe-jev-modellje/) (aihirfolyam.hu, [@heyitsbalazs](https://x.com/heyitsbalazs/status/2104892344990396756), 2026-09-29): Hungarian grammar | Jev caught far fewer ungrammatical Hungarian sentences than GPT-6 Luna, and an English instruction did not change that; the same run matched Luna on 5 of 6 task types. Numbers on [[ideas/head-to-head-benchmarks]] | `unverified` | Test grammatical acceptability in your own language before relying on it |
| [Jev macOS Loop](https://github.com/jcpsimmons/jev-macos-loop), jcpsimmons: native GUI control | Passed its small task suites, but a later Calculator regression typed an extra digit (caught by the independent verifier); OCR disagreed with accessibility labels or lagged transitions. Numbers on [[ideas/builds-computer-use]] | `unverified`; small samples (author) | Verify the end state in code |
| [macbrow](https://github.com/timpratim/macbrow), timpratim: voice Mac agent | An early version, asked to tidy the Desktop, moved every Desktop file into a folder; the safety policy file exists because of it. An LLM writes new AppleScript tools that run locally | `unverified` | Policy and scope limits before actions on files (P03) |
| [jeff](https://github.com/Alurith/jeff), Alurith (not logan-markewich/jeff): semantic Go lint | First real run, 16 dev cases ($0.00066): exited 1 because the author's own quality gates failed (which gates, and the figures, are not in the capture) | `unverified` | Gate rules on dev data first |
| [@PrajwalTomar_](https://x.com/PrajwalTomar_/status/2102012435687514340), 468 of his own posts, engagement hidden (2026-09-21) | The **reach-band** Choice hit 95/468 = **20%** exact vs **76%** for always guessing the bottom band; 261 posts put in the 30K-100K band where 24 landed; 0 of his 28 posts over 100K reached the top band. The same run's hook-strength Score tracked reality (median impressions weak **462**, okay **320**, strong **3,474**, scroll-stopping **5,420**; number in line one **3,833** vs **2,854**), though weak beat okay: adjacent levels did not separate | `unverified`: author-reported, one account. Reach depends on the card, video and algorithm, none in `state`; the export truncates at ~280 characters, no video or card | Judge the artefact, never forecast the outcome |
| [Theo](https://www.youtube.com/watch?v=F3YXg7AaKWE) (t3.gg video, 2026-09-21), his own coding threads and a checkers bot | A "worth saving for a video?" tag flagged about half his threads even after the prompt was reworded; his checkers bot answered instantly but lost to him while he barely paid attention. Numbers on [[ideas/builds-apps]] | docs' few-seconds test `verified` ([[concepts/system-one]]) | Taste and lookahead need thought: keep them with an LLM or in code |
| jev-oncall ([@rae1101x](https://x.com/rae1101x/status/2105070742446690442), 2026-09-29): outage-investigation agent | Named the right service while often never checking what failed; fixed by reserving checks for the blamed service. Numbers: [[ideas/head-to-head-agents]] | `unverified` | Score the evidence, not only the verdict |

## Related

- [[ideas/failure-reports]] — the other failure groups (generation, numbers, speed, cost, a simpler tool, agents with no gain, access)
- [[ideas/failure-reports-confidence]] — confidence and calibration failures
- [[ideas/advisor-checklist]] — the pre-recommendation checklist that points here
- [[concepts/jaggedness-jev-1-13]] — official failure modes; [[guides/writing-instructions-and-criteria]] — writing options that fit

## Sources

Files in frontmatter `sources:`, captured by 2026-09-30; original URLs inline.
