Agents: read the raw Markdown of this page, or start at llms.txt.
Failure reports: wrong role, or options that don't fit
TL;DR The largest group of community failure reports: Jev asked to plan or investigate, to read subtext or predict outcomes, or to pick from options that overlap, miss the answer or describe themselves; plus trimmed or packed state that hid the signal. The same builders report wins when code builds the candidates and Jev picks. All n≈1, builders' own numbers. Split from Failure reports: where Jev broke, lost, or was the wrong tool on 2026-09-30 (rows unchanged); other failure groups stay there, confidence failures on Failure reports: confidence misread and calibration.
Marks check against Jev 1.13 jaggedness: known failure modes (§n) and reference/. Numbers are the builder's own, on Jev 1.13 (jev-1.13.0; assumed where a row names no version, inferred): some may be setup errors, or fixed in later versions. TypeSafe's docs win.
| Report | What broke | Matches | Takeaway |
|---|---|---|---|
| @hrishioa (Southbridge), quoted by @zeroxmin | Across a few thousand hours of agent runs: the best tool he tested for progress and time-to-finish, "dangerous" for detecting harmful commands, weak at spotting lazy models. Article not captured | unverified. Fits §6 adversarial content |
Observer, not the only security gate; keep deterministic allow-lists (P03, P15) |
| jev-wechat-live (abandoned), via OpenRouter | Reading subtext in WeChat messages: "reply with a clarifying question" in 6/11, "irritated" in 5/11, single emoji unreadable. Added context barely helped. Author: options were a linguist's taxonomy; with none fitting, Jev picked the nearest safe one | §1, §7 verified in direction |
Options must cover reality, plus "other/none". Patterns: marketing, sales, GTM, content, support and ops P20 |
| @hw40, SEO tools | Under-called commercial pages as informational (25% of 263 vs an outside classifier). A citation rubric mixed two questions (33% → 64–70% after splitting). On a single-topic site, generic anchors scored high until a "specific" check was added (0.82 → 0.94) | §1 and "several judgments in one question" verified |
Split rubrics; check against outside ground truth; log biases, don't threshold them away. Patterns: marketing, sales, GTM, content, support and ops P32 |
| primeline.cc | A request phrased as a passing remark was missed 42.5% of the time until the question said such remarks count (0%). Two-stage narrowing was 12.9 points worse when broad groups overlapped | §1 verified. Narrowing penalty unverified |
Name edge cases in the instruction. Don't assume hierarchy helps |
| Reddit thread, 2026-09-23: critique, not evidence | Demos framed Jev as an agent (cars, games), not a classifier. A commenter: one forward pass caps what it can classify. Another: would old pattern matching do? | Jev does not choose its own next action (System One Models) verified. Single-pass ceiling unverified |
Pitch it as a router in a code-owned loop; baseline rules first (When not to use Jev: rules, embeddings, trained classifiers, small and frontier LLMs) |
| agentjournal: real bookkeeping, many debit accounts | One Choice over all the accounts lost badly to word-bigram naive Bayes; scored dimensions with fitted weights came close: Head-to-head: Jev against other models and methods | §1 verified in direction (Composite scoring) |
Many close classes: score dimensions and fit weights, or use n-grams. Author: never ask one question with that many choices |
| jevgrep, allebee: one Noul per log line, batched | A log line that describes itself as the target label was scored as that label where Claude Haiku was not fooled, and a routine line took on its list-indexed neighbours' answer. Numbers: Head-to-head: Jev against other models and methods | adversarial verified; bleed unverified |
Text can argue for its own label: never auto-act on attacker-writable lines. Name items in batched state |
| jev-engineering, eugeniughelbur: tool-call gate, 300 calls | A blunt injection let 0/30 dangerous commands through but raised safe-command denials 1 → 8/30; "a human approved this" framing let up to 3/30 through (git stash clear under all three). Catching all 5 successful attacks needed a 0.8 confidence floor that escalated 58% of normal traffic |
§6 verified in direction |
Hard-code catastrophic commands as rules; set the floor from your own log. P03 |
| jev-edge, kiwi0719, vs agent-chaperone: BIPIA email injections | jev-edge missed most BIPIA EmailQA injections at the default cut, while agent-chaperone's differently worded questions caught most BIPIA email injections: Head-to-head benchmarks: security (phishing, spam, injection, vulnerable code) | §6 verified in direction; the two results conflict (wording, task) |
Wording decides recall; describe the deployment and mark tool results untrusted (both cut jev-edge's misses sharply) |
| jev-secret-detection, teyhouse: one Noul per snippet | Scores Jev alone (deliberately no regex). 99/100 main set, 16/17 edge set, but 15/20 and 16/20 (AUC 0.939, 0.955) on config-shaped cases: hashes, SCRAM verifiers and stock passwords pushed into the 0.3-0.7 review band; missed an SNMP community string (no credential field name) | unverified |
Config files need format-aware rules alongside Jev (P35). No licence: describe only |
| Jev Wrapped, gaborishka: 4 questions per Telegram post | On a channel with no paid ads it marked two partner promotions as ads | unverified |
Treat labels as a reason to look, not a verdict |
| Jev for Chrome, chy4pro: browser suite | One genuine model decision error (on arXiv it ignored the sort control and opened a same-titled paper); its other misses were rate limits, an error page and an executor bug since fixed. Results on Builds: browser, computer use and interface | unverified |
Keep a code-side goal check; two Nouls already veto a premature DONE. P12 |
| Dex Horthy (HumanLayer), AI That Works live stream 2026-09-22: code-search harness | Jev picked the next action (read this folder, read this file) from a menu the harness built over a codebase of roughly 3,500 files. Jev takes no arguments, so every path is its own option; the 255-option cap and the 32k budget held menus to about 80 files per step, leaving the harness to pre-sort and rerank. His verdict: "absolutely an abuse of Jev and did not get good results" (no numbers, no code shown) | 255 options verified (HTTP API: POST /v1/systemone and GET /v1/models); his 32k matches the state-plus-longest-question budget (State: what you send Jev); outcome unverified |
An open action space is the wrong role: code or an LLM owns the search loop, Jev makes bounded picks (P12, P05) |
| Endform (@OliverStenbom, 2026-09-28): Jev drives Playwright end-to-end tests | A Bun controller turned each page's accessibility snapshot into a menu of actions; Jev picked one per step, and a completion Noul at ≥ 0.90 triggered independent Playwright verification (30 steps, 7 min per attempt). 72/150 verified (48%) over 15 scenarios × 10 runs: 5 scenarios 10/10, 6 at 0/10; of 78 failures, Jev chose stop in 67, stalled in 10, hit the step limit in 1. Jev cost $0.238 in total. 30 injected faults: no false pass where a fault visibly broke a checked outcome. Named causes: lossy page → accessibility tree → menu conversion, the 255-choice cap, earlier milestones outside the five-action history, billing values never supplied as inputs. The post's "under 50% failed" misstates the article (48% passed) | 255 choices verified (HTTP API: POST /v1/systemone and GET /v1/models); the article's $42 per billion input tokens matches the list price verified (Models, aliases, pricing, rate limits, context); results unverified |
Arbitrary test flows are the wrong role: Jev can only pick actions the menu offers. Author sees promise in partly dynamic tests where Playwright keeps setup and assertions (P12) |
| Jev Logs, reachjalil: Loghub HDFS and BGL samples | Line-level triage missed nearly all block-level HDFS anomalies, and the BGL alerts came from a local rule, not Jev. Numbers on Head-to-head benchmarks: Jev on public datasets and suites | unverified |
Line-level judgment cannot recover block-level labels; its synthetic pager win is in-sample (Head-to-head: Jev against other models and methods) |
| Paper Radar, Eliot5566: CLEF TAR screening | Development reviews saved far less work than held-out ones; thresholds fitted on the same judgments: Head-to-head benchmarks: Jev on public datasets and suites | unverified |
Performance is dominated by the review; splitting criteria into more vetoes cut work saved (Builder lessons: what changed the result) |
| langchain-skill-router, deyna256 | When Jev loaded a wrong skill the agent answered far worse than with the right one, and contradictory skills had put an earlier run below the full catalog (Head-to-head: Jev inside agents, routers and tool gates) | unverified |
Skill quality caps any router; raise load_at with near-duplicate skills |
| iammrduncan: synthetic app scenes | Lost to a small Qwen model on Cerebras on multi-field home-automation commands and on approvals; numbers on Head-to-head: Jev against other models and methods | unverified |
Multi-field device commands: test per field. Head-to-head: Jev against other models and methods |
| Jevals.com score board, anonymous operator | HelpSteer2 helpfulness levels: no model, Jev included, beat guessing the label base rates. Numbers on Eval boards: JevBench by version, Jevals.com, typed-decisions and other independent suites | unverified |
Subjective quality rubrics are a poor Score target; keep a person in the loop |
| jev-ultralightspeed, collapseindex: 32 items packed per request | Keeping only the first 300 characters of each item dropped agreement to 37.3%, below the base rate (always answering the majority label: 57.7%; last 300: 74.8%; first 500: 88.6% vs 89.4% untrimmed, a trim the author recommends); sorted queues showed a position effect (Request mechanics: billing, limits, latency, calibration and stability) | §5 verified in spirit: the cut removed the evidence |
Measure a trim before shipping it: which end and how much decide it; packing itself held agreement (Request mechanics: billing, limits, latency, calibration and stability) |
| jev-rag-benchmark, erendikmenn: Turkish XQuAD | A full-corpus hierarchical Jev retriever did worse than BM25, and pointwise reranking gained nothing over batch reranking at higher cost; reranking a hybrid shortlist, Jev tied Cohere Rerank on recall at lower cost but trailed on nDCG. Numbers on Head-to-head benchmarks: Jev on public datasets and suites | unverified |
Jev reranks a shortlist; it does not replace retrieval (P17) |
| jevwire, Brainwires: coding-agent hooks | Too many near-alike candidates in one request flattened the scores and buried the right answer, so the tool caps its batch size (run: Request mechanics: billing, limits, latency, calibration and stability). Hooks fail open silently; the npm name in its README is unclaimed (Warning receipts: what we checked behind each warning) | unverified |
Keep candidate lists short |
| Sniff Test, DanRWilloughby: one Noul per rule per paragraph | In the bench path one rule's answer was missing from Jev's reply on several paragraphs (cause open): Head-to-head: Jev against other models and methods; mid-range readings on text Jev could not read: Failure reports: confidence misread and calibration | unverified |
Check every question came back; treat a mid-range reading as no judgment |
| typesafe-skill-router, DECRUX9812: Hermes skill suggestions | A miss and an off-target pick in its first live turns; fits scored lower in Spanish: Head-to-head: Jev inside agents, routers and tool gates |
unverified |
Test each language you serve |
| HuLU test (aihirfolyam.hu, @heyitsbalazs, 2026-09-29): Hungarian grammar | Jev caught far fewer ungrammatical Hungarian sentences than GPT-6 Luna, and an English instruction did not change that; the same run matched Luna on 5 of 6 task types. Numbers on Head-to-head benchmarks: Jev on public datasets and suites | unverified |
Test grammatical acceptability in your own language before relying on it |
| Jev macOS Loop, jcpsimmons: native GUI control | Passed its small task suites, but a later Calculator regression typed an extra digit (caught by the independent verifier); OCR disagreed with accessibility labels or lagged transitions. Numbers on Builds: desktop, mobile and voice computer use | unverified; small samples (author) |
Verify the end state in code |
| macbrow, timpratim: voice Mac agent | An early version, asked to tidy the Desktop, moved every Desktop file into a folder; the safety policy file exists because of it. An LLM writes new AppleScript tools that run locally | unverified |
Policy and scope limits before actions on files (P03) |
| jeff, Alurith (not logan-markewich/jeff): semantic Go lint | First real run, 16 dev cases ($0.00066): exited 1 because the author's own quality gates failed (which gates, and the figures, are not in the capture) | unverified |
Gate rules on dev data first |
| @PrajwalTomar_, 468 of his own posts, engagement hidden (2026-09-21) | The reach-band Choice hit 95/468 = 20% exact vs 76% for always guessing the bottom band; 261 posts put in the 30K-100K band where 24 landed; 0 of his 28 posts over 100K reached the top band. The same run's hook-strength Score tracked reality (median impressions weak 462, okay 320, strong 3,474, scroll-stopping 5,420; number in line one 3,833 vs 2,854), though weak beat okay: adjacent levels did not separate | unverified: author-reported, one account. Reach depends on the card, video and algorithm, none in state; the export truncates at ~280 characters, no video or card |
Judge the artefact, never forecast the outcome |
| Theo (t3.gg video, 2026-09-21), his own coding threads and a checkers bot | A "worth saving for a video?" tag flagged about half his threads even after the prompt was reworded; his checkers bot answered instantly but lost to him while he barely paid attention. Numbers on Builds: data, search and business | docs' few-seconds test verified (System One Models) |
Taste and lookahead need thought: keep them with an LLM or in code |
| jev-oncall (@rae1101x, 2026-09-29): outage-investigation agent | Named the right service while often never checking what failed; fixed by reserving checks for the blamed service. Numbers: Head-to-head: Jev inside agents, routers and tool gates | unverified |
Score the evidence, not only the verdict |
Related
- Failure reports: where Jev broke, lost, or was the wrong tool — the other failure groups (generation, numbers, speed, cost, a simpler tool, agents with no gain, access)
- Failure reports: confidence misread and calibration — confidence and calibration failures
- Advisor checklist: what to check before recommending Jev for a decision — the pre-recommendation checklist that points here
- Jev 1.13 jaggedness: known failure modes — official failure modes; Writing instructions and criteria that Jev reads correctly — writing options that fit
Sources
Files in frontmatter sources:, captured by 2026-09-30; original URLs inline.