---
title: "Jev inside your agent harness: hooks, tools and fail-safes"
type: guide
source_tier: mixed
tags: [agents, harness, hooks, guardrails, routing, compaction]
created: 2026-10-05
updated: 2026-10-05
confidence: medium
sources:
  - raw/docs/introduction__coding-agents.md
  - raw/docs/models.md
  - raw/docs/api.md
  - raw/docs/confidence.md
  - raw/community/claude-code-docs-hooks.md
  - raw/community/claude-code-docs-plugins-reference.md
  - raw/community/claude-code-docs-mods-overview.md
  - raw/community/claude-code-docs-mods-events.md
  - raw/community/claude-code-docs-mods-reference.md
  - raw/community/openai-codex-docs-hooks.md
  - raw/community/cursor-docs-hooks.md
  - raw/community/opencode-docs-plugins.md
  - raw/x-repos/anomalyco__opencode__plugin-index.md
  - raw/x-repos/earendil-works__pi__docs-extensions.md
  - raw/x-repos/earendil-works__pi__extensions-types.md
  - raw/x-repos/earendil-works__pi__docs-virtual-models.md
  - raw/x-repos/earendil-works__pi__docs-compaction.md
  - raw/x-repos/earendil-works__pi__examples-jev-router.md
  - raw/x-repos/earendil-works__pi__ai-providers-typesafe.md
  - raw/x-repos/earendil-works__pi__ai-system-one-shared.md
  - raw/x-repos/disler__ten-levels-of-jev.md
  - raw/x-repos/disler__ten-levels-of-jev__apps-ten-levels-readme.md
  - raw/x-repos/disler__ten-levels-of-jev__extensions-jev-guard.md
  - raw/x-repos/disler__ten-levels-of-jev__extensions-jev-compact.md
  - raw/x-repos/disler__ten-levels-of-jev__extensions-ask-jev-file.md
  - raw/x-repos/disler__ten-levels-of-jev__extensions-ask-jev-files.md
  - raw/x-repos/disler__ten-levels-of-jev__extensions-ask-jev.md
  - raw/x-repos/disler__ten-levels-of-jev__extensions-report.md
  - raw/x-repos/disler__ten-levels-of-jev__src-core-client.md
  - raw/x-repos/disler__ten-levels-of-jev__src-core-types-helpers.md
  - raw/x-repos/disler__ten-levels-of-jev__src-levels-level06.md
  - raw/x-repos/disler__ten-levels-of-jev__src-levels-level07.md
  - raw/x-repos/disler__ten-levels-of-jev__src-levels-level08.md
  - raw/x-repos/disler__ten-levels-of-jev__src-levels-level09.md
  - raw/x-repos/disler__ten-levels-of-jev__src-levels-level10.md
  - raw/x-repos/disler__ten-levels-of-jev__skill-cookbook-production.md
  - raw/x-repos/disler__ten-levels-of-jev__web-lib-cost.md
  - raw/community/youtube-indydevdan-10-levels-of-jev.txt
  - raw/community/youtube-indydevdan-openrouter-state-of-models.txt
jev_version: "jev-1.13.0"
sdk_python: "0.7.2"
summary: "Wiring Jev into Claude Code, Codex, Cursor, Pi or opencode: which hook for which decision, a complete pre-tool gate, Jev as an agent tool, timeouts, fail direction, privacy."
---

# Jev inside your agent harness: hooks, tools and fail-safes

> **TL;DR** Keep the LLM as the agent's brain. Use Jev for the narrow, repeated decisions a harness makes around it: in **hooks the model never sees** (block a tool call, screen a tool result, route the prompt to a model, decide when to compact, check "done"), or as **tools the model calls** that answer a typed question about files or command output without loading them into its context. Each harness decides what happens when a hook errors or times out. So set your own Jev time budget, choose fail-open or fail-closed per hook on purpose, and remember that hook input is conversation text you are sending to TypeSafe or a gateway. Community tier: harness contracts come from each vendor's docs; designs and numbers come from builders, chiefly IndyDevDan's [ten-levels-of-jev](https://github.com/disler/ten-levels-of-jev).

## What "Jev in the harness" means (and does not)

TypeSafe's own position: Jev cannot replace the model behind Claude Code, Cursor or Codex, but it is "often exactly the right tool *inside* an agent or app you're building" ([[concepts/jev-with-coding-agents]]). The harness is the code around the LLM: the loop, the tool runner, permissions, compaction, model choice. Those are full of fixed-option decisions that today are hard-coded rules, a human prompt, or an extra LLM call. Those decisions are where Jev goes.

There are three shapes, in rising order of autonomy (the ladder of IndyDevDan's video, [transcript](https://www.youtube.com/watch?v=_U-O5lYhJ7Q), levels 6-10):

| Shape | Who decides to call Jev | Examples | Pattern |
|---|---|---|---|
| **Hook** (invisible to the model) | harness code, on an event | pre-tool gate, result screen, prompt router, compaction trigger, done-check | [[ideas/patterns-agents]] P02, P03, P04, P07, P08 |
| **Tool** (fixed questions) | the model, via a tool you register | "is this file about auth?" over one file or a glob | P36 (inferred fit) |
| **Agent-authored questions** | the model writes the question JSON | classify a test failure, score the risk of its own diff | [[ideas/patterns-coding-agents]] P11 |

## Step 1: pick the decisions

Good first candidates are decisions that fire often, have a fixed answer set, and whose input fits in `state`:

| Decision | Primitive | Typical action in code |
|---|---|---|
| Is this command read-only, reversible or irreversible? | `Choice` + a destructive-intent `Noul` | block, ask, allow |
| Does this write touch secrets? | `Noul` (+ file-kind `Choice`) | block |
| Does this tool output contain instructions aimed at the agent? | `Noul` | add a "treat as data" banner |
| Is this request simple or complex? | `Choice` | pick model or effort |
| Did the user switch tasks; is the agent at a clean boundary? | `Noul`s + a `Score` | suggest or request compaction |
| Is the agent's "done" backed by evidence in the transcript? | `Noul` | block the stop, rerun checks |

Keep counting, paths, size limits and anything exact in code; Jev handles the judgment ([[concepts/jaggedness-jev-1-13]]). Rules come first and Jev sees only what they cannot settle. Most community gates do this; see [[ideas/builds-gates-and-routers]] for why widening permissions needs hard bounds.

## Step 2: find the hook in your harness

Event names come from each vendor's docs (captured 2026-10-05). The mapping of decision to event is ours.

| Decision | Claude Code | Codex CLI | Cursor | Pi | opencode |
|---|---|---|---|---|---|
| Tool-call gate | `PreToolUse`; mod `tool.call` | `PreToolUse` | `preToolUse`, `beforeShellExecution` | `pi.on("tool_call")` | `tool.execute.before` |
| Replace the human approval | `PermissionRequest`; mod `tool.check` | `PermissionRequest` | `beforeShellExecution` | `tool_call` + `ctx.ui.confirm` | `permission.ask` |
| Screen tool output | `PostToolUse` | `PostToolUse` (block replaces the result) | `postToolUse` | `tool_result` | `tool.execute.after` |
| Model or effort routing | mod `turn.step`; `UserPromptSubmit` (hint only) | custom provider or proxy | none found | `registerVirtualModel` + `ctx.modelRegistry.classify` | `chat.params` |
| Done-check | `Stop` | `Stop` | `stop` | `agent_before_settle` | `session.idle` (observe only) |
| Compaction | `PreCompact` (can block); mod `session.compact` (skip) | `PreCompact` (can stop) | `preCompact` (observe only) | `session_before_compact` (cancel or supply) | `experimental.session.compacting` |
| Skill or context selection | `UserPromptSubmit` `additionalContext` | `UserPromptSubmit` | `beforeSubmitPrompt` (block only) | `before_agent_start`, `context` | `experimental.chat.system.transform` |

Contract details that change the design:

- **Claude Code** command hooks read JSON on stdin and decide with `hookSpecificOutput.permissionDecision` = `allow | deny | ask | defer`. Exit `2` blocks. **Exit 1, or a timed-out `PreToolUse` hook, does not block.** Default timeouts are 600 s for command hooks and 30 s on `UserPromptSubmit`; added context from a timed-out prompt hook is dropped. Plugins can store the key with `userConfig` `sensitive: true` ([Claude Code hooks](https://code.claude.com/docs/en/hooks), plugins reference).
- **Claude Code mods** (in-process JS hooks) are documented, need v2.1.287+, and are on by default. That version ignores `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS`, which older community READMEs ask you to set. A mod hook gets 10 s of its own time; a hook that throws or times out is skipped unless you attach `.catch` (mods overview, events and reference pages).
- **Codex CLI**: `PreToolUse` can deny or rewrite but does **not support `ask` yet**; the hook is marked failed and the call continues. Hooks must be trusted by hash in `/hooks`; the default timeout is 600 s ([Codex hooks](https://developers.openai.com/codex/hooks)).
- **Cursor** fails open unless a hook sets `failClosed: true`, but blocks on invalid JSON from a permission hook. `ask` is not enforced on `preToolUse`. Hook input includes `user_email` ([Cursor hooks](https://cursor.com/docs/agent/hooks)).
- **Pi**: a `tool_call` handler returns `{ block: true, reason }`. **A handler that throws blocks the tool** (fail-safe). Pi also ships Jev as a built-in classifier model (`typesafe` provider, `TYPESAFE_API_KEY`; a 2026-09-24 commit also serves it through OpenRouter and Cloudflare Workers AI) with an official [`jev-router.ts`](https://github.com/earendil-works/pi/blob/5b6c792b424e73edefbfa558b901bcd64788dad2/packages/coding-agent/examples/extensions/jev-router.ts) example; details on [[ideas/framework-integrations]].
- **opencode**: block in `tool.execute.before` by throwing; hooks run in load order; the docs give no timeout rule.

## Step 3: write the hook (a complete Claude Code pre-tool gate)

`.claude/settings.json`:

```json
{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [
          {
            "type": "command",
            "command": "python3 ${CLAUDE_PROJECT_DIR}/.claude/hooks/jev_bash_gate.py",
            "timeout": 10
          }
        ]
      }
    ]
  }
}
```

`.claude/hooks/jev_bash_gate.py` (needs `pip install typesafe-sdk` and `TYPESAFE_API_KEY`):

```python
import json
import os
import sys

from typesafe_sdk import Choice, Noul, RetryPolicy, TypeSafeClient, TypeSafeError

FAIL_CLOSED = os.environ.get("JEV_GATE_FAIL_CLOSED") == "1"


def decide(decision: str, reason: str) -> None:
    print(json.dumps({
        "hookSpecificOutput": {
            "hookEventName": "PreToolUse",
            "permissionDecision": decision,
            "permissionDecisionReason": reason,
        }
    }))
    sys.exit(0)


event = json.load(sys.stdin)
command = event.get("tool_input", {}).get("command", "")

try:
    with TypeSafeClient(model="jev-1.13.0") as client:
        result = client.system_one(
            state={"command": command[:4000], "cwd": event.get("cwd", "")},
            questions={
                "effect": Choice(
                    instructions="What happens to files, history or remote state if this shell command runs?",
                    criteria={
                        "read_only": "Only reads or lists; changes nothing.",
                        "reversible": "Changes things that can be restored or regenerated.",
                        "irreversible": "Deletes, overwrites, force-pushes or publishes something that cannot be restored.",
                    },
                ),
                "destructive": Noul(instructions="Is this command meant to destroy data or history?"),
            },
            retry=RetryPolicy(max_retries=1, timeout=4.0),
            timeout=3.0,
        )
except TypeSafeError as error:
    if FAIL_CLOSED:
        decide("ask", f"Jev gate unavailable ({type(error).__name__}); asking instead of allowing.")
    sys.exit(0)  # fail open: no decision, the normal permission flow continues

effect = result.choices["effect"]
destructive = result.nouls["destructive"].noul

if (effect.choice == "irreversible" and effect.confidence >= 0.6) or destructive >= 0.7:
    decide("deny", f"Jev: {effect.choice} (confidence {effect.confidence:.2f}), destructive {destructive:.2f}.")
if effect.choice != "read_only" or effect.confidence < 0.5:
    decide("ask", f"Jev: {effect.choice} (confidence {effect.confidence:.2f}).")
sys.exit(0)  # read-only and confident: no decision, normal flow
```

Notes on the sample:

- The thresholds are the starting values ten-levels-of-jev's Pi guard uses (block on `irreversible` at confidence ≥ 0.6 or the destructive `Noul` ≥ 0.7). The 0.5 floor follows the docs' own example ([[concepts/confidence]]). Tune them on your own logged commands.
- The gate only narrows: it denies or asks and never returns `allow`, so your permission rules stay in charge. An `allow` would widen permissions; read [[ideas/builds-gates-and-routers]] first.
- `RetryPolicy(timeout=4.0)` caps the whole call including retries, and `timeout=3.0` caps each attempt. The SDK defaults are a 30 s total retry budget and 10 s per attempt ([[reference/python-sdk-retries-errors]]), which can outlast a hook's timeout.
- The model is pinned to `jev-1.13.0` so a new release behind `jev-latest` cannot shift your thresholds ([[guides/production-operations]]).
- A `Noul` has no `confidence`; its probability is the number you threshold ([[concepts/noul]]).
- This is one process per call. To avoid start-up cost, community gates moved to a daemon or an in-process mod or extension ([[ideas/request-mechanics]]).

The same gate in Pi is a `pi.on("tool_call", ...)` handler that returns `{ block: true, reason }`. ten-levels-of-jev's [`jev-guard.ts`](https://github.com/disler/ten-levels-of-jev/blob/777adaf47d37ae0553220d35b2f15b3a3a063305/apps/ten-levels/extensions/jev-guard.ts) adds a write gate (paths outside the repo are refused in code with no Jev call; content is checked for secrets) and a `tool_result` injection screen. Because a throwing Pi handler blocks the tool, its try/catch is what makes it fail open.

## Step 4: expose Jev as a tool the agent calls

Hooks make Jev invisible. Tools let the agent ask it things. ten-levels-of-jev registers them with `pi.registerTool`:

- **`ask_jev_file_bool` / `_choice` / `_score`** (level 8): the tool's code reads the file and sends `{path, content}` as Jev `state`. The agent gets back only the typed answer, so the file never enters the LLM's context. The tool description tells the agent when to use it: "a judgment about what a file does"; read the file when you need to edit or quote it; use grep for exact lookups. A `choice` always gets an "other" option.
- **`ask_jev_files`** (level 9): the agent passes paths or globs plus question JSON. Code expands the globs, skips `node_modules`, `.git`, binaries and files outside the repo, then makes one Jev call per file, 16 in flight.
- **`ask_jev`** (level 10): the agent writes the questions itself (the schema is in the tool description) and supplies its own state, paths, or a command whose output becomes state. A system-prompt nudge (`before_agent_start`) tells it to prefer `ask_jev` for classifications, risk scores and yes/no checks. Over the size budget, the tool refuses and proposes a split instead of truncating.

The design point: the expensive model decides *what to ask*, and Jev answers over content the expensive model never reads. The repo prices one avoided read: input tokens times a frontier model's input price, against Jev's reported cost, one read only. Its README numbers are on [[ideas/builds-agents]]. Agents skip optional tools, though. If a check must run every time, put it in a hook, not a tool ([[ideas/head-to-head-agents]]).

## Step 5: budgets and limits

| Limit | Value | Consequence in a harness |
|---|---|---|
| Rate limits | 100K tokens/s and 80 requests/s, `429` on either ([[reference/models-and-pricing]]) | parallel tool calls, subagents and per-file fan-out multiply requests; batch questions into one request per event, cache, back off on `429` ([[reference/rate-limits-and-errors]]) |
| Context | 64k tokens per request; 32k for `state` plus the longest question | trim transcripts, diffs and files in code. ten-levels-of-jev budgets ~60k of state with one question in levels 8-10, beyond the 32k rule: `contradicts docs` |
| Price | $0.042 per million input tokens; output free | a gate on every tool call costs input tokens only; long transcripts dominate |
| Latency | "End-to-end 70 ms–500 ms" (TypeSafe, [[concepts/system-one]]) | measured hook latencies and cold starts: [[ideas/request-mechanics]], [[ideas/tools-guardrails]] |
| Availability | no uptime SLA ([[reference/legal-and-data]]) | decide the fail direction per hook (below) |

## Step 6: choose the fail direction on purpose

The harness decides what an error means unless your code decides first:

| Harness | Unhandled error or timeout in a gate |
|---|---|
| Claude Code command hook | does not block (only exit `2` or a JSON deny blocks) |
| Claude Code mod | hook skipped unless `.catch` handles it |
| Codex CLI | non-blocking for MCP-tool hook errors; `ask` unsupported |
| Cursor | open, unless `failClosed: true` |
| Pi `tool_call` | **blocks** (fail-safe) |
| opencode `tool.execute.before` | a throw **blocks** |

Most community gates fail open on purpose: cost and noise control should not stop work when Jev is down. That makes them **hygiene, not security** ([[ideas/builds-gates-and-routers]]). ten-levels-of-jev's own skill guide says "Do not turn an exception into permission to act," while its guard fails open on every hook. If a call must never run unchecked, fail closed (deny or ask) and keep hard rules final in the platform. A probability should never switch on a bypass mode ([[ideas/patterns-agents]] P03).

## Step 7: privacy, keys and logs

- **Hook input is conversation text.** Claude Code hooks receive the prompt, full tool inputs, tool results, the last assistant message and the transcript path. Cursor adds the user's email. Whatever you put in `state` goes to TypeSafe on your key, or to the gateway you route through. ten-levels-of-jev's extensions always call OpenRouter.
- **An in-process hook bypasses a proxy** set up for the main model; a redacting proxy does not see it. See the fast-jev-compaction receipt on [[ideas/warning-receipts]].
- **Logs keep what you sent.** ten-levels-of-jev writes the full state (file contents, command output) to stderr and the session file. Treat those as sensitive.
- **The agent can read its environment.** ten-levels-of-jev passes the agent only an allowlist of environment variables because, in one live run, the agent printed every variable it was given.
- TypeSafe's data terms (no training on Input; Telemetry carve-out) are on [[reference/legal-and-data]]. Look-alike gateways and replicas answering to `jev-latest` are on [[ideas/warnings]].

## Gotchas

- **A block message is an instruction, not a control.** In one ten-levels-of-jev run, an agent blocked from writing `.env` wrote it with a bash heredoc. The fix was a notice saying the block is final, plus gating every write path, not only the one the agent chose first.
- **`ask` is not universal.** Codex and Cursor `preToolUse` cannot ask a human, so a three-way gate collapses to allow or deny there.
- **Compaction replacement is not in Claude Code's mods reference.** It lists only `{ skip }` for `session.compact`. Community compaction mods that replace messages were built against early-access declarations (`unverified` on current builds).
- **Do not confuse two "ten levels".** The skill bundled in ten-levels-of-jev (`hyper-jev`) has a different set of levels 6-10 than the app and video.
- **Measure the agent, not the call.** Whole-agent A/Bs with and without Jev found gains, losses and ties ([[ideas/head-to-head-agents]]). IndyDevDan says Jev tools in his Pi agent save "about 20%" of token spend with Sonnet 5.5 (2026-10-05 video); that is his own estimate, method not shown, `unverified`.

## Ready-made integrations

Install before you build: Claude Code, Codex and Cursor gates on [[ideas/tools-guardrails]]; model, skill and compaction routers on [[ideas/tools-agent-routing]]; harness builds with results on [[ideas/builds-gates-and-routers]] and [[ideas/builds-agents]]. Vet code before running it, and check [[ideas/warnings]] for anything that asks for your key.

## Related

- [[concepts/jev-with-coding-agents]] — why Jev is not the agent's LLM, and where it fits instead
- [[reference/agent-skill]] — the TypeSafe skill that teaches your coding agent the Jev API
- [[ideas/patterns-agents]] — P02 routing, P03 tool gates, P04 loop control, P07 compaction, P08 skill selection
- [[ideas/patterns-coding-agents]] — P10 diff review, P11 test-output interpretation
- [[guides/agent-integration-playbook]] — building with Jev, end to end
- [[guides/production-operations]] — pinning, retries, monitoring, fallbacks
- [[cookbooks/llm-guardrails]] — the official guardrail cookbook

## Sources

- raw/docs/introduction__coding-agents.md (https://docs.typesafe.ai/introduction/coding-agents)
- raw/docs/models.md, raw/docs/api.md, raw/docs/confidence.md (TypeSafe docs)
- raw/community/claude-code-docs-hooks.md, claude-code-docs-plugins-reference.md, claude-code-docs-mods-overview.md, claude-code-docs-mods-events.md, claude-code-docs-mods-reference.md (Claude Code docs, 2026-10-05)
- raw/community/openai-codex-docs-hooks.md, raw/community/cursor-docs-hooks.md, raw/community/opencode-docs-plugins.md, raw/x-repos/anomalyco__opencode__plugin-index.md
- raw/x-repos/earendil-works__pi__* (Pi docs and source at 5b6c792b, incl. examples/extensions/jev-router.ts)
- raw/x-repos/disler__ten-levels-of-jev* (https://github.com/disler/ten-levels-of-jev at 777adaf)
- raw/community/youtube-indydevdan-10-levels-of-jev.txt (https://www.youtube.com/watch?v=_U-O5lYhJ7Q, 2026-09-28), raw/community/youtube-indydevdan-openrouter-state-of-models.txt (https://www.youtube.com/watch?v=8BD6w5wELRo, 2026-10-05)
