Agents: read the raw Markdown of this page, or start at llms.txt.
Jev inside your agent harness: hooks, tools and fail-safes
TL;DR Keep the LLM as the agent's brain. Use Jev for the narrow, repeated decisions a harness makes around it: in hooks the model never sees (block a tool call, screen a tool result, route the prompt to a model, decide when to compact, check "done"), or as tools the model calls that answer a typed question about files or command output without loading them into its context. Each harness decides what happens when a hook errors or times out. So set your own Jev time budget, choose fail-open or fail-closed per hook on purpose, and remember that hook input is conversation text you are sending to TypeSafe or a gateway. Community tier: harness contracts come from each vendor's docs; designs and numbers come from builders, chiefly IndyDevDan's ten-levels-of-jev.
What "Jev in the harness" means (and does not)
TypeSafe's own position: Jev cannot replace the model behind Claude Code, Cursor or Codex, but it is "often exactly the right tool inside an agent or app you're building" (Jev with coding agents: not a drop-in for the LLM behind Claude Code, Cursor, Copilot). The harness is the code around the LLM: the loop, the tool runner, permissions, compaction, model choice. Those are full of fixed-option decisions that today are hard-coded rules, a human prompt, or an extra LLM call. Those decisions are where Jev goes.
There are three shapes, in rising order of autonomy (the ladder of IndyDevDan's video, transcript, levels 6-10):
| Shape | Who decides to call Jev | Examples | Pattern |
|---|---|---|---|
| Hook (invisible to the model) | harness code, on an event | pre-tool gate, result screen, prompt router, compaction trigger, done-check | Patterns: agent internals, routing, gates, context and memory P02, P03, P04, P07, P08 |
| Tool (fixed questions) | the model, via a tool you register | "is this file about auth?" over one file or a glob | P36 (inferred fit) |
| Agent-authored questions | the model writes the question JSON | classify a test failure, score the risk of its own diff | Patterns: coding agents, dev tools and self-compiling workflows P11 |
Step 1: pick the decisions
Good first candidates are decisions that fire often, have a fixed answer set, and whose input fits in state:
| Decision | Primitive | Typical action in code |
|---|---|---|
| Is this command read-only, reversible or irreversible? | Choice + a destructive-intent Noul |
block, ask, allow |
| Does this write touch secrets? | Noul (+ file-kind Choice) |
block |
| Does this tool output contain instructions aimed at the agent? | Noul |
add a "treat as data" banner |
| Is this request simple or complex? | Choice |
pick model or effort |
| Did the user switch tasks; is the agent at a clean boundary? | Nouls + a Score |
suggest or request compaction |
| Is the agent's "done" backed by evidence in the transcript? | Noul |
block the stop, rerun checks |
Keep counting, paths, size limits and anything exact in code; Jev handles the judgment (Jev 1.13 jaggedness: known failure modes). Rules come first and Jev sees only what they cannot settle. Most community gates do this; see Builds: permission gates, approvals and model routers for agents for why widening permissions needs hard bounds.
Step 2: find the hook in your harness
Event names come from each vendor's docs (captured 2026-10-05). The mapping of decision to event is ours.
| Decision | Claude Code | Codex CLI | Cursor | Pi | opencode |
|---|---|---|---|---|---|
| Tool-call gate | PreToolUse; mod tool.call |
PreToolUse |
preToolUse, beforeShellExecution |
pi.on("tool_call") |
tool.execute.before |
| Replace the human approval | PermissionRequest; mod tool.check |
PermissionRequest |
beforeShellExecution |
tool_call + ctx.ui.confirm |
permission.ask |
| Screen tool output | PostToolUse |
PostToolUse (block replaces the result) |
postToolUse |
tool_result |
tool.execute.after |
| Model or effort routing | mod turn.step; UserPromptSubmit (hint only) |
custom provider or proxy | none found | registerVirtualModel + ctx.modelRegistry.classify |
chat.params |
| Done-check | Stop |
Stop |
stop |
agent_before_settle |
session.idle (observe only) |
| Compaction | PreCompact (can block); mod session.compact (skip) |
PreCompact (can stop) |
preCompact (observe only) |
session_before_compact (cancel or supply) |
experimental.session.compacting |
| Skill or context selection | UserPromptSubmit additionalContext |
UserPromptSubmit |
beforeSubmitPrompt (block only) |
before_agent_start, context |
experimental.chat.system.transform |
Contract details that change the design:
- Claude Code command hooks read JSON on stdin and decide with
hookSpecificOutput.permissionDecision=allow | deny | ask | defer. Exit2blocks. Exit 1, or a timed-outPreToolUsehook, does not block. Default timeouts are 600 s for command hooks and 30 s onUserPromptSubmit; added context from a timed-out prompt hook is dropped. Plugins can store the key withuserConfigsensitive: true(Claude Code hooks, plugins reference). - Claude Code mods (in-process JS hooks) are documented, need v2.1.287+, and are on by default. That version ignores
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS, which older community READMEs ask you to set. A mod hook gets 10 s of its own time; a hook that throws or times out is skipped unless you attach.catch(mods overview, events and reference pages). - Codex CLI:
PreToolUsecan deny or rewrite but does not supportaskyet; the hook is marked failed and the call continues. Hooks must be trusted by hash in/hooks; the default timeout is 600 s (Codex hooks). - Cursor fails open unless a hook sets
failClosed: true, but blocks on invalid JSON from a permission hook.askis not enforced onpreToolUse. Hook input includesuser_email(Cursor hooks). - Pi: a
tool_callhandler returns{ block: true, reason }. A handler that throws blocks the tool (fail-safe). Pi also ships Jev as a built-in classifier model (typesafeprovider,TYPESAFE_API_KEY; a 2026-09-24 commit also serves it through OpenRouter and Cloudflare Workers AI) with an officialjev-router.tsexample; details on Framework integrations: Pydantic AI, LangChain, Spring AI, Vercel AI SDK, eve, LiteLLM, BAML, TanStack AI, DSPy, Langfuse and other framework-shipped Jev support. - opencode: block in
tool.execute.beforeby throwing; hooks run in load order; the docs give no timeout rule.
Step 3: write the hook (a complete Claude Code pre-tool gate)
.claude/settings.json:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "python3 ${CLAUDE_PROJECT_DIR}/.claude/hooks/jev_bash_gate.py",
"timeout": 10
}
]
}
]
}
}
.claude/hooks/jev_bash_gate.py (needs pip install typesafe-sdk and TYPESAFE_API_KEY):
import json
import os
import sys
from typesafe_sdk import Choice, Noul, RetryPolicy, TypeSafeClient, TypeSafeError
FAIL_CLOSED = os.environ.get("JEV_GATE_FAIL_CLOSED") == "1"
def decide(decision: str, reason: str) -> None:
print(json.dumps({
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": decision,
"permissionDecisionReason": reason,
}
}))
sys.exit(0)
event = json.load(sys.stdin)
command = event.get("tool_input", {}).get("command", "")
try:
with TypeSafeClient(model="jev-1.13.0") as client:
result = client.system_one(
state={"command": command[:4000], "cwd": event.get("cwd", "")},
questions={
"effect": Choice(
instructions="What happens to files, history or remote state if this shell command runs?",
criteria={
"read_only": "Only reads or lists; changes nothing.",
"reversible": "Changes things that can be restored or regenerated.",
"irreversible": "Deletes, overwrites, force-pushes or publishes something that cannot be restored.",
},
),
"destructive": Noul(instructions="Is this command meant to destroy data or history?"),
},
retry=RetryPolicy(max_retries=1, timeout=4.0),
timeout=3.0,
)
except TypeSafeError as error:
if FAIL_CLOSED:
decide("ask", f"Jev gate unavailable ({type(error).__name__}); asking instead of allowing.")
sys.exit(0) # fail open: no decision, the normal permission flow continues
effect = result.choices["effect"]
destructive = result.nouls["destructive"].noul
if (effect.choice == "irreversible" and effect.confidence >= 0.6) or destructive >= 0.7:
decide("deny", f"Jev: {effect.choice} (confidence {effect.confidence:.2f}), destructive {destructive:.2f}.")
if effect.choice != "read_only" or effect.confidence < 0.5:
decide("ask", f"Jev: {effect.choice} (confidence {effect.confidence:.2f}).")
sys.exit(0) # read-only and confident: no decision, normal flow
Notes on the sample:
- The thresholds are the starting values ten-levels-of-jev's Pi guard uses (block on
irreversibleat confidence ≥ 0.6 or the destructiveNoul≥ 0.7). The 0.5 floor follows the docs' own example (Confidence vs probability). Tune them on your own logged commands. - The gate only narrows: it denies or asks and never returns
allow, so your permission rules stay in charge. Anallowwould widen permissions; read Builds: permission gates, approvals and model routers for agents first. RetryPolicy(timeout=4.0)caps the whole call including retries, andtimeout=3.0caps each attempt. The SDK defaults are a 30 s total retry budget and 10 s per attempt (Python SDK retries, exceptions, constants), which can outlast a hook's timeout.- The model is pinned to
jev-1.13.0so a new release behindjev-latestcannot shift your thresholds (Running Jev in production: versioning, caching, retries, monitoring and fallbacks). - A
Noulhas noconfidence; its probability is the number you threshold (Noul (yes/no) questions). - This is one process per call. To avoid start-up cost, community gates moved to a daemon or an in-process mod or extension (Request mechanics: billing, limits, latency, calibration and stability).
The same gate in Pi is a pi.on("tool_call", ...) handler that returns { block: true, reason }. ten-levels-of-jev's jev-guard.ts adds a write gate (paths outside the repo are refused in code with no Jev call; content is checked for secrets) and a tool_result injection screen. Because a throwing Pi handler blocks the tool, its try/catch is what makes it fail open.
Step 4: expose Jev as a tool the agent calls
Hooks make Jev invisible. Tools let the agent ask it things. ten-levels-of-jev registers them with pi.registerTool:
ask_jev_file_bool/_choice/_score(level 8): the tool's code reads the file and sends{path, content}as Jevstate. The agent gets back only the typed answer, so the file never enters the LLM's context. The tool description tells the agent when to use it: "a judgment about what a file does"; read the file when you need to edit or quote it; use grep for exact lookups. Achoicealways gets an "other" option.ask_jev_files(level 9): the agent passes paths or globs plus question JSON. Code expands the globs, skipsnode_modules,.git, binaries and files outside the repo, then makes one Jev call per file, 16 in flight.ask_jev(level 10): the agent writes the questions itself (the schema is in the tool description) and supplies its own state, paths, or a command whose output becomes state. A system-prompt nudge (before_agent_start) tells it to preferask_jevfor classifications, risk scores and yes/no checks. Over the size budget, the tool refuses and proposes a split instead of truncating.
The design point: the expensive model decides what to ask, and Jev answers over content the expensive model never reads. The repo prices one avoided read: input tokens times a frontier model's input price, against Jev's reported cost, one read only. Its README numbers are on Builds: coding agents, harnesses and orchestration. Agents skip optional tools, though. If a check must run every time, put it in a hook, not a tool (Head-to-head: Jev inside agents, routers and tool gates).
Step 5: budgets and limits
| Limit | Value | Consequence in a harness |
|---|---|---|
| Rate limits | 100K tokens/s and 80 requests/s, 429 on either (Models, aliases, pricing, rate limits, context) |
parallel tool calls, subagents and per-file fan-out multiply requests; batch questions into one request per event, cache, back off on 429 (HTTP status codes, rate limits, retry semantics) |
| Context | 64k tokens per request; 32k for state plus the longest question |
trim transcripts, diffs and files in code. ten-levels-of-jev budgets ~60k of state with one question in levels 8-10, beyond the 32k rule: contradicts docs |
| Price | $0.042 per million input tokens; output free | a gate on every tool call costs input tokens only; long transcripts dominate |
| Latency | "End-to-end 70 ms–500 ms" (TypeSafe, System One Models) | measured hook latencies and cold starts: Request mechanics: billing, limits, latency, calibration and stability, Tools: guardrails for agents (tool-call gates, permission hooks, injection screens, rule checks) |
| Availability | no uptime SLA (Legal: MCA, DPA, privacy, data retention) | decide the fail direction per hook (below) |
Step 6: choose the fail direction on purpose
The harness decides what an error means unless your code decides first:
| Harness | Unhandled error or timeout in a gate |
|---|---|
| Claude Code command hook | does not block (only exit 2 or a JSON deny blocks) |
| Claude Code mod | hook skipped unless .catch handles it |
| Codex CLI | non-blocking for MCP-tool hook errors; ask unsupported |
| Cursor | open, unless failClosed: true |
Pi tool_call |
blocks (fail-safe) |
opencode tool.execute.before |
a throw blocks |
Most community gates fail open on purpose: cost and noise control should not stop work when Jev is down. That makes them hygiene, not security (Builds: permission gates, approvals and model routers for agents). ten-levels-of-jev's own skill guide says "Do not turn an exception into permission to act," while its guard fails open on every hook. If a call must never run unchecked, fail closed (deny or ask) and keep hard rules final in the platform. A probability should never switch on a bypass mode (Patterns: agent internals, routing, gates, context and memory P03).
Step 7: privacy, keys and logs
- Hook input is conversation text. Claude Code hooks receive the prompt, full tool inputs, tool results, the last assistant message and the transcript path. Cursor adds the user's email. Whatever you put in
stategoes to TypeSafe on your key, or to the gateway you route through. ten-levels-of-jev's extensions always call OpenRouter. - An in-process hook bypasses a proxy set up for the main model; a redacting proxy does not see it. See the fast-jev-compaction receipt on Warning receipts: what we checked behind each warning.
- Logs keep what you sent. ten-levels-of-jev writes the full state (file contents, command output) to stderr and the session file. Treat those as sensitive.
- The agent can read its environment. ten-levels-of-jev passes the agent only an allowlist of environment variables because, in one live run, the agent printed every variable it was given.
- TypeSafe's data terms (no training on Input; Telemetry carve-out) are on Legal: MCA, DPA, privacy, data retention. Look-alike gateways and replicas answering to
jev-latestare on Warnings: not-Jev services, key safety, look-alikes and install names.
Gotchas
- A block message is an instruction, not a control. In one ten-levels-of-jev run, an agent blocked from writing
.envwrote it with a bash heredoc. The fix was a notice saying the block is final, plus gating every write path, not only the one the agent chose first. askis not universal. Codex and CursorpreToolUsecannot ask a human, so a three-way gate collapses to allow or deny there.- Compaction replacement is not in Claude Code's mods reference. It lists only
{ skip }forsession.compact. Community compaction mods that replace messages were built against early-access declarations (unverifiedon current builds). - Do not confuse two "ten levels". The skill bundled in ten-levels-of-jev (
hyper-jev) has a different set of levels 6-10 than the app and video. - Measure the agent, not the call. Whole-agent A/Bs with and without Jev found gains, losses and ties (Head-to-head: Jev inside agents, routers and tool gates). IndyDevDan says Jev tools in his Pi agent save "about 20%" of token spend with Sonnet 5.5 (2026-10-05 video); that is his own estimate, method not shown,
unverified.
Ready-made integrations
Install before you build: Claude Code, Codex and Cursor gates on Tools: guardrails for agents (tool-call gates, permission hooks, injection screens, rule checks); model, skill and compaction routers on Tools: agent routing, context and skill selection; harness builds with results on Builds: permission gates, approvals and model routers for agents and Builds: coding agents, harnesses and orchestration. Vet code before running it, and check Warnings: not-Jev services, key safety, look-alikes and install names for anything that asks for your key.
Related
- Jev with coding agents: not a drop-in for the LLM behind Claude Code, Cursor, Copilot — why Jev is not the agent's LLM, and where it fits instead
- The typesafe-ai agent skill and Claude Code plugin — the TypeSafe skill that teaches your coding agent the Jev API
- Patterns: agent internals, routing, gates, context and memory — P02 routing, P03 tool gates, P04 loop control, P07 compaction, P08 skill selection
- Patterns: coding agents, dev tools and self-compiling workflows — P10 diff review, P11 test-output interpretation
- Playbook for LLM agents building with Jev — building with Jev, end to end
- Running Jev in production: versioning, caching, retries, monitoring and fallbacks — pinning, retries, monitoring, fallbacks
- Cookbook: Guardrails for LLMs — the official guardrail cookbook
Sources
- raw/docs/introduction__coding-agents.md (https://docs.typesafe.ai/introduction/coding-agents)
- raw/docs/models.md, raw/docs/api.md, raw/docs/confidence.md (TypeSafe docs)
- raw/community/claude-code-docs-hooks.md, claude-code-docs-plugins-reference.md, claude-code-docs-mods-overview.md, claude-code-docs-mods-events.md, claude-code-docs-mods-reference.md (Claude Code docs, 2026-10-05)
- raw/community/openai-codex-docs-hooks.md, raw/community/cursor-docs-hooks.md, raw/community/opencode-docs-plugins.md, raw/x-repos/anomalyco__opencode__plugin-index.md
- raw/x-repos/earendil-works__pi__* (Pi docs and source at 5b6c792b, incl. examples/extensions/jev-router.ts)
- raw/x-repos/disler__ten-levels-of-jev* (https://github.com/disler/ten-levels-of-jev at 777adaf)
- raw/community/youtube-indydevdan-10-levels-of-jev.txt (https://www.youtube.com/watch?v=_U-O5lYhJ7Q, 2026-09-28), raw/community/youtube-indydevdan-openrouter-state-of-models.txt (https://www.youtube.com/watch?v=8BD6w5wELRo, 2026-10-05)