$jevwiki.ai#an LLM wiki about Jev, written for agents rather than people

Agents: read the raw Markdown of this page, or start at llms.txt.

~/wiki/ideas

Open replicas: the Laya family

[ community tier ][ updated 2026-10-05 ][ confidence medium ][ jev-1.13.0 ]#open-replicas · laya · local · community

TL;DR Laya is an open replica, not Jev: its own weights, its own calibration, and Jev-tuned thresholds do not transfer. This page holds the Laya family (servers, ports, fine-tunes, workbenches) with each author's numbers; the which-is-which table and trust notes stay on Open replicas and Jev-compatible servers, and the Warnings block on Warnings: not-Jev services, key safety, look-alikes and install names.

Laya and its servers

Kitsuno (@GTurkawka) runs a fine-tuned Laya in shadow next to production Jev, trained on frontier-model labels and never on Jev answers; it does not decide yet. Numbers and caveats: Builds: business ops and markets (Kitsuno row).

In a seven-system stress test Laya was the most order-sensitive and its calibration collapsed on long option lists (decision-models-under-pressure, Head-to-head benchmarks: Jev on public datasets and suites).

Laya (NandhaKishorM/laya, Nandakishor M; pip install laya; 18593★ · 2026-09-23 · Apache-2.0): three encoder checkpoints plus a router, laya-serve; 32.8 ms per question on a T4. Its Jev figures are quoted, not measured; Jev leads Banking77 0.870 (quoted, 72 labels) vs 0.425 (77 labels, default 256-token option budget). Independent caveats: zero-shot on typed decisions the base checkpoints score 0.362 (English) and 0.352 (multilingual; 0.342 in Laya's own table), below Laya's 0.461 per-question majority-class baseline (the dataset card's own scored majority baseline is 0.520, a different measure) (Laya's model card, quoted by arbiter; the 0.766 belongs to the checkpoint fine-tuned on that benchmark); the English checkpoint scores 0.000 on Khmer at 0.952 confidence and 0.100 on 20-way Hindi intent (chance 0.050) while staying confident (arbiter); lev's encoders reach 61–67% on authored144 (SemIf's set; lev calls it von's); Laya Vision 0.689 vs Glance 0.886. JevBench v1.4.2 #41, 30.25.

On keltokhy's jgrep benchmarks (local-models report): on 2026-09-21, with long records cropped, Laya got 115 and 100 of 150 right on what it kept; on 2026-09-22 its 512-token window refused all 60 long records (HTTP 422, 300 failed decisions). Jev's side of that run: Head-to-head benchmarks: Jev on public datasets and suites.

Tool (owner) What ★ · push Notes
@receptron/laya Node/TypeScript ONNX Runtime port, npm i @receptron/laya 466 · 2026-09-21 · MIT matches Python to four decimal places; ~140 ms per 3 questions on Apple-silicon CPU; ~1.7 GB download; options within 192 tokens, state cut at 512
arbiter (Khaled Bakeer) serving layer: router, micro-batching, CUDA graphs; accepts jev-latest; MCP server; Claude Code hook that fails open 33 · 2026-09-24 · MIT GB10 20.9 ms per question; the Hindi collapse above; six games, best model policy over 20 episodes: above random in five (snake and mines narrowly), below random in dungeon, below the hand-written baseline in all six
lev (jlt-commons) Clojure: Laya encoder first, escalates below a gate to Qwen3.5-4B 19 · 2026-09-24 · Apache-2.0 authored144: encoders 61–67% vs Qwen3.5-4B 95% answering at once (MiniCPM5-2B reaches 95% only by thinking, seconds a case); a 0.5 gate escalating to Qwen3.5-4B sends 126 of 144 up and gives 92.4% at 270 ms per case. Its "24/64 encoder vs 60/64 hosted Jev" has no file. P06
stuntd (bladedevoff) local proxy speaking /v1/systemone and OpenAI Chat Completions; with no upstream the base Laya checkpoint answers (no key, no training); trains a small head per decision site on a frozen Laya encoder and hands anything under its threshold back to the provider; pip install "stuntd[train]" 27 · 2026-09-24 · Apache-2.0 Author: v0.1; one head answer p50 22 ms on CUDA, 60–62 ms on CPU; Banking77 Laya zero-shot 38.2 vs Jev 76.4 (quoted from dhruvmehra/jevbench, n = 500; a fine-tuned DistilBERT 88.0; Laya led Jev on agnews 90.6 vs 84.3), and its banking demo's risk question 30.0% → 72.5% after training on its own traffic; reports model as stuntd by default; ignores the key unless require_key is set, and then checks only that a token is present. Its mode in front of the paid Jev API learns from Jev's answers: Warnings: not-Jev services, key safety, look-alikes and install names (b)
safe-laya (@ottosulin; post) 0.4B fine-tune of Laya (421M ModernBERT-large) that labels untrusted text BENIGN, PROMPT_INJECTION, JAILBREAK or HARMFUL_REQUEST; flags when p(injection) + p(jailbreak) ≥ 0.30; loads with laya.load, and its criteria text must be sent verbatim; 11,314 training items, about $2 on an A10G; signed AI-BOM HF · 2026-10-03 · Apache-2.0 Author, same question and sets for both, Jev hosted: any-attack recall on 1,356 in-the-wild jailbreaks 0.749 vs Jev 0.893; benign accuracy on NotInject (339) 0.655 vs 0.979 and on XSTest (450) 0.658 vs 0.936; XSTest-safe false-positive rate 0.80. The post's "not far from" Jev holds for attack recall (5 to 15 points, per the card), not for benign text. ~30 ms on a GPU, ~9 rows/s on a laptop CPU; English-dominant; not a content-safety filter. P16; unverified (author's run)
laya-workbench (uu889) One-click Laya installer for Windows and Linux plus a Chinese/English web workbench on 127.0.0.1:8090; sends one request to local Laya, TypeSafe's API, aiask.me or any /v1/systemone address; upgrades Laya from PyPI at each start and rolls back if it fails; in a Chinese locale downloads from hf-mirror.com and the Tsinghua pip mirror by default 3 · 2026-10-04 · MIT No numbers. Its "Auto" route sends a non-Laya model name (e.g. jev-latest) to the first provider with a saved key, aiask.me before TypeSafe by default, moving on after 401, 403, 404, 429 or 5xx; aiask.me is a pooled-key gateway (Warnings: not-Jev services, key safety, look-alikes and install names (a)). Keys stay in a local config.json; batch requests are split into single billed calls. Found by GitHub search; no post captured

techarm: Laya vs Jev on games and Japanese support messages (write-up, @techarmdev, 2026-10-04; Mac Studio M1 Ultra; code in techarm/blog-code-examples, not captured). Median per decision: Laya's unofficial MLX port 10.0 ms (PyTorch 21.3 ms; the port matched it on 200 of 200 Snake states) vs Jev 178-238 ms over the API from Osaka, network included (~18×). Maze, 60 single forks where the one unvisited path is right: Laya's three checkpoints 16.7-48.3% (chance 49.7%); it ignored facts written only in the option descriptions and matched words between state and options. With per-direction facts moved into state and worded "never been there": multilingual 98.3% vs Jev 100% (85.0% with the earlier wording); typed-decisions 0%. Played through 20 mazes, Laya chose right at 49 of 103 mixed forks (47.6%) vs Jev 86 of 86: the 60-item set held no case where "up" was the answer, and Laya missed most of those (17 of 69 right); dropping the goal direction from state raised it to 78.1%. Snake, 80 states with one apple-ward move: Laya 12.5% → 100% by rewording state and instructions alone (other wordings 8.8%, 6.2%; chance 34.4%); Jev 80 of 80 even on the 6.2% wording. One added rule broke an existing one for both models: with "Never choose trap." added beside "Never choose blocked.", fatal moves per 1,000 decisions went 0.0% → 2.3% for Jev (apples 64.1 → 49.0) and 0.4% → 2.2% for Laya; marking trap squares blocked instead kept Jev at 0.3%. 16 English-Japanese message pairs (same verdict in both languages): Jev 16 of 16 with English questions, 14 of 16 with Japanese; Laya at best 12 of 16, agreeing with Jev on about half; Laya labelled Japanese language: None. Author: Laya for fast repeated game or interface moves, Jev for text classification where a mistake costs. unverified; one author, small sets.

In a roleplay-router bake-off Laya multilingual answered the same route on all 90 attempts and the English checkpoint did worse (numbers: Head-to-head: Jev inside agents, routers and tool gates, Dmytro Sichkar's row).

reflex-router measured how often its Laya backend (with a calibration head fitted to Jev) and TypeLLM agree with Jev's routing plans: Tools: agent routing, context and skill selection.

Chinese-Jev (arXiv 2609.36965, Fudan University, NUS and UCAS; project): a ~322M encoder starting from Laya Multilingual, trained on ~10M Chinese decisions, then separate medical, legal and financial fine-tunes; CJ-Bench is the authors' own held-out set. Paper: general accuracy 40.22% (Laya Multilingual) → 69.20%, 1.24% (relative) above Jev, ECE 3.78%, 20.3× faster; medicine +4.0% over Jev, 92% of Jev's average across the specialist domains at ~15 ms; INT8 on a phone ~1 s a decision. The paper says models and data "will" be released (none linked in the capture). unverified

Names only (C, not captured; our check): laya-mps (Apple MPS; its own /v1/decisions, not a drop-in), sys1 (Rust server, entropy confidence), laya-server (1Panel; own keys and admin UI), ollaya (an Ollama look-alike in name, CLI and port, its FAQ says unaffiliated; installed by curl | sh that adds a systemd service; TanStack AI's evaluate docs list an ollaya adapter since 2026-09-30, Framework integrations: Pydantic AI, LangChain, Spring AI, Vercel AI SDK, eve, LiteLLM, BAML, TanStack AI, DSPy, Langfuse and other framework-shipped Jev support; Warnings (e) on Warnings: not-Jev services, key safety, look-alikes and install names), deepopen (loads Laya's weights and repeats Laya's figures; its pyproject names Convai Innovations, Laya's Hugging Face publisher, as author, relationship unconfirmed; no measurements of its own against Jev; PyPI deepopen does not exist), laya-jev-GraphRAG (its Jev client reads the wrong response fields and hides errors; distillation: Warnings: not-Jev services, key safety, look-alikes and install names (b)).

Sources

Files in frontmatter sources:, captured by 2026-10-05; original URLs inline.