Oct 10, 2026
Jev vs Laya, CLM-8B and the rest: a 2026 field guide to decision models
Eight System One-style decision models compared: licenses, hosting, cost crossovers, and the latency numbers we measured on the live Jev gateway.
The landscape at a glance
Since TypeSafe shipped the Jev decision model in September 2026, a small ecosystem of System One-style decision models has appeared — open weights, hosted APIs, and everything in between. We run a hosted gateway for Jev, and this guide is a field map of the alternatives: what each one is, what it costs to run, and when the hosted path wins.
- Jev (TypeSafe) — hosted API; calibrated probabilities per answer over the /v1/systemone contract. You call it through gateways like Optio.
- Laya (Convai Innovations) — Apache-2.0 encoder heads on ModernBERT/mmBERT; you host it on a CPU, GPU or Apple Silicon; a multilingual checkpoint covers more than English.
- OpenJev — open, self-hosted; the DiffusionGemma build answers questions about images as well as text.
- SemIf (TheoLeeCJ) — MIT; typed decisions read from a frozen open LLM's logits, with a temperature you fit per workload. (An earlier name for the project was OpenJev.)
- AnyJev (Nokia Applied Research) — Apache-2.0; training-free decisions from a base LLM's hidden states, served by you.
- Kev (Jared Palmer) — Apache-2.0 Qwen fine-tunes that speak the same /v1/systemone request, from a 0.8B Mac-sized checkpoint upward.
- CLM-8B (Contrastive-LM) — Apache-2.0 contrastive heads on a frozen Qwen3-8B encoder; scores candidates and ranks up to roughly a thousand of them.
- Liquid d1 (LiquidAI) — edge-sized decision models (600M omni with vision and audio, 3B text+vision) fine-tuned from LFM2.5 bases; GGUF builds for llama.cpp.
Real latency from live calls
We measured the hosted Jev gateway from a single client region on a business day: 16 successive live calls covering one-question and multi-question requests. The reported latency_ms clustered at 400–560 ms with a median near 445 ms; the first call after an idle period showed an occasional ~3 s cold start. Design your timeout from the p99 of your own region, not from any single measurement — ours included.
The cost crossover
Hosted credits are per-use: 1 credit per 1,000 input tokens, minimum 1 per request, output tokens free — $0.0001 scale on Starter, which is $10 for 100,000 credits that never expire. Self-hosting open models is compute plus ops: near-zero marginal cost past an always-on GPU, plus the labeling or calibration work your model family demands.
A working rule from teams comparing the two: below a few thousand decisions a day, hosted wins almost every time. The crossover sits above sustained tens of thousands per day, and only after your validation set is done. Data-residency requirements move the line — if the data cannot leave your machines, the hosted path is disqualified no matter what it costs.
Where the claims need care
- CLM-8B's publisher benchmarks zero-shot parity with Jev at up to 9× lower latency, 13× with state/action caching at roughly 1,000 candidates, and fine-tuned-verifier state of the art on DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%) at 4–6× Jev's speed. Real numbers — but the caching wins need warm caches you operate, and the SOTA lines need fine-tuning the heads on your task.
- Liquid d1's weights carry a license field of "other" on Hugging Face, not a standard open license. Read the actual terms before shipping a product on it.
- AnyJev's answers inherit your base LLM's blind spots: every base-model swap re-opens the calibration question.
- Laya's zero-shot base will not stand in for hosted Jev on your task until you fine-tune it on your own labels.
Whatever you pick, re-validate thresholds on your own ten real cases before switching models — publisher benchmarks are somebody else's task.
How to choose in one afternoon
- Need zero operations, keys, credits and a dashboard? Hosted Jev through a gateway.
- Need the weights inside your VPC, or ranking over a thousand candidates? CLM-8B or an open encoder family.
- Need vision or audio inputs on the edge? Liquid d1 is the small-model niche; Jev stays text.
- Already serving an open LLM around the clock? SemIf or AnyJev add typed decisions without a new vendor.
Is CLM-8B better than Jev?
On its publisher's benchmarks, parity with much lower latency — a claim about their tasks and hardware. For you, run ten real cases through both and compare the low-confidence rates.
Which decision model is cheapest?
Below a few thousand decisions a day, hosted credits from $10. At sustained high volume on hardware you already run, open weights approach zero marginal cost — after the ops and calibration work.
Is this guide affiliated with TypeSafe?
No. Optio is an independently operated gateway for calling Jev, not affiliated with TypeSafe or with any model covered here; the facts about third-party models come from their own documentation.
Try the Jev AI playground to send your first free decision after you finish this guide.