CLM-8B — the open decision model that scores instead of generating
CLM-8B (Contrastive-LM v0.1) is an Apache-2.0 decision model built on a frozen Qwen3-8B encoder: Noul, Choice and Score questions, candidate ranking, and published latency numbers several times lower than Jev on its own benchmarks. What it is, how to run it, and what to be skeptical about.
What CLM-8B is
CLM (Contrastive Language Model) is a decision model family released on Hugging Face around September 21, 2026 by the Contrastive-LM organization — the model card cites Jacky Kwok, Hangoo Kang, Tarun Suresh, Jon Saad-Falcon, Marco Pavone, Christopher Ré and Azalia Mirhoseini. CLM-v0.1-8B is two small projection heads (a state head and an action head) on top of a frozen Qwen3-8B encoder, trained with a bidirectional InfoNCE contrastive objective. It is Apache-2.0, and its tags on Hugging Face read text-ranking, verifier and reranker.
It scores; it never generates
Like Jev, CLM speaks the System One question vocabulary — Noul (yes/no), Choice (pick among labeled options) and Score (place on a scale) — plus a rank() call for free-form candidate ranking over large sets (the card cites up to ~1k candidates). The output is a probability distribution over the candidates you supplied. Nothing is generated, which is why temperature does not appear anywhere in its documentation.
The published numbers, read carefully
The model card claims zero-shot parity with Jev on computer-use, gaming and tool-calling workloads at up to 9× lower latency; with state and action caching and roughly a thousand candidates, 13× faster; and, once fine-tuned as a verifier, state-of-the-art results on DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%) at 4–6× Jev's speed. Two caveats before quoting any of that: the SOTA verifier numbers require fine-tuning the heads, and the caching speedups depend on embedding caches staying warm — which is your serving problem to run, not a property that ships with the weights.
How you run it
Install contrastive-lm with pip, serve the frozen Qwen3-8B encoder with vLLM (a --runner pooling server), then run clm-serve, which exposes an API and a local playground on port 8700. Community GGUF and MLX quantizations exist for Apple Silicon and llama.cpp. In practice you are operating an inference stack: GPU time, capacity, model versioning and on-call are yours, in exchange for owning the weights and the data path.
Limitations worth naming
CLM-8B is encoder-locked — it needs Qwen3-8B's last-token-pooled embeddings, so swapping the base means re-working the stack. It does not generate text. Its headline verifier results are fine-tuned, not zero-shot. A multimodal CLM-35B was announced as coming in early October. And like every decision model, its probabilities need validating on your own cases before anything auto-approves.
FAQ
- Is CLM-8B related to Jev?
- They target the same problem — typed decisions with bounded answer spaces — and the CLM publishers benchmarked against Jev, but they are independent models: CLM-8B is Apache-2.0 open weights on a frozen Qwen3-8B encoder that scores candidates locally; Jev is TypeSafe's hosted decision model.
- Can I run CLM-8B for free?
- The weights are free to download (Apache-2.0), and the 8B checkpoint runs on Apple Silicon via community MLX/GGUF builds. What is not free is operating it: an embedding-serving stack with vLLM plus clm-serve, capacity planning, and re-checking calibration when you change anything.
- Does Optio host CLM-8B?
- No. Optio calls Jev. This page exists because CLM-8B speaks the same decision vocabulary, and people comparing typed decision models deserve an honest description rather than a marketing page.
- CLM-8B vs Jev — which should I use?
- If the requirement is zero-operations hosted decisions with keys and credits: Jev through Optio. If you want the weights in your own environment, candidate ranking over large sets, or the lowest latency on hardware you control: CLM-8B. Compare both on your own ten cases before switching anything.