
Many LLM calls ask for a label, not prose: spam or not, which team, how urgent. TypeSafe’s Jev is a zero-shot classifier for these. The request carries text and a label set; Jev returns a distribution over the labels without generating tokens.
Three answer types:
- Noul: the probability of yes.
- Choice: a categorical distribution over up to 255 options.
- Score: a distribution over up to 10 ordered levels, plus its mean.
Typed answers let symbolic AI sit on top of neural AI. The network reads the text; ordinary code (rules, thresholds, expected-utility decisions) acts on its probabilities.
TypeSafe reports:
- 193.6× faster and 444.6× cheaper than frontier LLMs, gains it expects are “on the higher end”;
- 70–500 ms per call;
- $0.042 per million input tokens; output is free.
Quality still trails frontier LLMs; by TypeSafe’s notes, Jev counts unreliably. TypeSafe’s quality claim scores Jev against the average answer of two frontier LLMs: agreement, not accuracy. Its calibration claim, “higher confidence means higher accuracy”, describes monotonicity, which is weaker than calibration.
Typed. Probabilities. Feed. Code. Fast. Cheap. Not. Yet. Frontier.
References
- TypeSafe (2026). Introducing System One Models & Jev. Launch post, 2026-09-15.
- TypeSafe (2026). Documentation, accessed 2026-09-29: Score, Models, Jev 1.13 jaggedness.
- Calibrated Decisions, Part 1: Jev and the System One Model — the interface and these claims in depth.
- Calibrated Decisions, Part 2: Calibration and Proper Scoring Rules — why monotonicity is weaker than calibration.
- Calibrated Decisions, Part 3: Recalibration and Decisions — checking calibration on your own data and deciding with it.