Calibrated Decisions, Part 3: Recalibration and Decisions
Estimating the calibration curve on your own data, repairing it, and turning calibrated probabilities into actions
A model calibrated on its vendor’s data need not be calibrated on yours. LOESS calibration curves and their pitfalls, Platt, temperature, isotonic and beta recalibration on a train/calibration/test split, subgroup and domain-shift failures, how many labels a repair needs, and the expected-loss thresholds that make calibrated probabilities worth having.
Statistics
Machine Learning
Python
Author
Ravi Kalia
Published
September 27, 2026
Calibrated Decisions, Part 3: Recalibration and Decisions
A weather service can be well calibrated across a country and still be wrong about one valley. The same holds for a model sold to many customers: calibration measured on the vendor’s evaluation set says little about the tickets, claims or patients in yours. What you need is the calibration curve on your own labelled data, a repair when it is off, and a rule that turns the repaired probability into an action.
Part 1 described Jev and what TypeSafe claims for it. Part 2 defined calibration and the scores that measure it. This part estimates the calibration curve with LOESS, repairs it with a recalibration map fitted on held-out data, shows how shift breaks both, and derives the thresholds that make a calibrated probability worth having.
Statements about Jev carry a label. vendor means documented by TypeSafe. theory means standard statistics. inference means our reading of documented facts. speculation means a hypothesis about undisclosed internals.
1 Data
Every data set in this part is synthetic, drawn from the process Part 2 used, so the true probability is known and every estimate can be checked against it:
\[
X \sim N(0, 1), \qquad p^*(x) = \sigma(2x), \qquad Y \mid X = x \sim \operatorname{Bernoulli}\big(p^*(x)\big).
\]
It stands in for a labelled sample of your own traffic for one yes/no question, such as a Noul asking whether a ticket requests a refund.
The miscalibrated model is either the logit-scale distortion \(\hat p = \sigma(\alpha \operatorname{logit} p^*)\) with \(\alpha = 2.5\), or the overfitted logistic regression from Part 2 (1,000 training cases, 200 features, one of them informative).
A wrong recalibration here costs money in Section 6: refunds approved that should have been declined, and reviewers paid to look at cases the model could have settled.
Code
def draw(rng, n, slope=2.0, shift=0.0):"""X ~ N(0, 1), p*(x) = sigmoid(slope * x + shift), Y ~ Bernoulli(p*).""" x = rng.standard_normal(n) p_star = expit(slope * x + shift) y = (rng.random(n) < p_star).astype(float)return x, p_star, ydef distort(p, alpha):return expit(alpha * logit(np.clip(p, 1e-15, 1-1e-15)))def brier(p, y):return np.mean((p - y) **2)def log_loss(p, y): p = np.clip(p, EPS, 1- EPS)return-np.mean(y * np.log(p) + (1- y) * np.log(1- p))def ece(p, y, m=15): idx = np.minimum((p * m).astype(int), m -1) gap =sum(np.sum(idx == b) *abs(y[idx == b].mean() - p[idx == b].mean()) for b in np.unique(idx))return gap /len(p)def reliability_bins(p, y, m=10): idx = np.minimum((p * m).astype(int), m -1) rows = [(p[idx == b].mean(), y[idx == b].mean(), np.sum(idx == b)) for b in np.unique(idx)]return np.array(rows).Tdef loess_curve(p, y, frac=0.3, it=0):"""LOWESS of y on p: sorted x, fitted m(x). it=0 turns off robustness iterations.""" fit = lowess(y, p, frac=frac, it=it, delta=0.005, return_sorted=True) xs, keep = np.unique(fit[:, 0], return_index=True)return xs, fit[keep, 1]def cell(a, d=3):"""Median over seeds, with the 5-95% interval on a second, smaller line.""" a = np.asarray(a) lo, hi = np.quantile(a, [0.05, 0.95])returnf"{np.median(a):.{d}f}<br><span class='rng'>{lo:.{d}f}–{hi:.{d}f}</span>"def summary(a, d=3): a = np.asarray(a)returnf"{np.median(a):.{d}f} [{np.quantile(a, 0.05):.{d}f}, {np.quantile(a, 0.95):.{d}f}]"def md_table(df):return Markdown(df.to_markdown(index=False, disable_numparse=True))
2 LOESS calibration curves
The target is the calibration function of the model’s forecasts:
\[
m(p) = \mathbb E\big[Y \mid \hat p = p\big] = P\big(Y = 1 \mid \hat p = p\big).
\]
A calibrated model has \(m(p) = p\). A reliability diagram estimates \(m\) one bin at a time. LOESS estimates it as a smooth curve: at each point \(p_0\) it fits a straight line to the pairs \((\hat p_i, y_i)\) nearest \(p_0\), weighting close pairs more, and reads off the line’s height at \(p_0\) (Cleveland 1979; Austin & Steyerberg 2014 for its use on risk models).
The outcomes are 0 or 1, but a local average of zeros and ones is a local event frequency, which is what \(m\) is.
The span is the fraction of the data in each local fit. A small span follows noise; a large one flattens real structure.
Weights are tricube: \(w_i = \big(1 - (d_i / d_{\max})^3\big)^3\) for the distance \(d_i\) from \(p_0\).
Statistical subtlety
statsmodels.nonparametric.lowess runs three robustness iterations by default (it=3). Each iteration down-weights points with large residuals. With binary outcomes every residual is large, and the minority label in a neighbourhood gets treated as an outlier, so the curve is pulled toward whichever label is locally more common. For calibration, set it=0. The right panel of Figure 1 shows the damage.
Code
rng = np.random.default_rng(3)x, p_star, y = draw(rng, 2000)p_hat = distort(p_star, 2.5)xs0, m0 = loess_curve(p_hat, y, it=0)xs3, m3 = loess_curve(p_hat, y, it=3)boot = []for b inrange(200): i = rng.integers(0, len(y), len(y)) bx, bm = loess_curve(p_hat[i], y[i]) boot.append(np.interp(xs0, bx, bm))band = np.quantile(boot, [0.05, 0.95], axis=0)true_m = expit(logit(np.clip(xs0, 1e-15, 1-1e-15)) /2.5)# How far each version lands from the true curve at five points, over seeds.grid = np.array([0.05, 0.2, 0.5, 0.8, 0.95])err = {0: [], 3: []}for r inrange(100): _, ps_r, y_r = draw(np.random.default_rng(700+ r), 2000) ph_r = distort(ps_r, 2.5)for it in err: gx, gm = loess_curve(ph_r, y_r, it=it) err[it].append(np.interp(grid, gx, gm) - expit(logit(grid) /2.5))rmse = {it: np.sqrt(np.mean(np.square(err[it]), axis=0)) for it in err}
Figure 1: The overconfident forecaster (\(\alpha = 2.5\)), 2,000 cases. Left: outcomes jittered at 0 and 1, ten-bin reliability points, LOESS (span 0.3, no robustness iterations) with a 90% bootstrap band, and the true calibration curve. Right: the same data smoothed with statsmodels’ default of three robustness iterations.
Over 100 samples of 2,000 cases, the root-mean-square error of the LOESS curve against the true \(m(p)\) at \(p = 0.2\) is 0.028 with it=0 and 0.364 with it=3; at \(p = 0.8\) it is 0.026 against 0.363.
2.1 Reading the curve
Above the diagonal at \(p\), the event happens more often than the model says there; below it, less often. Whether that is over- or underconfidence depends on which side of 0.5 the point sits:
Table 1: How to read a calibration curve against the diagonal.
Curve at low \(p\)
Curve at high \(p\)
Shape
Diagnosis
above the diagonal
below the diagonal
flatter than the diagonal
overconfident
below the diagonal
above the diagonal
steeper than the diagonal
underconfident
above the diagonal
above the diagonal
shifted up
forecasts too low everywhere
below the diagonal
below the diagonal
shifted down
forecasts too high everywhere
Common misconception
“LOESS above the diagonal means underconfident.” It does only for \(p > 0.5\). An overconfident model sits above the diagonal at low forecasts: for the \(\alpha = 2.5\) forecaster, the true curve at a forecast of 0.04 is at 0.22. The diagnosis comes from the slope of the curve, not from which side of the diagonal one point is on.
2.2 Strengths and limits
LOESS has three advantages over bins theory:
a continuous estimate of \(m(p)\) at every forecast value;
local structure, such as a bump at 0.7, that a bin would average away;
no bin edges to choose.
Plain LOESS also has limits theory:
it need not stay inside \([0, 1]\), and it need not be monotone;
near 0 and 1 the local line extrapolates from one side, so the ends are its least reliable part;
the curve depends on the span, and the “right” span depends on \(n\) and on the curve;
fitted and evaluated on the same data, it can chase noise.
2.3 Diagnosis versus repair
Used diagnostically, LOESS answers “is the model calibrated on this data?” Used as a recalibration map, it becomes part of the model:
The second use needs the discipline of any model fit. Fit \(\widehat m\) on data the model was not trained on, evaluate the repaired probabilities on data \(\widehat m\) was not fitted on, or cross-fit: split the calibration data into folds, fit \(\widehat m\) on all folds but one, and apply it to the held-out fold.
3 Recalibration maps
A recalibration map \(g\) is a function from the raw probability (or logit) to a repaired probability, fitted on labelled data the model did not see. Five families cover most practice:
Table 2: Recalibration maps (Platt 1999; Guo et al. 2017; Zadrozny & Elkan 2002; Kull, Silva Filho & Flach 2017).
Map
Form
Parameters
Monotone?
Typical failure
Platt scaling
\(\sigma(a z + b)\), \(z = \operatorname{logit} \hat p\)
2
if \(a > 0\)
cannot bend: wrong when the distortion is not linear in the logit
Temperature scaling
\(\operatorname{softmax}(\mathbf z / T)\)
1
yes
one scale for every class; keeps the arg max
Isotonic regression
best non-decreasing step function
one per step
yes
steps and plateaus; outputs of exactly 0 or 1 on small samples
Beta calibration
\(\sigma\big(a \ln \hat p - b \ln(1 - \hat p) + c\big)\)
3
if \(a, b \ge 0\)
three parameters still limit the shape
LOESS, splines, GAMs
smooth function of \(\hat p\) or \(z\)
span or knots
not by default
needs smoothing choices and cross-fitting
Platt scaling is logistic regression on the model’s logit. For a binary model, temperature scaling is Platt scaling with \(b = 0\) and \(a = 1/T\).
Isotonic regression is fitted by the pool-adjacent-violators algorithm: sort by \(\hat p\), then merge neighbouring blocks until the block means never decrease. It suits a model whose ranking you trust and whose probabilities you do not.
Beta calibration contains the identity map, and it can fit shapes that Platt’s logistic curve cannot. A logistic GAM on the logit is the smooth, bounded alternative to LOESS.
No map wins everywhere. The next section measures three of them.
4 Experiment: train, calibrate, test
The experiment uses three disjoint splits for each seed:
Train (1,000 cases): fit the overfitted logistic regression. The \(\alpha = 2.5\) forecaster needs no training.
Calibrate (1,000 cases): fit Platt, isotonic and LOESS maps to the raw forecasts.
Test (20,000 cases): evaluate every version of the forecasts.
Code
def fit_platt(p, y): z = logit(np.clip(p, 1e-12, 1-1e-12)) fit = sm.Logit(y, sm.add_constant(z)).fit(disp=0)returnlambda s: expit(fit.params[0] + fit.params[1] * logit(np.clip(s, 1e-12, 1-1e-12)))def fit_isotonic(p, y): iso = IsotonicRegression(y_min=0, y_max=1, out_of_bounds="clip").fit(p, y)return iso.predictdef fit_loess(p, y, frac=0.3): xs, m = loess_curve(p, y, frac=frac)returnlambda s: np.clip(np.interp(s, xs, m), 0, 1)MAPS = {"Platt": fit_platt, "Isotonic": fit_isotonic, "LOESS": fit_loess}
Code
def run_calmap(model, seed): rng = np.random.default_rng(seed)if model =="mle": d =200 X = rng.standard_normal((1000, d)) y_tr = (rng.random(1000) < expit(2* X[:, 0])).astype(float) fit = sm.Logit(y_tr, sm.add_constant(X)).fit(disp=0, maxiter=100)ifnot fit.mle_retvals["converged"]:returnNone beta = fit.params Xc, Xt = rng.standard_normal((1000, d)), rng.standard_normal((20_000, d)) pc, pt = expit(beta[0] + Xc @ beta[1:]), expit(beta[0] + Xt @ beta[1:]) xc, xt = Xc[:, 0], Xt[:, 0]else: xc, xt = rng.standard_normal(1000), rng.standard_normal(20_000) pc, pt = distort(expit(2* xc), 2.5), distort(expit(2* xt), 2.5) yc = (rng.random(1000) < expit(2* xc)).astype(float) yt = (rng.random(20_000) < expit(2* xt)).astype(float) out = {("Raw", "test"): pt, ("Raw", "cal"): pc, ("Truth", "test"): expit(2* xt)}for name, fitter in MAPS.items(): g = fitter(pc, yc) out[(name, "test")], out[(name, "cal")] = g(pt), g(pc)return out, yc, ytcalmap = {m: {} for m in ("alpha", "mle")}examples, skipped = {}, 0for model in calmap:for r inrange(R): run = run_calmap(model, 30_000+ r)if run isNone: skipped +=1continue out, yc, yt = runfor (name, split), p in out.items(): yy = yt if split =="test"else yc calmap[model].setdefault((name, split), []).append((brier(p, yy), log_loss(p, yy), ece(p, yy))) examples.setdefault(model, (out, yt))calmap = {m: {k: np.array(v) for k, v in d.items()} for m, d in calmap.items()}def calmap_table(model): rows = []for name in ("Raw", "Platt", "Isotonic", "LOESS", "Truth"): t = calmap[model][(name, "test")] c = calmap[model].get((name, "cal")) rows.append({"Forecasts": name if name !="Truth"else"True $p^*$","Brier (test)": cell(t[:, 0]),"Log loss (test)": cell(t[:, 1]),"ECE (test)": cell(t[:, 2]),"Log loss (calibration data)": cell(c[:, 1]) if c isnotNoneelse"—", })return md_table(pd.DataFrame(rows))
Code
calmap_table("alpha")
Table 3: Recalibrating the overconfident forecaster (\(\alpha = 2.5\)). Maps fitted on 1,000 calibration cases; test metrics on 20,000 fresh cases; 200 seeds, median with 5–95% interval beneath.
Forecasts
Brier (test)
Log loss (test)
ECE (test)
Log loss (calibration data)
Raw
0.169 0.165–0.172
0.601 0.586–0.615
0.117 0.113–0.122
0.595 0.537–0.662
Platt
0.152 0.150–0.154
0.464 0.457–0.470
0.015 0.008–0.028
0.460 0.434–0.492
Isotonic
0.154 0.151–0.157
0.473 0.465–0.484
0.027 0.017–0.040
0.442 0.414–0.474
LOESS
0.153 0.151–0.156
0.469 0.462–0.476
0.025 0.017–0.038
0.460 0.430–0.490
True \(p^*\)
0.152 0.149–0.154
0.463 0.456–0.469
0.008 0.006–0.011
—
Code
calmap_table("mle")
Table 4: Recalibrating the overfitted logistic regression (200 features, one informative). Same splits and seeds.
Figure 2: One seed of the overfitted logistic regression on its 20,000 test cases: raw forecasts and the three recalibration maps, ten-bin reliability curves. The maps were fitted on a separate 1,000 calibration cases.
The regression’s Newton fit converged on 200 of 200 seeds, and its table uses those. What the two tables show:
All three maps cut test ECE and test log loss for both models. Platt does best on the \(\alpha = 2.5\) forecaster, where the distortion is linear in the logit and Platt is the correct model for it.
No map recovers the true \(p^*\) for the overfitted regression. Recalibration fixes the calibration term; the 199 noise features have already damaged the ranking, and a monotone map cannot restore resolution.
On its own calibration data, isotonic regression has an ECE of zero on every seed, up to rounding error (largest value 1.4e-16). Every pooled block of the fit has one fitted value equal to the mean of its labels, and a block’s cases all land in the same bin, so every bin’s forecast and frequency agree.
On the regression model, isotonic regression’s log loss on its own calibration data is lower than on test data by a median of 0.026 per seed. For Platt’s two-parameter map the gap is -0.002, as close to zero as for the unfitted raw forecasts (-0.005). The more flexible the map, the more its own fit flatters it.
Common misconception
“The recalibrated model has an ECE of 0.000.” On the data the map was fitted to, isotonic regression always does. Evaluate a recalibration map on data it has not seen.
5 Domain shift
Suppose Jev’s probabilities are calibrated on the distribution TypeSafe evaluates on. That does not imply
Calibration is a property of a model on a distribution, and four versions of it differ in what they condition on theory (Van Calster et al. 2016; Hébert-Johnson et al. 2018):
Marginal (mean) calibration: the average forecast equals the event rate.
Calibration in the sense of Part 2: \(P(Y = 1 \mid \hat p = p) = p\) for every \(p\).
Subgroup calibration:\(P(Y = 1 \mid \hat p = p, G = g) = p\) for every group \(g\) you care about.
Conditional (strong) calibration:\(P(Y = 1 \mid X = x) = \hat p(x)\) for every input, which makes \(\hat p\) the true probability.
Each implies every condition earlier in the list, with two provisos: subgroup calibration implies calibration only when the groups cover every case, and conditional calibration implies subgroup calibration only for groups that the model’s input determines. A group the model cannot see can be miscalibrated even when the forecast is the true probability given all the inputs the model sees, as the next example shows. Only conditional calibration survives a change in the input distribution with the relationship between \(X\) and \(Y\) held fixed; the others can fail when the mix of inputs moves (Ovadia et al. 2019 measured the effect on neural networks).
5.1 Subgroup calibration example
Two groups of equal size share the input distribution but differ in base rate: \(p^*_A(x) = \sigma(2x + 0.8)\) and \(p^*_B(x) = \sigma(2x - 0.8)\). A model that does not see the group, and outputs the average of the two, is calibrated overall. It is even conditionally calibrated given \(x\), the only input it has: \(\hat p(x) = P(Y = 1 \mid X = x)\).
Code
rng = np.random.default_rng(7)n =200_000x = rng.standard_normal(n)g = rng.random(n) <0.5p_true = expit(2* x + np.where(g, 0.8, -0.8))y = (rng.random(n) < p_true).astype(float)p_hat =0.5* expit(2* x +0.8) +0.5* expit(2* x -0.8)sub = {"All cases": np.ones(n, bool), "Group A": g, "Group B": ~g}sub_ece = {k: ece(p_hat[s], y[s]) for k, s in sub.items()}
Figure 3: A model that averages over two groups, 200,000 cases. Its forecasts are calibrated on all cases together, too low for group A and too high for group B.
The ECE is 0.002 on all cases and 0.120 and 0.119 within the groups. A calibration check on pooled data cannot see this; one run per group can.
5.2 Recalibrating to a shifted domain
The second experiment changes the relationship between \(X\) and \(Y\). A model calibrated on the source, \(\hat p = \sigma(2x)\), is deployed where the truth is different, and a recalibration map is fitted on \(n\) labelled cases from the new domain. Two kinds of shift:
Logit-linear shift, \(p^*_T(x) = \sigma(1.2x - 1)\): a lower base rate and a weaker signal. Platt’s two-parameter map is the correct model for it.
Label-noise shift, \(p^*_T(x) = 0.15 + 0.7\,\sigma(2x)\): 15% of labels in the new domain are flipped at random. No logistic curve in the logit fits it.
Code
SIZES = [50, 100, 300, 1000, 3000, 10_000]SHIFTS = {"Logit-linear shift": lambda x: expit(1.2* x -1.0),"Label-noise shift": lambda x: 0.15+0.7* expit(2* x),}shift = {}for label, truth in SHIFTS.items(): res = {k: {n: [] for n in SIZES} for k in MAPS} raw, best = [], []for r inrange(R): rng = np.random.default_rng(40_000+ r) xt = rng.standard_normal(20_000) pt_true = truth(xt) yt = (rng.random(20_000) < pt_true).astype(float) pt = expit(2* xt) floor = log_loss(pt_true, yt) raw.append(log_loss(pt, yt) - floor)for n in SIZES: xc = rng.standard_normal(n) yc = (rng.random(n) < truth(xc)).astype(float)for k, fitter in MAPS.items(): res[k][n].append(log_loss(fitter(expit(2* xc), yc)(pt), yt) - floor) shift[label] = (res, np.array(raw))
Code
fig, axes = plt.subplots(1, 2, figsize=(7.6, 3.4), sharey=True)for ax, (label, (res, raw)) inzip(axes, shift.items()):for (k, colour) inzip(MAPS, (PURPLE, TEAL, BLUE)): q = np.array([np.quantile(res[k][n], [0.05, 0.5, 0.95]) for n in SIZES]) ax.fill_between(SIZES, q[:, 0], q[:, 2], color=colour, alpha=0.12, lw=0) ax.plot(SIZES, q[:, 1], "o-", ms=3.5, lw=1.4, color=colour, label=k) ax.axhline(np.median(raw), color=INK, lw=1, ls=":") ax.set(xscale="log", yscale="log", title=label, xlabel="Labelled cases from the new domain")axes[0].set_ylabel("Excess log loss")axes[0].legend(frameon=False)plt.tight_layout()plt.show()def excess(label, k, n):return np.median(shift[label][0][k][n])def stays_ahead(label, a, b):"""Clause saying from which n map a keeps a lower median excess than map b."""for i, n inenumerate(SIZES):ifall(excess(label, a, m) < excess(label, b, m) for m in SIZES[i:]):returnf"stays ahead of {b} from {n:,} labels"returnf"does not stay ahead of {b} at any size up to {SIZES[-1]:,}"
Figure 4: Excess test log loss over the true probabilities after recalibrating on \(n\) labelled cases from the new domain; median and 5–95% band over 200 seeds, on log axes. The dotted line is the source model with no recalibration. On a finite test set the excess can dip below zero; the log scale cuts those seeds off the bottom of the Platt band.
With no recalibration the source model’s excess log loss is 0.145 under the logit-linear shift and 0.054 under label noise.
Under the logit-linear shift, Platt’s map with 100 labels brings the excess to 0.0074. Isotonic regression is still at 0.0101 with 1,000 labels.
Under label noise, Platt levels off near 0.0035 because it has the wrong shape. LOESS stays ahead of Platt from 1,000 labels and reaches 0.0003 at 10,000; isotonic regression stays ahead of Platt from 10,000 labels.
A two-parameter map wins when its shape is right and data are scarce; a flexible map wins when data are plentiful and the shape is unknown. Which regime you are in is an empirical question about your data.
5.3 Implications for Jev
TypeSafe states that “Jev is not fine-tuned or LoRA-adapted with customer data” and that “the same weights serve every account” vendor. Any calibration on your distribution is therefore something you check and, if needed, repair yourself inference.
The documentation says English is where accuracy is best and asks users to “test on your own content” for other languages vendor. Language is a subgroup to calibrate on separately.
TypeSafe advises pinning a versioned model ID if you have tuned thresholds against it vendor. A recalibration map is tuned in the same sense and belongs to one version.
The documentation lists a Noul and its negation on one ticket at 0.72 and 0.47, which sum to 1.19, and says “P(noul) and 1 - P(not noul) may not be directly comparable” vendor. Calibrate each question in the wording you deploy, and do not derive one polarity from the other.
A calibration check for one Noul question looks like this. The code is not run here: it needs an API key and a labelled sample of your own tickets. The SDK names follow TypeSafe’s Python documentation.
import numpy as npfrom typesafe_sdk import Noul, TypeSafeClientfrom sklearn.isotonic import IsotonicRegressionclient = TypeSafeClient(model="jev-1.13.0") # pin the version the map is fitted toquestion = {"refund": Noul(instructions="Does this message request a refund?")}def jev_probability(ticket):return client.system_one(ticket, question).nouls["refund"].noul# tickets, labels: a random sample of your own traffic, labelled by people.p = np.array([jev_probability(t) for t in tickets])y = np.array(labels, dtype=float)cal, test = np.arange(len(y)) %2==0, np.arange(len(y)) %2==1# 1. Diagnose on one half: LOESS with it=0, reliability bins, Brier and log loss.# 2. Fit the map on that half; evaluate it on the other half only.g = IsotonicRegression(y_min=0, y_max=1, out_of_bounds="clip").fit(p[cal], y[cal])p_repaired = g.predict(p[test])
How many labels the diagnosis needs follows from the binomial standard error: a bin of \(n_b\) cases estimates its frequency to about \(\sqrt{p(1-p)/n_b}\). At \(p = 0.9\), a bin of 100 cases gives a standard error of 0.03, so a gap of 0.05 is at the edge of detection. Checking the region above a 0.9 threshold therefore needs a few hundred labelled cases in that region, not a few hundred overall.
6 Decisions
A probability is worth calibrating because actions depend on it. Suppose a refund-routing system receives \(P(\text{warranted}) = 0.999\) for one ticket and \(0.71\) for another. The first can be approved automatically at almost any cost of error; the second cannot, unless mistakes are cheap. Statistical decision theory makes that precise.
6.1 Expected loss
With actions \(a\), outcomes \(y\) and a loss \(L(a, y)\), the Bayes action minimises expected loss under the predictive distribution:
The rule uses \(P(y \mid x)\) as a probability. If the model’s number is not one, because it is miscalibrated, the rule optimises the wrong expectation.
6.2 Two actions: a cost-derived threshold
Approve or decline a refund. Approving an unwarranted refund costs \(c_{\mathrm{FP}}\); declining a warranted one costs \(c_{\mathrm{FN}}\); correct actions cost nothing. With \(p = P(\text{warranted} \mid x)\):
\[
p > p_0 = \frac{c_{\mathrm{FP}}}{c_{\mathrm{FP}} + c_{\mathrm{FN}}}.
\]
This is Elkan’s (2001) threshold. With \(c_{\mathrm{FP}} = 10\) and \(c_{\mathrm{FN}} = 40\), \(p_0 = 0.2\): declining a warranted refund is four times worse, so approval starts at a low probability. The threshold is only as good as the probability it is applied to.
6.3 Three actions: human review
Add a third action, send to a person, at cost \(c_{\mathrm{H}}\) whatever the outcome. Review beats both automatic actions when \(c_{\mathrm{H}} < (1 - p)\,c_{\mathrm{FP}}\) and \(c_{\mathrm{H}} < p\,c_{\mathrm{FN}}\), that is, for
\[
\frac{c_{\mathrm{H}}}{c_{\mathrm{FN}}} < p < 1 - \frac{c_{\mathrm{H}}}{c_{\mathrm{FP}}}.
\]
This is Chow’s (1970) reject option in cost form. When the interval is empty, review is never worth its cost and the two-action threshold applies.
With \(c_{\mathrm{FP}} = 10\), \(c_{\mathrm{FN}} = 40\) and \(c_{\mathrm{H}} = 3\), the Bayes rule declines below 0.075, approves above 0.70, and sends the cases between to review. Applied to the process in Section 1, computed by quadrature over \(X\) rather than by sampling, with costs in the same units as \(c_{\mathrm{FP}}\), \(c_{\mathrm{FN}}\) and \(c_{\mathrm{H}}\):
Calibrated probabilities: 2293 per 1,000 cases, with 44% handled automatically.
The overconfident forecaster (\(\alpha = 2.5\)), same thresholds: 3085 per 1,000 cases, with 74% automated. It automates more and pays for it in wrong approvals and declines.
After Platt recalibration on 1,000 labelled cases: 2297 [2294, 2313] per 1,000 cases over 200 seeds.
The widget runs the same computation. Change the costs and the distortion, and compare the cost of acting on raw and on recalibrated probabilities.
At the default costs, raise \(\alpha\) from 1 to 4. The raw forecaster’s cost rises and its automation rate climbs; the recalibrated cost does not move.
At the same costs, lower \(\alpha\) below 1. The underconfident forecaster sends too many cases to review, and it also costs more than the recalibrated one.
Raise the review cost until the review band closes. The rule collapses to the single threshold \(c_{\mathrm{FP}} / (c_{\mathrm{FP}} + c_{\mathrm{FN}})\).
6.4 Calibration versus ranking
Cost-derived thresholds, expected-utility decisions, abstention and escalation all read the number as a probability. They need calibration.
A fixed review budget, such as “the 50 least certain tickets each day”, needs only the order of the cases, which a monotone miscalibration preserves.
Resource allocation and model routing sit between the two. Sending a case to an expensive model when \(P(\text{correct})\) is below a cost-derived line is a threshold decision; sending the bottom decile is a ranking decision.
TypeSafe’s routing guidance applies thresholds to its confidence field, a statistic of the spread of the distribution rather than a probability vendor. Part 1 converts those thresholds to probabilities. A cost-derived threshold belongs on a calibrated probability theory.
7 Synthesis
Jev combines older ideas: a foundation model’s reading of text, a classifier’s closed output space, a forecaster’s probabilities, a decision theorist’s loss, and a typed programming interface. What may be new is building the model, the API and the training objective around one map:
The first map is useful to software only if its probabilities mean what they say on the software’s own data. The tools for checking that are the ones in this series, and they do not depend on how the model was trained.
8 Open questions
The architecture questions are in Part 1. On calibration:
What does RLCD optimise?
Does a proper scoring rule enter the objective directly?
How does TypeSafe evaluate calibration: which metric, which binning, which sample sizes?
Calibration over which distribution of tasks, questions and inputs?
How well does calibration hold up under domain shift, across languages, and across question wordings?
How much does task-specific recalibration improve Jev on real workloads?
9 Constraints
All data is synthetic, with one binary outcome. Choice and Score answers need classwise or top-label maps and ordinal checks.
The decision example has three actions with fixed costs. Real costs vary by case, and review has limited capacity.
None of the experiments call Jev. Its calibration on any real workload remains untested in public.
Calibrate. Locally. Repair. On. Unseen. Data. Then. Let. Costs. Decide.