Calibrated confidence for zero-shot LLMs

How llmex turns a causal LM’s logits into probabilities you can actually filter and trust.

Zero-shot with an instruct-tuned LM is a fantastic way to prototype an NLP pipeline. In fifteen lines of Python you can classify tickets, extract entities, verify claims, or rank passages — no fine-tuning, no annotated data, no task-specific head.

The default recipe for turning that free-text output into something a downstream system can consume is structured extraction: hand the model a JSON schema in the prompt and ask it to emit conforming JSON. That works, but it has three costs you pay every call:

You could try to squeeze a confidence out by prompting: “On a scale of 0 to 1, how confident are you?”. Don’t. Self-reported confidence from a decoder LM is only weakly correlated with correctness, saturates near 1.0, and is trained-to-please rather than trained-to-be-honest. When you actually measure it on REBEL relation extraction, the model reports a mean confidence of 0.58 on a task where it gets 7/84 = 8.3 % right — off by a factor of seven, in the wrong direction.

llmex fixes both. Every function returns a value and a confidence in [0, 1] derived from the model’s own logits (not from anything the model says about itself), and can be calibrated to match empirical accuracy with a 50-example fit. The schema stays short — you pass it as a Python object, not as prompt text — and structural correctness comes for free because llmex reads probabilities over the candidates you provided, one at a time, instead of generating a JSON blob and hoping.

Standard zero-shot Prompt "Classify: 'I love it'" Causal LM Free-form reply "positive" ? — no confidence llmex — logit scoring over candidates Prompt + candidates choices=["pos","neg","neu"] Causal LM Softmax over candidates {"pos":0.91,"neg":0.05,"neu":0.04} value + confidence
Instead of parsing a text reply, llmex reads the model's own probability distribution over the labels you supplied.

The core idea in one paragraph

An instruct-tuned causal LM, given a prompt like "Sentiment: 'I love it'. Answer:\n", computes a probability distribution over its ~150k-token vocabulary for the very next token. score_choices picks out just the tokens that spell your candidate labels, renormalises to a proper distribution, and returns it. One forward pass, N candidates, a valid probability distribution.

from llmex import classify
classify("I love this product!", ["positive", "negative", "neutral"], model, tokenizer)
# {"positive": 0.9134, "negative": 0.0512, "neutral": 0.0354}

That single trick powers classify, mcqa, verify, rank, disambiguate, and every “score the alternatives” step in extract / extract_entities / analyze / extract_relations.

Two edge cases, handled automatically

First-token collisions

The tokeniser doesn’t respect your label semantics. "Sports" and "Sci/Tech" both start with the token "S"; the digits "1" and "10" start with the token "1". A naïve first-token softmax over such labels is meaningless — it can’t distinguish them.

score_choices detects a collision → falls back to full-span scoring Candidates "Sports" → ["S", "ports"] "Sci/Tech" → ["S", "ci", "/Tech"] First-token clash on "S" next-token softmax can't separate the two labels Fall back to score_span geometric mean of P(token | prefix) across the full label span Cost: only the colliding candidates pay for extra forward passes. Non-colliding labels stay on the fast 1-pass path. Length-normalised so a 3-token label doesn't lose to a 1-token label just for being longer.
The tokeniser doesn't respect label boundaries — llmex detects and works around it per-label, not per-batch.

Multi-label ≠ softmax

If a review can talk about both battery and screen, softmax over labels is the wrong model — it forces the probability mass to sum to 1 across labels that aren’t mutually exclusive. classify(multilabel=True) asks one independent yes/no question per label and returns P(yes). Real probabilities. 0.5 is a meaningful threshold.

Confidence for extraction: recognition prompts

For open-ended tasks — NER, ABSA, relations, structured JSON — the classic approach is: generate the span, then teacher-force it back through the model and read the joint probability of the tokens. That works, but the numbers you get are useless in practice: entity text confidences of 0.004, aspect confidences of 0.0001. Sort-by-confidence returns garbage.

llmex switches to a recognition prompt instead. After the generation pass finds a candidate span, it re-prompts the model with a yes/no question — "Is [Elon Musk] the PERSON entity in this text?" — and reads P(yes) via the same collision-safe logit-scoring path.

extract_entities() — generate → recognise Stage 1 — greedy generation (1 pass) source "Elon Musk announced…" Model emits spans + labels candidate spans [Elon Musk/PERSON, Tesla/ORG] Stage 2 — per-span recognition (yes/no scoring) prompt (repeated per span, KV-cached prefix) "Is 'Elon Musk' the PERSON entity in the text?" "Answer yes or no:" score_choices(["yes","no"]) P(yes) = 0.963 ← text_confidence label_confidence via softmax Per-entity output: confidence = geometric mean(text_conf, label_conf) — one number for sort / filter, either weak component pulls it down.
The recognition prompt gives well-separated span confidences (0.5–1.0 range) instead of teacher-forcing crush (near 0). Same trick for ABSA, relations, and per-field extract().

The result: extract_entities() gives you span confidences you can actually filter on. The KV prefix cache reuses the source-text prefix across all recognition prompts, so the second stage is nearly free even on CPU.

Making the numbers trustworthy: temperature scaling

Even with recognition prompts, the raw probabilities coming out of an instruct-tuned LM are systematically overconfident. A 0.95 score on a task with 8 % accuracy tells you the model likes to sound sure. Ranking is fine (AUC > 0.5), but the numbers themselves lie.

Fit a single scalar T on a small labelled calibration set (~50 examples) and apply softmax(logits / T) at inference. That’s temperature scaling. It is monotonic — argmax, accuracy, and AUC are exactly preserved — so it can only help.

REBEL relation extraction — reliability diagram (want: dots on the diagonal) Before (raw) ECE = 0.496 0 0.5 1 0 0.5 1 confidence bin accuracy After (T=4.0) ECE = 0.089 0 0.5 1 0 0.5 1 confidence bin Argmax, accuracy, and AUC are unchanged by T-scaling — only the probability changes. Mean conf drops from 0.579 to 0.172, matching true accuracy 0.083.
Fitting a single scalar T on ~50 labelled examples took REBEL relation confidences from wildly overconfident to nearly perfectly calibrated. Same trick works for every task in llmex.
from llmex import fit_temperature, classify

# 50 labelled examples drawn from your task distribution
T = fit_temperature(cls_calib, choices, model, tokenizer)   # e.g. 1.7

# Same call, one extra kwarg
scores = classify(text, choices, model, tokenizer, temperature=T)

Two things to watch for. First, the calibration set has to include hard cases — wrong or uncertain predictions. Feed fit_temperature only examples the model got right, and NLL fitting picks T < 1 (sharpening) and the overconfidence stays. Stratify by prediction confidence when you build the set. Second, per-task T values in the wild span roughly 0.7 to 5+; there’s no universal default, so fit per task and per model.

If the fitted T saturates at the top of the grid, extend it:

T = fit_temperature(calib, choices, model, tokenizer, t_grid=[2, 5, 8, 12, 16, 20])

When to trust the score

A pre-calibration rule of thumb:

RangeMeaning
0.90 – 1.00Model was highly certain
0.70 – 0.90Confident, but alternatives considered
0.50 – 0.70Moderate — worth reviewing
< 0.50Low — may be unreliable

After temperature scaling on a set that included hard cases, those thresholds correspond to genuine probabilities and you can pick decision thresholds on principled grounds — target ECE, precision at a coverage rate, or a conformal-prediction-set width. Composing per-check confidences into a per-record verdict? Use the geometric mean, not the arithmetic mean or raw product — it penalises any weak link and stays comparable across records with different numbers of checks.

What llmex is not

It doesn’t fine-tune. It doesn’t call an API. It doesn’t quantise. It just reads logits from a local causal LM and turns them into probability distributions with the correct shape for each task. The whole library is small enough to read in an afternoon.

If your workflow currently parses free-text LM output with a regex and hopes for the best, classify(), extract(), and a 50-sample calibration pass will give you a triage queue you can actually work.


Open source publication coming soon. We’re preparing llmex for public release — watch this space.