Calibrated confidence for zero-shot LLMs
How llmex turns a causal LM’s logits into probabilities you can actually filter and trust.
Zero-shot with an instruct-tuned LM is a fantastic way to prototype an NLP pipeline. In fifteen lines of Python you can classify tickets, extract entities, verify claims, or rank passages — no fine-tuning, no annotated data, no task-specific head.
The default recipe for turning that free-text output into something a downstream system can consume is structured extraction: hand the model a JSON schema in the prompt and ask it to emit conforming JSON. That works, but it has three costs you pay every call:
- The schema eats context. A modest Pydantic model burns 300–500 tokens of prompt before the input starts. On long documents that’s context you could have spent on the actual source.
- Parsing fails. Trailing commas, missing required keys, hallucinated fields, quoted numbers — common enough that most production wrappers ship a JSON-repair library and a retry loop.
- The returned JSON has no confidences. You get
category: electronicsand no signal about whether the model was 80 % sure or 30 % sure. Every downstream consumer has to treat every field as ground truth.
You could try to squeeze a confidence out by prompting: “On a scale of 0 to 1, how confident are you?”. Don’t. Self-reported confidence from a decoder LM is only weakly correlated with correctness, saturates near 1.0, and is trained-to-please rather than trained-to-be-honest. When you actually measure it on REBEL relation extraction, the model reports a mean confidence of 0.58 on a task where it gets 7/84 = 8.3 % right — off by a factor of seven, in the wrong direction.
llmex fixes both. Every function returns a value and a confidence in [0, 1] derived from the model’s own logits (not from anything the model says about itself), and can be calibrated to match empirical accuracy with a 50-example fit. The schema stays short — you pass it as a Python object, not as prompt text — and structural correctness comes for free because llmex reads probabilities over the candidates you provided, one at a time, instead of generating a JSON blob and hoping.
The core idea in one paragraph
An instruct-tuned causal LM, given a prompt like "Sentiment: 'I love it'. Answer:\n", computes a probability distribution over its ~150k-token vocabulary for the very next token. score_choices picks out just the tokens that spell your candidate labels, renormalises to a proper distribution, and returns it. One forward pass, N candidates, a valid probability distribution.
from llmex import classify
classify("I love this product!", ["positive", "negative", "neutral"], model, tokenizer)
# {"positive": 0.9134, "negative": 0.0512, "neutral": 0.0354}
That single trick powers classify, mcqa, verify, rank, disambiguate, and every “score the alternatives” step in extract / extract_entities / analyze / extract_relations.
Two edge cases, handled automatically
First-token collisions
The tokeniser doesn’t respect your label semantics. "Sports" and "Sci/Tech" both start with the token "S"; the digits "1" and "10" start with the token "1". A naïve first-token softmax over such labels is meaningless — it can’t distinguish them.
Multi-label ≠ softmax
If a review can talk about both battery and screen, softmax over labels is the wrong model — it forces the probability mass to sum to 1 across labels that aren’t mutually exclusive. classify(multilabel=True) asks one independent yes/no question per label and returns P(yes). Real probabilities. 0.5 is a meaningful threshold.
Confidence for extraction: recognition prompts
For open-ended tasks — NER, ABSA, relations, structured JSON — the classic approach is: generate the span, then teacher-force it back through the model and read the joint probability of the tokens. That works, but the numbers you get are useless in practice: entity text confidences of 0.004, aspect confidences of 0.0001. Sort-by-confidence returns garbage.
llmex switches to a recognition prompt instead. After the generation pass finds a candidate span, it re-prompts the model with a yes/no question — "Is [Elon Musk] the PERSON entity in this text?" — and reads P(yes) via the same collision-safe logit-scoring path.
The result: extract_entities() gives you span confidences you can actually filter on. The KV prefix cache reuses the source-text prefix across all recognition prompts, so the second stage is nearly free even on CPU.
Making the numbers trustworthy: temperature scaling
Even with recognition prompts, the raw probabilities coming out of an instruct-tuned LM are systematically overconfident. A 0.95 score on a task with 8 % accuracy tells you the model likes to sound sure. Ranking is fine (AUC > 0.5), but the numbers themselves lie.
Fit a single scalar T on a small labelled calibration set (~50 examples) and apply softmax(logits / T) at inference. That’s temperature scaling. It is monotonic — argmax, accuracy, and AUC are exactly preserved — so it can only help.
from llmex import fit_temperature, classify
# 50 labelled examples drawn from your task distribution
T = fit_temperature(cls_calib, choices, model, tokenizer) # e.g. 1.7
# Same call, one extra kwarg
scores = classify(text, choices, model, tokenizer, temperature=T)
Two things to watch for. First, the calibration set has to include hard cases — wrong or uncertain predictions. Feed fit_temperature only examples the model got right, and NLL fitting picks T < 1 (sharpening) and the overconfidence stays. Stratify by prediction confidence when you build the set. Second, per-task T values in the wild span roughly 0.7 to 5+; there’s no universal default, so fit per task and per model.
If the fitted T saturates at the top of the grid, extend it:
T = fit_temperature(calib, choices, model, tokenizer, t_grid=[2, 5, 8, 12, 16, 20])
When to trust the score
A pre-calibration rule of thumb:
| Range | Meaning |
|---|---|
| 0.90 – 1.00 | Model was highly certain |
| 0.70 – 0.90 | Confident, but alternatives considered |
| 0.50 – 0.70 | Moderate — worth reviewing |
| < 0.50 | Low — may be unreliable |
After temperature scaling on a set that included hard cases, those thresholds correspond to genuine probabilities and you can pick decision thresholds on principled grounds — target ECE, precision at a coverage rate, or a conformal-prediction-set width. Composing per-check confidences into a per-record verdict? Use the geometric mean, not the arithmetic mean or raw product — it penalises any weak link and stays comparable across records with different numbers of checks.
What llmex is not
It doesn’t fine-tune. It doesn’t call an API. It doesn’t quantise. It just reads logits from a local causal LM and turns them into probability distributions with the correct shape for each task. The whole library is small enough to read in an afternoon.
If your workflow currently parses free-text LM output with a regex and hopes for the best, classify(), extract(), and a 50-sample calibration pass will give you a triage queue you can actually work.
Open source publication coming soon. We’re preparing llmex for public release — watch this space.