Calibration and scoring¶
Certify thresholds on your own labelled decisions, and score results against labels.
The file format both read is one JSON object per line —
{"context": ..., "decision": <spec>, "label": ...} — with the labels
The certificate lists.
from mimir import Mimir
from mimir.core.labels import read_labelled
from mimir.evaluation.bench import bench
from mimir.runtime.calibrate import calibrate
model = Mimir.from_pretrained("Mythologic/MIMIR-1")
labelled = read_labelled("labelled.jsonl")
calibration = calibrate(model, labelled, risk=0.01, confidence=0.95)
results = [model.decide_uncertified(item.context, item.decision) for item in labelled]
report = bench(results, [item.label for item in labelled])
calibrate replaces the release thresholds with ones certified on your data at one risk
level; calibration.types reports each decision type's records, taken count and
certified threshold. Certifiable types are binary, categorical, multilabel, ranking and
ordinal. bench reports accuracy, coverage and realised risk per spec type,
each with a 95% Wilson interval. The mimir calibrate and mimir bench commands do the
same over files; see Command line.
mimir.runtime.calibrate ¶
Custom policies certified on labelled decisions (mimir calibrate).
A custom policy copies the release policy's calibration, gate and conformal scores and replaces
its thresholds with ones certified on the caller's data at a single risk level. Each decision
type is one test, and the tests split 1 - confidence equally (Bonferroni). The fingerprint is
set to the local model, runtime and hardware, so the policy loads only in that configuration.
TypeCalibration
dataclass
¶
Certification outcome for one model decision type.
needed is the fewest accepted decisions that could certify the risk with zero errors.
calibrate ¶
calibrate(
engine: Mimir,
labelled: Sequence[LabelledDecision],
*,
risk: float,
confidence: float,
batch_size: int | None = None,
) -> Calibration
Certify thresholds on labelled and return the custom policy with a per-type report.
Raises:
| Type | Description |
|---|---|
PolicyError
|
the engine has no release policy to start from, or no labelled decision is of a certifiable type. |
mimir.evaluation.bench ¶
Accuracy, coverage and realised risk of results against labels (mimir bench).
Per spec type:
accuracy: correct answers among all records, ignoring the policy;coverage: records not deferred;risk: wrong answers among records not deferred.
Each rate has a 95% Wilson interval. estimate also reports the mean absolute error.
bench ¶
Score results against their labels, grouped by spec type.
mimir.core.labels ¶
Labelled decisions, the input format of mimir bench and mimir calibrate.
A JSONL file with one {"context": ..., "decision": <spec>, "label": ...} object per line.
Labels by spec type:
| Spec | Label |
|---|---|
choice |
an option id, or null when no option applies |
multi_choice |
the list of option ids that apply |
yes_no |
true or false |
verify |
supported, contradicted or not_enough_information |
rank |
the id of the best candidate, or the list of ids that are equally best |
rate |
a level id |
estimate |
a number within [low, high] |
LabelError ¶
Bases: ValueError
Invalid labelled file: bad JSON, an invalid spec, or a label that does not fit it.
is_correct ¶
Return whether the answer matches the label.
For rank, the top candidate must be one of the labelled best. For estimate, the label
must lie in the prediction interval.
read_labelled ¶
Read a labelled JSONL file. Errors report the line number and field.