Skip to content

Calibration and scoring

Certify thresholds on your own labelled decisions, and score results against labels. The file format both read is one JSON object per line — {"context": ..., "decision": <spec>, "label": ...} — with the labels The certificate lists.

from mimir import Mimir
from mimir.core.labels import read_labelled
from mimir.evaluation.bench import bench
from mimir.runtime.calibrate import calibrate

model = Mimir.from_pretrained("Mythologic/MIMIR-1")
labelled = read_labelled("labelled.jsonl")
calibration = calibrate(model, labelled, risk=0.01, confidence=0.95)
results = [model.decide_uncertified(item.context, item.decision) for item in labelled]
report = bench(results, [item.label for item in labelled])

calibrate replaces the release thresholds with ones certified on your data at one risk level; calibration.types reports each decision type's records, taken count and certified threshold. Certifiable types are binary, categorical, multilabel, ranking and ordinal. bench reports accuracy, coverage and realised risk per spec type, each with a 95% Wilson interval. The mimir calibrate and mimir bench commands do the same over files; see Command line.

mimir.runtime.calibrate

Custom policies certified on labelled decisions (mimir calibrate).

A custom policy copies the release policy's calibration, gate and conformal scores and replaces its thresholds with ones certified on the caller's data at a single risk level. Each decision type is one test, and the tests split 1 - confidence equally (Bonferroni). The fingerprint is set to the local model, runtime and hardware, so the policy loads only in that configuration.

TypeCalibration dataclass

Certification outcome for one model decision type.

needed is the fewest accepted decisions that could certify the risk with zero errors.

calibrate

calibrate(
    engine: Mimir,
    labelled: Sequence[LabelledDecision],
    *,
    risk: float,
    confidence: float,
    batch_size: int | None = None,
) -> Calibration

Certify thresholds on labelled and return the custom policy with a per-type report.

Raises:

Type Description
PolicyError

the engine has no release policy to start from, or no labelled decision is of a certifiable type.

mimir.evaluation.bench

Accuracy, coverage and realised risk of results against labels (mimir bench).

Per spec type:

  • accuracy: correct answers among all records, ignoring the policy;
  • coverage: records not deferred;
  • risk: wrong answers among records not deferred.

Each rate has a 95% Wilson interval. estimate also reports the mean absolute error.

bench

bench(
    results: Sequence[DecisionResult],
    labels: Sequence[Label],
) -> BenchReport

Score results against their labels, grouped by spec type.

mimir.core.labels

Labelled decisions, the input format of mimir bench and mimir calibrate.

A JSONL file with one {"context": ..., "decision": <spec>, "label": ...} object per line. Labels by spec type:

Spec Label
choice an option id, or null when no option applies
multi_choice the list of option ids that apply
yes_no true or false
verify supported, contradicted or not_enough_information
rank the id of the best candidate, or the list of ids that are equally best
rate a level id
estimate a number within [low, high]

LabelError

Bases: ValueError

Invalid labelled file: bad JSON, an invalid spec, or a label that does not fit it.

is_correct

is_correct(result: DecisionResult, label: Label) -> bool

Return whether the answer matches the label.

For rank, the top candidate must be one of the labelled best. For estimate, the label must lie in the prediction interval.

read_labelled

read_labelled(path: Path) -> list[LabelledDecision]

Read a labelled JSONL file. Errors report the line number and field.