Skip to content

Core

The contract every surface shares: decision specs, contexts, results, decision tools, and tool-call checks. It depends only on Pydantic, so a project can build requests and read results without installing ONNX Runtime.

Decider

Decider is the interface both Mimir and MimirClient implement. It defines:

  • decide(context, spec, risk) and decide_many(pairs, risk) — the primary decision methods, over any spec type.
  • One shortcut per spec: choose, yes_no, verify, rank, rate, estimate.
  • Async counterparts for every method: adecide, adecide_many, achoose, and so on.
  • decide_uncertified(context, spec) — the raw model answer with no policy applied: no certificate, no deferral gate.
  • tool(name, spec, description) — bind a spec to a name to create a decision tool.
  • tool_call_check(rules, tools) — create a check that gates an agent's tool calls against rules.
  • info() — the loaded model's id, revision, variant, runtime fingerprint, certified risk levels, and input limits.

mimir.core.decider

The Decider base class, implemented by Mimir (local) and MimirClient (HTTP).

Subclasses implement _run and info. The public methods, shortcuts and async variants are defined here.

Decider

Bases: ABC

Abstract base for decision engines.

info abstractmethod

info() -> ModelInfo

Return the loaded model, runtime and certified risk levels.

decide

decide(
    context: ContextLike,
    spec: Choice,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> ChoiceResult
decide(
    context: ContextLike,
    spec: MultiChoice,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> MultiChoiceResult
decide(
    context: ContextLike,
    spec: YesNo,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> YesNoResult
decide(
    context: ContextLike,
    spec: Verify,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> VerifyResult
decide(
    context: ContextLike,
    spec: Rank,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> RankResult
decide(
    context: ContextLike,
    spec: Rate,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> RateResult
decide(
    context: ContextLike,
    spec: Estimate,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> EstimateResult
decide(
    context: ContextLike,
    spec: DecisionSpec,
    *,
    risk: float = DEFAULT_RISK,
    alpha: float | None = None,
) -> DecisionResult

Make a certified decision.

risk must be one of info().risk_levels. alpha is the miscoverage of the conformal prediction set; None uses the release default.

decide_many

decide_many(
    items: Items,
    *,
    risk: float = DEFAULT_RISK,
    alpha: float | None = None,
    batch_size: int | None = None,
) -> list[DecisionResult]

Make certified decisions for many (context, spec) pairs, returned in order.

Requests are sorted by length and packed into batches up to the release's token budget. batch_size additionally caps the number of requests per batch.

decide_uncertified

decide_uncertified(
    context: ContextLike, spec: Choice
) -> ChoiceResult
decide_uncertified(
    context: ContextLike, spec: MultiChoice
) -> MultiChoiceResult
decide_uncertified(
    context: ContextLike, spec: YesNo
) -> YesNoResult
decide_uncertified(
    context: ContextLike, spec: Verify
) -> VerifyResult
decide_uncertified(
    context: ContextLike, spec: Rank
) -> RankResult
decide_uncertified(
    context: ContextLike, spec: Rate
) -> RateResult
decide_uncertified(
    context: ContextLike, spec: Estimate
) -> EstimateResult
decide_uncertified(
    context: ContextLike, spec: DecisionSpec
) -> DecisionResult

Return the raw model answer, without calibration, gating or certificate.

decide_uncertified_many

decide_uncertified_many(
    items: Items, *, batch_size: int | None = None
) -> list[DecisionResult]

Return raw model answers for many (context, spec) pairs, batched as decide_many.

adecide async

adecide(
    context: ContextLike,
    spec: Choice,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> ChoiceResult
adecide(
    context: ContextLike,
    spec: MultiChoice,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> MultiChoiceResult
adecide(
    context: ContextLike,
    spec: YesNo,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> YesNoResult
adecide(
    context: ContextLike,
    spec: Verify,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> VerifyResult
adecide(
    context: ContextLike,
    spec: Rank,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> RankResult
adecide(
    context: ContextLike,
    spec: Rate,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> RateResult
adecide(
    context: ContextLike,
    spec: Estimate,
    *,
    risk: float = ...,
    alpha: float | None = ...,
) -> EstimateResult
adecide(
    context: ContextLike,
    spec: DecisionSpec,
    *,
    risk: float = DEFAULT_RISK,
    alpha: float | None = None,
) -> DecisionResult

Async version of decide.

adecide_many async

adecide_many(
    items: Items,
    *,
    risk: float = DEFAULT_RISK,
    alpha: float | None = None,
    batch_size: int | None = None,
) -> list[DecisionResult]

Async version of decide_many.

adecide_uncertified async

adecide_uncertified(
    context: ContextLike, spec: Choice
) -> ChoiceResult
adecide_uncertified(
    context: ContextLike, spec: MultiChoice
) -> MultiChoiceResult
adecide_uncertified(
    context: ContextLike, spec: YesNo
) -> YesNoResult
adecide_uncertified(
    context: ContextLike, spec: Verify
) -> VerifyResult
adecide_uncertified(
    context: ContextLike, spec: Rank
) -> RankResult
adecide_uncertified(
    context: ContextLike, spec: Rate
) -> RateResult
adecide_uncertified(
    context: ContextLike, spec: Estimate
) -> EstimateResult
adecide_uncertified(
    context: ContextLike, spec: DecisionSpec
) -> DecisionResult

Async version of decide_uncertified.

adecide_uncertified_many async

adecide_uncertified_many(
    items: Items, *, batch_size: int | None = None
) -> list[DecisionResult]

Async version of decide_uncertified_many.

tool

tool(
    name: str,
    spec: DecisionSpec,
    description: str,
    *,
    risk: float = DEFAULT_RISK,
    alpha: float | None = None,
) -> DecisionTool

Create a DecisionTool that applies spec to any context it is called with.

tool_call_check

tool_call_check(
    rules: str | Sequence[str],
    *,
    question: str = DEFAULT_QUESTION,
    tools: Iterable[str] | None = None,
    risk: float = DEFAULT_RISK,
) -> ToolCallCheck

Create a ToolCallCheck deciding pending calls of tools (every tool if None) against rules.

mimir.core.decisions

Decision specs: the question and the options to decide between.

Spec Model decision type Answer
Choice categorical (binary for two options) an option id, or None
MultiChoice multilabel the option ids that apply
YesNo binary True or False
Verify categorical over three verdicts a Verdict
Rank ranking candidate ids, best first
Rate ordinal a level id
Estimate continuous a number in [low, high]

Options are a sequence of texts or a mapping of id to text. Results use the ids. YesNo and Verify use fixed option texts matching the training data.

Choice

Bases: _Spec

Pick one option, or none.

MultiChoice

Bases: _Spec

Pick every option that applies. The answer may be empty.

YesNo

Bases: _Spec

Answer a yes/no question.

Verify

Bases: BaseModel

Check whether the context supports or contradicts a claim.

Rank

Bases: _Spec

Rank candidates from best to worst.

Rate

Bases: _Spec

Rate on an ordered scale. Levels are given from lowest to highest.

Estimate

Bases: _Spec

Estimate a value in [low, high].

mimir.core.context

Decision inputs: passages, tables and typed fields.

Every entry point accepts a ContextLike:

  • str: one passage;
  • Sequence[str]: one passage per string;
  • Mapping: a JSON state, flattened into fields with dotted keys;
  • Context or JsonState.

Numbers and dates in table cells and fields are typed; their original text is kept.

Passage

Bases: _Frozen

A passage of text with an optional title.

Cell

Bases: _Frozen

A table cell with its parsed number and date, if any.

parse classmethod

parse(raw: str) -> Self

Create a cell from raw text, parsing its number and date.

Table

Bases: _Frozen

A table with an optional caption and header.

from_rows classmethod

from_rows(
    rows: Sequence[Sequence[CellValue]],
    *,
    header: Sequence[str] = (),
    caption: str = "",
) -> Self

Create a table from raw values.

Strings are parsed for a number and a date. Numbers keep their value, integers and decimals their exact text. Dates are typed; a datetime at midnight is its date. Booleans read true or false, and None is a blank cell.

from_dataframe classmethod

from_dataframe(
    frame: DataFrame | DataFrame, *, caption: str = ""
) -> Self

Create a table from a pandas or polars DataFrame: its columns are the header, its index is dropped, and missing values are blank cells. Cells convert as in from_rows.

Field

Bases: _Frozen

A typed field of a structured state, keyed by its dotted path.

from_json classmethod

from_json(
    value: JsonValue, prefix: str = ""
) -> tuple[Field, ...]

Flatten a JSON value into fields, e.g. {"a": {"b": [1]}} gives key a.b[0].

Kinds: a list of strings is one list field, booleans and null are category, numbers are number, date strings are datetime and other strings are text.

Context

Bases: _Frozen

The input of a decision. May be empty, in which case only the question is read.

coerce classmethod

coerce(value: ContextLike) -> Context

Convert a ContextLike to a Context.

JsonState

Bases: _Frozen

A structured state in a JSON request body: {"state": {...}}.

field_date

field_date(value: str) -> datetime | None

Parse a datetime field value: a date in a parse_date format, or ISO 8601.

mimir.core.results

Decision results.

answer is the model's prediction. status is the policy's verdict on it:

  • DECIDED: the answer is an option and is certified.
  • ABSTAINED: the answer is "none of the options" and is certified.
  • DEFERRED: not certified; deferral.reason gives the cause.

Results from decide_uncertified have raw model probabilities, a status taken from the answer alone, and no certificate, deferral or prediction set.

relevant_context is sorted by relevance, highest first. Relevance is the share of the model's evidence attention on each part of the context, not a causal attribution.

Deferral

Bases: _Frozen

Reason a decision was deferred.

  • no_certified_threshold: no threshold is certified for this decision type and risk.
  • out_of_distribution: gate_p_value is at or below the policy's gate level.
  • below_threshold: confidence is below threshold.

Certificate

Bases: _Frozen

The certified threshold a decision was checked against.

On records held-out decisions, taken reached the threshold and errors of those were wrong. The exact binomial test of errors out of taken at rate risk gives p_value: the error rate of taken decisions is at most risk with probability confidence. coverage is taken / records, with a 95% Wilson interval.

ContextRelevance

Bases: _Frozen

Relevance of one part of the context.

index is the position of the passage, table or field in the context. row is the row position for table_row and None otherwise. A table part is the caption and header.

OptionSet

Bases: _Frozen

A conformal prediction set. abstain is true when "none of the options" is included.

mimir.core.tools

Named decision tools for agents.

A DecisionTool fixes the question and options, so the caller supplies only the context. Framework adapters and the MCP server expose these tools. ToolDefinitions declares them, as a tools file does.

ToolArguments

Bases: BaseModel

Arguments of a decision tool call.

DecisionTool dataclass

A decision spec bound to a name. tool(context) returns its result.

agent_description property

agent_description: str

The description an agent reads: description, then how to act on the status.

input_schema property

input_schema: dict[str, JsonValue]

JSON Schema of the tool's arguments.

output_schema property

output_schema: dict[str, JsonValue]

JSON Schema of the tool's result.

call_with

call_with(
    arguments: Mapping[str, object],
) -> DecisionResult

Decide on a tool call's arguments, validated against input_schema.

acall_with async

acall_with(
    arguments: Mapping[str, object],
) -> DecisionResult

call_with, awaitable.

ToolDefinition

Bases: BaseModel

One declared tool: the arguments of Decider.tool.

ToolDefinitions

Bases: BaseModel

The tools a server exposes, in the order it lists them. Names are unique.

bind

bind(decider: Decider) -> tuple[DecisionTool, ...]

Return one DecisionTool per definition, answered by decider.

mimir.core.checks

Tool-call checks: whether an agent's pending tool call may run under written rules.

The model reads each rule as a passage and the call as fields (tool, arguments.<name>), and answers a yes/no question about it. A certified yes allows the call, a certified no denies it, and a deferral or an abstention escalates it to a person.

CheckOutcome dataclass

The permission for one pending call, and the decision it was read from. result is None for a tool the check does not apply to.

reason property

reason: str

One sentence for the agent or the approver, naming the tool and the confidence.

ToolCallCheck dataclass

A yes/no decision on pending tool calls, made against rules.

tools names the tools it applies to; None applies it to every tool. Calls to other tools are allowed without a decision.

applies_to

applies_to(tool: str) -> bool

Whether calls to tool are decided rather than allowed outright.

context

context(
    tool: str, arguments: Mapping[str, JsonValue]
) -> Context

The context the model reads for one pending call.