Core¶
The contract every surface shares: decision specs, contexts, results, decision tools, and tool-call checks. It depends only on Pydantic, so a project can build requests and read results without installing ONNX Runtime.
Decider¶
Decider is the interface both Mimir and MimirClient implement. It defines:
decide(context, spec, risk)anddecide_many(pairs, risk)— the primary decision methods, over any spec type.- One shortcut per spec:
choose,yes_no,verify,rank,rate,estimate. - Async counterparts for every method:
adecide,adecide_many,achoose, and so on. decide_uncertified(context, spec)— the raw model answer with no policy applied: no certificate, no deferral gate.tool(name, spec, description)— bind a spec to a name to create a decision tool.tool_call_check(rules, tools)— create a check that gates an agent's tool calls against rules.info()— the loaded model's id, revision, variant, runtime fingerprint, certified risk levels, and input limits.
mimir.core.decider ¶
The Decider base class, implemented by Mimir (local) and MimirClient (HTTP).
Subclasses implement _run and info. The public methods, shortcuts and async variants are
defined here.
Decider ¶
Bases: ABC
Abstract base for decision engines.
info
abstractmethod
¶
Return the loaded model, runtime and certified risk levels.
decide ¶
decide(
context: ContextLike,
spec: Choice,
*,
risk: float = ...,
alpha: float | None = ...,
) -> ChoiceResult
decide(
context: ContextLike,
spec: MultiChoice,
*,
risk: float = ...,
alpha: float | None = ...,
) -> MultiChoiceResult
decide(
context: ContextLike,
spec: YesNo,
*,
risk: float = ...,
alpha: float | None = ...,
) -> YesNoResult
decide(
context: ContextLike,
spec: Verify,
*,
risk: float = ...,
alpha: float | None = ...,
) -> VerifyResult
decide(
context: ContextLike,
spec: Rank,
*,
risk: float = ...,
alpha: float | None = ...,
) -> RankResult
decide(
context: ContextLike,
spec: Rate,
*,
risk: float = ...,
alpha: float | None = ...,
) -> RateResult
decide(
context: ContextLike,
spec: Estimate,
*,
risk: float = ...,
alpha: float | None = ...,
) -> EstimateResult
decide(
context: ContextLike,
spec: DecisionSpec,
*,
risk: float = DEFAULT_RISK,
alpha: float | None = None,
) -> DecisionResult
Make a certified decision.
risk must be one of info().risk_levels. alpha is the miscoverage of the conformal
prediction set; None uses the release default.
decide_many ¶
decide_many(
items: Items,
*,
risk: float = DEFAULT_RISK,
alpha: float | None = None,
batch_size: int | None = None,
) -> list[DecisionResult]
Make certified decisions for many (context, spec) pairs, returned in order.
Requests are sorted by length and packed into batches up to the release's token budget.
batch_size additionally caps the number of requests per batch.
decide_uncertified ¶
decide_uncertified(
context: ContextLike, spec: Choice
) -> ChoiceResult
decide_uncertified(
context: ContextLike, spec: MultiChoice
) -> MultiChoiceResult
decide_uncertified(
context: ContextLike, spec: YesNo
) -> YesNoResult
decide_uncertified(
context: ContextLike, spec: Verify
) -> VerifyResult
decide_uncertified(
context: ContextLike, spec: Rank
) -> RankResult
decide_uncertified(
context: ContextLike, spec: Rate
) -> RateResult
decide_uncertified(
context: ContextLike, spec: Estimate
) -> EstimateResult
Return the raw model answer, without calibration, gating or certificate.
decide_uncertified_many ¶
Return raw model answers for many (context, spec) pairs, batched as decide_many.
adecide
async
¶
adecide(
context: ContextLike,
spec: Choice,
*,
risk: float = ...,
alpha: float | None = ...,
) -> ChoiceResult
adecide(
context: ContextLike,
spec: MultiChoice,
*,
risk: float = ...,
alpha: float | None = ...,
) -> MultiChoiceResult
adecide(
context: ContextLike,
spec: YesNo,
*,
risk: float = ...,
alpha: float | None = ...,
) -> YesNoResult
adecide(
context: ContextLike,
spec: Verify,
*,
risk: float = ...,
alpha: float | None = ...,
) -> VerifyResult
adecide(
context: ContextLike,
spec: Rank,
*,
risk: float = ...,
alpha: float | None = ...,
) -> RankResult
adecide(
context: ContextLike,
spec: Rate,
*,
risk: float = ...,
alpha: float | None = ...,
) -> RateResult
adecide(
context: ContextLike,
spec: Estimate,
*,
risk: float = ...,
alpha: float | None = ...,
) -> EstimateResult
adecide(
context: ContextLike,
spec: DecisionSpec,
*,
risk: float = DEFAULT_RISK,
alpha: float | None = None,
) -> DecisionResult
Async version of decide.
adecide_many
async
¶
adecide_many(
items: Items,
*,
risk: float = DEFAULT_RISK,
alpha: float | None = None,
batch_size: int | None = None,
) -> list[DecisionResult]
Async version of decide_many.
adecide_uncertified
async
¶
adecide_uncertified(
context: ContextLike, spec: Choice
) -> ChoiceResult
adecide_uncertified(
context: ContextLike, spec: MultiChoice
) -> MultiChoiceResult
adecide_uncertified(
context: ContextLike, spec: YesNo
) -> YesNoResult
adecide_uncertified(
context: ContextLike, spec: Verify
) -> VerifyResult
adecide_uncertified(
context: ContextLike, spec: Rank
) -> RankResult
adecide_uncertified(
context: ContextLike, spec: Rate
) -> RateResult
adecide_uncertified(
context: ContextLike, spec: Estimate
) -> EstimateResult
Async version of decide_uncertified.
adecide_uncertified_many
async
¶
Async version of decide_uncertified_many.
tool ¶
tool(
name: str,
spec: DecisionSpec,
description: str,
*,
risk: float = DEFAULT_RISK,
alpha: float | None = None,
) -> DecisionTool
Create a DecisionTool that applies spec to any context it is called with.
tool_call_check ¶
tool_call_check(
rules: str | Sequence[str],
*,
question: str = DEFAULT_QUESTION,
tools: Iterable[str] | None = None,
risk: float = DEFAULT_RISK,
) -> ToolCallCheck
Create a ToolCallCheck deciding pending calls of tools (every tool if None)
against rules.
mimir.core.decisions ¶
Decision specs: the question and the options to decide between.
| Spec | Model decision type | Answer |
|---|---|---|
Choice |
categorical (binary for two options) | an option id, or None |
MultiChoice |
multilabel | the option ids that apply |
YesNo |
binary | True or False |
Verify |
categorical over three verdicts | a Verdict |
Rank |
ranking | candidate ids, best first |
Rate |
ordinal | a level id |
Estimate |
continuous | a number in [low, high] |
Options are a sequence of texts or a mapping of id to text. Results use the ids.
YesNo and Verify use fixed option texts matching the training data.
Choice ¶
Bases: _Spec
Pick one option, or none.
MultiChoice ¶
Bases: _Spec
Pick every option that applies. The answer may be empty.
YesNo ¶
Bases: _Spec
Answer a yes/no question.
Verify ¶
Bases: BaseModel
Check whether the context supports or contradicts a claim.
Rank ¶
Bases: _Spec
Rank candidates from best to worst.
Rate ¶
Bases: _Spec
Rate on an ordered scale. Levels are given from lowest to highest.
Estimate ¶
Bases: _Spec
Estimate a value in [low, high].
mimir.core.context ¶
Decision inputs: passages, tables and typed fields.
Every entry point accepts a ContextLike:
str: one passage;Sequence[str]: one passage per string;Mapping: a JSON state, flattened into fields with dotted keys;ContextorJsonState.
Numbers and dates in table cells and fields are typed; their original text is kept.
Passage ¶
Bases: _Frozen
A passage of text with an optional title.
Cell ¶
Bases: _Frozen
A table cell with its parsed number and date, if any.
parse
classmethod
¶
Create a cell from raw text, parsing its number and date.
Table ¶
Bases: _Frozen
A table with an optional caption and header.
from_rows
classmethod
¶
from_rows(
rows: Sequence[Sequence[CellValue]],
*,
header: Sequence[str] = (),
caption: str = "",
) -> Self
Create a table from raw values.
Strings are parsed for a number and a date. Numbers keep their value, integers and
decimals their exact text. Dates are typed; a datetime at midnight is its date.
Booleans read true or false, and None is a blank cell.
from_dataframe
classmethod
¶
Create a table from a pandas or polars DataFrame: its columns are the header, its
index is dropped, and missing values are blank cells. Cells convert as in from_rows.
Field ¶
Bases: _Frozen
A typed field of a structured state, keyed by its dotted path.
from_json
classmethod
¶
from_json(
value: JsonValue, prefix: str = ""
) -> tuple[Field, ...]
Flatten a JSON value into fields, e.g. {"a": {"b": [1]}} gives key a.b[0].
Kinds: a list of strings is one list field, booleans and null are category,
numbers are number, date strings are datetime and other strings are text.
Context ¶
JsonState ¶
Bases: _Frozen
A structured state in a JSON request body: {"state": {...}}.
field_date ¶
Parse a datetime field value: a date in a parse_date format, or ISO 8601.
mimir.core.results ¶
Decision results.
answer is the model's prediction. status is the policy's verdict on it:
DECIDED: the answer is an option and is certified.ABSTAINED: the answer is "none of the options" and is certified.DEFERRED: not certified;deferral.reasongives the cause.
Results from decide_uncertified have raw model probabilities, a status taken from the answer
alone, and no certificate, deferral or prediction set.
relevant_context is sorted by relevance, highest first. Relevance is the share of the
model's evidence attention on each part of the context, not a causal attribution.
Deferral ¶
Bases: _Frozen
Reason a decision was deferred.
no_certified_threshold: no threshold is certified for this decision type and risk.out_of_distribution:gate_p_valueis at or below the policy's gate level.below_threshold:confidenceis belowthreshold.
Certificate ¶
Bases: _Frozen
The certified threshold a decision was checked against.
On records held-out decisions, taken reached the threshold and errors of those were
wrong. The exact binomial test of errors out of taken at rate risk gives p_value:
the error rate of taken decisions is at most risk with probability confidence.
coverage is taken / records, with a 95% Wilson interval.
ContextRelevance ¶
Bases: _Frozen
Relevance of one part of the context.
index is the position of the passage, table or field in the context. row is the row
position for table_row and None otherwise. A table part is the caption and header.
OptionSet ¶
Bases: _Frozen
A conformal prediction set. abstain is true when "none of the options" is included.
mimir.core.tools ¶
Named decision tools for agents.
A DecisionTool fixes the question and options, so the caller supplies only the context.
Framework adapters and the MCP server expose these tools. ToolDefinitions declares them, as
a tools file does.
ToolArguments ¶
Bases: BaseModel
Arguments of a decision tool call.
DecisionTool
dataclass
¶
ToolDefinition ¶
Bases: BaseModel
One declared tool: the arguments of Decider.tool.
ToolDefinitions ¶
Bases: BaseModel
The tools a server exposes, in the order it lists them. Names are unique.
bind ¶
bind(decider: Decider) -> tuple[DecisionTool, ...]
Return one DecisionTool per definition, answered by decider.
mimir.core.checks ¶
Tool-call checks: whether an agent's pending tool call may run under written rules.
The model reads each rule as a passage and the call as fields (tool, arguments.<name>),
and answers a yes/no question about it. A certified yes allows the call, a certified no denies
it, and a deferral or an abstention escalates it to a person.
CheckOutcome
dataclass
¶
The permission for one pending call, and the decision it was read from. result is
None for a tool the check does not apply to.
reason
property
¶
One sentence for the agent or the approver, naming the tool and the confidence.
ToolCallCheck
dataclass
¶
A yes/no decision on pending tool calls, made against rules.
tools names the tools it applies to; None applies it to every tool. Calls to other tools
are allowed without a decision.