Decisions
Decisions let a deployed agent answer a yes/no gate, pick one option from a set, or place something on a rubric, and receive a calibrated probability for each answer instead of generated text. Agent steps call the typed ctx.sapiom.decisions.evaluate capability; Sapiom supplies the authenticated cloud connection when the deployed agent runs.
The capability does not generate content, run tools, or parse a reply. You define the possible answers up front; the result tells you how likely each one is.
Choose the question type
Section titled “Choose the question type”| What the agent needs to know | Question type | Answer |
|---|---|---|
| Whether a condition holds | noul | noul: probability of “yes”, 0–1 |
| Which one of a named set applies | choice | choice, probabilities, confidence |
| How far along an ordered rubric something sits | score | score, legend, probabilities, confidence |
noul is the yes/no question type: its answer is a single probability that the answer is yes, with no separate confidence. To check several labels that can apply at once, ask one noul per label rather than a choice.
If the answer is free text, a schema-shaped object, or the result of a tool loop, use llm.run or models.run instead. See Choose a call surface.
Evaluate questions over a state
Section titled “Evaluate questions over a state”Every call takes one state and a map of named questions. All questions are evaluated against the same state in parallel, so ask every independent question in one call:
const result = await ctx.sapiom.decisions.evaluate({ state: { subject: input.subject, message: input.body, customerTier: input.tier, }, questions: { urgent: { type: "noul", instructions: "Does `message` need a same-day response?", }, team: { type: "choice", instructions: "Which team should handle `message`?", criteria: { shipping: "Delivery delays, lost or damaged parcels", billing: "Charges, refunds, and invoices", account: "Login, password, and profile changes", other: null, }, }, frustration: { type: "score", instructions: "How frustrated is the customer?", criteria: [ "Neutral or friendly tone", "Mildly annoyed, still cooperative", "Angry, threatening to cancel or escalate", ], }, },});
if (result.answers.urgent.noul > 0.8) { await escalate(result.answers.team.choice);}
return { team: result.answers.team.choice, teamConfidence: result.answers.team.confidence, frustration: result.answers.frustration.score, inputTokens: result.usage.inputTokens,};answers is typed by the questions you passed: result.answers.team.choice is one of "shipping" | "billing" | "account" | "other", and a noul question never has a probabilities field. Question names are for your code only; the model never sees them.
state and instructions accept a string, an object, or an array. Prefer named fields when the context has several parts, and refer to them by name in instructions.
Question rules
Section titled “Question rules”noultakes optionalcriteria: { true?: string; false?: string }describing what each answer means.choicerequirescriteriawith 2–32 options. Each value is a description ornullwhen the name is self-explanatory. Include a no-match option such asother: nullwhen nothing may fit.scorerequirescriteriaas 2–10 level descriptions, ordered lowest first. Each level should describe a concrete situation, not a number.- A call carries 1–64 questions, each name at most 128 characters, and the serialized request must be at most 256,000 bytes.
The request is validated before any metered work runs. An invalid question fails the whole call with a 400 and a message naming the question, for example question 'team' choice criteria must have at least two options.
Read the answers
Section titled “Read the answers”| Field | Meaning |
|---|---|
noul | Probability of “yes”. A value near 0.5 means undecided, not “moderately”. |
choice | The highest-probability option. |
probabilities | Every option or level → its probability; the values sum to 1. |
confidence | 0–1, how concentrated the distribution is. It is not a measure of correctness. |
score | Probability-weighted average across the levels, from 0 to levels - 1. See below. |
legend | Level index (as a string) → the level description you passed. |
Because score is an average, it can land on a level the model gave little weight. An answer split evenly between level 0 and level 2 has score: 1, even though level 1 may have almost no probability. Read probabilities and confidence before acting on score, and do not round it to a level.
Look up probabilities by option name or level index. Keys are not guaranteed to come back in the order you listed the criteria.
Set thresholds in code and keep them explicit. A choice with confidence: 0.4 is a signal to route to review, not a reason to trust the top option.
HTTP request and response
Section titled “HTTP request and response”Deployed agents should call the typed client, but the same capability is available to any authenticated caller at POST /v1/capabilities/decisions.evaluate:
curl -sS https://api.sapiom.ai/v1/capabilities/decisions.evaluate \ -H "x-api-key: $SAPIOM_API_KEY" -H 'content-type: application/json' \ -d '{ "state": { "message": "My parcel was due Monday and still has not arrived." }, "questions": { "urgent": { "type": "noul", "instructions": "Does `message` need a same-day response?" }, "team": { "type": "choice", "instructions": "Which team should handle `message`?", "criteria": { "shipping": "Delivery issues", "billing": "Charges and refunds", "other": null } } } }'{ "answers": { "urgent": { "type": "noul", "noul": 0.36 }, "team": { "type": "choice", "choice": "shipping", "probabilities": { "billing": 0.01, "shipping": 0.97, "other": 0.02 }, "confidence": 0.96 } }, "usage": { "inputTokens": 409, "outputTokens": 70 }, "cost": { "estimateUsd": 0.0000357, "currency": "USD", "reference": "txn_01J…", "isEstimate": true, "source": "quote" }}A score answer looks like { "type": "score", "score": 1.26, "legend": { "0": "…", "1": "…", "2": "…" }, "probabilities": { "0": 0, "1": 0.74, "2": 0.26 }, "confidence": 0.61 }.
The response contains answers, token usage, and, when available, cost. It carries no model or provider identity. cost is a pre-call estimate sized from the request and is usually higher than the settled charge. It is omitted when no estimate was available, never reported as zero. When reference is present, use it to look up the settled transaction.
Test the behavior locally
Section titled “Test the behavior locally”Local Run replaces decisions.evaluate with a deterministic, shape-correct answer for every question and applies the same request validation as the cloud. Override the path to exercise a specific branch:
{ "version": 1, "steps": { "triage-ticket": { "decisions.evaluate": { "answers": { "urgent": { "type": "noul", "noul": 0.95 }, "team": { "type": "choice", "choice": "billing", "probabilities": { "shipping": 0.05, "billing": 0.9, "account": 0.03, "other": 0.02 }, "confidence": 0.87 }, "frustration": { "type": "score", "score": 1.8, "legend": { "0": "Neutral or friendly tone", "1": "Mildly annoyed, still cooperative", "2": "Angry, threatening to cancel or escalate" }, "probabilities": { "0": 0.05, "1": 0.1, "2": 0.85 }, "confidence": 0.8 } }, "usage": { "inputTokens": 1, "outputTokens": 0 } } } }}Test the undecided branch (noul near 0.5, low confidence) with a separate fixture. After each run, assert the terminal output and require both unusedStubs and stubWarnings to be empty. A passing Local Run proves the step’s thresholds against fixture data, not the calibration of a real answer.
Manage usage and result quality
Section titled “Manage usage and result quality”Calls are metered on input tokens: state is counted once per call, and each question’s instructions and criteria add to the total. A failed call is not retried automatically, and a retry is a new metered call. Group every independent question over the same state into one call, and make a second call only when a question depends on an earlier answer. Keep state to the fields the questions need.
Ambiguous criteria produce flat distributions. Describe each option or level as a concrete situation, add a no-match option to every choice, and read confidence before acting on choice or score.
Use the signed-in capability catalog for current availability, limits, and pricing rather than copying a per-token rate into the agent.
© 2026 Sapiom, Inc.