Decisions API
Minimum-discount routing: Prefix the path with a min{N} segment (e.g. /min30/v1/decisions) to require marketplace seller offers to meet a minimum estimated buyer discount before routing. See Minimum-Discount Routing.
Ask ordered classification and scoring questions about shared evidence, and get calibrated probabilities back. The endpoint speaks OpenAI's public Decisions API — point the OpenAI SDK at Surplus and call client.decisions.create — and keeps accepting the JEV format (state plus a map of questions) that earlier Surplus integrations use. Both formats route through the marketplace: API-key auth, x402 payment challenges, seller selection, usage logging and accounting work as on every other endpoint.
Basic request
questions is an array, answered in order. Three question types:
predicate— a yes/no question; the answer is aprobability.choice— pick one ofchoices; values are typed (the string"true"and the booleantrueare different choices).score— rate against orderedlevels; the answer is a probability-weighted level index.
name is optional; unnamed questions are answered as name: null. input is a string, or user messages whose content is input_text and input_image parts.
curl $SURPLUS_BASE_URL/decisions \ -H "Authorization: Bearer $SURPLUS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-luna", "input": "Card charged twice for the June invoice; refund requested 38 hours ago.", "questions": [ { "type": "predicate", "name": "urgent", "instructions": "Should on-call be paged within the hour?" }, { "type": "choice", "name": "queue", "instructions": "Which team owns it?", "choices": [{ "value": "billing" }, { "value": "support" }] }, { "type": "score", "name": "anger", "instructions": "How upset is the customer?", "levels": [{ "label": "calm" }, { "label": "annoyed" }, { "label": "furious" }] } ] }'
from openai import OpenAI client = OpenAI(base_url="$SURPLUS_BASE_URL", api_key=SURPLUS_API_KEY) decision = client.decisions.create( model="gpt-6-luna", input="Card charged twice for the June invoice; refund requested 38 hours ago.", questions=[ {"type": "predicate", "name": "urgent", "instructions": "Should on-call be paged within the hour?"}, {"type": "choice", "name": "queue", "instructions": "Which team owns it?", "choices": [{"value": "billing"}, {"value": "support"}]}, ], ) print(decision.answers[0].probability)
import OpenAI from 'openai'; const client = new OpenAI({ baseURL: '$SURPLUS_BASE_URL', apiKey: process.env.SURPLUS_API_KEY }); const decision = await client.decisions.create({ model: 'gpt-6-luna', input: 'Card charged twice for the June invoice; refund requested 38 hours ago.', questions: [{ type: 'predicate', name: 'urgent', instructions: 'Should on-call be paged within the hour?' }], }); console.log(decision.answers[0]);
Response
{model, answers, usage} — one answer per question, in question order, each carrying the question's name. A question the model declines comes back as {"type": "refusal", "name": …}; a refusal is still a served (and billed) answer.
usage is the provider's own count. Decisions are billed on input tokens only; output is free.
{
"model": "gpt-6-luna",
"answers": [
{ "type": "predicate", "name": "urgent", "probability": 0.82 },
{ "type": "choice", "name": "queue", "choice": "billing",
"probabilities": [{ "value": "billing", "probability": 0.91 }, { "value": "support", "probability": 0.09 }],
"confidence": 0.84 },
{ "type": "score", "name": "anger", "score": 1.2,
"probabilities": [
{ "value": 0, "label": "calm", "probability": 0.1 },
{ "value": 1, "label": "annoyed", "probability": 0.6 },
{ "value": 2, "label": "furious", "probability": 0.3 }
],
"confidence": 0.7 }
],
"usage": {
"input_tokens": 412,
"input_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 },
"output_tokens": 0,
"output_tokens_details": { "reasoning_tokens": 0 },
"total_tokens": 412
}
}Images
Add input_image parts with a base64 data URL (PNG, JPEG, WebP or GIF) and an optional detail (low, high, auto, original). Hosted image URLs and file ids are not supported. At most 128 images per request.
{
"model": "gpt-6-luna",
"input": [{ "role": "user", "content": [
{ "type": "input_text", "text": "Is the package in this photo damaged?" },
{ "type": "input_image", "image_url": "data:image/png;base64,iVBORw0KGgo…", "detail": "low" }
] }],
"questions": [{ "type": "predicate", "instructions": "Is there visible damage?" }]
}The JEV format
The earlier Surplus format still works: state (a string, object or array) and questions as a map keyed by your own ids, with noul / choice / score questions and criteria. Answers come back as a map keyed by those ids. On models whose sellers speak this format (such as jev-1.13) the request and the answer are passed through unchanged.
When a JEV request goes to a model whose sellers speak the OpenAI format (e.g. gpt-6-luna), Surplus translates it only if nothing is lost: noul without criteria, choice criteria that map values to text descriptions, and score criteria that are a list of labels. A state object is sent as JSON text. Anything else is refused with 400 decisions_mapping_lossy and the reason — send the OpenAI format instead. Translated responses carry X-SI-Adapted-Response: answers; a refusal appears as {"type": "refusal"}, a score answer carries its legend but not the per-level probabilities, and usage is the OpenAI seller's.
Routing, privacy and limits
- Provider pins.
provider,providers_blocked,provider_orderandsi_routework as on chat completions. From the Python SDK pass them throughextra_body; from the Node SDK put them in the request object. - End-user ids. On OpenAI-format sellers,
safety_identifier(and a translated JEV request'suser) reach the seller only as a stable one-way digest. Translated fields are listed inX-SI-Adapted-Params. - Access. In some environments
/v1/decisionsis limited to pilot accounts (403 buyer_not_in_allowlist). - Limits. Up to 64 questions, 64 choices or levels per question, names up to 128 characters, 128 images, and a 10 MB body.
- Idempotency.
Idempotency-Keyis not used on this endpoint (as on the other inference endpoints). - Errors. A malformed request is a
400whoseerror.paramnames the field. A body that mixes the two formats (statewith an array of questions, orinput/safety_identifierwith a map) is400 decisions_dialect_mismatch. An OpenAI-format request for a model whose sellers accept only the JEV format is400 decisions_format_unsupported.
Pricing
gpt-6-luna decisions list at $0.10 per 1M input tokens; output tokens are free. The provider adds a fixed overhead per request and per question, so even a short request bills a few hundred input tokens.