Letter forcing.
Options are presented as A–Z lettered choices and the answer is constrained to a single first token. Its log-probabilities become your per-option distribution.
Documentation / 2BA.AI
The service facts, every client setup and the decision-mode API reference. Plain text, no login — one amber model behind OpenAI- and Anthropic-compatible endpoints.
General information · Client setup · Decision mode · 100% EU hosted · Zero prompt logging
01 / General information
One API key from your dashboard authenticates every endpoint. Everything below is current and verifiable.
https://api.2ba.ai/v1 — OpenAI compatiblehttps://api.2ba.ai — Claude Code appends /v1/messagesamber — text and vision input02 / Client setup
The installer pairs your browser, creates your API key and configures the tools it finds. Per-tool steps live in the client guides; the essentials:
curl -fsSL https://2ba.ai/install.sh | sh does the whole pass.https://api.2ba.ai/v1, model amber.amber.ANTHROPIC_BASE_URL=https://api.2ba.ai and ANTHROPIC_AUTH_TOKEN set to your key.2ba provider with npm @ai-sdk/openai in ~/.config/opencode/opencode.json.apiBase: https://api.2ba.ai/v1.aider --model openai/amber with OPENAI_API_BASE=https://api.2ba.ai/v1.03 / Chat API
Point any OpenAI client at the base URL with your 2BA key. amber takes text and images, streams tokens, and speaks the same wire format your tools already expect.
$ curl https://api.2ba.ai/v1/chat/completions \
-H "Authorization: Bearer your_2ba_api_key" \
-H "Content-Type: application/json" \
-d '{ "model": "amber",
"messages": [{ "role": "user", "content": "Hello!" }] }'
POST /v1/chat/completions — OpenAI compatible. Anthropic clients use https://api.2ba.ai; Claude Code appends /v1/messages.amber — text and image input via standard content parts."stream": true for server-sent events in the OpenAI format.04 / Decision mode · how it works
POST /v1/choices turns the served model into a typed classifier: your state and 2–26 options go in, one chosen option and per-option probabilities come out — no answer sentence is written. The gateway sends the options as lettered choices with thinking disabled and reads the first content token’s letter log-probs; each rotation emits a single constrained letter token, which is what the usage fields count.
Options are presented as A–Z lettered choices and the answer is constrained to a single first token. Its log-probabilities become your per-option distribution.
Each decision runs 1–5 cyclic rotations of the option list (default 3) and aggregates the per-rotation distributions, cancelling position and letter bias. Rotations run concurrently, so latency tracks one call.
label_mass is the probability the model answered the question asked. Below 0.5 the choice is null and abstained is true — route those to review instead of guessing.
05 / Request
Authentication is the same API key as every other 2BA endpoint. The state can be in any language; it is passed through verbatim.
$ curl -X POST https://api.2ba.ai/v1/choices \
-H "Authorization: Bearer your_2ba_api_key" \
-H "Content-Type: application/json" \
-d '{
"state": "Customer was charged twice for the same invoice.",
"question": "Which queue should handle this request?",
"options": [
{"id": "billing", "description": "Invoices, payments, refunds"},
{"id": "technical", "description": "Bugs, outages, system errors"},
{"id": "sales", "description": "Pricing and account questions"}
]
}'
statequestionoptionsids. description is optional and falls back to the id.rotationsinstructions06 / Response
The response carries the aggregated distribution, the abstain signal, and the per-rotation tally that tells you when the model was guessing.
{
"model": "amber",
"choice": "billing",
"probabilities": {"billing": 0.86, "technical": 0.11, "sales": 0.03},
"confidence": 0.86,
"label_mass": 0.97,
"vote_split": {"billing": 3},
"abstained": false,
"rotations": 3,
"usage": {"prompt_tokens": 612, "completion_tokens": 3}
}
probabilitieslabel_masschoice is null and abstained is true.vote_splitusage07 / Examples
Any HTTP client works; the endpoint is plain JSON in, plain JSON out.
# pip install requests
import requests
API_KEY = "your_2ba_api_key" # from your 2BA dashboard
d = requests.post("https://api.2ba.ai/v1/choices",
headers={"Authorization": "Bearer " + API_KEY},
json={
"state": "Customer was charged twice for the same invoice.",
"question": "Which queue should handle this request?",
"options": [
{"id": "billing", "description": "Invoices, payments, refunds"},
{"id": "technical", "description": "Bugs, outages, system errors"},
{"id": "sales", "description": "Pricing and account questions"}],
}).json()
print(d["choice"], d["vote_split"], d["abstained"])
// Standard library only.
package main
import (
"encoding/json"
"fmt"
"net/http"
"strings"
)
func main() {
body := `{"state": "Customer was charged twice for the same invoice.",
"question": "Which queue should handle this request?",
"options": [
{"id": "billing", "description": "Invoices, payments, refunds"},
{"id": "technical", "description": "Bugs, outages, system errors"},
{"id": "sales", "description": "Pricing and account questions"}]}`
req, err := http.NewRequest("POST", "https://api.2ba.ai/v1/choices", strings.NewReader(body))
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer your_2ba_api_key")
req.Header.Set("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
var d struct {
Choice string `json:"choice"`
VoteSplit map[string]int `json:"vote_split"`
Abstained bool `json:"abstained"`
}
if err := json.NewDecoder(res.Body).Decode(&d); err != nil {
panic(err)
}
fmt.Println(d.Choice, d.VoteSplit, d.Abstained)
}
// Node 18+, Deno and Bun: fetch is built in.
const res = await fetch("https://api.2ba.ai/v1/choices", {
method: "POST",
headers: {
Authorization: "Bearer your_2ba_api_key",
"Content-Type": "application/json",
},
body: JSON.stringify({
state: "Customer was charged twice for the same invoice.",
question: "Which queue should handle this request?",
options: [
{ id: "billing", description: "Invoices, payments, refunds" },
{ id: "technical", description: "Bugs, outages, system errors" },
{ id: "sales", description: "Pricing and account questions" },
],
}),
});
const decision = await res.json();
console.log(decision.choice, decision.vote_split, decision.abstained);
08 / Measured
Product categorization over a fixed 36-item dataset: one run per configuration, no retries, no prompt tuning per item.
01 / Product categorization
Items assigned the correct category on the same fixed dataset. Higher is better.
86%correct
81%correct
A decision costs ~4 output tokens and ~150 ms in steady state, because no prose is generated. Throughput scales with your plan’s capacity, not with answer length.
Internal evaluation, single run per configuration on a 36-item dataset. Not a general-purpose benchmark; validate on your own data before routing production traffic. Coding benchmarks live on the quality page.
09 / Limits & caveats
First-token log-probs sit near 1.0 for correct and incorrect answers alike. Treat confidence as a ranking, not a calibration; use vote_split for uncertainty.
The upstream does not support logit_bias. The letter contract is enforced by the prompt, which measurements show the template honours when thinking is disabled.
The option set is capped at 26 (A–Z). For larger sets, compose hierarchically: choose a branch first, then a leaf.
If every rotation fails upstream, the request fails with 502. Otherwise failed rotations are dropped and the reduced count is reported in rotations.
10 / Your states are the next test
Pair your browser, create an API key and send your first decision.