Skip to content

Documentation / 2BA.AI

Install, configure,
decide.

The service facts, every client setup and the decision-mode API reference. Plain text, no login — one amber model behind OpenAI- and Anthropic-compatible endpoints.

General information · Client setup · Decision mode · 100% EU hosted · Zero prompt logging

01 / General information

The service,
in plain facts.

One API key from your dashboard authenticates every endpoint. Everything below is current and verifiable.

Base URL · chat completions
https://api.2ba.ai/v1 — OpenAI compatible
Base URL · Anthropic messages
https://api.2ba.ai — Claude Code appends /v1/messages
Default model
amber — text and vision input
Capacity
4,500 requests per rolling 5-hour window
Concurrency
1–2 agents running simultaneously
Price
€20 per month, flat — cancel anytime
Privacy
Zero prompt logging · 100% EU hosted · GDPR native

02 / Client setup

Every tool,
one command.

The installer pairs your browser, creates your API key and configures the tools it finds. Per-tool steps live in the client guides; the essentials:

Getting started
curl -fsSL https://2ba.ai/install.sh | sh does the whole pass.
VS Code
OpenAI-compatible provider pointed at https://api.2ba.ai/v1, model amber.
Cursor
Settings → Models: override the OpenAI base URL, model name amber.
Claude Code
ANTHROPIC_BASE_URL=https://api.2ba.ai and ANTHROPIC_AUTH_TOKEN set to your key.
OpenCode
2ba provider with npm @ai-sdk/openai in ~/.config/opencode/opencode.json.
Windsurf
Run OpenCode or Aider in the integrated terminal.
Continue.dev
Custom model provider with apiBase: https://api.2ba.ai/v1.
Aider
aider --model openai/amber with OPENAI_API_BASE=https://api.2ba.ai/v1.
Chat API basics
Completions and streaming against the same base URL and key.

03 / Chat API

Chat completions,
plain OpenAI.

Point any OpenAI client at the base URL with your 2BA key. amber takes text and images, streams tokens, and speaks the same wire format your tools already expect.

POST /v1/chat/completionsbash
$ curl https://api.2ba.ai/v1/chat/completions \
  -H "Authorization: Bearer your_2ba_api_key" \
  -H "Content-Type: application/json" \
  -d '{ "model": "amber",
        "messages": [{ "role": "user", "content": "Hello!" }] }'
Endpoint
POST /v1/chat/completions — OpenAI compatible. Anthropic clients use https://api.2ba.ai; Claude Code appends /v1/messages.
Model
amber — text and image input via standard content parts.
Streaming
Set "stream": true for server-sent events in the OpenAI format.
Quota
Chat and agents draw on the 4,500-request rolling 5-hour window; decision mode is metered per request on its own cycle.

04 / Decision mode · how it works

One forward pass.
A letter, read precisely.

POST /v1/choices turns the served model into a typed classifier: your state and 2–26 options go in, one chosen option and per-option probabilities come out — no answer sentence is written. The gateway sends the options as lettered choices with thinking disabled and reads the first content token’s letter log-probs; each rotation emits a single constrained letter token, which is what the usage fields count.

01 / Letters, not prose

Letter forcing.

Options are presented as A–Z lettered choices and the answer is constrained to a single first token. Its log-probabilities become your per-option distribution.

02 / Rotations

Bias averaged out.

Each decision runs 1–5 cyclic rotations of the option list (default 3) and aggregates the per-rotation distributions, cancelling position and letter bias. Rotations run concurrently, so latency tracks one call.

03 / Abstain

Off-question answers rejected.

label_mass is the probability the model answered the question asked. Below 0.5 the choice is null and abstained is true — route those to review instead of guessing.

05 / Request

A state, a question,
your options.

Authentication is the same API key as every other 2BA endpoint. The state can be in any language; it is passed through verbatim.

POST /v1/choicesbash
$ curl -X POST https://api.2ba.ai/v1/choices \
  -H "Authorization: Bearer your_2ba_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer was charged twice for the same invoice.",
    "question": "Which queue should handle this request?",
    "options": [
      {"id": "billing", "description": "Invoices, payments, refunds"},
      {"id": "technical", "description": "Bugs, outages, system errors"},
      {"id": "sales", "description": "Pricing and account questions"}
    ]
  }'
state
Required, nonempty. The text to classify; any language, passed through verbatim.
question
Optional. Defaults to a generic instruction; use it to steer what the decision is about.
options
2–26 entries with unique nonempty ids. description is optional and falls back to the id.
rotations
1–5, default 3. Cyclic rotations of the option list; more rotations cost more prefill but do not add latency.
instructions
Optional extra sentence appended to the decision-engine system prompt.

06 / Response

A choice,
with its evidence.

The response carries the aggregated distribution, the abstain signal, and the per-rotation tally that tells you when the model was guessing.

200 OKjson
{
  "model": "amber",
  "choice": "billing",
  "probabilities": {"billing": 0.86, "technical": 0.11, "sales": 0.03},
  "confidence": 0.86,
  "label_mass": 0.97,
  "vote_split": {"billing": 3},
  "abstained": false,
  "rotations": 3,
  "usage": {"prompt_tokens": 612, "completion_tokens": 3}
}
probabilities
Aggregated per-option distribution: the mean of per-rotation normalized letter log-probs, with letters missing from the top-k floored.
label_mass
Mean probability that the first token was a letter at all. Below 0.5 the model did not answer the question asked; choice is null and abstained is true.
vote_split
Per-rotation argmax tally. This is the honest uncertainty signal: probabilities are sharply peaked on wrong answers too, while rotation disagreement identifies borderline states. A 2:1 split deserves a look even when confidence is high.
usage
Token counts summed across rotations; metered as one gateway request against your quota.

07 / Examples

One call,
from any stack.

Any HTTP client works; the endpoint is plain JSON in, plain JSON out.

decide.pypython
# pip install requests
import requests

API_KEY = "your_2ba_api_key"  # from your 2BA dashboard

d = requests.post("https://api.2ba.ai/v1/choices",
    headers={"Authorization": "Bearer " + API_KEY},
    json={
        "state": "Customer was charged twice for the same invoice.",
        "question": "Which queue should handle this request?",
        "options": [
            {"id": "billing", "description": "Invoices, payments, refunds"},
            {"id": "technical", "description": "Bugs, outages, system errors"},
            {"id": "sales", "description": "Pricing and account questions"}],
    }).json()

print(d["choice"], d["vote_split"], d["abstained"])
main.gogo
// Standard library only.
package main

import (
	"encoding/json"
	"fmt"
	"net/http"
	"strings"
)

func main() {
	body := `{"state": "Customer was charged twice for the same invoice.",
		"question": "Which queue should handle this request?",
		"options": [
			{"id": "billing", "description": "Invoices, payments, refunds"},
			{"id": "technical", "description": "Bugs, outages, system errors"},
			{"id": "sales", "description": "Pricing and account questions"}]}`
	req, err := http.NewRequest("POST", "https://api.2ba.ai/v1/choices", strings.NewReader(body))
	if err != nil {
		panic(err)
	}
	req.Header.Set("Authorization", "Bearer your_2ba_api_key")
	req.Header.Set("Content-Type", "application/json")

	res, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer res.Body.Close()

	var d struct {
		Choice    string         `json:"choice"`
		VoteSplit map[string]int `json:"vote_split"`
		Abstained bool           `json:"abstained"`
	}
	if err := json.NewDecoder(res.Body).Decode(&d); err != nil {
		panic(err)
	}
	fmt.Println(d.Choice, d.VoteSplit, d.Abstained)
}
decide.tstypescript
// Node 18+, Deno and Bun: fetch is built in.
const res = await fetch("https://api.2ba.ai/v1/choices", {
  method: "POST",
  headers: {
    Authorization: "Bearer your_2ba_api_key",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    state: "Customer was charged twice for the same invoice.",
    question: "Which queue should handle this request?",
    options: [
      { id: "billing", description: "Invoices, payments, refunds" },
      { id: "technical", description: "Bugs, outages, system errors" },
      { id: "sales", description: "Pricing and account questions" },
    ],
  }),
});

const decision = await res.json();
console.log(decision.choice, decision.vote_split, decision.abstained);

08 / Measured

Classified correctly,
at classifier cost.

Product categorization over a fixed 36-item dataset: one run per configuration, no retries, no prompt tuning per item.

Evaluation / 36 itemsCompleted

01 / Product categorization

86% correct.

Items assigned the correct category on the same fixed dataset. Higher is better.

2BA decision mode amber, rotations 3, zero reasoning tokens

86%correct

Jev external decision API, as configured by its own docs

81%correct

A decision costs ~4 output tokens and ~150 ms in steady state, because no prose is generated. Throughput scales with your plan’s capacity, not with answer length.

Internal evaluation, single run per configuration on a 36-item dataset. Not a general-purpose benchmark; validate on your own data before routing production traffic. Coding benchmarks live on the quality page.

09 / Limits & caveats

Read this part
before you ship.

Probabilities are peaked.

First-token log-probs sit near 1.0 for correct and incorrect answers alike. Treat confidence as a ranking, not a calibration; use vote_split for uncertainty.

No logit_bias.

The upstream does not support logit_bias. The letter contract is enforced by the prompt, which measurements show the template honours when thinking is disabled.

26 options per request.

The option set is capped at 26 (A–Z). For larger sets, compose hierarchically: choose a branch first, then a leaf.

Upstream failures.

If every rotation fails upstream, the request fails with 502. Otherwise failed rotations are dropped and the reduced count is reported in rotations.

10 / Your states are the next test

Put the classifier to work.

Pair your browser, create an API key and send your first decision.

Terminal — secure installerCopy