One key, the standard chat-completions request format, five effort levels. Tokens used through the API are billed to your key, separately from your plan.
Open Avalos AI, go to Settings → Developer · API and choose Create API key. Keys start with avl_live_. API keys come with a paid plan or prepaid credits; there are no free developer keys . Treat the key like a password: keep it on your server or in an environment variable, never in a web page or a public repository.
export AVALOS_API_KEY="avl_live_..."
Send the key in the Authorization header as a bearer token:
Authorization: Bearer avl_live_...
The header X-Avalos-Key: avl_live_... is accepted as well, and it is what the Python SDK sends.
POST https://avalos.ai/v1/chat/completions takes a list of messages (roles system, user, assistant) and an effort. The answer is in choices[0].message.content.
curl https://avalos.ai/v1/chat/completions \
-H "Authorization: Bearer $AVALOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"effort": "star",
"messages": [
{"role": "system", "content": "Answer in one short paragraph."},
{"role": "user", "content": "Explain CRDTs."}
]
}'
import os, requests
r = requests.post(
"https://avalos.ai/v1/chat/completions",
headers={"Authorization": "Bearer " + os.environ["AVALOS_API_KEY"]},
json={
"effort": "star",
"messages": [
{"role": "system", "content": "Answer in one short paragraph."},
{"role": "user", "content": "Explain CRDTs."},
],
},
timeout=600,
)
r.raise_for_status()
print(r.json()["choices"][0]["message"]["content"])
effort | Tier | What it does |
|---|---|---|
quick | Base-T1 | Direct answer from the model's own knowledge, no tools. |
star | Star (T2) | Direct answer on GPU, with tools such as web search. Default. |
galaxy | Galaxy (T3) | Reasons step by step before answering. |
cosmos | Cosmos (T4) | Up to three independent answers; when they can be compared, the one most of them agree on is returned. |
quantum | Quantum (T5) | Up to five independent answers, voted the same way; it stops early when the first answers agree. Slowest by design. |
Higher effort uses more tokens and takes longer. Allow long timeouts (up to 10 minutes) for cosmos and quantum.
GET https://avalos.ai/v1/models returns the model id for clients that require a model field. The tier is always chosen with effort (or tier: T1–T5), never by the model name.
curl https://avalos.ai/v1/models \
-H "Authorization: Bearer $AVALOS_API_KEY"
# {"object": "list", "data": [{"id": "avalos-sulphur-c", "object": "model", ...}]}
import os, requests
r = requests.get("https://avalos.ai/v1/models",
headers={"Authorization": "Bearer " + os.environ["AVALOS_API_KEY"]}, timeout=30)
r.raise_for_status()
for m in r.json()["data"]:
print(m["id"])
avalos_sdk.py is a single file that uses only the Python standard library (Python 3.8 or newer). It covers chat, the Coder Engine and image questions.
Download avalos_sdk.py Open Avalos AIfrom avalos_sdk import Avalos
ai = Avalos() # reads AVALOS_API_KEY, or ~/.avalos/config.json
print(ai.chat("Explain CRDTs in two sentences", effort="star"))
# full response dict, same shape as the REST call above
r = ai.completions([{"role": "user", "content": "Hello"}], effort="galaxy")
# Coder Engine: returns the edited file; it never writes to your disk
fixed = ai.code("app.py", open("app.py").read(), "fix the off-by-one in the loop")
# ask about an image
print(ai.vision("photo.png", "What colour is the car?"))
Errors raise AvalosError with the HTTP status and the server's message. Set AVALOS_BASE to point the SDK at a different host.
| Status | Meaning |
|---|---|
401 | Missing or unknown key. |
402 | The key has no token balance left. |
429 | Too many requests. Wait for the time in Retry-After, then retry. |
5xx | Temporary server problem. Retry with a backoff. |