Every call has a request ID and one final billed amount. LLM requests return the ID in the x-apimodels-request-id response header, and non-streaming requests also carry x-apimodels-cost; for image, video and audio tasks the ID is the task_id returned at creation. GET /v1/records/{id} with any of these IDs returns exactly what that request took from your balance. Built for customers who resell us or need to reconcile each of their own charges against ours.
Every successful response from the LLM endpoints (/v1/chat/completions, /v1/messages, the native Gemini /v1beta, /v1/responses and /v1/completions) carries the x-apimodels-request-id header; its value is the task_id of that request. Non-streaming responses also carry x-apimodels-cost: the amount actually deducted, in USD, after every discount on your account, written only once the deduction has completed. Streaming responses carry the ID only, because the amount is known when the last chunk settles; look the record up with the ID.
HTTP/2 200
content-type: application/json
x-apimodels-request-id: cmumilhfd0013mg012q3qobr9
x-apimodels-cost: 0.0001
access-control-expose-headers: x-apimodels-request-id, x-apimodels-costOne exception: when a non-streaming LLM request runs past about 95 seconds, the gateway sends the response headers early to keep the connection alive, so the two headers cannot carry values; the same two values are then added to the response body under an apimodels field. The OpenAI and Anthropic SDKs ignore the unknown field.
{
"id": "chatcmpl-…",
"choices": [ … ],
"usage": { … },
"apimodels": { "request_id": "cmumilhfd0013mg012q3qobr9", "cost": 0.0001 }
}For image, video and audio tasks the ID is the data.taskId returned at creation. The completion callback (callback_url) also includes credits in data, equal to what the record lookup returns.
{
"code": 200,
"msg": "success",
"data": {
"taskId": "cmum4dq3z00dwlm01wugboo5z",
"state": "completed",
"credits": 0.338,
"resultUrls": ["https://r2.apimodels.app/videos/cmum4dq3z00dwlm01wugboo5z.mp4"]
}
}When calling from a browser, both headers are exposed through Access-Control-Expose-Headers, so response.headers.get() can read them.
GET /api/v1/records/{task_id}
Authorization: Bearer <API_KEY>cURL
curl https://api.apimodels.app/v1/records/cmumilhfd0013mg012q3qobr9 \
-H "Authorization: Bearer $API_KEY"Python
import requests, os
# 1) make the call and keep the request id from the response headers
r = requests.post(
"https://api.apimodels.app/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
json={"model": "space-bunny-alpha", "messages": [{"role": "user", "content": "Hi"}]},
)
request_id = r.headers["x-apimodels-request-id"]
print("charged now:", r.headers.get("x-apimodels-cost")) # non-streaming only
# 2) any time later: the authoritative billed amount for that one request
rec = requests.get(
f"https://api.apimodels.app/v1/records/{request_id}",
headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
).json()["data"]
print(rec["settled"], rec["credits"], rec["currency"])Node.js
const headers = { Authorization: "Bearer " + process.env.API_KEY };
// 1) make the call and keep the request id from the response headers
const r = await fetch("https://api.apimodels.app/v1/chat/completions", {
method: "POST",
headers: { ...headers, "Content-Type": "application/json" },
body: JSON.stringify({ model: "space-bunny-alpha", messages: [{ role: "user", content: "Hi" }] }),
});
const requestId = r.headers.get("x-apimodels-request-id");
console.log("charged now:", r.headers.get("x-apimodels-cost")); // non-streaming only
// 2) any time later: the authoritative billed amount for that one request
const rec = await fetch("https://api.apimodels.app/v1/records/" + requestId, { headers });
const { data } = await rec.json();
console.log(data.settled, data.credits, data.currency);{
"code": 200,
"msg": "success",
"data": {
"task_id": "cmumilhfd0013mg012q3qobr9",
"model": "space-bunny-alpha",
"type": "LANGUAGE_MODEL",
"state": "completed",
"settled": true,
"credits": 0.0001,
"currency": "USD",
"usage": {
"input_tokens": 165,
"output_tokens": 6,
"cached_input_tokens": 149,
"cache_creation_tokens": 0
},
"created_at": 1790676571513,
"completed_at": 1790676572799
}
}| Field | Required | Type | Description |
|---|---|---|---|
| data.task_id | — | string | The request ID, identical to the header value or the taskId returned at creation |
| data.model | — | string | Public model name, e.g. space-bunny-alpha |
| data.type | — | string | LANGUAGE_MODEL / TEXT_TO_IMAGE / IMAGE_TO_VIDEO / … as on the call record |
| data.state | — | string | pending (still running) / completed / failed |
| data.settled | — | boolean | true once the amount is final; while false, credits is null — poll again later |
| data.credits | — | number | null | USD actually taken from the balance, after account discounts and per-model prices; 0 for failed requests |
| data.currency | — | string | Always "USD" |
| data.usage | — | object | LLM only: input_tokens, output_tokens, cached_input_tokens, cache_creation_tokens; reasoning tokens are inside output_tokens |
| data.created_at | — | number | Creation time, Unix milliseconds |
| data.completed_at | — | number | null | Completion time, Unix milliseconds; null until done |
Read x-apimodels-request-id from the response headers (LLM) or keep the task_id returned at creation (image, video, audio), then call GET https://api.apimodels.app/v1/records/{id} with your API key. data.credits is the USD amount actually deducted after all discounts; failed requests are 0. Non-streaming LLM responses also carry the amount directly in x-apimodels-cost.
The request has not finished settling. Streaming LLM calls settle on the last chunk, and some video models settle on the actual output length. Poll the record again after completion; once settled is true the amount is final and will not change.
No. Streaming responses carry only x-apimodels-request-id, because the amount is known when the stream ends. Query GET /v1/records/{id} after the final chunk. Non-streaming responses carry both headers; a non-streaming LLM request that runs past about 95 seconds gets the same two values in an apimodels field in the response body instead, because the headers were already sent to keep the connection alive.
No. A request ID that belongs to a different account returns 404, exactly like a missing one; the API never confirms whether such an ID exists. The lookup itself is free and is not counted as a call.