GPT-5.6 ships in three tiers — Sol (frontier), Terra (mini) and Luna (nano) — and the price gap between them is 20x, so picking the tier matters more than picking the provider. This page puts OpenAI list prices next to what APIMODELS charges, tier by tier, and is explicit about where going direct still wins.
One thing most comparison articles still get wrong: OpenAI cut the GPT-5.6 list prices on 2026-08-21. Sol went from $5/$30 to $4/$20, Terra from $2.5/$15 to $2/$12, and Luna from $1/$6 to $0.20/$1.20 per 1M input/output tokens. Any page still quoting the old ladder — and several major aggregator pages were, weeks later — overstates every discount it claims. The figures below were re-verified on 2026-09-05 against live marketplace listings.
APIMODELS charges $1.324 input / $6.618 output / $0.132 cached input on Sol, which is about 67% below the $4/$20 list. Terra is $0.551 / $3.309 / $0.055, about 72% below its $2/$12 list. Luna is $0.16 / $0.96 / $0.016, 20% below its $0.20/$1.20 list — a much thinner margin than the other two, because OpenAI prices the nano tier close to the floor and there is simply less room underneath it. We would rather say that plainly than quote a headline discount that only applies to one tier.
Reliability, from production rather than a marketing claim: over the last 30 days of external traffic (internal accounts excluded), Sol served 37,198 calls at a 99.7% success rate with a median of 9.1 seconds and p95 of 53.1 seconds. Terra ran 3,963 calls at 99.9%, Luna 4,920 at 99.0%. Failed requests are never billed.
Two contract details worth knowing before you budget. Requests above 272,000 input tokens are billed at 2x input and 1.5x output — that is OpenAI’s own long-context step, and we pass it through rather than hiding it. And the reasoning-depth suffixes (-low / -medium / -high / -xhigh / -max / -ultra) are no longer served: only the three base ids exist, and a suffixed id returns a clear 404 instead of being silently mapped to a different model.
Access is the other half. APIMODELS needs no OpenAI account and no organization verification, works from mainland China, bills a single USD balance across every model, and takes Stripe, PayPal and Alipay. Both /v1/chat/completions and /v1/responses are supported, so existing OpenAI SDK code runs after changing the base URL and key.
Price per 1M tokens — OpenAI list vs APIMODELS, verified 2026-09-05
input output cached in vs list
GPT-5.6 Sol
OpenAI (list) $4.00 $20.00 $0.40
APIMODELS $1.324 $6.618 $0.132 -67%
GPT-5.6 Terra
OpenAI (list) $2.00 $12.00 --
APIMODELS $0.551 $3.309 $0.055 -72%
GPT-5.6 Luna
OpenAI (list) $0.20 $1.20 $0.02
APIMODELS $0.16 $0.96 $0.016 -20%
OpenAI cut this ladder on 2026-08-21 (Sol was $5/$30, Terra $2.5/$15,
Luna $1/$6). Pages still quoting the old numbers overstate their discount.
Long context: above 272,000 input tokens a request bills at
2x input / 1.5x output on all three tiers, on both sides.cURL
# Drop-in OpenAI endpoint — only the base URL and key change
curl https://api.apimodels.app/v1/chat/completions \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"messages": [{"role": "user", "content": "Explain prompt caching in two sentences."}]
}'Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_APIMODELS_KEY",
base_url="https://api.apimodels.app/v1",
)
# gpt-5.6 is an alias for gpt-5.6-sol; use -terra or -luna for the cheaper tiers
r = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Summarise this changelog."}],
)
print(r.choices[0].message.content)
print(r.usage) # cached tokens show up here and bill at the cached rateOpenAI list prices per 1M tokens are $4 / $20 for Sol, $2 / $12 for Terra and $0.20 / $1.20 for Luna, after the 2026-08-21 cut. On APIMODELS the same three tiers are $1.324 / $6.618, $0.551 / $3.309 and $0.16 / $0.96 — about 67%, 72% and 20% below list respectively. Cached input is $0.132, $0.055 and $0.016.
Sol for maximum reasoning and long-horizon agent work. Terra is the everyday default — exactly half of Sol’s output rate and about 40% of its input rate, and it handles chat, coding and tool-calling agents well. Luna is for high-volume, cost-sensitive jobs like bulk classification, extraction and rewriting, at about an eighth of Sol’s input rate. All three share one OpenAI-compatible endpoint, so switching is a one-string change.
Because OpenAI prices the nano tier close to the floor. Luna’s list price is $0.20 / $1.20 per 1M tokens, so there is far less room underneath it than under Sol’s $4 / $20. We would rather state the real per-tier number than advertise a single headline discount that only holds on the top tier. If per-token cost is the only thing that matters for your workload, Luna direct from OpenAI is close enough that the deciding factor should be access and billing, not price.
Sign up at apimodels.app, create an API key, and point the OpenAI SDK at https://api.apimodels.app/v1 — no OpenAI account, no organization verification, reachable from mainland China. Both /v1/chat/completions and /v1/responses are supported. Payment works with Stripe, PayPal and Alipay, and new accounts on consumer email domains get $0.10 in free credit to test with.
Yes. Cached input tokens bill at $0.132 (Sol), $0.055 (Terra) and $0.016 (Luna) per 1M — a tenth of the standard input rate on each tier. Cache hits appear in the usage object of every response, so you can verify what you were charged for rather than taking it on trust.
Three honest cases. First, the Batch API: OpenAI discounts asynchronous batch jobs by 50%, which we do not resell — for large offline workloads that can be cheaper than any gateway. Second, enterprise compliance paperwork (DPAs, SOC 2 chains, data residency commitments) that only a direct OpenAI or Azure relationship provides. Third, priority service tiers and dedicated capacity, which are account-level products rather than per-request options.