fal.ai is an inference platform: next to a gallery of more than 1,000 hosted image, video, audio and 3D models, it sells serverless GPU deployment, dedicated GPUs and training endpoints. APIMODELS does a narrower job. It calls well-known commercial models such as GPT Image 2, Nano Banana Pro, Seedance, Kling, Veo, Claude and GPT through one key, and prices them per image, per second or per million tokens, with nothing to deploy. If your workload is “call the best current model and get a file back”, the two overlap heavily. If it is “host and tune my own model”, that is fal’s job and not ours.
On the models both platforms sell, the gap is widest on Google and OpenAI models. GPT Image 2 at 1024×1024 and medium quality is $0.025 here against $0.053 on fal, which defaults to high quality at $0.211. An 8-second 1080p Veo 3.1 clip with audio is $1.00 here against $3.20 at fal’s $0.40 per second. Nano Banana Pro is $0.08 against $0.15 per image, and Seedance 2.0 at 720p is $0.197 per second against $0.303. Every fal price on this page was read from fal’s own model pages on 6 October 2026.
We judge a gateway on stability first and price second, so here is what you can check. Many models here run on more than one upstream route and fail over automatically when one errors or is overloaded. A failed generation is never charged. Image and video jobs accept a callback_url, retried three times and then in the background for up to 30 minutes, and GET /v1/records/{id} returns the exact amount any single call cost. There is no organization verification, the site is reachable from mainland China, and you can pay by card, PayPal, Alipay or USDT.
Everything in the fal column comes from fal’s own pricing pages, docs and terms, checked on 2026-10-06. The sources are listed at the end of this page.
| APIMODELS | fal.ai | |
|---|---|---|
| What it is | Multi-model API gateway for hosted commercial models | Inference platform: hosted models plus serverless GPUs, dedicated compute and training |
| Model coverage | 149 models: 37 image, 40 video, 49 LLM, 21 audio, 2 embedding (October 2026) | 1,000+ image, video, audio and 3D models; LLMs through an OpenRouter-backed router endpoint |
| Billing unit | Per image, per second of video, per 1M tokens, per 1K characters of speech | Per output (image, second, token); GPU-seconds for models without a fixed price |
| Failed requests | Never charged | HTTP 5xx not charged; HTTP 422 errors, including content-policy rejections, may be charged |
| API shape | REST; OpenAI-compatible chat and image endpoints, native Anthropic Messages | fal queue API and SDKs (Python, JS, Swift, Java, Kotlin, Dart); OpenAI-compatible only for the LLM router |
| Async jobs | Create and poll, or callback_url (retried for up to 30 minutes) | Queue and poll, or webhook_url (ED25519-signed) |
| Payment | Card (Stripe), PayPal, Alipay, USDT; top-ups from $10 | Card or ACH, in US dollars |
| Invoices | Automatic for card payments, with your company tax ID if added | Invoice-based billing for higher-volume customers |
| Verification | None: email sign-up, no organization or ID check | None stated for model APIs; serverless deployments need per-account approval |
| Free credit | $0.10 for consumer-email sign-ups, usable on the API | Free credits work only in Sandbox and Playground, not through the API |
| Custom models and training | Not offered | fal Serverless, dedicated GPUs (H100 from $2.49/hr), LoRA training ($2 per FLUX LoRA run) |
USD, checked on 2026-10-06. fal prices come from each model’s page on fal.ai; APIMODELS prices are what our billing charges. Seedance on APIMODELS is billed by tokens, so its per-second figure is the 16:9 equivalent. Veo clips here are always 8 seconds with an audio track, so the fal column multiplies its with-audio per-second rate by eight.
| Model and unit | APIMODELS | fal.ai | Lower price |
|---|---|---|---|
| GPT Image 2, 1024×1024, medium quality, per image | $0.025 | $0.053 | APIMODELS |
| GPT Image 2, 4K (3840×2160), medium quality, per image | $0.05 | $0.101 | APIMODELS |
| Nano Banana Pro (Gemini 3 Pro Image), 1K–2K / 4K, per image | $0.08 / $0.13 | $0.15 / $0.30 | APIMODELS |
| Seedance 2.0, 720p / 1080p, per second | $0.197 / $0.492 | $0.303 / $0.682 | APIMODELS |
| Veo 3.1, 8 s, 1080p, with audio, per clip | $1.00 | $3.20 ($0.40/s) | APIMODELS |
| Veo 3.1 Fast, 8 s, 1080p, with audio, per clip | $0.07 | $1.20 ($0.15/s) | APIMODELS |
| Grok Imagine Video 1.5, 480p / 720p / 1080p, per second | $0.045 / $0.0529 / $0.0882 | $0.08 / $0.14 / $0.25 | APIMODELS |
Choose fal if you need to run your own model. fal Serverless deploys your code with fal deploy and bills runner time per second, fal rents dedicated GPUs (H100 from $2.49 an hour on its pricing page), and it has training endpoints such as FLUX LoRA fast training at $2 per run. APIMODELS cannot host your weights and offers no training or fine-tuning.
Choose fal for 4K Veo, the long tail and real-time work. fal serves Veo 3.1 at 4K, while our Veo tiers stop at 1080p. It lists more than 1,000 models, including 3D and many open-weight image models, supports streaming and WebSocket inference, and its enterprise plan adds private model hosting, SOC 2 and SSO.
Replicate suits teams that want to package their own model with Cog and run it next to 50,000+ community models; its official commercial models are priced per output.
OpenRouter suits LLM-heavy workloads that want each model routed across several providers; it also runs an async video API with 29 models.
Kie.ai suits buyers who need its Suno V6 and Runway endpoints.
The model makers’ own APIs (Google, OpenAI, ByteDance) suit teams that need the compliance paperwork and account-level products that only a direct contract provides. Each of the first three has its own comparison page, linked at the bottom.
Moving a fal integration takes three changes: the endpoint, the model id and the result field. fal calls a model by its endpoint id through fal_client or https://queue.fal.run/<endpoint-id>. Here every video model sits behind POST https://api.apimodels.app/v1/video/generations and every image model behind /v1/images/generations, and the model is a field in the JSON body: fal-ai/veo3.1 becomes veo-3.1, fal-ai/nano-banana-pro becomes gemini-3-pro-image.
Prompt, duration, aspect ratio and image URLs carry over, but field names differ between model families, so check each model’s docs page. Replace webhook_url with callback_url, or poll GET /v1/video/generations?task_id=… until state is completed. The file URLs come back in resultUrls and stay downloadable for 30 days.
LLM code only needs a new base URL: point the OpenAI SDK at https://api.apimodels.app/v1, or the Anthropic SDK at https://api.apimodels.app, and use ids such as claude-sonnet-5-5 or gpt-6.1-sol. GET https://api.apimodels.app/v1/models lists every id without a key.
| Fact | Source |
|---|---|
| GPT Image 2 per-image table | fal.ai/models/openai/gpt-image-2 |
| Nano Banana Pro | fal.ai/models/fal-ai/nano-banana-pro |
| Seedance 2.0 | fal.ai/models/bytedance/seedance-2.0/text-to-video |
| Veo 3.1 and Veo 3.1 Fast | fal.ai/models/fal-ai/veo3.1 and /veo3.1/fast |
| Grok Imagine Video 1.5 | fal.ai/models/xai/grok-imagine-video/v1.5/text-to-video |
| Billing, failed requests, free credits | fal.ai/docs/documentation/model-apis/pricing, /faq and /sandbox |
| Payment and credit expiry | fal.ai/legal/terms-of-service |
| Serverless, GPUs, LoRA training | fal.ai/docs/documentation/serverless/pricing, fal.ai/pricing, fal.ai/models/fal-ai/flux-lora-fast-training |
| APIMODELS prices | apimodels.app/pricing and each model page |
cURL
# Before (fal): POST https://queue.fal.run/fal-ai/veo3.1 with "Authorization: Key $FAL_KEY"
# After (APIMODELS): one endpoint for every video model; the model is a body field
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1",
"prompt": "Slow dolly shot of a lighthouse at dawn, waves breaking below",
"aspect_ratio": "16:9",
"callback_url": "https://your-app.example.com/hooks/video"
}'
# → {"data": {"taskId": "..."}} 8 s, 1080p, with audio, $1.00
# no webhook? poll: curl "https://api.apimodels.app/v1/video/generations?task_id=TASK_ID" -H "Authorization: Bearer $APIMODELS_API_KEY"Python
import requests, time
BASE = "https://api.apimodels.app/v1"
H = {"Authorization": "Bearer YOUR_APIMODELS_KEY"}
# Before (fal): fal_client.subscribe("fal-ai/nano-banana-pro", arguments={"prompt": ...})
# After: same model, called as gemini-3-pro-image ($0.08 at 1K/2K, $0.13 at 4K)
task = requests.post(f"{BASE}/images/generations", headers=H, json={
"model": "gemini-3-pro-image",
"prompt": "Studio photo of a ceramic teapot on linen, soft window light",
"aspect_ratio": "16:9",
"resolution": "2k",
}).json()["data"]["taskId"]
while True:
r = requests.get(f"{BASE}/images/generations", headers=H,
params={"task_id": task}).json()["data"]
if r["state"] in ("completed", "failed"):
print(r.get("resultUrls"), r.get("failMsg"))
break
time.sleep(3)Yes, on every model in the price table on this page. Checked on 6 October 2026: GPT Image 2 at 1024×1024 and medium quality is $0.025 here and $0.053 on fal; Nano Banana Pro is $0.08 against $0.15 per image; Seedance 2.0 at 720p is $0.197 against $0.303 per second; an 8-second 1080p Veo 3.1 clip with audio is $1.00 against $3.20; and Grok Imagine Video 1.5 at 720p is $0.0529 against $0.14 per second. GPT Image 2 at 4K and medium quality is $0.05 here and $0.101 on fal. The size of the gap varies by model, so compare the ones you actually use.
We do not publish an uptime figure, so here is what can be verified instead. Many models on APIMODELS run on more than one upstream route, and when one route errors or is overloaded the request fails over to another automatically. A generation that fails is never charged. fal’s rule is different: it does not charge HTTP 5xx failures, but its docs say HTTP 422 errors, including content-policy rejections, may still be billed. Async jobs here accept a callback_url; if your endpoint does not answer, delivery is retried three times with backoff and then in the background for up to 30 minutes. Every call’s final charge can be looked up with GET /v1/records/{id}. One gap to know about: our callbacks are not signed yet, while fal signs its webhooks, so put a secret token in your callback URL and check it on receipt.
For most integrations it is an afternoon of renaming, not a rewrite. Image and video calls move from fal’s queue (fal_client or https://queue.fal.run/<endpoint-id>) to two REST endpoints, /v1/images/generations and /v1/video/generations, with the model named in the JSON body: gemini-3-pro-image instead of fal-ai/nano-banana-pro, veo-3.1 instead of fal-ai/veo3.1. Prompt, duration, aspect ratio and image URLs carry over, though field names differ by model family, so check each model’s docs page. Replace webhook_url with callback_url, or poll with the task_id until state is completed, and read the files from resultUrls. LLM code only needs the OpenAI or Anthropic SDK pointed at our base URL. The one thing that does not migrate is anything you deployed yourself on fal Serverless.
APIMODELS takes cards through Stripe, PayPal, Alipay and USDT on TRC-20 or BEP-20, with top-ups from $10; the $50, $100, $500 and $1,000 packages add a 2–5% bonus. Card payments are invoiced automatically, and if you tick the business option at checkout, your company name and tax number appear on every later invoice, automatic top-ups included. New accounts on consumer email domains get $0.10 to test with on the API. fal’s terms list payment by card or ACH in US dollars, it offers invoice-based billing to higher-volume customers, and purchased credits expire 365 days after purchase. fal’s free credits can be spent only in its Sandbox and Playground, not through the API.
Yes. apimodels.app and api.apimodels.app are reachable from mainland China without a proxy, the site and docs are in Chinese as well as English, and you can pay with Alipay, which bills the CNY equivalent at the order-time exchange rate. The same key works for Claude, GPT and Gemini text models and for Veo, Kling, Seedance and GPT Image, so a team in China does not need a separate overseas account with each model maker. fal’s published restrictions cover sanctioned regions, and it bills by card or ACH in US dollars, so for teams in China the practical difference is mostly payment and network access rather than the models themselves.
No. APIMODELS serves hosted models only, so you cannot upload weights, train a LoRA or deploy a container. If that is part of your workflow, keep fal (Serverless, dedicated GPUs, training endpoints such as FLUX LoRA fast training at $2 per run) or Replicate (Cog and Deployments) for those jobs, and use APIMODELS for the commercial models in the price table above; running both is common. Many models here accept reference images instead (Nano Banana Pro takes several, Seedance 2.0 up to nine), which covers a lot of what people reach for a LoRA to do, such as keeping a character or a product consistent across shots.