Replicate, part of Cloudflare since December 2025, is two products in one: a marketplace of 50,000+ community models billed by GPU time, and more than 100 “official” models such as Nano Banana Pro, Veo 3.1, Kling and Seedance billed per output. APIMODELS competes only with the second. If you call commercial models and never deploy your own, the decision comes down to price, API shape and how each platform behaves when something fails, and that is what this page compares.
Nano Banana Pro is $0.08 per 1K–2K image here against $0.15 on Replicate, and $0.13 against $0.30 at 4K. An 8-second Veo 3.1 clip with audio is $1.00 here against $3.20 at Replicate’s $0.40 per second, and Veo 3.1 Fast is $0.07 against $1.20. Kling 3.0 pro with audio is $0.24 per second against $0.336, and GPT Image 2 at medium quality is $0.025 (1K) and $0.03 (2K) against $0.047. All Replicate figures were read from its model pages on 6 October 2026.
Two practical differences show up in production. Replicate removes API outputs after one hour by default, so files must be copied out straight away; result files here stay downloadable for 30 days. And Replicate signs you in through GitHub only, while APIMODELS takes any email, needs no organization verification, is reachable from mainland China, and accepts card, PayPal, Alipay and USDT. On both platforms failed runs of hosted models are not charged.
The Replicate column comes from its model pages, docs, pricing page and terms, checked on 2026-10-06; sources are at the end of this page.
| APIMODELS | Replicate | |
|---|---|---|
| What it is | Multi-model API gateway for hosted commercial models | Model platform: community models on GPU time, official models per output, custom deployments |
| Model coverage | 149 models: 37 image, 40 video, 49 LLM, 21 audio, 2 embedding (October 2026) | 50,000+ models, of which more than 100 are official per-output models |
| Billing | Fixed price per image, per second or per 1M tokens | Official models per output; other models per second of GPU time (H100 $5.49/hr) |
| Failed requests | Never charged | Not charged on public and official models; billed on private models and deployments |
| API shape | OpenAI-compatible chat and image endpoints, native Anthropic Messages; async video with callback_url | Predictions API, async with webhooks or Prefer: wait (up to 60 s); no OpenAI-compatible endpoint documented |
| Output retention | 30 days | API outputs removed after 1 hour by default |
| Sign-up and verification | Any email; no organization or ID verification | GitHub sign-in; no ID verification mentioned |
| Payment | Card (Stripe), PayPal, Alipay, USDT; top-ups from $10 | Prepaid credit by card or bank transfer (valid 1 year, non-refundable), or billing in arrears |
| Custom models | Not offered | Cog packaging, private models, Deployments with min / max instances, fine-tuning |
USD, checked on 2026-10-06, against Replicate’s official-model prices. Veo clips here are 8 seconds at 1080p with audio. Replicate charges GPT Image 2 one price per quality level at any size and lists Grok Imagine Video 1.5 as an image-to-video preview at 480p and 720p. On LLMs, Replicate lists Claude Sonnet 5 at Anthropic’s $2 / $10 per million tokens against $1.60 / $8 here, and had no listing for Claude Sonnet 5.5, Opus 5.5 or GPT-6 on the check date.
| Model and unit | APIMODELS | Replicate | Lower price |
|---|---|---|---|
| Nano Banana Pro, 1K–2K / 4K, per image | $0.08 / $0.13 | $0.15 / $0.30 | APIMODELS |
| GPT Image 2, medium quality, 1K / 2K, per image | $0.025 / $0.03 | $0.047 | APIMODELS |
| GPT Image 2, high quality, 1K / 2K, per image | $0.07 / $0.10 | $0.128 | APIMODELS |
| Kling 3.0 standard / pro, with audio, per second | $0.18 / $0.24 | $0.252 / $0.336 | APIMODELS |
| Veo 3.1, 8 s with audio, per clip | $1.00 | $3.20 ($0.40/s) | APIMODELS |
| Veo 3.1 Fast, 8 s with audio, per clip | $0.07 | $1.20 ($0.15/s) | APIMODELS |
| Grok Imagine Video 1.5, 480p / 720p, per second (on Replicate: preview, image-to-video only) | $0.045 / $0.0529 | $0.08 / $0.08 | APIMODELS |
Choose Replicate when you need to run a model that is not a commercial API. You can package any model with Cog, keep it private, serve it through Deployments with minimum and maximum instance counts, and fine-tune FLUX for about $1.46 per run in Replicate’s documented example. APIMODELS cannot host your weights.
Choose Replicate for the long tail of open models: 50,000+ community models, from niche upscalers to research checkpoints, all behind the same predictions API, with SDKs for Node.js, Python, Swift and Go and an MCP server.
fal.ai suits teams that want hosted models and serverless GPU deployment from one vendor, plus Veo 3.1 at 4K.
OpenRouter suits LLM-first workloads that need provider routing and zero-data-retention controls.
Kie.ai suits buyers who use its Suno V6 toolset and Runway endpoints. Each of the three has its own comparison page, linked at the bottom.
Swapping replicate.run(...) for APIMODELS is one HTTP request per job. Replicate takes the model as a path (owner/name) plus an input object; APIMODELS takes the model as a body field, such as veo-3.1, kling-v3, seedance-2.0 or gemini-3-pro-image, next to the prompt and settings, at POST /v1/video/generations or /v1/images/generations.
Replicate’s webhook becomes callback_url. Its Prefer: wait mode maps to the OpenAI-style image call: send an OpenAI size such as 1024x1024 and no callback_url, and the response carries the image. For video you poll GET /v1/video/generations?task_id=… until state is completed. Because results stay up for 30 days, a copy-within-the-hour step that exists only because of Replicate’s retention can be relaxed.
LLM calls move to the OpenAI SDK with base_url https://api.apimodels.app/v1, or the Anthropic SDK with https://api.apimodels.app, using ids such as claude-sonnet-5 or claude-sonnet-5-5.
| Fact | Source |
|---|---|
| Image prices | replicate.com/google/nano-banana-pro; replicate.com/openai/gpt-image-2 |
| Video prices | replicate.com/kwaivgi/kling-v3-video; /google/veo-3.1; /google/veo-3.1-fast; /xai/grok-imagine-video-1.5 |
| LLM prices | replicate.com/anthropic/claude-sonnet-5 |
| Billing, failed runs, prepaid credit | replicate.com/docs/topics/billing; /docs/topics/billing/prepaid-credit; replicate.com/pricing; replicate.com/terms |
| API, webhooks, retention | replicate.com/docs/topics/predictions/create-a-prediction; /docs/topics/predictions/data-retention |
| Cog, Deployments, fine-tuning | replicate.com/docs/get-started/deploy-a-custom-model; /docs/topics/deployments; /docs/get-started/fine-tune-with-flux |
| Cloudflare acquisition, model count | replicate.com/blog/replicate-cloudflare; blog.cloudflare.com/why-replicate-joining-cloudflare |
| APIMODELS prices | apimodels.app/pricing and each model page |
cURL
# Before (Replicate): replicate.run("google/veo-3.1", input={"prompt": ...})
# After (APIMODELS): model goes in the body; 8 s, 1080p, with audio, $1.00 per clip
curl -X POST https://api.apimodels.app/v1/video/generations \
-H "Authorization: Bearer $APIMODELS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1",
"prompt": "A paper boat drifting down a rain-soaked street at dusk",
"aspect_ratio": "9:16",
"images": ["https://example.com/first-frame.jpg"]
}'
# images is optional: [first frame] or [first frame, last frame]
# poll: curl "https://api.apimodels.app/v1/video/generations?task_id=TASK_ID" -H "Authorization: Bearer $APIMODELS_API_KEY"Python
import requests, time
BASE = "https://api.apimodels.app/v1"
H = {"Authorization": "Bearer YOUR_APIMODELS_KEY"}
# Before (Replicate): replicate.run("kwaivgi/kling-v3-video", input={...})
task = requests.post(f"{BASE}/video/generations", headers=H, json={
"model": "kling-v3",
"prompt": "A barista pours latte art in a sunlit cafe, ambient chatter",
"mode": "pro", # std or pro
"duration": "10",
"sound": "on", # pro + audio: $0.24 per second
}).json()["data"]["taskId"]
while True:
s = requests.get(f"{BASE}/video/generations", headers=H,
params={"task_id": task}).json()["data"]
if s["state"] in ("completed", "failed"):
print(s.get("resultUrls")) # downloadable for 30 days
break
time.sleep(5)Yes, on every model in the price table on this page. Checked on 6 October 2026: Nano Banana Pro is $0.08 per 1K–2K image here and $0.15 on Replicate ($0.13 against $0.30 at 4K); an 8-second Veo 3.1 clip with audio is $1.00 against $3.20, and Veo 3.1 Fast $0.07 against $1.20; Kling 3.0 with audio is $0.18 / $0.24 per second (standard / pro) against $0.252 / $0.336; GPT Image 2 at medium quality is $0.025 at 1K and $0.03 at 2K against a flat $0.047; and Grok Imagine Video 1.5 is $0.045 / $0.0529 per second at 480p / 720p against $0.08 for Replicate’s image-to-video preview. Replicate prices its official models per output, so compare the exact resolution and quality you use.
We do not publish an uptime figure, so compare the rules instead. On APIMODELS a failed generation is never charged, many models fail over automatically between more than one upstream route, and an async job retries its callback for up to 30 minutes if your server is down. Replicate’s docs say failed runs on public and official models are not charged, while failed and cancelled runs on private models and deployments are; it offers webhooks and a synchronous Prefer: wait mode of up to 60 seconds. Retention matters too: Replicate removes API outputs after an hour by default, so a missed webhook can mean a lost file, whereas results here stay downloadable for 30 days and any task can be re-queried by its task_id.
Replace each replicate.run or predictions call with one POST: video models to /v1/video/generations and image models to /v1/images/generations, with the model as a body field rather than a path, so google/veo-3.1 becomes "model": "veo-3.1" and kwaivgi/kling-v3-video becomes "model": "kling-v3". Move the input fields up to the top level of the JSON body and check each model’s docs page for exact names, since they differ by model family. Swap webhook for callback_url, or poll with the returned task_id until state is completed, then read resultUrls. For a synchronous image, send an OpenAI-style size and no callback_url. LLM code points the OpenAI SDK at https://api.apimodels.app/v1. Custom models packaged with Cog cannot move here.
Replicate asks you to buy credit upfront by card or bank transfer; purchased credit is valid for one year and is not refundable, and auto-reload has a $5 minimum threshold and a $15 minimum reload. Larger accounts can be billed in arrears. Community models are billed by GPU time (an H100 is $5.49 an hour), official models per output. APIMODELS is prepaid too: top-ups from $10 by Stripe card, PayPal, Alipay or USDT, with a 2–5% bonus on packages from $50, automatic invoices for card payments, and one fixed price per image, per second or per million tokens for every model, so a job’s cost is known before you send it. New accounts on consumer email domains get $0.10 to test with.
No. APIMODELS serves hosted commercial and open models chosen by us, so there is no way to upload weights, package a container or fine-tune. If you depend on Cog, private models or Deployments with warm instances, keep those on Replicate and route the commercial models (Veo, Kling, Nano Banana Pro, GPT Image 2, Claude) through APIMODELS. Several models here cover cases that used to need a custom checkpoint: Nano Banana Pro and GPT Image 2 edit from reference images, Seedance 2.0 takes up to nine reference images plus reference video and audio, and Kling Motion Control transfers motion from a source clip.
Not on APIMODELS: sign up with any email address, create a key in the console, and call the API; there is no organization, ID or face verification, and the service is reachable from mainland China with Alipay accepted. Replicate’s sign-in page offers GitHub only, and its pages mention no ID verification. For teams whose members do not all have GitHub accounts, or that pay through Alipay, that difference decides it. For teams already on GitHub and Cloudflare, Replicate fits into what they have.