
fish-s2.1-proFish Audio S2.1 Pro is a text-to-speech model that covers 83 languages from a single model id, and its distinguishing feature on this platform is that voice cloning costs nothing: you send a clean sample of at least ten seconds, get a voice id back in under twenty seconds, and pass that id as reference_id on every later request. We created three clone slots in a row while measuring and were charged zero for all of them, and the provider documents no cap on how many voices an account may store. Output comes back as mp3, wav, pcm or opus, with speed adjustable between 0.5 and 2.0. The one thing to plan around is the billing unit: this model is priced per 1,000 UTF-8 bytes of input text rather than per 1,000 characters like every other speech model here, and a Chinese character is three bytes while a Latin letter is one. A thousand English characters costs $0.03; a thousand Chinese characters is three thousand bytes and costs $0.09. Budget from bytes, not from a character count. There is also an evaluation tier, model id fish-s2.1-pro-free, that runs the identical model at $0.005 per 1,000 bytes; it is best-effort rather than guaranteed, took 41 seconds against 17 on the same passage in our measurement, and its requests may be used by the provider to improve their models, so use it to audition voices and prototype rather than for anything a user waits on.
Unlimited voice slots; a clone trains in under 20 seconds and is usable immediately
Not per character — Chinese text costs three times Latin text of the same length
No per-language endpoint or premium for non-English speech
fish-s2.1-pro for production; fish-s2.1-pro-free at a sixth of the price, slower and without guarantees
Fish Audio S2.1 Pro is a Audio & Speech API provided by Fish Audio. Fish Audio S2.1 Pro is a text-to-speech model that covers 83 languages from a single model id, and its distinguishing feature on this platform is that voice cloning costs nothing: you send a clean sample of at least ten seconds, get a voice id back in under twenty seconds, and pass that id as reference_id on every later request. We created three clone slots in a row while measuring and were charged zero for all of them, and the provider documents no cap on how many voices an account may store. Output comes back as mp3, wav, pcm or opus, with speed adjustable between 0.5 and 2.0. The one thing to plan around is the billing unit: this model is priced per 1,000 UTF-8 bytes of input text rather than per 1,000 characters like every other speech model here, and a Chinese character is three bytes while a Latin letter is one. A thousand English characters costs $0.03; a thousand Chinese characters is three thousand bytes and costs $0.09. Budget from bytes, not from a character count. There is also an evaluation tier, model id fish-s2.1-pro-free, that runs the identical model at $0.005 per 1,000 bytes; it is best-effort rather than guaranteed, took 41 seconds against 17 on the same passage in our measurement, and its requests may be used by the provider to improve their models, so use it to audition voices and prototype rather than for anything a user waits on. Through APIMODELS platform, you can access this model via a unified API with transparent pay-as-you-go pricing. Current pricing: standard, per 1K UTF-8 bytes: $0.03, evaluation tier, per 1K UTF-8 bytes: $0.005.
Generate professional-grade voiceovers for videos, animations, and ads with diverse voice options.
Quickly produce podcast audio content with support for multi-character dialogue.
Convert text content into natural, fluid speech for audiobook production.
AI-powered multilingual dubbing and translation to help content reach global audiences.
Fish Audio S2.1 Pro is available through APIMODELS at: standard, per 1K UTF-8 bytes: $0.03, evaluation tier, per 1K UTF-8 bytes: $0.005. Billing is pay-as-you-go — you only pay for what you generate.
Sign up at APIMODELS, get your API key, and call our unified API endpoint. We provide detailed API documentation with code examples in cURL, Python, and Node.js.
APIMODELS offers the same Fish Audio S2.1 Pro model through our aggregation platform. We provide a unified API interface so you do not need separate accounts for each provider - one API key to access all models.
Because that is how the upstream model charges, and passing the same unit through keeps the two sides from drifting. A Latin letter is one UTF-8 byte, a Chinese character is three, and an emoji is four. So 1,000 English characters is 1,000 bytes and costs $0.03, while 1,000 Chinese characters is 3,000 bytes and costs $0.09. Every other speech model on apimodels.app bills per 1,000 characters, so if you are switching from one of those, re-estimate CJK workloads on bytes rather than assuming the character count carries over.
No. Cloning is free and we found no cap on how many voices an account can hold. We created three clone slots back to back while measuring and each one was billed at zero. Send a clean sample of at least ten seconds, get a voice id back in roughly five to twenty seconds, then pass that id as reference_id on any later synthesis request. You only pay for the speech you generate afterwards, at the normal per-byte rate.
Use the free tier (model id fish-s2.1-pro-free) to audition voices and build prototypes, and the standard tier for anything a user waits on. They run the same model, so quality is identical, but the free tier is best-effort: the same 2,000-byte passage took 41 seconds there against 17 seconds on the standard tier in our measurement. The free tier also carries no data agreement, meaning the provider may use those requests to improve their models, and its low price has a published end date that has already moved four times.
On APIMODELS, Fish Audio S2.1 Pro runs alongside 60+ models on one API key and one balance, so choosing is about fit, not lock-in. It supports 83 Languages, Free Voice Cloning, mp3 / wav / pcm / opus, Speed 0.5-2.0, and you can weigh it on price and capability against other Audio & Speech models, then switch by changing a single model-name string — no new account or integration. Browse every Audio & Speech option with live pricing at apimodels.app/models.
Fish Audio S2.1 Pro supports: 83 Languages, Free Voice Cloning, mp3 / wav / pcm / opus, Speed 0.5-2.0. See the APIMODELS docs for full parameters and call examples.
Yes. APIMODELS exposes Fish Audio S2.1 Pro through a single unified API and one key — no separate provider accounts, and no need to handle each provider's regional network access yourself.
We support Stripe (Visa, Mastercard, and other international cards) and Alipay. Credits are available instantly after payment.