
fish-s2.1-proFish Audio S2.1 Pro は、ひとつのモデル ID で 83 言語をカバーする音声合成モデルです。このプラットフォームでの最大の利点は、音声クローンが無料である点です。十秒以上のクリアなサンプルを送ると二十秒以内にボイス ID が返り、以降はそれを reference_id として渡すだけで同じ声で話します。計測中に三つのクローンを続けて作成しましたが、いずれも課金はゼロで、提供元はアカウントごとの音声数の上限を文書化していません。出力は mp3、wav、pcm、opus の四形式、速度は 0.5 から 2.0 まで調整できます。事前に押さえておくべきなのは課金単位です。このモデルは入力テキスト 1000 文字ではなく 1000 UTF-8 バイト単位で課金され、漢字は 3 バイト、ラテン文字は 1 バイトです。英語 1000 文字は $0.03、中国語 1000 文字は 3000 バイトで $0.09 になります。文字数ではなくバイト数で見積もってください。
Unlimited voice slots; a clone trains in under 20 seconds and is usable immediately
Not per character — Chinese text costs three times Latin text of the same length
No per-language endpoint or premium for non-English speech
fish-s2.1-pro for production; fish-s2.1-pro-free at a sixth of the price, slower and without guarantees
Fish Audio S2.1 Pro は Fish Audio の音声合成 API です。Fish Audio S2.1 Pro は、ひとつのモデル ID で 83 言語をカバーする音声合成モデルです。このプラットフォームでの最大の利点は、音声クローンが無料である点です。十秒以上のクリアなサンプルを送ると二十秒以内にボイス ID が返り、以降はそれを reference_id として渡すだけで同じ声で話します。計測中に三つのクローンを続けて作成しましたが、いずれも課金はゼロで、提供元はアカウントごとの音声数の上限を文書化していません。出力は mp3、wav、pcm、opus の四形式、速度は 0.5 から 2.0 まで調整できます。事前に押さえておくべきなのは課金単位です。このモデルは入力テキスト 1000 文字ではなく 1000 UTF-8 バイト単位で課金され、漢字は 3 バイト、ラテン文字は 1 バイトです。英語 1000 文字は $0.03、中国語 1000 文字は 3000 バイトで $0.09 になります。文字数ではなくバイト数で見積もってください。 APIMODELS のプラットフォーム経由なら、統一 API と明朗な従量課金でこのモデルを呼び出せます。 現在の料金: standard, per 1K UTF-8 bytes: $0.03, evaluation tier, per 1K UTF-8 bytes: $0.005。
動画、アニメーション、広告向けにプロ品質のナレーションを、豊富な声質から選んで生成します。
複数話者の掛け合いにも対応し、ポッドキャストの音声を素早く仕上げます。
テキストを自然で流れのよい読み上げに変換し、オーディオブックとして出せます。
多言語の吹き替えと翻訳で、コンテンツを世界の視聴者に届けます。
Fish Audio S2.1 Pro は APIMODELS 経由で standard, per 1K UTF-8 bytes: $0.03, evaluation tier, per 1K UTF-8 bytes: $0.005 で利用できます。課金は従量制で、生成した分だけの支払いです。
APIMODELS に登録して API キーを取得し、統一エンドポイントを呼ぶだけです。cURL / Python / Node.js のサンプルを含む詳しいドキュメントを用意しています。
APIMODELS は同じ Fish Audio S2.1 Pro を集約プラットフォーム経由で提供します。統一された API インターフェースなので、プロバイダごとにアカウントを作る必要はありません。キー 1 本ですべてのモデルに届きます。
Because that is how the upstream model charges, and passing the same unit through keeps the two sides from drifting. A Latin letter is one UTF-8 byte, a Chinese character is three, and an emoji is four. So 1,000 English characters is 1,000 bytes and costs $0.03, while 1,000 Chinese characters is 3,000 bytes and costs $0.09. Every other speech model on apimodels.app bills per 1,000 characters, so if you are switching from one of those, re-estimate CJK workloads on bytes rather than assuming the character count carries over.
No. Cloning is free and we found no cap on how many voices an account can hold. We created three clone slots back to back while measuring and each one was billed at zero. Send a clean sample of at least ten seconds, get a voice id back in roughly five to twenty seconds, then pass that id as reference_id on any later synthesis request. You only pay for the speech you generate afterwards, at the normal per-byte rate.
Use the free tier (model id fish-s2.1-pro-free) to audition voices and build prototypes, and the standard tier for anything a user waits on. They run the same model, so quality is identical, but the free tier is best-effort: the same 2,000-byte passage took 41 seconds there against 17 seconds on the standard tier in our measurement. The free tier also carries no data agreement, meaning the provider may use those requests to improve their models, and its low price has a published end date that has already moved four times.
APIMODELS では Fish Audio S2.1 Pro が 60 以上のモデルと同じ API キー・同じ残高の上に並んでいるので、選択は「相性」の問題であって「囲い込み」の問題ではありません。83 Languages、Free Voice Cloning、mp3 / wav / pcm / opus、Speed 0.5-2.0 に対応しており、他の音声合成モデルと価格・機能を並べて評価できます。乗り換えはモデル名の文字列を 1 つ書き換えるだけ。新しいアカウントも追加の実装も要りません。音声合成の選択肢と最新価格は apimodels.app/models で確認できます。
Fish Audio S2.1 Pro は次に対応しています: 83 Languages、Free Voice Cloning、mp3 / wav / pcm / opus、Speed 0.5-2.0。パラメータの全一覧と呼び出し例は APIMODELS のドキュメントをご覧ください。
はい。APIMODELS は Fish Audio S2.1 Pro を単一の統一 API とキー 1 本で提供します。プロバイダごとのアカウントも、各社の地域ごとのネットワーク経路を自分で面倒みる必要もありません。
Stripe(Visa、Mastercard などの国際カード)と Alipay に対応しています。支払い後、残高はすぐ反映されます。