Fish Audio S1
Audio Available Full comparison ↗ fish-audio/s1 · by Fish Audio
· neural-tts
Pricing — 1 offering(s)
Text characters
- $15.00 / 1M characters
Current
2026-07-30 → present
Fish Audio's native billing unit is UTF-8 BYTES, not characters — the pricing page states "$15.00 / M UTF-8…
Fish Audio's native billing unit is UTF-8 BYTES, not characters — the pricing page states "$15.00 / M UTF-8 bytes". For English and other Latin-script text 1 character = 1 byte, so the rate is equivalent; for CJK and other multi-byte scripts, actual cost per character can run up to 3-4x this rate. Recorded here as per_1m_characters (closest schema unit) since the registry has no byte-based unit; effective_from is the date this entry was sourced, not a confirmed price-change date.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Fish Audio S1 fits into a cost-aware routing setup
See how →Capability profile
How Fish Audio S1 rates across core capability dimensions, with the task-level evidence behind each rating.
0.8% WER / 0.4% CER per Fish Audio's own figures; #1 ranking on the independent TTS-Arena2 leaderboard at time of writing — one of the stronger independently-corroborated naturalness signals in this registry. Cross-checked 2026-08-19 (SCO-462 sweep) against a second, separate independent arena: the Artificial Analysis Text to Speech Arena (artificialanalysis.ai/text-to-speech/leaderboard/provider-voice) lists this model as "OpenAudio S1" — confirmed to be the same model (OpenAudio is Fish Audio's research-lab brand for this exact release, per openaudio.com/blogs/s1) — at #42 of 98 voices tracked, Elo 1079. This is a more modest, mid-table result than the TTS-Arena2 "#1" claim — the two are different arenas with different voter pools and snapshot dates, so this doesn't contradict the TTS-Arena2 figure, but it's a meaningfully less flattering independent data point worth recording alongside it rather than only keeping the stronger claim.
13 languages — narrower than S2 Pro/S2.1 Pro (80+/83) and below ElevenLabs Multilingual v2 (29) or Eleven v3 (70+).
Voice cloning from as little as 10-15 seconds of reference audio, plus access to Fish Audio's community voice library (2M+ voices claimed platform-wide, not S1-specific).
No S1-specific latency figures published. S2 Pro/S2.1 Pro carry an explicit 100ms time-to-first-audio guarantee that S1 does not — Fish Audio's docs don't state whether S1 shares comparable latency.
Zero-shot voice cloning from ~10-15 seconds of reference audio, per Fish Audio's platform-wide cloning claim.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| Word error rate (vendor-reported) | 0.8 % | Fish Audio's own reported WER. Not independently re-measured by Modelglass. | source ↗ |
| TTS-Arena2 ranking (independent, community leaderboard) | — | Ranked #1 at time of writing per Fish Audio's own docs — an independent/community leaderboard, not a vendor-run benchmark, though the ranking snapshot itself wasn't independently re-verified by Modelglass. | — |
Operator guidance
Fish Audio's own docs label S1 "Previous Model" and recommend s2.1-pro for all new projects — S1 is kept available only "for existing integrations," priced identically to S2 Pro/S2.1 Pro. Use S2.1 Pro instead for anything new; S1 is documented here for operators already integrated against it.
Use cases
- Existing integrations already built on S1 (Fish Audio's own guidance — new projects should use S2.1 Pro instead)
- Cost-sensitive TTS with strong naturalness at S2/S2.1 Pro's same price point
- Expressive narration via 64+ emotion/tone markers using (parenthesis) syntax
Limitations
- Fish Audio itself recommends against using S1 for new projects — kept only for existing integrations
- Narrower language support (13) than S2 Pro/S2.1 Pro (80+/83) or ElevenLabs' models
- No published time-to-first-audio guarantee, unlike S2 Pro/S2.1 Pro's 100ms SLA
- Fish Audio's native billing unit is UTF-8 bytes, not characters — see registry pricing notes