Stable Audio 2.0
Audio Deprecated Full comparison ↗ stability-ai/stable-audio-2 · by Stability AI
· Diffusion transformer (DiT)
Pricing — 1 offering(s)
Per generation (Professional plan)
- $0.20 / generation
Current
2024-04-01 → present
Professional plan (USD $20/month, 100 credits). Each generation = 1 credit. Free tier: 20 generations/month…
Professional plan (USD $20/month, 100 credits). Each generation = 1 credit. Free tier: 20 generations/month. SCO-617: citation URL updated — the old bare-domain stability.ai/pricing now redirects to a 404 (stability.ai/brand-studio-plans); this file was the one entry in the repo not yet migrated to platform.stability.ai/pricing, the URL already used by its sibling Stable Audio entries. Historical pricing figure unchanged, not re-verified (deprecated model).
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how Stable Audio 2.0 fits into a cost-aware routing setup
See how →Capability profile
How Stable Audio 2.0 rates across core capability dimensions, with the task-level evidence behind each rating.
44.1kHz stereo output with high fidelity. Particularly strong on instrumental music and sound design; one of the highest-quality open audio generation models.
Broad style coverage from text descriptions. Strong on electronic, ambient, cinematic, and classical. Text conditioning produces detailed style differentiation.
Limited vocal generation capability; designed primarily for instrumental music and sound design. Occasional vocal artefacts from text prompts that imply vocals.
Up to 3 minutes of audio per generation — significantly longer than MusicGen (30s) or Suno's 2-minute ceiling. Timing conditioning ensures consistent structure throughout.
No stem export. Single mixed-down audio file output.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| CLAP score and FAD (Stability AI internal) | — | Stability AI reports improved CLAP and FAD scores vs Stable Audio 1.0. No independent third-party benchmark published comparing against Suno v4 or Udio on a standardised dataset. | — |
Operator guidance
DEPRECATED as of mid-2026 — superseded by Stable Audio 2.5 and 3.0, which are now the active products on platform.stability.ai/pricing. Use stable-audio-2-5-stability-ai for the same per-generation price ($0.20) with improved coherence and audio conditioning, or stable-audio-3-0-stability-ai for 6-minute generation at $0.26.
Use cases
- Long-form background music (up to 3 minutes) for video and podcasts
- Sound design and ambient audio for games and installations
- High-fidelity instrumental generation at 44.1kHz
- Applications where output duration precision is required (timing conditioning)
Limitations
- Deprecated — no longer listed on platform.stability.ai/pricing as of 2026-06-26
- Limited vocal generation — primarily instrumental and sound design
- No stem export