MusicGen Large
Audio Available Full comparison ↗ meta/musicgen-large · by Meta AI
· transformer-decoder
Pricing — 1 offering(s)
Per generation (Replicate-hosted)
- $0.042 / generation
Historical
2023-08-01 → 2026-09-19
Approximate cost for a 30-second clip on Replicate (GPU compute pricing; varies by output duration and server…
Approximate cost for a 30-second clip on Replicate (GPU compute pricing; varies by output duration and server load). Self-hosted compute cost will differ. SCO-608 (2026-09-20) — source.url now 404s: Replicate consolidated the standalone "musicgen-large" model page into a single unified model page with a version selector (see the 2026-09-20 entry below).
- $0.10 / generation
Current
2026-09-20 → present
SCO-608: dead-URL fix + price re-verification. Replicate retired the standalone /meta/musicgen-large page…
SCO-608: dead-URL fix + price re-verification. Replicate retired the standalone /meta/musicgen-large page (confirmed genuine 404, not bot-blocking) and consolidated Large + Melody variants into one model page (/meta/musicgen) with an in-page version selector; same page is mirrored at /facebookresearch/musicgen. Live page states "approximately $0.10 to run on Replicate, or 10 runs per $1, but this varies depending on your inputs" on Nvidia A100 80GB hardware, ~72s per prediction — a flat approximate-per-run rate that does not differ between the Large and Melody variants shown on the page. Real increase from the prior $0.042/30s-clip estimate, not a data error. SCO-672: marked estimated. Replicate's figure is a rolling per-run estimate ("approximately ... varies depending on your inputs"), not a list price: $0.10 on 2026-09-20, $0.036 on 2026-09-30. It isn't monitored as a price.
Showing the active price and any recorded history. Full pricing history is available via the paid API — see API docs.
See how MusicGen Large fits into a cost-aware routing setup
See how →Capability profile
How MusicGen Large rates across core capability dimensions, with the task-level evidence behind each rating.
Good audio quality for an open-weights model; not at the level of Suno v4 or Stable Audio 2.0 on instrumental fidelity. No post-processing artifacts from a well-configured setup.
Strong genre coverage from training data. Text conditioning works well for mainstream and classical genres; niche sub-genres are less reliable.
No vocal generation. Pure instrumental output only — a fundamental architectural limitation, not a configuration option.
Default generation limit of 30 seconds per call. Longer audio requires multiple generations and stitching.
No stem export. Mixed-down audio output only.
Benchmarks
| Benchmark | Score | Config | Source |
|---|---|---|---|
| FAD (Fréchet Audio Distance, published in paper) | — | Copet et al. report FAD and KL-divergence scores on MusicCaps dataset. MusicGen Large outperforms prior open models (MusicLM, Riffusion). Exact scores in the paper. | source ↗ |
Operator guidance
Choose MusicGen Large when open-weights ownership matters (CC BY-NC 4.0 allows commercial use with attribution), or when self-hosting eliminates per-call costs at scale. For production quality with vocals, use Suno v4. For longer-form high-fidelity instrumentals, use Stable Audio 2.0. For melody conditioning specifically, MusicGen's melody-conditioning feature is unique among major models.
Use cases
- Background music generation in applications where open-weights licensing matters
- Research and experimentation on music generation pipelines
- Self-hosted music generation to eliminate per-call API costs at scale
- Melody-conditioned generation: constraining output to follow a reference melody
Limitations
- Instrumental only — no vocal generation
- 30-second generation ceiling per call; longer clips require stitching
- CC BY-NC 4.0 licence — commercial use requires attribution; no-commercial clause applies to the model weights
- Requires GPU infrastructure for reasonable generation speed (no first-party hosted API)