← All models

Groq

Available

groq.com ↗

Coverage across verticals

Language 1 model

1 previous-gen

Language pricing page ↗

Inference provider (LPU hardware), not a model creator. Serves open-weight models (Llama, etc.) at low latency with per-1M-token input/output pricing.

Audio 1 model

1 ungraded

Audio pricing page ↗

Inference provider (LPU hardware), not a model creator. Hosts Whisper Large v3 (and a cheaper Turbo variant) for speech-to-text at very low latency — 217x real-time per Groq's own benchmark. Billed per hour of audio processed, with a 10-second minimum per request.

Release cadence

2 models tracked across all verticals. Earliest release 2023-11, most recent 2024-12.

Derived from Modelglass's per-vertical provider registries — image/llm/video/audio's registry/providers/*.yaml records, plus science/coding/agentic's benchmark registries, where a provider is identified by the prefix before "/" in each model's model_id rather than a registry record of its own. Where the same company uses a different provider slug or model_id prefix across verticals (e.g. Google DeepMind vs. Google Cloud, or Qwen's "alibaba" model_id prefix in the coding registry vs. its "qwen" provider slug elsewhere), coverage is merged onto a single profile per an explicit, human-reviewed alias list (registry/provider-aliases.yaml) — everything else reflects exact-slug or exact-prefix matches only.