Groq
AvailableCoverage across verticals
1 previous-gen
Language pricing page ↗Inference provider (LPU hardware), not a model creator. Serves open-weight models (Llama, etc.) at low latency with per-1M-token input/output pricing.
1 ungraded
Audio pricing page ↗Inference provider (LPU hardware), not a model creator. Hosts Whisper Large v3 (and a cheaper Turbo variant) for speech-to-text at very low latency — 217x real-time per Groq's own benchmark. Billed per hour of audio processed, with a 10-second minimum per request.
Release cadence
2 models tracked across all verticals. Earliest release 2023-11, most recent 2024-12.
Derived from Modelglass's per-vertical provider registries — image/llm/video/audio's registry/providers/*.yaml records, plus science/coding/agentic's benchmark registries, where a provider is identified by the prefix before "/" in each model's model_id rather than a registry record of its own. Where the same company uses a different provider slug or model_id prefix across verticals (e.g. Google DeepMind vs. Google Cloud, or Qwen's "alibaba" model_id prefix in the coding registry vs. its "qwen" provider slug elsewhere), coverage is merged onto a single profile per an explicit, human-reviewed alias list (registry/provider-aliases.yaml) — everything else reflects exact-slug or exact-prefix matches only.