What is Groq?
Groq is an inference provider that runs models on Language Processing Units — purpose-built silicon for sequential token generation rather than general-purpose GPUs. Depending on the model it serves between roughly 276 and 1,500+ tokens per second, against the 40–100 typical of GPU-based providers. It hosts open-weight models including Llama 4 and Llama 3.x, Mixtral, Gemma, DeepSeek R1 distills and Whisper.