Shipping
Cerebras Cloud API
Cloud-based inference service with OpenAI API compatibility, supporting open models including GLM, OpenAI, Qwen, and Llama. Enables deployment across cloud, on-premise, and on-device with in-region inference.
Cerebras Cloud Inference
Cloud-based AI inference service supporting open models (GLM, OpenAI, Qwen, Llama) with API key access. Delivers up to 15x faster inference than GPUs using Cerebras' Wafer-Scale Engine.
Cerebras Dedicated Cloud
Private cloud API/endpoint for scaling custom models on dedicated capacity, enabling fine-tuning and optimization for specific use cases.
Cerebras Inference (CS-3)
Ultra-fast AI inference engine running on the Cerebras Wafer-Scale Engine, delivering up to 2,000 tokens per second and supporting multimodal models including Gemma 4. Designed for low-latency enterprise AI applications.
Cerebras Inference API
Public inference endpoints hosting open-source language models with support for rate-limited free access and dedicated endpoints for production SLAs.
Cerebras On-Prem Deployment
On-premises deployment option for full control over models, data, and infrastructure in private data centers or private clouds.
Cerebras Training & Fine-tuning
Platform enabling model fine-tuning and pre-training on the same infrastructure as inference. Allows optimization of models with custom data for specific use cases.
Dedicated Endpoints
Production inference endpoints on Cerebras offering reserved capacity, higher throughput, production SLAs, and support for additional model families.
GLM-4.7
Frontier intelligence model now available on Cerebras infrastructure, delivering record-speed inference with capability comparable to Claude Sonnet 4.5 but 20x faster.
Kimi K2.6
Trillion-parameter model for enterprise inference, bringing large-scale reasoning capabilities to production workloads with Cerebras' high-speed inference infrastructure.
OpenAI GPT OSS
A 120 billion parameter open-source GPT model optimized for production inference on Cerebras infrastructure, delivering ~3000 tokens/second throughput.
OpenAI GPT OSS 120B
A 120 billion parameter open-source GPT model available on Cerebras public endpoints, delivering ~3000 tokens/second for inference workloads.
Preview
Gemma 4 31B
A 31 billion parameter preview model hosted on Cerebras delivering ~1850 tokens/second, intended for evaluation purposes.
Z.ai GLM 4.7
A 355 billion parameter preview model available on Cerebras delivering ~1000 tokens/second, intended for evaluation purposes.