All vendors
    Other

    Cerebras Inference

    Foundation model providers, model access platforms, inference, model development and deployment platforms.

    Shipping

    Cloud-based inference service with OpenAI API compatibility, supporting open models including GLM, OpenAI, Qwen, and Llama. Enables deployment across cloud, on-premise, and on-device with in-region inference.

    Cloud-based AI inference service supporting open models (GLM, OpenAI, Qwen, Llama) with API key access. Delivers up to 15x faster inference than GPUs using Cerebras' Wafer-Scale Engine.

    Cerebras Dedicated Cloud

    Shipping

    Private cloud API/endpoint for scaling custom models on dedicated capacity, enabling fine-tuning and optimization for specific use cases.

    Ultra-fast AI inference engine running on the Cerebras Wafer-Scale Engine, delivering up to 2,000 tokens per second and supporting multimodal models including Gemma 4. Designed for low-latency enterprise AI applications.

    Public inference endpoints hosting open-source language models with support for rate-limited free access and dedicated endpoints for production SLAs.

    Cerebras On-Prem Deployment

    Shipping

    On-premises deployment option for full control over models, data, and infrastructure in private data centers or private clouds.

    Platform enabling model fine-tuning and pre-training on the same infrastructure as inference. Allows optimization of models with custom data for specific use cases.

    Production inference endpoints on Cerebras offering reserved capacity, higher throughput, production SLAs, and support for additional model families.

    GLM-4.7

    Shipping

    Frontier intelligence model now available on Cerebras infrastructure, delivering record-speed inference with capability comparable to Claude Sonnet 4.5 but 20x faster.

    Kimi K2.6

    Shipping

    Trillion-parameter model for enterprise inference, bringing large-scale reasoning capabilities to production workloads with Cerebras' high-speed inference infrastructure.

    A 120 billion parameter open-source GPT model optimized for production inference on Cerebras infrastructure, delivering ~3000 tokens/second throughput.

    A 120 billion parameter open-source GPT model available on Cerebras public endpoints, delivering ~3000 tokens/second for inference workloads.

    Preview

    A 31 billion parameter preview model hosted on Cerebras delivering ~1850 tokens/second, intended for evaluation purposes.

    A 355 billion parameter preview model available on Cerebras delivering ~1000 tokens/second, intended for evaluation purposes.

    Keep going — across the app