Cerebras Systems, Inc. is an American semiconductor, computer systems, and artificial intelligence hardware company founded in 2015 by former SeaMicro and AMD executives Andrew Feldman, Gary Lauterbach, Sean Lie, Michael James, and Jean-Philippe Fricker. Headquartered in Sunnyvale, California, Cerebras is renowned for creating the world's first and only commercial 'wafer-scale' processors—most notably the Wafer-Scale Engine 3 (WSE-3), a single 300mm silicon wafer packing 4 trillion transistors, 900,000 cores, and 44 gigabytes of on-chip SRAM. In 2026, Cerebras surpassed an annualized revenue run-rate of $250 million ($250M+ ARR) at a private valuation exceeding $7.0 billion, filing for a landmark Nasdaq initial public offering (IPO) under the leadership of CEO Andrew Feldman.
Cerebras Systems, Inc.: Key Facts & Operational Metrics
| Company Name | Cerebras Systems, Inc. |
|---|---|
| Founded | 2015 |
| Founders | Andrew Feldman, Gary Lauterbach, Sean Lie, Michael James, Jean-Philippe Fricker |
| Headquarters | Sunnyvale, California, United States |
| Industry | Semiconductors, Artificial Intelligence Supercomputing & Cloud Inference |
| Chief Executive Officer | Andrew Feldman |
| Chief Technology Officer | Gary Lauterbach |
| Chief Hardware Architect | Sean Lie |
| Employees | Approximately 450 personnel |
| Annualized Revenue (ARR) | $250M+ ARR (2026 Run-Rate) |
| Valuation | $7.0 billion+ (IPO S-1 Registration) |
| Core Silicon | Wafer-Scale Engine 3 (4 Trillion Transistors, 900,000 Cores, 44GB SRAM) |
| Flagship Products | CS-3 Supercomputer, Cerebras Inference Cloud, Condor Galaxy |
| Key Partners & Customers | TSMC, G42, Qualcomm, Argonne National Laboratory, Mayo Clinic |
| Website | cerebras.ai |
- Financial and operational metrics verified via official SEC Form S-1 registration statements and quarterly disclosures
- Chip architectural specifications verified via Hot Chips IEEE presentations and TSMC technical documentation
- Inference performance benchmarks independently verified by Artificial Analysis and open-source researchers
- For informational purposes only - not financial advice
For more than seventy years, the global semiconductor industry operated under an immutable manufacturing dogma: silicon wafers are sliced into hundreds of tiny individual dies. When Robert Noyce and Jack Kilby co-invented the integrated circuit, chips were microscopic. As lithography advanced, chips grew larger, but manufacturing physics dictated a strict ceiling: because microscopic dust specks and crystal lattice defects inevitably occur during fabrication, larger chips have lower manufacturing yields. If a chip is too big, defects are guaranteed, and the entire chip must be thrown away. In the 1980s, legendary computer architect Gene Amdahl raised hundreds of millions of dollars to build a 'wafer-scale' computer at Trilogy Systems; the attempt failed catastrophically, bankrupting the company and convincing generations of venture capitalists that building a computer chip the size of an entire silicon wafer was physically impossible.
In 2015, five veteran Silicon Valley hardware engineers led by Andrew Feldman and Gary Lauterbach founded Cerebras Systems with the express purpose of defying that seventy-year dogma. By designing proprietary hardware-level defect tolerance into the silicon itself, Cerebras succeeded where Gene Amdahl failed—fabricating the Wafer-Scale Engine, a single continuous square of silicon measuring 46,225 square millimeters containing 4 trillion transistors. By eliminating the copper cables, optical transceivers, and off-chip memory bottlenecks that plague NVIDIA GPU clusters, Cerebras engineered a supercomputing monolith that shatters global AI speed records, charting a path toward a historic initial public offering.
What Does Cerebras Systems Do?
Cerebras Systems designs, manufactures, and deploys wafer-scale computing systems engineered specifically for artificial intelligence model training and real-time inference:
- Wafer-Scale Engine 3 (WSE-3): Built on TSMC's 5-nanometer process, the WSE-3 is the largest computer chip in the world. Measuring 56 times larger than the largest GPU die, it packs 4 trillion transistors, 900,000 independent AI cores, 44 gigabytes of ultra-fast on-chip SRAM, and delivers 125 petaflops of peak computational power.
- CS-3 AI Supercomputer: The turnkey systems cabinet that houses the WSE-3. Standing roughly 26 inches tall (15 rack units), the CS-3 delivers closed-loop liquid cooling, uniform power distribution of over 20 kilowatts, and internal 1.2 Terabits per second I/O, replacing an entire row of traditional GPU servers with a single compact system.
- Cerebras Inference Cloud: A high-throughput, serverless developer cloud that delivers the fastest AI inference on the planet. By keeping entire foundation models in on-chip SRAM with 21 Petabytes per second of memory bandwidth, Cerebras generates over 1,800 to 2,100 tokens per second on Llama 3.1 70B—up to 20x faster than NVIDIA H100 clouds.
- Condor Galaxy Supercomputers: Multi-exoflop supercomputing constellations built in partnership with G42, interconnecting dozens of CS-3 systems to train frontier foundation models, such as Jais (Arabic LLM) and Med42 (clinical AI).
- CSoft Compiler Platform: Turnkey software compiler suite that maps standard PyTorch and TensorFlow neural networks directly onto the 900,000 cores without requiring developers to write complex distributed MPI or Megatron code.
How Does Cerebras Make Money?
Cerebras operates a high-margin semiconductor, enterprise systems, and recurring cloud software business model:
- Supercomputer Hardware Sales (CS-3 Systems): Sovereign nation-states, defense agencies, and enterprise conglomerates purchase turnkey CS-3 systems cabinets. Individual supercomputer hardware configurations sell for several million dollars per unit, generating substantial upfront hardware revenues and ongoing multi-year maintenance service contracts.
- Turnkey Multi-Exoflop Clusters (G42 Strategic Contracts): Multi-hundred-million-dollar mega-contracts to construct and manage massive supercomputing constellations like Condor Galaxy 1, 2, and 3, comprising dozens of interconnected wafer-scale engines.
- Cerebras Inference Cloud Token Consumption: Software developers, AI voice agent startups, and enterprise applications pay usage-based fees per million tokens consumed on the Cerebras Inference Cloud, generating rapidly compounding, software-grade recurring revenue.
Cerebras Financials & Revenue Trajectory
Cerebras has executed an explosive revenue acceleration curve as disclosed in its public SEC filings:
- 2022: Annual revenue stood at roughly $24 million as CS-2 deployments commenced.
- 2023: Revenue surged over 450% to $136.2 million, driven by massive supercomputing contracts with UAE-based G42.
- 2024: Revenue reached $136 million in the first half of the year alone, expanding toward an annualized run-rate exceeding $250 million ($250M+ ARR) ahead of its public IPO filing at a valuation over $7.0 billion.
Cerebras maintains gross margins exceeding 50% on hardware systems and over 75% on its inference cloud services. Backed by premier venture capital firms including Foundation Capital, Eclipse Ventures, and Alpha Wave, Cerebras represents one of the most commercially significant pure-play semiconductor IPOs of the decade.
Origins: Overcoming Gene Amdahl's Wafer-Scale Ghost
The quest for wafer-scale computing is one of the great tragic epics of computer engineering. In 1980, Gene Amdahl—the legendary chief designer of the IBM System/360—founded Trilogy Systems with $230 million in capital to build a wafer-scale computer. The project collapsed in ruin because silicon manufacturing defects were inevitable: if even one defect occurred on a wafer, the entire chip was ruined. For thirty-five years, 'wafer-scale' was considered a career-ending fool's errand.
In 2015, Andrew Feldman, Gary Lauterbach, and Sean Lie—having just sold micro-server maker SeaMicro to AMD—decided that deep learning presented a unique historical opportunity. Unlike general-purpose CPUs, deep learning neural networks are inherently parallel and fault-tolerant. Sean Lie conceived a brilliant hardware innovation: Algorithmic Defect Tolerance. Cerebras designed the wafer with 1.5% redundant cores and a reconfigurable 2D mesh network. When a wafer is powered on, custom firmware tests every single core, identifies the microscopic silicon defects, and automatically routes communication around the flawed cores. By making defects irrelevant, Cerebras turned the 50-year-old impossibility of wafer-scale computing into an operational reality.
The Physics of Speed: 21 Petabytes/sec Memory Bandwidth
To understand why Cerebras crushes NVIDIA GPUs on inference speed, one must understand the 'Memory Wall'. In traditional GPUs (such as the NVIDIA H100), the GPU compute die is physically separated from its high-bandwidth memory (HBM3) stacks by a silicon interposer. While HBM3 is fast, the physical bus connecting the memory to the compute cores maxes out at approximately 3.35 Terabytes per second.
In large language model inference, every single generated token requires reading every single parameter of the model from memory. As a result, GPUs spend the vast majority of their time waiting for data to travel across the memory bus. Cerebras eliminated this physical separation entirely by integrating 44 Gigabytes of ultra-dense SRAM directly onto the silicon wafer itself. Because the memory is located micrometers away from the 900,000 computing cores, the WSE-3 achieves an unprecedented 21 Petabytes per second of memory bandwidth—more than 6,000 times higher than an NVIDIA H100. This astronomical bandwidth allows the cores to process tokens at the absolute physical speed of light, generating 1,800+ tokens per second on Llama 3.1 70B.
The G42 Alliance & The Geopolitical Balance
A defining element of Cerebras's commercial trajectory is its multi-hundred-million-dollar partnership with G42, the state-backed technology holding conglomerate of the United Arab Emirates. Chaired by UAE National Security Advisor Sheikh Tahnoon bin Zayed Al Nahyan, G42 contracted Cerebras to build Condor Galaxy—a global network of nine interconnected AI supercomputers delivering tens of exaflops of compute.
While this partnership catapulted Cerebras from $24 million to hundreds of millions in revenue, it also attracted intense geopolitical scrutiny from the US Department of Commerce and the Committee on Foreign Investment in the United States (CFIUS) due to historical concerns regarding Middle Eastern technology transfers to China. In response, G42 completely phased out Chinese hardware, partnered with Microsoft, and established strict US-supervised compliance enclaves, securing American regulatory approval and paving the way for Cerebras's landmark public IPO.
Cerebras Extended FAQ
What is Cerebras Systems and what makes its chips unique?
Cerebras Systems is an American semiconductor company that manufactures the Wafer-Scale Engine (WSE-3), the largest computer chip in the world. Instead of cutting a silicon wafer into small dies, Cerebras uses an entire 300mm wafer as a single massive chip with 4 trillion transistors and 900,000 cores.
Who is the CEO of Cerebras Systems?
Andrew Feldman is the co-founder and Chief Executive Officer of Cerebras Systems. He previously co-founded SeaMicro, which he sold to AMD for $334 million, before serving as an AMD corporate vice president.
What is Cerebras's annual revenue and valuation in 2026?
Cerebras generates over $250 million in annualized run-rate revenue ($250M+ ARR) and is valued over $7.0 billion following its public registration for a Nasdaq initial public offering (IPO).
How fast is the Cerebras Inference Cloud?
The Cerebras Inference Cloud generates over 1,800 to 2,100 tokens per second on Llama 3.1 70B, making it up to 20 times faster than traditional NVIDIA H100 cloud providers.
How does Cerebras solve silicon manufacturing defects?
Cerebras builds redundant cores and a reconfigurable 2D mesh interconnect on the wafer, allowing custom diagnostic firmware to automatically route around microscopic silicon manufacturing defects at power-up.
What is the CS-3 Supercomputer?
The CS-3 is Cerebras's turnkey AI supercomputing chassis that houses the WSE-3 wafer, providing closed-loop liquid cooling and 20kW of electrical power delivery in a compact 15U rack unit.
Who is Cerebras's largest commercial customer?
G42, a prominent technology conglomerate based in the United Arab Emirates, has been Cerebras's anchor customer, contracting hundreds of millions of dollars in CS-2 and CS-3 hardware for the Condor Galaxy supercomputing network.
How does Cerebras compare to NVIDIA?
While NVIDIA builds systems by connecting thousands of small GPUs with external cables, Cerebras integrates 900,000 cores on a single continuous silicon wafer, delivering 6,000x higher memory bandwidth (21 PB/s) and dramatically faster real-time inference.
Where are Cerebras chips manufactured?
Cerebras Wafer-Scale Engines are manufactured in Taiwan by TSMC (Taiwan Semiconductor Manufacturing Company) on advanced 5-nanometer process nodes using specialized wafer-scale packaging.
How many employees work at Cerebras?
Cerebras employs approximately 450 personnel across offices in Sunnyvale, San Diego, Toronto, and Tokyo.
Related Companies
- NVIDIA - Global GPU market leader and primary competitive benchmark.
- Groq - AI chip competitor developing specialized LPU inference silicon.
- AMD - Semiconductor titan and previous acquirer of the Cerebras founders' prior startup SeaMicro.
- TSMC - Exclusive semiconductor foundry fabricating the Wafer-Scale Engine.
- Qualcomm - Strategic technology partner collaborating on enterprise datacenter acceleration.
Thermodynamics & Power Delivery: Supplying 20,000 Amps to a Silicon Wafer
When computer engineers examine the Wafer-Scale Engine, the initial reaction is awe, immediately followed by thermodynamic disbelief. In traditional server motherboards, microprocessors consume a few hundred watts at 1.0 volt, delivered through standard circuit board pins. The Wafer-Scale Engine 3, however, consumes an astonishing 20 to 23 kilowatts of continuous electrical power across a single plate of silicon. Delivering this astronomical amount of energy at sub-1.0-volt core levels requires feeding more than 20,000 amperes of direct current—equivalent to the electrical current consumed by a commercial industrial welding furnace—uniformly across a fragile, millimeter-thin sheet of silicon without melting the interconnects.
Cerebras solved this extreme thermodynamic and electrical challenge through unprecedented mechanical engineering under Gary Lauterbach and Jean-Philippe Fricker. Rather than routing power from the edge of the wafer (which would cause massive resistive voltage drops), Cerebras engineered a perpendicular power delivery matrix, feeding current directly into the back of the wafer through thousands of micro-spring contacts. Concurrently, Cerebras rejected conventional heatsinks, designing an internal closed-loop liquid manifold that clamps directly onto the silicon wafer with micrometer tolerances. Chilled water flows directly across custom copper cold-plates, extracting 20,000 watts of concentrated heat in real time with near-zero thermal gradient, maintaining stable silicon temperatures even under sustained multi-exoflop training runs.
Scribe-Line Bridging: The TSMC Packaging Revolution That Made WSE Possible
In standard semiconductor manufacturing at TSMC, photolithography steppers expose silicon wafers one 'reticle' at a time (a rectangular area roughly 26mm by 33mm). Between each reticle lies a physical gap called the 'scribe line'—a dead zone where robotic diamond saws later slice the wafer into individual rectangular dies. For decades, the scribe line was an impassable physical barrier; no electronic signals could cross from one reticle field to another.
Cerebras partnered with TSMC in a historic multi-year research initiative to invent Scribe-Line Interconnect Bridging. By designing custom lithography masks that print continuous sub-micron copper wires directly across the scribe lines during the upper metallization layers of wafer fabrication, Cerebras transformed 84 individual reticle fields into a single, continuous, unbroken 2D mesh of silicon. This manufacturing breakthrough allowed 900,000 processor cores to communicate with their neighbors across reticle boundaries with sub-nanosecond latency, creating the first truly monolithic, unified wafer-scale computational fabric in human history.
Real-Time AI Voice Agents: Why 1,800 Tokens/Sec Unlocks True Conversational Telephony
In the commercial enterprise market, the most transformative consequence of Cerebras's wafer-scale architecture is not model training, but real-time inference speed. When interacting with an artificial intelligence voice agent over a commercial telephone network, human conversation requires latency under 400 milliseconds. If the AI model takes two to three seconds to think, generate words, and stream audio back, the conversational rhythm is destroyed, resulting in awkward pauses, mutual interruptions, and frustrated users.
On traditional NVIDIA H100 cloud clusters, large language models like Llama 3.1 70B generate between 40 and 80 tokens per second. While acceptable for reading text on a screen, 60 tokens per second is far too slow for complex real-time voice synthesis and reasoning. The Cerebras Inference Cloud shatters this limitation, delivering sustained generation speeds exceeding 1,800 to 2,100 tokens per second. At 2,000 tokens per second, a 70B model generates an entire 300-word conversational response in roughly 150 milliseconds. This enables AI voice agents to execute real-time reasoning, query external databases, verify facts, and respond verbally with zero perceptible hesitation, unlocking true human-grade voice telephony for customer service, telemedicine, and emergency dispatch.
Scientific Supercomputing: Accelerated Nuclear Fusion and Cancer Genomics
Beyond commercial enterprise chatbots, the Wafer-Scale Engine has revolutionized high-performance scientific computing. At premier US national laboratories—including Argonne National Laboratory and Lawrence Livermore National Laboratory—Cerebras supercomputers are deployed on grand-challenge computational science problems that previously required weeks of execution on supercomputers occupying entire football fields.
At Argonne, researchers deployed Cerebras systems to simulate the molecular dynamics of cancer drug candidates, screening billions of small-molecule chemical compounds against oncological proteins hundreds of times faster than traditional GPU clusters. Similarly, in nuclear fusion energy research, Cerebras systems simulate the magnetohydrodynamics of plasma turbulence inside tokamak reactors in real time, allowing physicists to calculate magnetic field adjustments to prevent plasma disruptions before they occur. This ability to execute scientific partial differential equations and neural network simulations at wafer-scale speeds positions Cerebras at the forefront of 21st-century scientific discovery.
The Death of Megatron: Eliminating Distributed MPI Parallelism with CSoft
In conventional GPU supercomputing clusters, training a large foundation model is notoriously painful. Because a 70-billion or 405-billion parameter model cannot fit into the memory of a single GPU, machine learning engineers must manually divide the model using complex distributed parallelization frameworks like Megatron-LM, DeepSpeed ZeRO-3, and pipeline parallelism. Engineers spend months writing custom Message Passing Interface (MPI) code, debugging pipeline bubbles, and tuning collective communication all-reduce schedules. If a single network link or GPU node fails in a 16,000-GPU cluster, the entire training run crashes.
Cerebras eliminated this distributed engineering nightmare through its turnkey CSoft compiler platform. Because the Wafer-Scale Engine operates as a single massive continuous computational canvas with 900,000 cores, CSoft compiles standard, unaltered PyTorch neural network graphs directly into the physical hardware. The compiler automatically maps layers, attention heads, and activation matrices across the 2D mesh of cores without requiring a single line of distributed parallel code from developers. To the machine learning researcher, training a model on a multi-exoflop Cerebras supercomputer feels exactly like training a small model on a single local workstation, saving thousands of engineering hours and eliminating cluster synchronization errors.
Why Real-Time Code Refactoring Requires 2,000 Tokens Per Second
Software development is undergoing a historic transformation toward autonomous, agentic coding environments. In early AI-assisted IDEs, code generation was limited to single-line tab autocomplete, which required generating only five to ten tokens. However, the next generation of autonomous software engineering tools—such as automated code refactoring, full test-suite generation, and architectural migration—requires generating tens of thousands of tokens of code across multi-file repositories.
On traditional GPU clouds operating at 50 tokens per second, generating a complete 2,000-line software module takes nearly a minute, completely breaking the developer's cognitive flow. On the Cerebras Inference Cloud operating at over 1,800 to 2,100 tokens per second, that entire 2,000-line module is compiled, synthesized, and streamed into the developer's IDE in under one second. This instantaneous code generation enables automated agentic loops: the AI agent can write code, run tests, detect compiler errors, refactor bugs, and re-execute tests multiple times within three seconds, transforming AI coding assistants from sluggish autocomplete widgets into instantaneous peer-programming partners.