Groq was founded in 2016 in Mountain View, California by Jonathan Ross and Douglas Wightman. Ross was a pioneering systems engineer at Google who co-created the original Tensor Processing Unit (TPU v1), the custom ASIC that powered Google's search algorithms and AlphaGo. While at Google, Ross realized that the semiconductor industry was trapped in a GPU paradigm: processors were burdened by complex dynamic branch prediction, speculative execution hardware, and slow external memory buses that created massive latency bottlenecks. Ross believed that the future of computing was sequential dataflow streaming, where a smart compiler dictates the exact physical path of every piece of data through silicon. Leaving Google in 2016, Ross and Wightman founded Groq to build a processor from the software down. They spent nearly eight years in quiet research, engineering a single-core streaming processor equipped with 230MB of on-chip SRAM and a deterministic compiler. When generative large language models emerged in 2023, the world suddenly faced a severe inference latency crisis. When Groq demonstrated Llama 2 and Llama 3 running at over 800 tokens per second in early 2024, the tech industry was stunned, instantly propelling Groq into the forefront of global AI infrastructure. In early 2016, Jonathan Ross walked away from Google despite having just co-founded and built the Google TPU, one of the most successful silicon projects in the company's history. Meeting in Silicon Valley coffee shops with Douglas Wightman, Ross sketched the architecture for a processor that had zero runtime hardware schedulers—a chip where every single transistor was dedicated to pure arithmetic computation, directed with clock-cycle precision by a mathematical compiler. When the first Tensor Streaming Processors returned from the foundry and powered on with 100% deterministic execution on the very first day, the founders knew they had invented the blueprint for the future of artificial intelligence.