The Inference Layer Splinters: Purpose-Built Chips Challenge Nvidia's Universal GPU
Hot Chips 2026 was the most consequential single week in AI chip competition since Nvidia's H100 launch. OpenAI Jalapeño, Google TPU 8i, Microsoft Maia 200, SambaNova SN50, and Intel Crescent Island all presented production-track inference architectures on the same stage, each optimised for workload-specific metrics that Nvidia's general-purpose GPU handles adequately but not optimally. The economic logic is now clear: inference is a volume workload with tight cost-per-token economics, and every token served on proprietary silicon rather than Nvidia hardware is a margin point that stays inside the hyperscaler's P&L. OpenAI's Jalapeño — deployed at scale in its own infrastructure, not a research exercise — is the sharpest expression of this logic. Its 700W TDP against Nvidia's 1,400W GB300 directly addresses the energy cost constraint that is becoming the binding limit on inference fleet economics.
Nvidia's response is visible in two moves disclosed at the same conference: the integration of acquired Groq LPU technology for the decode phase of inference in heterogeneous Vera Rubin clusters, and SpaceX AI's commitment to deploy standalone Vera CPUs for Grok's agentic workloads. These are defensive architecture moves — Nvidia using its acquisition to protect latency benchmarks against exactly the claims OpenAI is making. The broader signal is structural: Nvidia's training infrastructure lead remains substantial and is not under near-term threat, but the inference segment — now the dominant and fastest-growing AI compute workload — is being carved up from below by hyperscalers who control their own silicon and pay no GPU margins on it.