Google Cloud TPU Inference
Google CloudGoogle Cloud TPU Inference serves models on Google's Tensor Processing Units (TPUs), spanning the v5e, v6e (Trillium), and Ironwood (TPU7x) generations. The Pathways distributed runtime integrates with the JetStream serving engine to partition models across multi-host TPU pods, with disaggregated serving for large language models. Google offers it as a managed service under AI Hypercomputer, deployed through GKE.
Verified 2026-08-09 via the Google Cloud AI Hypercomputer inference blog post and the TPU7x (Ironwood) documentation.
Openness
1 high confidence- source
- closed
- license
- proprietary
- runtime
- Pathways(Google-internal distributed runtime, not open sourced)
- hardware-lock
- Google TPU-only
- managed
- Google Cloud/GKE
- note
- JetStream(separate, Apache-2.0 OSS) is the open serving frontend but the scored product is the proprietary Pathways/TPU runtime
The named product (Pathways / TPU Runtime managed inference) is proprietary, TPU-locked, and cloud-delivered. JetStream (open) is a distinct artifact and not what this SKU is; scored closed.
- https://cloud.google.com/blog/products/compute/ai-hypercomputer-inference-updates-for-google-cloud-tpu-and-gpu recorded 2026-08-09
Available for the first time for Google Cloud customers, Google's Pathways runtime is now integrated into JetStream, enabling multi-host inference and disaggregated serving
Adoption
3 low confidencePowers Google's own internal and Cloud TPU inference through AI Hypercomputer and GKE, but no standalone usage figure is published for the Pathways runtime offering. Level 3 reflects reported managed-cloud traction, supported by the Osmos customer account of a scaled Trillium/JetStream production deployment.
- https://cloud.google.com/blog/products/compute/ai-hypercomputer-inference-updates-for-google-cloud-tpu-and-gpu recorded 2026-08-09
Osmos is building the world's first AI Data Engineer...We have vLLM and JetStream in scaled production deployment on Trillium and are able to achieve industry leading performance
Capability
4 medium confidenceHigh-throughput multi-host TPU inference: Google reports 1,703 tokens per second on Llama 3.1 405B running on Trillium, and roughly 3x the inference per dollar of TPU v5e, with multi-host disaggregated serving. Strong, but TPU-locked, and the figures are Google's own rather than an independent MLPerf submission for the runtime itself.
- https://cloud.google.com/blog/products/compute/ai-hypercomputer-inference-updates-for-google-cloud-tpu-and-gpu recorded 2026-08-09
With multi-host inference, JetStream achieves 1703 token/s on Llama 3.1 405B on Trillium. This translates to three times more inference per dollar compared to TPU v5e.
Verified 2026-08-09