AI Potluck
Model components / Inference code

Google Cloud TPU Inference

Google Cloud

Google Cloud TPU Inference serves models on Google's Tensor Processing Units (TPUs), spanning the v5e, v6e (Trillium), and Ironwood (TPU7x) generations. The Pathways distributed runtime integrates with the JetStream serving engine to partition models across multi-host TPU pods, with disaggregated serving for large language models. Google offers it as a managed service under AI Hypercomputer, deployed through GKE.

Verified 2026-08-09 via the Google Cloud AI Hypercomputer inference blog post and the TPU7x (Ironwood) documentation.

Openness

1 high confidence
1.0
source
closed
license
proprietary
runtime
Pathways(Google-internal distributed runtime, not open sourced)
hardware-lock
Google TPU-only
managed
Google Cloud/GKE
note
JetStream(separate, Apache-2.0 OSS) is the open serving frontend but the scored product is the proprietary Pathways/TPU runtime

The named product (Pathways / TPU Runtime managed inference) is proprietary, TPU-locked, and cloud-delivered. JetStream (open) is a distinct artifact and not what this SKU is; scored closed.

Adoption

3 low confidence
3.0

Powers Google's own internal and Cloud TPU inference through AI Hypercomputer and GKE, but no standalone usage figure is published for the Pathways runtime offering. Level 3 reflects reported managed-cloud traction, supported by the Osmos customer account of a scaled Trillium/JetStream production deployment.

Capability

4 medium confidence
4.0

High-throughput multi-host TPU inference: Google reports 1,703 tokens per second on Llama 3.1 405B running on Trillium, and roughly 3x the inference per dollar of TPU v5e, with multi-host disaggregated serving. Strong, but TPU-locked, and the figures are Google's own rather than an independent MLPerf submission for the runtime itself.

Verified 2026-08-09