We hereby declare September to be scalability month! As the world prepares for a surge of agentic fleets, we are shoring up our AI infrastructure and orchestration offerings to gracefully — and quickly — respond to that demand, all while maintaining workload isolation and security, and keeping costs in check. Read on to learn how these enhancements manifest across Google Cloud’s compute, network, storage, and orchestration offerings, plus new ways customers are using Google Cloud AI infrastructure, and third-party industry validation of our strategy. Product, technology, and tools updates Google Kubernetes Engine updates: The GKE team is all about improving the scalability of the platform, and in September, those improvements came in many shapes and sizes: New feature: Need an execution runtime with higher density for your agentic workloads? We engineered the new open-source GKE Agent Substrate to run millions of sandboxes with 10x higher density than standard container runtimes. Agent Substrate also delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation. Product update: GKE now has scale-to-zero capabilities built-in. No need to configure complex components to scale your workloads down, thanks to the HPA with the Autoscaling Metric and support for KEP-2021, which do the job for you, out of the box. Read the blog to learn more. Product update: Further, the GKE HPA (with the above-mentioned Autoscaling Metric) now lets you scale up and down based on custom PromQL metrics, in addition to standard metrics, allowing you to trigger workloads according to conditions that are meaningful and unique to your business. Read more here. New feature: Yet another scalability feature is GKE Pod snapshots, which lets you save the running state of your workload, including CPU and GPU memory, and restore it on demand. According to internal tests, GKE Pod snapshots can reduce AI inference start-up by as much as 89%. Learn more here. New migration tool: Finally, if you’ve always wanted to migrate your container workloads from AWS EKS to GKE but feared a daunting, high-friction engineering endeavor, we’ve just launched GKE agentic migration, a purpose-built agent plugin that replaces brittle, ad-hoc prompting with an AI-assisted migration pipeline protected by deterministic guardrails. Designed as a compilation of agent skills and a local Model Context Protocol (MCP) server, it uses AI to translate complex AWS EKS IaC and Kubernetes manifests directly into GKE landing zones. Get started with the onboarding guide. Feature updates: Reinforcement learning (RL) and evaluation workloads are a beast: In a standard agentic RL loop, an LLM policy generates actions like code snippets on GPUs and executes them inside isolated CPU sandboxes to observe a reward signal. However, when scaling up this loop to support tens of thousands of parallel rollouts, infrastructure bottlenecks emerge, for instance idle accelerators, image cardinality, and a saturated control plane. To help, we developed GKE Agent Sandbox optimized for RL, plus an Agent Sandbox RL orchestration SDK and native integrations for popular RL gyms and harnesses. All are now generally available, and you can learn more here. Storage updates: AI trains and creates lots of data, and that data has to live somewhere — in block storage systems, file systems, object stores and databases. We announced enhancements to our storage portfolio to help this critical layer roll with the agentic punches: Product update: Filestore agent volumes offer high-performance, elastic, persistent file storage for agentic workloads. Thanks to its tight integration with GKE Agent Substrate and GKE Agent Sandbox, Filestore agent volumes automatically allocates and attaches a dedicated, isolated file workspace to GKE agent sandboxes in milliseconds. Request access to the preview here. New product: If you run generative AI and RAG data layers —