Enhanced Dataflow Features for Streaming and Machine Learning Workloads
The world of artificial intelligence is advancing rapidly, and at Google Cloud, we’re dedicated to providing top-notch infrastructure to support your AI and ML workloads. One essential component of Google Cloud’s AI stack is Dataflow, which allows you to create batch and streaming pipelines for a wide range of analytics and AI applications. We’re thrilled to introduce a host of new features and capabilities that offer you more options, improved accessibility, and enhanced efficiency for running your batch and streaming ML workloads.
When it comes to hardware, we understand that not all ML workloads are the same. That’s why we’re expanding our hardware offerings to provide you with the flexibility to select the best accelerators for your specific requirements.
For instance, we’re continuously adding the latest GPUs to our lineup, including support for H100 and H100 Mega GPUs. These cutting-edge GPUs can greatly accelerate your AI inference workloads, allowing leading businesses to power innovative customer experiences, such as document translation for threat intelligence platform provider Flashpoint and at-scale podcast previews for media provider Spotify.
In addition to GPUs, our support for Tensor Processing Units (TPUs) like V5E, V5P, and V6E offers a powerful and cost-effective solution for large-scale ML tasks. These TPUs enable state-of-the-art ML builders to efficiently handle high-volume, low-latency machine learning inference workloads within their Dataflow jobs.
Accessing the necessary hardware resources when you need them is crucial to keeping your ML projects on track. That’s why we’ve introduced new ways to consume accelerators, making it easier than ever to obtain the resources you require.
You can now reserve GPUs and TPUs for your Dataflow jobs, ensuring that you have the necessary resources available when you need them, particularly for critical workloads that can’t afford to wait for resources to become available.
Moreover, our flex-start GPU provisioning, powered by Dynamic Workload Scheduler (DWS), helps streamline the process for batch jobs with flexible start times. By queuing your job and automatically starting it as soon as the required GPUs become available, we eliminate the need for manual resubmissions, mitigate stockout risk, and boost developer productivity.
Efficiency is key when it comes to running AI workloads, so we’re continuously working on enhancements to help you optimize your investments in AI. Features like ML-aware streaming and right fitting are designed to maximize efficiency and utilization of resources.
By making our streaming engine ML-aware, we’ve enabled smarter decision-making for executing your streaming ML pipelines. GPU-based autoscaling, for example, leverages GPU-related signals to enhance efficiency in streaming ML jobs by fine-tuning horizontal autoscaling decisions.
Additionally, right fitting allows you to use heterogeneous resource pools for different stages of your Dataflow pipelines, optimizing resource allocation and reducing costs by utilizing specialized hardware for compute-intensive stages and generic workers for less resource-intensive tasks.
We’re incredibly excited about the ML capabilities of Dataflow and the possibilities they unlock for our customers. Start using Dataflow today and leverage these new features to tackle your most challenging ML tasks.

