Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
By: Tasnuva Chowdhury , Tadashi Maeno , Fatih Furkan Akman and more
Potential Business Impact:
Predicts computer needs for science experiments.
The collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.
Similar Papers
Machine learning-based cloud resource allocation algorithms: a comprehensive comparative review
Distributed, Parallel, and Cluster Computing
Makes computers use cloud power smarter and cheaper.
Predicting the Performance of Scientific Workflow Tasks for Cluster Resource Management: An Overview of the State of the Art
Distributed, Parallel, and Cluster Computing
Helps computers guess how long jobs will take.
Artificial Intelligence for Cost-Aware Resource Prediction in Big Data Pipelines
Distributed, Parallel, and Cluster Computing
Saves money by guessing computer needs.