Workflow performance improvement using model-based scheduling over multiple clusters and clouds

详细信息查看全文

作者：Ketan Maheshwari^a ; ^{ketan@anl.gov" class="auth_mail" title="E-mail the corresponding author}Author Vitae ; Eun-Sung Jung^a ; ^{esjung@mcs.anl.gov" class="auth_mail" title="E-mail the corresponding author}Author Vitae ; Jiayuan Meng^b ; ^{meng.jiayuan@gmail.com" class="auth_mail" title="E-mail the corresponding author}Author Vitae ; Vitali Morozov^b ; ^{morozov@anl.gov" class="auth_mail" title="E-mail the corresponding author}Author Vitae ; Venkatram Vishwanath^b ; ^{venkatv@mcs.anl.gov" class="auth_mail" title="E-mail the corresponding author}Author Vitae ; Rajkumar Kettimuthu^a ; ^{kettimut@mcs.anl.gov" class="auth_mail" title="E-mail the corresponding author}Author Vitae
关键词：System modeling ; Workflow ; Optimization ; Swift ; Clouds
刊名：Future Generation Computer Systems
出版年：2016
出版时间：January 2016
年：2016
卷：54
期：Complete
页码：206-218
全文大小：1891 K

文摘

In recent years, a variety of computational sites and resources have emerged, and users often have access to multiple resources that are distributed. These sites are heterogeneous in nature and performance of different tasks in a workflow varies from one site to another. Additionally, users typically have a limited resource allocation at each site capped by administrative policies. In such cases, judicious scheduling strategy is required in order to map tasks in the workflow to resources so that the workload is balanced among sites and the overhead is minimized in data transfer. Most existing systems either run the entire workflow in a single site or use naïve approaches to distribute the tasks across sites or leave it to the user to optimize the allocation of tasks to distributed resources. This results in a significant loss in productivity. We propose a multi-site workflow scheduling technique that uses performance models to predict the execution time on resources and dynamic probes to identify the achievable network throughput between sites. We evaluate our approach using real world applications using the Swift parallel and distributed execution framework. We use two distinct computational environments-geographically distributed multiple clusters and multiple clouds. We show that our approach improves the resource utilization and reduces execution time when compared to the default schedule.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700