Benchmarking Spark Machine Learning Using BigBench

详细信息查看全文

关键词：Collaborative filtering using machine learning ; Predicting accuracy of data sets ; Visualization of bigbench machine learning queries using SPSS
刊名：Lecture Notes in Computer Science
出版年：2017
出版时间：2017
年：2017
卷：10080
期：1
页码：45-60
丛书名：Performance Evaluation and Benchmarking. Traditional - Big Data - Interest of Things
ISBN：978-3-319-54334-5
卷排序：10080

文摘

Databases such as dashDB are adding High Speed Connectors for Spark to efficiently extract large volumes of data. This allows them to be combined with other unstructured data sources and perform Machine Learning (ML) on top of it. Machine Learning is a key ingredient for such use cases. In order to assess performance of the data connectors and machine language frameworks, we sought benchmarks that have the ability to scale the size of datasets to very large volumes and apply Machine Learning algorithms. After exploring several options, we found BigBench to be a good fit. In this paper, we talk about our experiences of using BigBench with special focus on its 5 Machine Learning queries and their default implementation in Spark. We discuss on how we could improve effectiveness of BigBench for benchmarking Machine Learning by avoiding bias and inclusion of real time analytics. We also think that there is scope for improving the coverage of Machine Learning by adding more use cases like Collaborative Filtering. Lastly, we share some interesting visualization of 4 ML queries using SPSS Modeler and our experiments on different Clustering and Classification algorithms.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700