Experimental Evaluation of Big Data Analytical Tools

Due to the extensive use of SQL, the number of SQL-on-Hadoop systems has significantly increased, transforming Big Data Analytics in a more accessible practice and allowing users to perform ad-hoc querying and interactive analysis. Therefore, it is of upmost importance to understand these querying tools and the specific contexts in which each one of them can be used to accomplish specific analytical needs. Due to the high number of available tools, this work performs a performance evaluation, using the well-known TPC-DS benchmark, of some of the most popular Big Data Analytical tools, analyzing in more detail the behavior of Drill, Hive, HAWQ, Impala, Presto, and Spark.