- Conference Article
- 10.1109/iccca66364.2025.11325411
Query Performance Analysis using Apache Pig and Hive in Hadoop Environment
- Nov 28, 2025
- Shashi Shekhar Kumar + 4 more +4
Query execution is a challenging task in the big data environment, as the execution might not always be easy to handle. Some big data frameworks provide efficient ways to execute queries with the most optimized approach. These queries are based on the various operators used with Apache Pig and Apache Hive. Since these are used for performing query execution inside the Hadoop ecosystem, which contains large volumes of data in a distributed environment, both are built on top of the Map Reduce framework for query processing. Consequently, Apache Hive focuses on SQL-like queries, while Apache Pig uses a programmatic approach for executing Pig Latin script. In this research work, we propose a query execution-based approach using various operators of Pig and Hive. We wrote simple and complex queries based on an IoT dataset and analyzed the performance in terms of execution time for both components. The comparative analysis shows that simple and complex queries can be efficiently handled using the Apache Hadoop ecosystem. The hive performed better than the pig for various types of queries, while system resources differ based on types of queries.
Read more