Data lakes serve as repositories that house millions of tables covering a wide range of subjects. Data scientists use tables from data lakes to support decision-making processes, train machine learning models, perform statistical analysis, and more. However, the enormous size and heterogeneity of data lakes, together with missing or unreliable metadata, make it challenging for data scientists to find suitable tables for their analyses. Furthermore, after table discovery, another challenge is to integrate the discovered tables together to get a unified view of data, enabling queries that go beyond a single table. In this thesis, we introduce novel techniques for table discovery and integration tasks. First, we present SANTOS which helps data scientists discover tables from a data lake that union with their query table and possibly extend it with more rows. Contrary to existing techniques that only use the semantics of columns to perform union search, we improve the search accuracy by also using semantic relationships between pairs of columns in a table. Consequently, we present a new notion of unionability that considers relationships between columns, together with the semantics of columns, in a principled way. We show that SANTOS outperforms a state-of-the-art union search technique that uses a wide variety of column-based semantics, including word embeddings and regular expressions. Our empirical analysis is over real data lake benchmarks. Second, we present ALITE (Align and Integrate), the first proposal for scalable integration of tables that may have been discovered using join, union, or related table search. ALITE aligns the columns of discovered tables using holistic schema matching and then applies Full Disjunction, an associative version of outer join, to get an integrated table. ALITE relaxes previous assumptions required for full disjunction that tables share common attribute names (which completely determine the join columns), that the tables are complete (without null values), and that the integration have an acyclic join pattern. Our contributions include a new hierarchical-clustering-based holistic schema matching algorithm that does not require reliable metadata and performs well on tables from real open data. Furthermore, we also present an algorithm that adopts complementation and subsumption operators to compute the Full Disjunction. We show that our algorithm is faster in practice on real tables than previous implementations of Full Disjunction. Third, we describe DIALITE, a novel table discovery pipeline that allows users to discover, integrate, and analyze open data tables. DIALITE is flexible such that the user can easily add and compare additional discovery and integration algorithms. Lastly, we identify and address the new problem of searching for diverse tuples from a data lake that can be unioned with the given query table. We propose the DUST system as a solution to the problem. DUST uses a novel embedding model to represent unionable tuples that outperforms other tuple representation models by at least 15 % when representing unionable tuples. Within DUST, we present an efficient algorithm that diversifies a set of candidate unionable tuples. Using real data lake benchmarks, we show that our diversification algorithm is more than six times faster than the most efficient diversification baseline. We also show that it is more effective in diversifying unionable tuples than existing diversification algorithms. As part of our contributions, we have released new open benchmarks for table discovery and integration tasks to further research in this area.--Author's abstract
Read more