• Home
  • Search
  • Automating Data Science Pipelines with Tensor Completion
  • Open Access IconOpen Access
  • Cite Icon3
  • https://doi.org/10.1109/bigdata62323.2024.10825934Copy DOI Icon

Automating Data Science Pipelines with Tensor Completion

  • Dec 15, 2024
  • Shaan Pakala +7 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Hyperparameter optimization is an essential component in many data science pipelines and typically entails exhaustive time and resource-consuming computations in order to explore the combinatorial search space. Similar to this problem, other key operations in data science pipelines exhibit the exact same properties. Important examples are: neural architecture search, where the goal is to identify the best design choices for a neural network, and query cardinality estimation, where given different predicate values for a SQL query the goal is to estimate the size of the output. In this paper, we abstract away those essential components of data science pipelines and we model them as instances of tensor completion, where each variable of the search space corresponds to one mode of the tensor, and the goal is to identify all missing entries of the tensor, corresponding to all combinations of variable values, starting from a very small sample of observed entries. In order to do so, we first conduct a thorough experimental evaluation of existing state-of-the-art tensor completion techniques and introduce domain-inspired adaptations (such as smoothness across the discretized variable space) and an ensemble technique which is able to achieve state-of-the-art performance. We extensively evaluate existing and proposed methods in a number of datasets generated corresponding to (a) hyperparameter optimization for nonneural network models, (b) neural architecture search, and (c) variants of query cardinality estimation, demonstrating the effectiveness of tensor completion as a tool for automating data science pipelines. Furthermore, we release our generated datasets and code in order to provide benchmarks for future work on this topic.

Similar Papers
  • Research Article
  • Citations43

Improving reproducibility of data science pipelines through transparent provenance capture

  • Aug 01, 2020
  • Proceedings of the VLDB Endowment
  • Lukas Rupprecht +4
  • PDF
  • Conference Article
  • Citations61

The art and practice of data science pipelines

  • May 21, 2022
  • Sumon Biswas +2
  • Research Article
  • Citations2

Behavior Matters: An Alternative Perspective on Promoting Responsible Data Science

  • May 02, 2025
  • Proceedings of the ACM on Human-Computer Interaction
  • Ziwei Dong +4
  • Research Article
  • Citations17

Exploiting Global Low-Rank Structure and Local Sparsity Nature for Tensor Completion

  • Jul 24, 2018
  • IEEE Transactions on Cybernetics
  • Yong Du +6
  • Research Article
  • Citations8

Supporting Better Insights of Data Science Pipelines with Fine-grained Provenance

  • Apr 10, 2024
  • ACM Transactions on Database Systems
  • Adriane Chapman +3
  • Conference Instance
  • Citations33

Proceedings of the 2022 International Conference on Management of Data

  • Jun 10, 2022
  • Peter Boncz +99
  • Research Article
  • Citations15

Xel: A cloud-agnostic data platform for the design-driven building of high-availability data science services

  • Mar 17, 2023
  • Future Generation Computer Systems
  • J Armando Barron-Lugo +4
  • Research Article
  • Citations4

Large scale analysis of the SARS-CoV-2 main protease reveals marginal presence of nirmatrelvir-resistant SARS-CoV-2 Omicron mutants in Ontario, Canada, December 2021–September 2023

  • Oct 03, 2024
  • Canada Communicable Disease Report
  • Venkata Duvvuri +21
  • PDF
  • Research Article
  • Citations4

VeridicalFlow: a Python package for building trustworthy data science pipelines with PCS

  • Jan 12, 2022
  • The Journal of Open Source Software
  • James Duncan +4
  • PDF
  • Research Article
  • Citations37

Classification Framework of the Bearing Faults of an Induction Motor Using Wavelet Scattering Transform-Based Features.

  • Nov 19, 2022
  • Sensors
  • Rafia Nishat Toma +7
  • Book Chapter
  • Citations3

Crowdsourcing and Human‐in‐the‐Loop forIoT

  • Mar 06, 2020
  • The Internet of Things
  • Luis‐Daniel Ibáñez +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.