• Home
  • Search
  • Distributed frequent hierarchical pattern mining for robust and efficient large-scale association discovery
  • Cite Icon1
  • https://doi.org/10.32469/10355/63867Copy DOI Icon

Distributed frequent hierarchical pattern mining for robust and efficient large-scale association discovery

  • May 1, 2017
  • Michael Phinney
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Frequent pattern mining is a classic data mining technique, generally applicable to a wide range of application domains, and a mature area of research. The fundamental challenge arises from the combinatorial nature of frequent itemsets, scaling exponentially with respect to the number of unique items. Apriori-based and FPTree-based algorithms have dominated the space thus far. Initial phases of this research relied on the Apriori algorithm and utilized a distributed computing environment; we proposed the Cartesian Scheduler to manage Apriori's candidate generation process. To address the limitation of bottom-up frequent pattern mining algorithms such as Apriori and FPGrowth, we propose the Frequent Hierarchical Pattern Tree (FHPTree): a tree structure and new frequent pattern mining paradigm. The classic problem is redefined as frequent hierarchical pattern mining where the goal is to detect frequent maximal pattern covers. Under the proposed paradigm, compressed representations of maximal patterns are mined using a top-down FHPTree traversal, FHPGrowth, which detects large patterns before their subsets, thus yielding significant reductions in computation time. The FHPTree memory footprint is small; the number of nodes in the structure scales linearly with respect to the number of unique items. Additionally, the FHPTree serves as a persistent, dynamic data structure to index frequent patterns and enable efficient searches. When the search space is exponential, efficient targeted mining capabilities are paramount; this is one of the key contributions of the FHPTree. This dissertation will demonstrate the performance of FHPGrowth, achieving a 300x speed up over state-of-the-art maximal pattern mining algorithms and approximately a 2400x speedup when utilizing FHPGrowth in a distributed computing environment. In addition, we allude to future research opportunities, and suggest various modifications to further optimize the FHPTree and FHPGrowth. Moreover, the methods we offer will have an impact on other data mining research areas including contrast set mining as well as spatial and temporal mining.

Similar Papers
  • Research Article
  • Citations19

Closed frequent similar pattern mining: Reducing the number of frequent similar patterns without information loss

  • Dec 09, 2017
  • Expert Systems with Applications
  • Ansel Y Rodríguez-González +5
  • Conference Article
  • Citations6

An improved algorithm for frequent patterns mining problem

  • May 01, 2010
  • Thanh-Trung Nguyen
  • Research Article

Research on Combining Pattern Mining and Evolutionary Algorithm for Critical Node Detection Problems

  • Jan 01, 2022
  • International Journal of Frontiers in Engineering Technology
  • Hongyuan Ding +3
  • Conference Article

Assessing the Impact of a New Released Medicine Towards Medication Strategy Using Graph Based Visualization

  • Oct 01, 2019
  • Purnomo Husnul Khotimah +5
  • Research Article
  • Citations5

Prefix-트리를 이용한 동적 가중치 빈발 패턴 탐색 기법

  • Aug 31, 2010
  • The KIPS Transactions:PartD
  • Byeong-Soo Jeong +1
  • Research Article
  • Citations21

Handling Dynamic Weights in Weighted Frequent Pattern Mining

  • Nov 01, 2008
  • IEICE Transactions on Information and Systems
  • C F Ahmed +3
  • Conference Article
  • Citations33

Musk: Uniform Sampling of k Maximal Patterns

  • Apr 30, 2009
  • Mohammad Al Hasan +1
  • Research Article
  • Citations6

Efficient Top-K Identical Frequent Itemsets Mining without Support Threshold Parameter from Transactional Datasets Produced by IoT-Based Smart Shopping Carts

  • Oct 21, 2022
  • Sensors (Basel, Switzerland)
  • Saif Ur Rehman +4
  • PDF
  • Research Article
  • Citations25

Computational annotation of UTR cis-regulatory modules through Frequent Pattern Mining

  • Jun 01, 2009
  • BMC Bioinformatics
  • Antonio Turi +5
  • Research Article
  • Citations1

MINING TOP-K FREQUENT SEQUENTIAL PATTERN IN ITEM INTERVAL EXTENDED SEQUENCE DATABASE

  • Nov 23, 2018
  • Journal of Computer Science and Cybernetics
  • Duong Huy Tran +3
  • Book Chapter
  • Citations1

Using Non Boolean Similarity Functions for Frequent Similar Pattern Mining

  • Jan 01, 2010
  • Ansel Y Rodríguez-González +3
  • Research Article
  • Citations44

Similarity analysis of frequent sequential activity pattern mining

  • Sep 22, 2018
  • Transportation Research Part C: Emerging Technologies
  • Zhenyu Shou +1
  • Book Chapter
  • Citations7

Constrained Frequent Pattern Mining from Big Data Via Crowdsourcing

  • Aug 17, 2018
  • Calvin S H Hoi +2
  • Conference Article
  • Citations9

Incremental Mining of Frequent Query Patterns from XML Queries for Caching

  • Dec 01, 2006
  • Proceedings
  • Guoliang Li +4
  • Research Article
  • Citations21

Identifying Environmental and Human Factors Associated With Tick Bites using Volunteered Reports and Frequent Pattern Mining

  • May 13, 2016
  • Transactions in GIS
  • Irene Garcia‐Martí +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.