- Research Article
16
- 10.1016/j.procs.2016.03.128
Integration of Spark framework in Supply Chain Management
- Jan 01, 2016
- Procedia Computer Science
- Harjeet Singh Jaggi + 1 more +1
Integration of Spark framework in Supply Chain Management
The Smith-Waterman (SW) algorithm is universally used for a database search owing to its high sensitively. The widespread impact of the algorithm is reflected in over 8000 citations that the algorithm has received in the past decades. However, the algorithm is prohibitively high in terms of time and space complexity, and so poses significant computational challenges. Apache Spark is an increasingly popular fast big data analytics engine, which has been highly successful in implementing large-scale data-intensive applications on commercial hardware. This paper presents the first ever reported system that implements the SW algorithm on Apache Spark based distributed computing framework, with a couple of off-the-shelf workstations, which is named as SparkSW. The scalability and load-balancing efficiency of the system are investigated by realistic ultra-large database from the state-of-the-art UniRef100. The experimental results indicate that 1) SparkSW is load-balancing for parallel adaptive on workloads and scales extremely well with the increases of computing resource, 2) SparkSW provides a fast and universal option high sensitively biological sequence alignments. The success of SparkSW also reveals that Apache Spark framework provides an efficient solution to facilitate coping with ever increasing sizes of biological sequence databases, especially generated by second-generation sequencing technologies.
Integration of Spark framework in Supply Chain Management
Integration of Spark framework in Supply Chain Management
Abstract 6391: Scalable colorectal cancer (CRC) diagnosis model with arterial and unenhanced phases of KRAS mutational status using Apache Spark
The exponential increase in Medical data and Computer-aided diagnostic tools has made it easier for Machine Learning (ML) Researchers to extract valuable insights from medical images resulting in better patient outcomes during Medical Summarization. According to current statistics, 8.8 million people died every year due to cancer death. It is therefore imperative to use these available image datasets for better diagnosis, increased patient outcomes and recovery from deadly diseases. As a result, this study proposes a Colorectal Cancer classification model for detecting KRAS mutation status of Patients through their Radiomic images using the transfer learning of a deep learning pipeline on Apache Spark. The datasets will be curated from National Cancer Institute (NCI), features extraction or classification of Computed tomography(CT) images into CT- based handcrafted radiomic signatures with a Convolutional Neural Network(CNN) using the Apache Spark framework, the system uploads the segmented scanning images to the High Distributed File System (HDFS), Confidence Interval(CI), Area Under Curve(AUC), Precision and Recall for Validation cohort, Docker for app deployment and containerization of the model developed for reproducibility, transfer learning and model reuse. The result would be a real-time model using Deep learning pipelines with Apache Spark and Keras Tensor flow for KRAS Mutation detection using Colorectal Cancer images. In conclusion, this research project would produce a scalable and reproducible model for the faster diagnosis of Colorectal Cancer through the KRAS mutation status of CRC Patients. Citation Format: Mary Adetutu Adewunmi. Scalable colorectal cancer (CRC) diagnosis model with arterial and unenhanced phases of KRAS mutational status using Apache Spark [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 6391.
Read moreDI-Mondrian: Distributed improved Mondrian for satisfaction of the L-diversity privacy model using Apache Spark
DI-Mondrian: Distributed improved Mondrian for satisfaction of the L-diversity privacy model using Apache Spark
Shared Memory Based RDD Data Sharing on Spark
Apache Spark is an increasingly popular fast big data analytics engine, which focuses on a large-scale data processing. Currently, Spark's memory management mainly oriented to single application, and do not directly support the typical scenarios of multiple data processing applications. In this paper, we propose an extension of Apache Spark, called Shared Memory Spark (SMSpark). SMSpark introduces shared memory based on RDD data sharing between applications. We conducted experiment in a cluster using two typical applications. Experimental results show that compared to Spark, SMSpark gains better performance.
Read moreAn Overview of Hardware-Based Acceleration of Biological Sequence Alignment
Efficient biological sequence (proteins or DNA) alignment is an important and challenging task in bioinformatics. It is similar to string matching in the context of biological data and is used to infer the evolutionary relationship between a set of protein or DNA sequences. An accurate alignment can provide valuable information for experimentation on the newly found sequences. It is indispensable in basic research as well as in practical applications such as pharmaceutical development, drug discovery, disease prevention and criminal forensics. Many algorithms and methods, such as, dot plot (Gibbs & McIntyre, 1970),Needleman-Wunsch (N-W) (Needleman & Wunsch, 1970), Smith-Waterman (S-W) (Smith & Waterman, 1981), FASTA (Pearson & Lipman, 1985), BLAST (Altschul et al., 1990), HMMER (Eddy, 1998) and ClustalW (Thompson et al., 1994) have been proposed to perform and accelerate sequence alignment activities. An overview of these methods is given in (Hasan et al., 2007). Out of these, S-W algorithm is an optimal sequence alignment method, but its computational cost makes it inappropriate for practical purposes. To develop efficient and optimal sequence alignment solutions, the S-W algorithm has recently been implemented on emerging accelerator platforms such as Field Programmable Gate Arrays (FPGAs), Cell Broadband Engine (Cell/B.E.) and Graphics Processing Units (GPUs) (Buyukkur & Najjar, 2008; Hasan et al., 2010; Liu et al., 2009; 2010; Lu et al., 2008). This chapter aims at providing a broad overview of sequence alignment in general with particular emphasis on the classification and discussion of available methods and their comparison. Further, it reviews in detail the acceleration approaches based on implementations on different platforms and provides a comparison considering different parameters. This chapter is organized as follows: The remainder of this section gives a classification, discussion and comparison of the available methods and their hardware acceleration. Section 2 introduces the S-W algorithm which is the focus of discussion in the succeeding sections. Section 3 reviews CPU-based acceleration. Section 4 provides a review of FPGA-based acceleration. Section 5 overviews GPU-based acceleration. Section 6 presents a comparison of accelerations on different platforms, whereas Section 7 concludes the chapter.
Read moreEnhancing Spam Email Detection with Machine Learning: A Comparative Study of Logistic Regression and Naive Bayes Using Apache Spark
The spread of spam emails presents serious problems for both email security and user experience. This research aims to develop an effective spam email classification system utilizing machine learning techniques, specifically Logistic Regression and Naive Bayes, within the Apache Spark framework. The methodology encompasses a thorough preprocessing of the Enron email dataset. This process involves several critical steps: text cleaning to remove irrelevant information, tokenization to break down the text into individual words, removal of stop words to eliminate common but uninformative words, and text feature extraction using Term Frequency-Inverse Document Frequency (TF-IDF) to quantify the importance of terms within the dataset. The study is conducted on a subset of the Enron email dataset, comprising 11,029 emails, with 2,996 labeled as spam. Experimental results demonstrate that the Naive Bayes model outperforms Logistic Regression, achieving higher accuracy and F1 score. This finding underscores the robustness of Naive Bayes in spam email classification, highlighting its potential for enhancing email security by effectively filtering spam.
Read moreBuilding a knowledge graph by using cross-lingual transfer method and distributed MinIE algorithm on apache spark
The simplest and effective way to store human knowledge through centuries was using text. Along with the advancement of technology nowadays, the volume of text has grown to be larger and larger. To extract useful information from this amount of text becomes an exceptionally complex task. As an effort to solve that problem, in this paper, we present a pipeline to extract core knowledge from large quantity text using distributed computing. The components of our pipeline are systems that were known to yield good results. The outputs of our proposed system are stored in a knowledge graph. A knowledge graph is a graph for storing knowledge in the form of triples (head, relation, tail). Some of the existing knowledge graphs in the world are Google knowledge graph, YAGO, DBLP, or DBpedia. These knowledge graphs have one thing in common—they are in English. The English language is studied by many researchers in the world and it had become a rich-resource language (with many natural language processing tools and data set). Vietnamese, on the other hand, is a low-resource language. Therefore, we use cross-lingual transfer method to build a Vietnamese knowledge graph. Firstly, we collect data in form of text about Vietnam tourism, which was written mostly in Vietnamese, using Google search and Wikipedia. In the next step, we translate them into English with Google Translate and use English Natural Language Processing tools like Stanford Parser, Co-referencing, ClausIE, MinIE to extract useful triples from this text. Lastly, the triples are translated back to Vietnamese to build a Vietnam tourism knowledge graph. Since we are working with massive text, we develop a distributed algorithm to extract triples from sentences of massive text. This is a distributed version of MinIE, which was originally developed for a single machine model. In Apache Spark framework, we divide massive text into many smaller parts and move them to the worker nodes with distributed MinIE function. Spark distributed MinIE will extract the triples of sentences in the local text of this worker node in parallel. Finally, the result of worker nodes will be sent back to the master node for building the knowledge graph. We conduct experiments with the distributed MinIE on spark cluster to prove the outperformance of our proposed algorithm.
Read moreSearch for the Optimal Investor's Portfolio Using Spark Apache
According to the Markovitic, any investor must base its choice exclusively on the expected profitability and standard deviation when the portfolio is selected. Thus, having appreciated the various portfolio combinations, it must choose the “best”, based on the ratio of the expected profitability and the standard deviation of these portfolios. At the same time, the ratio of the profitability of the portfolio remains usual: the higher the yield, the higher the risk. The choice of the optimal portfolio is carried out taking into account two options for its orientation or on the priority receipt of income at the expense of interest and dividends, or on the increase in the course value of securities. The establishment of a combination of risk and profitability of the portfolio for the enterprise is achieved if it takes into account the rule that more income brings security, the greater the potential risk it has. Every financial asset trading system contains an algorithm inside. This algorithm helps to analyze the historical data of assets to collect an optimal portfolio for an investor. Nowadays those systems must be able to handle a vast amount of data. To be able to store and process big volumes of data frameworks, such as Apache Spark was developed. This particular framework allows handling the data distributed over many storages. However, an appropriate algorithm must be picked, because not all of them can be distributed through clusters successfully. This work focuses on two approaches for finding an optimal portfolio using risk metrics. One approach is the direct method, which maximizes expected returns concerning some level of risk, and another one minimizes the approximation of risk metric using Monte-Carlo sampling. Both of them lead to different optimization problems, solving which an optimal portfolio can be found. During this work, those approaches would be described, as well as the technics used to solve related optimization problems. Those approaches were implemented using the Apache Spark framework as a part of this work, and their efficiency will be compared through numerical experiments presented.
Read moreSpark-DIY: A Framework for Interoperable Spark Operations with High Performance Block-Based Data Models
This work was partially funded by the Spanish Ministry of Economy, Industry and Competitiveness under the grant TIN2016-79637-P ”Towards Unification of HPC and Big Data Paradigms”; the Spanish Ministry of Education under the FPU15/00422 Training Program for Academic and Teaching Staff Grant; the Advanced Scientific Computing \nResearch, Office of Science, U.S. Department of Energy, under Contract DE-AC02-06CH11357; and by DOE with agreement No. DE-DC000122495, program manager Laura Biven.
Read moreResearch on the Technical Path to Optimize the Reconstruction of Large-Scale Historical Scenes in Historical Documentaries Using Parallel Distributed Computing Techniques
The continuous development of distributed computing technology provides a new path for high-precision reconstruction of historical scenes. In this paper, we propose a technical optimization framework based on parallel distributed computing for the problems of low efficiency and lack of clarity in the reconstruction of large-scale historical scenes in historical documentaries. By integrating digital elevation model (DEM) and 3D Gaussian sputtering (3DGS) methods, combined with the DAG scheduling mechanism of Apache Spark framework and multi-factor weighted resilient distributed dataset (RDD) caching strategy, the reconstruction speed and rendering quality of the scene are significantly improved. The experiments show that the parallel computing nodes are set to 15, the nodes adopt an interval of 1500m, and the number of single file computations is 11 times to obtain higher modeling efficiency. The ambiguity of scene reconstruction with the introduction of parallel distributed computing is reduced to less than 17%, and the average ambiguity is lower than 16%. The values of three evaluation indexes are better than those of the comparison algorithms.
Read moreNovel Apache Spark based Algorithm to Solve Dirichlet Problem for Poisson Equation in 3D Computational Domain
Parallel computations are essential tool in solving large-scale computationally demanding problems. Due to large diversity and heterogeneity of the currently available parallel processing techniques and paradigms it is usually difficult to find the right solution that will perform well according to every performance metric. As one of the recent developments in parallel computing Apache Spark framework allows to process petabyte-scale data and possesses properties such as fault tolerance, scalability, load balancing and mechanisms of in memory computations across nodes of the cluster. All of these features are attractive for high performance scientific computing. It has been shown that Apache Spark outperforms Hadoop implementation of some machine learning algorithms by orders of magnitude. Since Hadoop platform is not well suited for iterative computing, typical for many computational problems, in this study we investigate performance characteristics of Apache Spark on scientific computing problems, particularly for solving Dirichlet problem for Poisson's equation. An algorithm for solving Dirichlet problem for Poisson's equation is described and analyzed and compared to optimized Hadoop-based implementations. Apache Spark uses new distributed data structure called RDD. Presented algorithm consists of operations on RDD such as mapping, grouping and partitioning. The benefits and drawbacks of the algorithm as well as applicability for stencil type computations are discussed and analyzed.
Read moreTwo Layer Hybrid Scheme of IMO and PSO for Optimization of Local Aligner: COVID-19 as a Case Study
Nowadays, meta-heuristic algorithm (MA) succeeded in optimizing many engineering problems. Ions motion optimization (IMO) algorithm is a MA that inspired its search strategy from ions attraction based on force law. IMO has good exploration capability but poor exploitation of the search space. The performance of IMO was tested for implementing fragmented local aligner technique (FLAT) which is a local aligner method for finding the longest common consecutive subsequence (LCCS) between pair of biological sequences. Due to the huge length of sequences FLAT based on IMO produce poor results due to the poor exploitation which need to be enhanced by adding particle swarm optimization (PSO) algorithm which has efficient exploitation capability. The enhanced version of IMO (IMO-PSO)was merged as two layer (bottom layer for exploration using IMO and the upper layer exploit the best solution founded from the bottom layer). This hybrid scheme increase the diversity of solutions which increase the quality of solutions. FLAT based on IMO-PSO was tested on real biological sequences gathered from NCBI versus IMO and the standard local alignment algorithm. Besides, COVID-19 was analyzed against other viruses to detect the LCCS between it. FLAT based on IMO-PSO produced an enhancement of the performance of IMO for finding LCCS between biological sequences.
Read moreHigh-Performance Computing for Satellite Image Processing Using Apache Spark
High-Performance Computing is the aggregate computing application that solves computational problems that are either huge or time-consuming for traditional computers. This technology is used for processing satellite images and analysing massive data sets quickly and efficiently. Parallel processing and distributed computing methods are very important to process satellite images quickly and efficiently. Parallel Computing is a computation type in which multiple processors execute multiple tasks simultaneously to rapidly process data using shared memory. In this, we process satellite image parallel in a single computer. In distributed computing, we use multiple systems to process satellite images quickly. With the help of VMware, we are creating a different operating system (like Linux, windows etc.) as a worker. In this project we are using cluster formation for connecting master and slave: apache spark is one of the important concepts in this project. Apache spark is one of the frameworks and Resilient Distributed Datasets are one of the concepts in the spark, we are using RDD for dividing dataset on the different node of the cluster.
Read moreBSCSO-STNN: A Big Data-Driven IoT Intrusion Detection Model
The rapid expansion of the Internet of Things (IoT) and Big Data (BD) has led to security challenges. Securing IoT-BD against cyberattacks is necessary. An increasing number of applications are being implemented on BD platforms due to the rapid proliferation of data on the Internet. As the volume of data increases, the possibility of intrusions on the platform correspondingly increases. Conventional Intrusion Detection Systems (IDS) are ineffective for managing the extensive volume of historical data and unable to fulfil the security demands of BD platforms. This research aims to propose a novel intrusion detection model using Binary Sand Cat Swarm Optimization and Spatiotemporal Transformer Neural Network (BSCSO-STNN) model to address these issues. The CIC-IoT-23 and Bot-IoT datasets are collected and applied to train the model for evaluation. The developed BSCSO-STNN model is deployed in an Apache Spark (APS) framework. The datasets are initially preprocessed in this framework with data cleaning, oversampling, label encoding, and normalization. After preprocessing, the data is applied to the BSCSO for feature selection. Using the selected features, the STNN model performs binary and multiclass classification for both datasets. The BSCSO-STNN model attained 99.08% accuracy, 98.78% detection rate, 99.02% precision, and 98.94% F1-score using the CIC-IoT-23 dataset. The model attained 99.04% accuracy, 98.81% detection rate, 98.97% precision, and 98.95% F1-score for the BoT-IoT dataset in multiclass classification. The developed model outperformed all the current models in this research and demonstrated its accuracy in detecting intrusions.
Read moreWaveform Mapping and Time-Frequency Processing of DNA and Protein Sequences
Current state-of-the-art approaches for biological sequence querying and alignment require preprocessing and lack robustness to repetitions in the sequence. In addition, these approaches do not provide much support for efficiently querying subsequences, a process that is essential for tracking localized database matches. We propose a query-based alignment method for biological sequences that first maps sequences to time-domain waveforms before processing the waveforms for alignment in the time-frequency plane. The mapping uses waveforms, such as Gaussian functions, with unique sequence representations in the time-frequency plane. The proposed alignment method employs a robust querying algorithm that utilizes a time-frequency signal expansion whose basis function is matched to the basic waveform in the mapped sequences. The resulting WAVEQuery approach was demonstrated for both deoxyribonucleic acid (DNA) and protein sequences using the matching pursuit decomposition as the signal basis expansion. We specifically evaluated the alignment localization of WAVEQuery over repetitive database segments, and we demonstrated its operation in real-time without preprocessing. We also demonstrated that WAVEQuery significantly outperformed the biological sequence alignment method BLAST for queries with repetitive segments for DNA sequences. A generalized version of the WAVEQuery approach with the metaplectic transform is also described for protein sequence structure prediction.
Read more