- Research Article
35
- 10.1016/j.jalgor.2004.04.001
On external-memory MST, SSSP and multi-way planar graph separation
- May 21, 2004
- Journal of Algorithms
- Lars Arge + 2 more +2
On external-memory MST, SSSP and multi-way planar graph separation
We implement a promising algorithm for sparse-matrix sparse-vector multiplication (SpMSpV) on the GPU. An efficient k-way merge lies at the heart of finding a fast parallel SpMSpV algorithm. We examine the scalability of three approaches -- no sorting, merge sorting, and radix sorting -- in solving this problem. For breadth-first search (BFS), we achieve a 1.26x speedup over state-of-the-art sparse-matrix dense-vector (SpMV) implementations. The algorithm seems generalize able for single-source shortest path (SSSP) and sparse-matrix sparse-matrix multiplication, and other core graph primitives such as maximal independent set and bipartite matching.
On external-memory MST, SSSP and multi-way planar graph separation
On external-memory MST, SSSP and multi-way planar graph separation
Techniques for Practical Parallel BFS and SSSP
Breadth-first search (BFS) and single source shortest paths (SSSP) are two fundamental graph problems with countless real-world applications. It is of major interest to develop efficient parallel algorithms and implementations to solve these problems on modern multicore architectures. The challenge is that these computations are often irregular, and there is no known algorithm that guarantees polylogarithmic depth for all graphs. For BFS, this paper describe several performance engineering methods to develop a practical parallel implementation for all classes of real-world graphs. For SSSP, we introduce a unique contraction based preprocessing method which significantly speeds up queries on high diameter graphs, like road networks. Our method is generic and can be used with any algorithm for SSSP queries.
Read moreI/O-Optimal Algorithms for Outerplanar Graphs
We present linear-I/O algorithms for fundamental graph problems on embedded outerplanar graphs. We show that breadth-first search, depth-first search, single-source shortest paths, triangulation, and computing an � separator of size O(1/� )t akeO(scan(N)) I/Os on embedded outerplanar graphs. We also show that it takes O(sort(N)) I/Os to test whether a given graph is outerplanar and to compute an outerplanar embedding of an outerplanar graph, thereby providing O(sort(N))-I/O algorithms for the above problems if no embedding of the graph is given. As all these problems have linear-time algorithms in internal memory, a simple simulation technique can be used to improve the I/O-complexity of our algorithms from O(sort(N)) to O(perm(N)). We prove matching lower bounds for embedding, breadth-first search, depth-first search, and singlesource shortest paths if no embedding is given. Our algorithms for the above problems use a simple linear-I/O time-forward processing algorithm for rooted trees whose vertices are stored in preorder.
Read moreComparison of MSP-EXP432P401R Program Performances by Using Different Development Environments
Texas Instruments development kits have a wide application in practical and scientific experiments due to their small size, processing power, available booster packs, and compatibility with different environments. The most popular integrated development environments for programming these development kits are Energia and Code Composer Studio. Unfortunately, there are no existing studies that compare the benefits and drawbacks of these environments and their performances. Conversely, the performances of the FreeRTOS environment are well-explored, making it a suitable baseline for embedded systems execution. In this paper, we performed the experimental evaluation of the performance of Texas Instruments MSP-EXP432P401R when using Energia, Code Composer Studio, and FreeRTOS for program execution. Three different sorting algorithms (bubble sort, radix sort, merge sort) and three different search algorithms (binary search, random search, linear search) were used for this purpose. The results show that Energia sorting algorithms outperform other environments with a maximum of 400 elements. On the other hand, FreeRTOS search algorithms far outperform other environments with a maximum of <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{2 5 5, 0 0 0}$</tex> elements (whereas this maximum was <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{1 0, 0 0 0}$</tex> elements for other environments). Code Composer Studio resulted in the largest processing time, which indicates that the lowlevel registry editing performed in this environment leads to significant performance issues.
Read moreA Comparative Study on MMDBM Classifier Incorporating Various Sorting Procedure
Classification is one of the most important methods in data mining. These methods are used to extract meaningful information from large database which can be effectively used for predicting unknown class. The classifier based on decision tree is called Mixed Mode Data Base Miner (MMDBM) which is tested with different sorting techniques (merge sort, quick sort, radix sort) to compare the processing time to SLIQ classifier. In this paper, we carried out a comparative study on MMDBM classifier incorporating various sorting procedure by using Blood Pressure (BP) database, and finally the proposed method MMDBM classifier is one of the best classifiers among SLIQ supervised learning method. This proposed method achieved less processing time and higher rate of accuracy. Keywords: Classification, Data Mining, Java Sorting
Read moreH-BFT: A fast Breadth-First Traversal algorithm for Sparse graphs and its GPU implementation
With Moore's law in effect, as the complexity of digital electronic circuits increases, the amount of time spent by the electronic design automation (EDA) tools to design such circuits also increases. It brings us to the point, where we need to improve the performance of EDA algorithms to fulfil the present and the future requirements of the EDA industry. Out of many algorithms used by these tools, Breadth First Traversal (BFT) is one of the most commonly used algorithms to traverse the gates of electronic circuits. In this paper, we present a new simple, fast and parallelizable BFT algorithm for sparse graphs, named H-BFT. We show that the CPU implementation of H-BFT is about 75× faster than the CPU implementation of the state of the art, the Sparse Matrix-Vector Product (SMVP) based BFT. Further, with the new features introduced by NVIDIA in their GPUs, we have accelerated both the state of the art SMVP based BFT implementation and our new H-BFT implementation. The best speedups we achieved via these accelerations are 180× and 25× for the SMVP-BFT and H-BFT respectively.
Read moreOrdered vs. unordered
Outside of computational science, most problems are formulated in terms of irregular data structures such as graphs, trees and sets. Unfortunately, we understand relatively little about the structure of parallelism and locality in irregular algorithms. In this paper, we study multiple algorithms for four such problems: discrete-event simulation, single-source shortest path, breadth-first search, and minimal spanning trees. We show that the algorithms can be classified into two categories that we call unordered and ordered, and demonstrate experimentally that there is a trade-off between parallelism and work efficiency: unordered algorithms usually have more parallelism than their ordered counterparts for the same problem, but they may also perform more work. Nevertheless, our experimental results show that unordered algorithms typically lead to more scalable implementations, demonstrating that less work-efficient irregular algorithms may be better for parallel execution.
Read morePopular Conjectures Imply Strong Lower Bounds for Dynamic Problems
We consider several well-studied problems in dynamic algorithms and prove that sufficient progress on any of them would imply a breakthrough on one of five major open problems in the theory of algorithms: 1) Is the 3SUM problem on n numbers in O(n2 -- aepsi;) time for some aepsi; > 0? 2) Can one determine the satisfiability of a CNF formula on n variables and poly n clauses in O((2 -- aepsi;)npolyn) time for some aepsi; > 0? 3) Is the All Pairs Shortest Paths problem for graphs on n vertices in O(n3 -- aepsi;) time for some aepsi; > 0? 4) Is there a linear time algorithm that detects whether a given graph contains a triangle? 5) Is there an O(n3 -- aepsi;) time combinatorial algorithm for n × n Boolean matrix multiplication? The problems we consider include dynamic versions of bipartite perfect matching, bipartite maximum weight matching, single source reachability, single source shortest paths, strong connectivity, subgraph connectivity, diameter approximation and some nongraph problems such as Pagh's problem defined in a recent paper by pa#x0103;traa#x015F;cu [STOC 2010].
Read moreMaximum weight bipartite matching in matrix multiplication time
Maximum weight bipartite matching in matrix multiplication time
How Well do CPU, GPU and Hybrid Graph Processing Frameworks Perform?
The importance of high-performance graph processing to solve big data problems targeting high-impact applications is greater than ever before. Recent graph processing frameworks target different hardware platforms (e.g., shared memory systems, accelerators such as GPUs, and distributed systems) and differ with respect to the programming model they adopt (e.g., based on linear algebra formulations of graph algorithms or enabling direct access to the graph structure). To better understand the impact of these choices, this paper, presents a comparative study of five state-of-the-art graph processing frameworks: two CPU-only frameworks - GraphMat and Galois, two GPU-based frameworks - Nvgraph and Gunrock; and Totem, a hybrid (CPU+GPU) framework. We use three popular graph algorithms (PageRank, Single Source Shortest Path, and Breadth-First Search), and massive scale graphs with up to billions of edges. Our evaluation focuses on three performance metrics: (i) execution time, (ii) scalability and (iii) energy consumption.
Read moreGraVF: A vertex-centric distributed graph processing framework on FPGAs
FPGAs are promising platforms to efficiently execute distributed graph algorithms. Unfortunately, they are notoriously hard to program, especially when the problem size and system complexity increases. In this paper, we propose GraVF, a high-level design framework for distributed graph processing on FPGAs. It leverages the vertex-centric paradigm, which is naturally distributed and requires the user to define only very small kernels and their associated message semantics for the target application. The user design may subsequently be elaborated and compiled to the target system automatically by the framework. To demonstrate the flexibility and capabilities of the proposed framework, 4 graph algorithms with distinct requirements have been implemented, namely breadth-first search, PageRank, single source shortest path, and connected component. Results show that the proposed framework is capable of producing FPGA designs with performance comparable to similar custom designs while requiring only minimal input from the user.
Read moreA fast and simple randomized parallel algorithm for the maximal independent set problem
A fast and simple randomized parallel algorithm for the maximal independent set problem
Deploying Graph Algorithms on GPUs: An Adaptive Solution
Thanks to their massive computational power and their SIMT computational model, Graphics Processing Units (GPUs) have been successfully used to accelerate a wide variety of regular applications (linear algebra, stencil computations, image processing and bioinformatics algorithms, among others). However, many established and emerging problems are based on irregular data structures, such as graphs. Examples can be drawn from different application domains: networking, social networking, machine learning, electrical circuit modeling, discrete event simulation, compilers, and computational sciences. It has been shown that irregular applications based on large graphs do exhibit runtime parallelism; moreover, the amount of available parallelism tends to increase with the size of the datasets. In this work, we explore an implementation space for deploying a variety of graph algorithms on GPUs. We show that the dynamic nature of the parallelism that can be extracted from graph algorithms makes it impossible to find an optimal solution. We propose a runtime system able to dynamically transition between different implementations with minimal overhead, and investigate heuristic decisions applicable across algorithms and datasets. Our evaluation is performed on two graph algorithms: breadth-first search and single-source shortest paths. We believe that our proposed mechanisms can be extended and applied to other graph algorithms that exhibit similar computational patterns.
Read moreA fast parallel algorithm for the maximal independent set problem
A parallel algorithm is presented that accepts as input a graph G and produces a maximal independent set of vertices in G . On a P-RAM without the concurrent write or concurrent read features, the algorithm executes in O ((log n ) 4 ) time and uses O (( n /(log n )) 3 ) processors, where n is the number of vertices in G . The algorithm has several novel features that may find other applications. These include the use of balanced incomplete block designs to replace random sampling by deterministic sampling, and the use of a “dynamic pigeonhole principle” that generalizes the conventional pigeonhole principle.
Read moreFine-Grained Task Migration for Graph Algorithms Using Processing in Memory
Graphs are used in a wide variety of applicationdomains, from social science to machine learning. Graphalgorithms present large numbers of irregular accesses with littledata reuse to amortize the high cost of memory accesses, requiring high memory bandwidth. Processing in memory (PIM) implemented through 3D die-stacking can deliver this highmemory bandwidth. In a system with multiple memory moduleswith PIM, the in-memory compute logic has low latency and highbandwidth access to its local memory, while accesses to remotememory introduce high latency and energy consumption. Ideally, in such a system, computation and data are partitioned amongthe PIM devices to maximize data locality. But the irregularmemory access patterns present in graph applications make itdifficult to guarantee that the computation in each PIM devicewill only access its local data. A large number of remote memoryaccesses can negate the benefits of using PIM. In this paper, we examine the feasibility and potential of finegrainedwork migration to reduce remote data accesses insystems with multiple PIM devices. First, we propose a datadrivenimplementation of our study algorithms: breadth-firstsearch (BFS), single source shortest path (SSSP) and betweennesscentrality (BC) where each PIM has a queue where the verticesthat it needs to process are held. New vertices that need to beprocessed are enqueued at the PIM device co-located with thememory that stores those vertices. Second, we propose hardwaresupport that takes advantage of PIM to implement highlyefficient queues that improve the performance of the queuingframework by up to 16.7%. Third, we develop a timing model forthe queueing framework to explore the benefits of workmigration vs. remote memory accesses. And, finally, our analysisusing the above framework shows that naive task migration canlead to performance degradations and identifies trade-offsamong data locality, redundant computation, and load balanceamong PIM devices that must be taken into account to realize thepotential benefits of fine-grain task migration.
Read more