- Research Article
36
- 10.1016/j.is.2017.10.007
Scalable and queryable compressed storage structure for raster data
- Oct 18, 2017
- Information Systems
- Susana Ladra + 2 more +2
Scalable and queryable compressed storage structure for raster data
The objective is to develop a data structure that is capable of handling large data volumes and offers support for querying, analysis and validation. Based on earlier results (i.e. the full decomposition of space, the use of a TEN structure and applying Poincare simplicial homology as mathematical foundation) a simplicial complex-based TEN structure is developed. Applying simplicial homology offers full control over orientation of simplexes and enables one to derive substantial parts of the TEN structure, instead of explicitly store the entire network. The described data structure is developed as a DBMS data structure and the usage of views, function based indexes and 3D R-trees result in a compact topological 3D data structure. Theoretical aspects of this approach are described earlier [12, 17,18]. This paper describes both theory and implementation of the approach.
Scalable and queryable compressed storage structure for raster data
Scalable and queryable compressed storage structure for raster data
Optimizing computational high-order schemes in finite volume simulations using unstructured mesh and topological data structures
Optimizing computational high-order schemes in finite volume simulations using unstructured mesh and topological data structures
Read moreOut-of-core build of a topological data structure from polygon soup
Many solid modeling applications require information not only about the geometry of an object but also about its topology. Most interchange formats do not provide this information, which the application must then derive as it builds its own topological data structure from unordered, "polygon soup" input. For very large data sets, the topological data structure itself can be bigger than core memory, so that a naive algorithm for building it that doesn't take virtual memory access patterns into account can become prohibitively slow due to thrashing. In this paper, we describe a new out-of-core algorithm that can build a topological data structure efficiently from very large data sets, improving performance by two orders of magnitude over a naive approach.
Read moreUsing Compressed Suffix-Arrays for a compact representation of temporal-graphs
Using Compressed Suffix-Arrays for a compact representation of temporal-graphs
A memory-efficient data structure representing exact-match overlap graphs with application for next-generation DNA assembly
Exact-match overlap graphs have been broadly used in the context of DNA assembly and the shortest super string problem where the number of strings n ranges from thousands to billions. The length ℓ of the strings is from 25 to 1000, depending on the DNA sequencing technologies. However, many DNA assemblers using overlap graphs suffer from the need for too much time and space in constructing the graphs. It is nearly impossible for these DNA assemblers to handle the huge amount of data produced by the next-generation sequencing technologies where the number n of strings could be several billions. If the overlap graph is explicitly stored, it would require Ω(n(2)) memory, which could be prohibitive in practice when n is greater than a hundred million. In this article, we propose a novel data structure using which the overlap graph can be compactly stored. This data structure requires only linear time to construct and and linear memory to store. For a given set of input strings (also called reads), we can informally define an exact-match overlap graph as follows. Each read is represented as a node in the graph and there is an edge between two nodes if the corresponding reads overlap sufficiently. A formal description follows. The maximal exact-match overlap of two strings x and y, denoted by ov(max)(x, y), is the longest string which is a suffix of x and a prefix of y. The exact-match overlap graph of n given strings of length ℓ is an edge-weighted graph in which each vertex is associated with a string and there is an edge (x, y) of weight ω=ℓ-|ov(max)(x, y)| if and only if ω ≤ λ, where |ov(max)(x, y)| is the length of ov(max)(x, y) and λ is a given threshold. In this article, we show that the exact-match overlap graphs can be represented by a compact data structure that can be stored using at most (2λ-1)(2⌈logn⌉+⌈logλ⌉)n bits with a guarantee that the basic operation of accessing an edge takes O(log λ) time. We also propose two algorithms for constructing the data structure for the exact-match overlap graph. The first algorithm runs in O(λℓnlogn) worse-case time and requires O(λ) extra memory. The second one runs in O(λℓn) time and requires O(n) extra memory. Our experimental results on a huge amount of simulated data from sequence assembly show that the data structure can be constructed efficiently in time and memory. Our DNA sequence assembler that incorporates the data structure is freely available on the web at http://www.engr.uconn.edu/~htd06001/assembler/leap.zip
Read moreDDSM: Design-Oriented Dual-Scale Shape-Material Model for Lattice Material Components
This paper proposes a new CAD model for the design of lattice material components. The CAD model better captures the user’s design intent and provides a dual-scale framework to represent the geometry and material distribution. Conventional CAD model formats based on B-Rep generate millions of data files, which also makes design intent and material information missing. In the present work, a new shape-material model for lattice material components is proposed. At the macroscopic scale, a compact face-based non-manifold topological data structure is proposed to express the lattice shape-material information without ambiguity. At the microscopic scale, implicit function is adopted for the representation of lattice material components. Numerical experiments verify that the proposed CAD model provides a powerful support for design intent with minor space costs. Meanwhile, the representation method supports solid modeling queries of geometric and material information on each scale.
Read moreSFCM: A Fuzzy Clustering Algorithm of Extracting the Shape Information of Data
Topological data analysis is a new theoretical trend using topological techniques to mine data. This approach helps determine topological data structures. It focuses on investigating the global shape of data rather than on local information of high-dimensional data. The Mapper algorithm is considered as a sound representative approach in this area. It is used to cluster and identify concise and meaningful global topological data structures that are out of reach for many other clustering methods. In this article, we propose a new method called the Shape Fuzzy C-Means (SFCM) algorithm, which is constructed based on the Fuzzy C-Means algorithm with particular features of the Mapper algorithm. The SFCM algorithm can not only exhibit the same clustering ability as the Fuzzy C-Means but also reveal some relationships through visualizing the global shape of data supplied by the Mapper. We present a formal proof and include experiments to confirm our claims. The performance of the enhanced algorithm is demonstrated through a comparative analysis involving the original algorithm, Mapper, and the other fuzzy set based improved algorithm, F-Mapper, for synthetic and real-world data. The comparison is conducted with respect to output visualization in the topological sense and clustering stability.
Read moreRe-designing Compact-structure based Forwarding for Programmable Networks
Forwarding packets based on networking names is essential for network protocols on different layers, where the `names' could be addresses, packet/flow IDs, and content IDs. For long there have been efforts using dynamic and compact data structures for fast and memory-efficient forwarding. In this work, we identify that the recently developed programmable network paradigm has the potential to further reduce the time/memory complexity of forwarding structures by separating the data plane and control plane. This work presents the new designs of network forwarding structures under the programmable network paradigm, applying three typical dynamic and compact data structures: Bloom filters, Cuckoo hashing, and Othello hashing. We conduct careful analyses and experiments in real networks of these forwarding methods for multiple performance metrics, including lookup throughput, memory footprint, construction time, dynamic updates, and lookup errors. The results give rich insights on designing forwarding algorithms with dynamic and compact data structures. In particular, the new designs based on Cuckoo hashing and Othello hashing show significant advantages over the extensively studied Bloom filter based methods, in all situations discussed in this paper.
Read moreHigh-Performance IPv6 Lookup with Real-Time Updates Using Hierarchical-Balanced Search Tree
Packet forwarding is the foundation of network communications. Internet routers forward packets by matching the packet destination address against the prefixes in the forwarding information base. Due to the emergence of cloud computing and network function virtualization, more and more applications require high-bandwidth low-latency IPv6 networks. However, most of the current IP lookup algorithms in practice are designed for IPv4 addresses and hard to scale to IPv6 addresses. In this paper, we first propose a novel data structure called hierarchical balanced search tree (Hi-BST). The compact and scalable data structure built can be stored in on-chip memory. On the basis of Hi-BST, we present a fast IPv6 lookup algorithm and prefix update algorithms. In the worst case, our algorithm can achieve fast lookup and update with the time that is proportional to the logarithm of the number of prefixes. Experiments using real FIBs give an integrated evaluation. Compared with the state-of-the-art algorithms, our algorithm achieves a fast lookup with realtime updates, which is 1.33 ~6.78 times that of the comparison algorithms respectively. Moreover, Hi-BST saves 4.3% 79.0% memory footprint of the comparison algorithms respectively.
Read moreEfficient Data Structure and Highly Scalable Algorithm for Total-Viewshed Computation
This paper presents an efficient and highly scalable algorithm, designed from scratch, to calculate total-viewshed in large high-resolution digital elevation models (DEMs) without restrictions as to whether or not the observer is linked to the ground. The keys to the high efficiency of the proposed method are: 1) the selection of a reliable sampling to represent the subareas of study; 2) the use of a compact and stable data structure to store the calculated data; and 3) the high reutilization of data and calculation between the large number of viewpoints. The obtained results demonstrate that the proposed algorithm is the fastest over the most commonly used GIS-software showing very similar numerical accuracy.
Read moreGapped Indexing for Consecutive Occurrences
The classic string indexing problem is to preprocess a string S into a compact data structure that supports efficient pattern matching queries. Typical queries include existential queries (decide if the pattern occurs in S), reporting queries (return all positions where the pattern occurs), and counting queries (return the number of occurrences of the pattern). In this paper we consider a variant of string indexing, where the goal is to compactly represent the string such that given two patterns \(P_1\) and \(P_2\) and a gap range \({[}\alpha , \beta ]\) we can quickly find the consecutive occurrences of \(P_1\) and \(P_2\) with distance in \({[}\alpha , \beta ]\), i.e., pairs of subsequent occurrences with distance within the range. We present data structures that use linear space and query time \({\widetilde{O}}(|P_1|+|P_2|+n^{2/3})\) for existence and counting and \({\widetilde{O}}(|P_1|+|P_2|+n^{2/3}\hbox {occ}^{1/3})\) for reporting. We complement this with a conditional lower bound based on the set intersection problem showing that any solution using \({\widetilde{O}}(n)\) space must use \({\widetilde{\Omega }}(|P_1| + |P_2| + \sqrt{n})\) query time. To obtain our results we develop new techniques and ideas of independent interest including a new suffix tree decomposition and hardness of a variant of the set intersection problem.
Read moreRun-Length Encoding in a Finite Universe
Text compression schemes and compact data structures usually combine sophisticated probability models with basic coding methods whose average codeword length closely match the entropy of known distributions. In the frequent case where basic coding represents run-lengths of outcomes that have probability p, i.e. the geometric distribution \(\Pr (i)=p^i(1-p)\), a Golomb code is an optimal instantaneous code, which has the additional advantage that codewords can be computed using only an integer parameter calculated from p, without need for a large or sophisticated data structure. Golomb coding does not, however, gracefully handle the case where run-lengths are bounded by a known integer n. In this case, codewords allocated for the case \(i>n\) are wasted. While negligible for large n, this makes Golomb coding unattractive in situations where n is recurrently small, e.g., when representing many short lists of integers drawn from limited ranges, or when the range of n is narrowed down by a recursive algorithm.
Read moreHierarchical digital differential analyzer for efficient ray-marching in OpenVDB
VDB[Museth 2013] is a compact data structure and toolset developed at DreamWorks Animation for high-resolution volumetric effects typically encountered in movie production. Since its open source release in 2012, as OpenVDB1, it has been adopted by major third-party renders and VFX tools, including Houdini by SideFX, RenderMan by Pixar, Arnold by Solid Angle, and RealFlow by Next Limit. Thus, it should come as no surprise that it is highly desirable to have efficient algorithms for ray-marching of sparse VDB volumes. The core problem is therefore, how to best utilize the hierarchical data structure of VDB for applications that require efficient ray-marching. As we will demonstrate, the solution is a novel hierarchical digital differential analyzer.
Read moreK2-treaps to represent and query data warehouses into main memory
In this paper we propose the use of the compact data structure k2-treap to process data cubes of Data Warehouses (DWs) into main memory. Compact data structures are data structures that allow compacting the data without losing the capacity of querying them in their compact form. A DW is a data repository to store historical data for decision support, and consists of dimensions and facts. The former are an abstract concept that groups data with a similar meaning, they are modelled as hierarchies of levels, which contain elements. The latter are quantitative data associated to dimensions. A data cube is a typical way to retrieve facts at different levels of granularity (through navigation on dimensions hierarchies). A DW can store terabytes of data, thus the efficient processing of data cubes is key in OLAP (On-line Analytical Processing). We show that by using a compact representation of data cubes and bitmaps to represent dimensions we are able to improve the use of space in main memory, and achieve better performance for query processing.
Read moreCONSERT: Constructing optimal name-based routing tables
CONSERT: Constructing optimal name-based routing tables