- Research Article
25
- 10.1016/j.jcss.2012.09.010
Multi-route query processing and optimization
- Oct 18, 2012
- Journal of Computer and System Sciences
- Rimma V Nehme + 4 more +4
Multi-route query processing and optimization
This paper addresses the problem of optimizing multiple distributed stream queries that are executing simultaneously in distributed data stream systems. We argue that the static query optimization approach of "plan, then deployment" is inadequate for handling distributed queries involving multiple streams and node dynamics faced in distributed data stream systems and applications. Thus, the selection of an optimal execution plan in such dynamic and networked computing systems must consider operator ordering, reuse, network placement, and search space reduction. We propose to use hierarchical network partitions to exploit various opportunities for operator-level reuse while utilizing network characteristics to maintain a manageable search space during query planning and deployment. We develop top-down, bottom-up, and hybrid algorithms for exploiting operator-level reuse through hierarchical network partitions. Formal analysis is presented to establish the bounds on the search space and suboptimality of our algorithms. We have implemented our algorithms in the IFLOW system, an adaptive distributed stream management system. Through simulations and experiments using a prototype deployed on Emulab, we demonstrate the effectiveness of our framework and our algorithms.
Multi-route query processing and optimization
Multi-route query processing and optimization
An Ant Colony Optimization Based Approach for Binary Search
Search is considered to be an important functionality in a computational system. Search techniques are applied in file retrievals and indexing. Though there exists various search techniques, binary search is widely used in many applications due to its advantage over other search techniques namely linear and hash search. Binary search is easy to implement and is used to search for an element in a large search space. The worst case time complexity of binary search is O (log 2 n) where n is the number of elements (search space) in the array. However, in binary search, searching is performed on the entire search space. The complexity of binary search may be further reduced if the search space is reduced. This paper proposes an Ant Colony Optimization based Binary Search (ACOBS) algorithm to find an optimal search space for binary search. ACOBS algorithm categorizes the search space and the key element is searched only in a specific category where the key element can exist thereby reducing the search space. The time complexity of ACOBS algorithm is O (log 2 c) where c is the number of elements in the reduced search space and c < n. The proposal is best suited for real time applications where searching is performed on a large domain.KeywordsAnt colony optimizationBinary searchOptimal search space
Read morePC-SSRDE: A paradigm crossover-based differential evolution algorithm with search space reduction
PC-SSRDE: A paradigm crossover-based differential evolution algorithm with search space reduction
Generating Query Plans for Distributed Query Processing Using Genetic Algorithm
Query Processing is a key determinant in the overall performance of distributed databases. It requires processing of data at their respective sites and transmission of the same between them. These together constitute a distributed query processing strategy (DQP). DQP aims to arrive at an efficient query processing strategy for a given query. This strategy involves generation of efficient query plans for a distributed query. In case of distributed relational queries, the number of possible query plans grows exponentially with an increase in the number of relations accessed by the query. This number increases further when the relations, accessed by the query, have replicas at different sites. Such a large search space renders it infeasible to find optimal query plans. This paper presents a query plan generation algorithm that attempts to generate optimal query plans, for a given query, using genetic algorithm. The query plans so generated involve fewer sites, thus leading to efficient query processing. Further, experimental results show that the proposed algorithm converges quickly towards optimal query plans for an observed crossover and mutation probability.
Read moreEfficient Intelligent Backtracking Using Linear Programming
Intelligent backtracking is a technique used in constraint programming for reducing search in solving combinatorial feasibility problems. The technique uses information derived from small sets of infeasible constraints discovered in one part of the search space to avoid searching other, similar, regions. It is often able to reduce the size of the search space significantly. For many problems, however, the computational effort required to achieve this reduction in search space is prohibitive. We introduce an algorithm that uses intelligent backtracking inside a linear-programming based branch-and-bound framework. We show that minimal infeasible sets can immediately be deduced from the dual extreme ray associated with the infeasible linear program. This allows us to obtain the reduction in search space associated with intelligent backtracking, without paying the large computational cost. We show the implementation of our intelligent backtracking approach as a branch-and-cut algorithm, and present computational results.
Read moreMulti-Level Search Space Reduction Framework for Face Image Database
In face recognition, searching and retrieval of relevant images from a large database form a major task. Recognition time is greatly related to the dimensionality of the original data and the number of training samples. This demands the selection of discriminant features that produce similar results as the entire set and a reduced search space. To address this issue, a Multi-Level Search Space Reduction framework for large scale face image database is proposed. The proposed approach identifies discriminating features and groups face images sharing similar properties using feature-weighted Fuzzy C-Means approach. A hierarchical tree model is then constructed inside every cluster based on the discriminating features which enables a branch based selection, thereby reducing the search space. The proposed framework is tested on three benchmark and two self-created databases. The experimental results show that the proposed method achieved an average accuracy of 93% and an average search time reduction of 66% compared to existing approaches for search space reduction of face recognition.
Read moreLAMBDA-SEARCH IN GAME TREES – WITH APPLICATION TO GO
This paper proposes a new method for searching two-valued (binary) game trees in games like chess or Go. Lambda-search uses null-moves together with different orders of threat-sequences (so-called lambda-trees), focusing the search on threats and threat-aversions, but still guaranteeing to find the mini-max value (provided that the game-rules allow passing or zugzwang is not a motive). Using negligible working memory in itself, the method seems able to offer a large relative reduction in search space over standard alpha-beta comparable to the relative reduction in search space of alpha-beta over minimax, among other things depending upon how non-uniform the search tree is. Lambda-search is compared to other resembling approaches, such as null-move pruning and proof-number search, and it is explained how the concept and context of different orders of lambda-trees may ease and inspire the implementation of abstract game-specific knowledge. This is illustrated on open-space Go block tactics, distinguishing between different orders of ladders, and offering some possible grounding work regarding an abstract formalization of the concept of relevancy-zones (zones outside of which added stones of any colour cannot change the status of the given problem).
Read moreA new framework for improving MPPT algorithms through search space reduction
A new framework for improving MPPT algorithms through search space reduction
Identification of top-K nodes in large networks using Katz centrality
Network theory concepts form the core of algorithms that are designed to uncover valuable insights from various datasets. Especially, network centrality measures such as Eigenvector centrality, Katz centrality, PageRank centrality etc., are used in retrieving top-K viral information propagators in social networks,while web page ranking in efficient information retrieval, etc. In this paper, we propose a novel method for identifying top-K viral information propagators from a reduced search space. Our algorithm computes the Katz centrality and Local average centrality values of each node and tests the values against two threshold (constraints) values. Only those nodes, which satisfy these constraints, form the search space for top-K propagators. Our proposed algorithm is tested against four datasets and the results show that the proposed algorithm is capable of reducing the number of nodes in search space at least by 70%. We also considered the parameter (alpha and beta) dependency of Katz centrality values in our experiments and established a relationship between the alpha values, number of nodes in search space and network characteristics. Later, we compare the top-K results of our approach against the top-K results of degree centrality.
Read moreMulti-objective parametric query optimization
Classical query optimization compares query plans according to one cost metric and associates each plan with a constant cost value. In this paper, we introduce the Multi-Objective Parametric Query Optimization (MPQ) problem where query plans are compared according to multiple cost metrics and the cost of a given plan according to a given metric is modeled as a function that depends on multiple parameters. The cost metrics may for instance include execution time or monetary fees; a parameter may represent the selectivity of a query predicate that is unspecified at optimization time. MPQ generalizes parametric query optimization (which allows multiple parameters but only one cost metric) and multi-objective query optimization (which allows multiple cost metrics but no parameters). We formally analyze the novel MPQ problem and show why existing algorithms are inapplicable. We present a generic algorithm for MPQ and a specialized version for MPQ with piecewise-linear plan cost functions. We prove that both algorithms find all relevant query plans and experimentally evaluate the performance of our second algorithm in a Cloud computing scenario.
Read moreDynamic multi-objective optimization applied to a solar-geothermal multi-generation system for hydrogen production, desalination, and energy storage
Dynamic multi-objective optimization applied to a solar-geothermal multi-generation system for hydrogen production, desalination, and energy storage
Read moreProcessing and optimizing main memory spatial-keyword queries
Important cloud services rely on spatial-keyword queries, containing a spatial predicate and arbitrary boolean keyword queries. In particular, we study the processing of such queries in main memory to support short response times. In contrast, current state-of-the-art spatial-keyword indexes and relational engines are designed for different assumptions. Rather than building a new spatial-keyword index, we employ a cost-based optimizer to process these queries using a spatial index and a keyword index. We address several technical challenges to achieve this goal. We introduce three operators as the building blocks to construct plans for main memory query processing. We then develop a cost model for the operators and query plans. We introduce five optimization techniques that efficiently reduce the search space and produce a query plan with low cost. The optimization techniques are computationally efficient, and they identify a query plan with a formal approximation guarantee under the common independence assumption. Furthermore, we extend the framework to exploit interesting orders. We implement the query optimizer to empirically validate our proposed approach using real-life datasets. The evaluation shows that the optimizations provide significant reduction in the average and tail latency of query processing: 7- to 11-fold reduction over using a single index in terms of 99th percentile response time. In addition, this approach outperforms existing spatial-keyword indexes, and DBMS query optimizers for both average and high-percentile response times.
Read moreBug Detection Using Particle Swarm Optimization with Search Space Reduction
A bug detection tool is an important tool in software engineering development. Many research papers have proposed techniques for detecting software bug, but there are certain semantic bugs that are not easy to detect. In our views, a bug can occur from incorrect logics that when a program is executed with a particular input, the program will behave in unexpected ways. In this paper, we propose a method and tool for software bugs detection by finding such input that causes an unexpected output guided by the fitness function. The method uses a Hierarchical Similarity Measurement Model (HSM) to help create the fitness function to examine a program behavior. Its tool uses Particle Swarm Optimization (PSO) with Search Space Reduction (SSR) to manipulate input by contracting and eliminating unfavorable areas of input search space. The programs under experiment were selected from four different domains such as financial, decision support system, algorithms and machine learning. The experimental result shows a significant percentage of success rate up to 93% in bug detection, compared to an estimated success rate of 28% without SSR.
Read moreFinding Image Semantics from a Hierarchical Image Database Based on Adaptively Combined Visual Features
Correlating image semantics with its low level features is a challenging task. Although, humans are adept in distinguishing object categories, both in visual as well as in semantic space, but to accomplish this computationally is yet to be fully explored. The learning based techniques do minimize the semantic gap, but unlimited possible categorization of objects in real world is a major challenge to these techniques. This work analyzes and utilizes the strength of a semantically categorized image database to assign semantics to query images. Semantics based categorization of images would result in image hierarchy. The algorithms proposed in this work exploit visual image descriptors and similarity measures in the context of a semantically categorized image database. A novel 'Branch Selection Algorithm' is developed for a highly categorized and dense image database, which drastically reduces the search space. The search space so obtained is further reduced by applying any one of the four proposed 'Pruning Algorithms'. Pruning algorithms maintain accuracy while reducing the search space. These algorithms use an adaptive combination of multiple visual features of an image database to find semantics of query images. Branch Selection Algorithm tested on a subset of 'ImageNet' database reduces search space by 75%. The best pruning algorithm further reduces this search space by 26% while maintaining 95% accuracy.
Read moreSearch space reduction through clustering in test generation
An important factor of performance in test generation is the number of decisions in the search space. The more decisions are made in order to find a test the larger is the search space. This paper introduces an algorithm for the reduction of the number of decisions. Unlike classical test generation methods which perform their search on a netlist of gates, this algorithm works at a higher level. The gates of the circuit to test are partitioned into clusters. The search performed at the cluster level becomes a progressive translation of a set of value assignments to another set of value assignments that steps over cluster boundaries. To be able to step over cluster boundaries sensitization and propagation conditions are used. Experimental data demonstrate the effectiveness of the algorithm in reducing the number of decisions.
Read more