- Research Article
45
- 10.1016/s0169-023x(01)00007-6
Materialized view selection under the maintenance time constraint
- Apr 04, 2001
- Data & Knowledge Engineering
- Weifa Liang + 2 more +2
Materialized view selection under the maintenance time constraint
On-Line Analytical Processing (OLAP) tools are frequently used in business, science and health to extract useful knowledgefrom massive databases. An important and hard optimization problem in OLAP data warehouses is the view selection problem, consisting of selecting a set of aggregate views of the data for speeding up future query processing. A common variant of the view selection problem addressed in the literature minimizes the sum of maintenance cost and query time on the view set. Converting what is inherently an optimization problem with multiple conflicting objectives into one with a single objective ignores the need and value of a variety of solutions offering various levels of trade-off between the objectives. We apply two non-elitist multiobjective evolutionary algorithms (MOEAs) to view selection under a size constraint. Our emphasis is to determine the suitability of the combination of MOEAs with constraint handling to the view selection problem, compared to a widely used greedy algorithm. We observe that the evolutionary process mimics that of the greedy in terms of the convergence process in the population. The MOEAs are competitive with the greedy on a variety of problem instances, often finding solutions dominating it in a reasonable amount of time.
Materialized view selection under the maintenance time constraint
Materialized view selection under the maintenance time constraint
Learning View Selection for 3D Scenes
Efficient 3D space sampling to represent an underlying 3D object/scene is essential for 3D vision, robotics, and beyond. A standard approach is to explicitly sample a dense collection of views and formulate it as a view selection problem, or, more generally, a set cover problem. In this paper, we introduce a novel approach that avoids dense view sampling. The key idea is to learn a view prediction network and a trainable aggregation module that takes the predicted views as input and outputs an approximation of their generic scores (e.g., surface coverage, viewing angle from surface normals). This methodology allows us to turn the set cover problem (or multi-view representation optimization) into a continuous optimization problem. We then explain how to effectively solve the induced optimization problem using continuation, i.e., aggregating a hierarchy of smoothed scoring modules. Experimental results show that our approach arrives at similar or better solutions with about 10 x speed up in running time, comparing with the standard methods.
Read moreA Constraint Optimization Method for Large-Scale Distributed View Selection
View materialization is a commonly used technique in many data-intensive systems to improve the query performance. Increasing need for large-scale data processing has led to investigating the view selection problem in distributed complex scenarios where a set of cooperating computer nodes may share data and issue numerous queries. In our work, the view selection and data placement problem is studied given a limited amount of resources e.g. storage space capacity per computer node and maximum view maintenance cost. We also consider the IO and CPU costs for each computer node as well as the network bandwidth. To address this problem, we have proposed a constraint programming approach which is based on constraint reasoning to tackle problems that aim to satisfy a set of constraints. Then, we have designed a set of efficient heuristics that result in a drastic reduction in the solution space so that the problem becomes solvable for complex scenarios consisting of realistically large numbers of sites, queries and views. Our experimental study shows that our approach performs consistently better compared to a practical approach designed for large-scale distributed environments which uses a genetic algorithm to compute which view has to be materialized at what computer node.
Read moreA case for dynamic view management
Materialized aggregate views represent a set of redundant entities in a data warehouse that are frequently used to accelerate On-Line Analytical Processing (OLAP). Due to the complex structure of the data warehouse and the different profiles of the users who submit queries, there is need for tools that will automate and ease the view selection and management processes. In this article we present DynaMat, a system that manages dynamic collections of materialized aggregate views in a data warehouse. At query time, DynaMat utilizes a dedicated disk space for storing computed aggregates that are further engaged for answering new queries. Queries are executed independently or can be bundled within a multiquery expression. In the latter case, we present an execution mechanism that exploits dependencies among the queries and the materialized set to further optimize their execution. During updates, DynaMat reconciles the current materialized view selection and refreshes the most beneficial subset of it within a given maintenance window. We show how to derive an efficient update plan with respect to the available maintenance window, the different update policies for the views and the dependencies that exist among them.
Read moreTemplate-Based Bitmap View Selection for Optimizing Queries Over Tree Data
Developing and exploiting flexible techniques for optimizing the evaluation of queries over loosely structured data (e.g. tree or graph databases) is of crucial importance for modern database applications. In this context, we consider a new type of views which can be materialized as compressed bitmaps over tree data. We introduce the concept of view structural template to define classes of views. We then define and address a novel view selection problem (called view class selection (VCS) problem) where the goal is to select classes of bitmap views in order to optimize the overall evaluation cost of all tree pattern queries (TPQs) that can be issued against a database while satisfying a space constraint and ensuring that all the TPQs can be answered using exclusively the materialized views. We show that the VCS problem is NP-hard and we design two heuristic greedy algorithms which iteratively generate new batches of candidate view classes and make them available for selection. Each algorithm uses a different view class expansion technique to enable the systematic generation of candidate view classes from classes with smaller templates. We run extensive experiments to evaluate both the effectiveness of the algorithms and their efficiency on real, benchmark and synthetic datasets. Our algorithms are able to suggest high quality selections of view classes in a reasonable amount of time.
Read moreAn MDA approach for developing secure OLAP applications: Metamodels and transformations
Decision makers query enterprise information stored in Data Warehouses (DW) by using tools (such as On-Line Analytical Processing (OLAP) tools) which employ specific views or cubes from the corporate DW or Data Marts, based on multidimensional modelling. Since the information managed is critical, security constraints have to be correctly established in order to avoid unauthorized access. In previous work we defined a Model-Driven based approach for developing a secure DW repository by following a relational approach. Nevertheless, it is also important to define security constraints in the metadata layer that connects the DW repository with the OLAP tools; that is, over the same multidimensional structures that end users manage. This paper incorporates a proposal for developing secure OLAP applications within our previous approach: it improves a UML profile for conceptual modelling; it defines a logical metamodel for OLAP applications; and it defines and implements transformations from conceptual to logical models, as well as from logical models to secure implementation in a specific OLAP tool (SQL Server Analysis Services).
Read moreMatérialisation de vues dans les entrepôts de données. Une approche dynamique
Nowadays, materialized views are increasingly being supported by a variety of commercial DBMS to speed up query response time. This technique is also very useful in dataware-housing for optimizing OLAP queries. However, most existing view selection methods are of static nature. Moreover, none of these methods have considered the problem of dematerializing the previously materialized views. This paper deals with the problem of dynamic view selection and the pending issue of removing materialized views in order to replace the less beneficial views with more beneficial views. More precisely, we have designed and implemented a view selection method, including a polynomial algorithm, to decide which views to materialize according to statistic metadata.
Read moreAn application for multidimensional analysis of the Web site traffic
Considers the use of online analytical processing (OLAP) tools to provide fast, interactive analysis of Web-site traffic. For the purpose of analysing our own hierarchically organized Web site, the Directory of Croatian WWW Servers, an OLAP application (OLAWEB) has been developed. We describe data extraction from existing server access log files, the data transformation necessary to prepare data for multidimensional analysis, and the storage of the Web-site traffic data. The basic characteristics of the OLAWEB application are explained. Using this application, the Web-site administrator and other users are able to get fast answers to many unpredictable and complex questions that could not be answered by available statistical tools.
Read moreModel for identifying regions’ potential for clustering using information extraction tools and techniques
Decision-makers adopt regional industrial clusters as a development tool for competitive advantages in globalization. Identifying clustering potential within regions serves as the foundation for effective policy formulation. However, a standard systematic approach for assessing regional clustering potential is lacking. Current methods combine quantitative analyses of employment statistics to evaluate sectoral density with qualitative assessments of cluster dynamics. Advances in information technology have accelerated alongside globalization, with exponential growth in data volume and analytical capabilities. Data science processes this information for knowledge discovery, with increasing use of data mining, SQL queries, and online analytical processing (OLAP) tools in clustering analysis. This study proposes a model for detecting regional industrial clustering potential by integrating OLAP technology with data mining through an online analytical mining (OLAM) framework. This approach addresses limitations in traditional quantitative methods, such as location quotient and three-star analysis, which rely on threshold values and single regional parameters. The OLAM mechanism combines OLAP’s multidimensional analysis with data mining’s pattern recognition, thus enabling nuanced identification of clustering potential. It eliminates threshold dependencies and single-parameter evaluations, offering an alternative to conventional techniques. Integration into institutional systems can replace time-consuming traditional methods with a dynamic framework for multidimensional reporting.
Read moreMultiobjective evolutionary algorithms: analyzing the state-of-the-art.
Solving optimization problems with multiple (often conflicting) objectives is, generally, a very difficult goal. Evolutionary algorithms (EAs) were initially extended and applied during the mid-eighties in an attempt to stochastically solve problems of this generic class. During the past decade, a variety, of multiobjective EA (MOEA) techniques have been proposed and applied to many scientific and engineering applications. Our discussion's intent is to rigorously define multiobjective optimization problems and certain related concepts, present an MOEA classification scheme, and evaluate the variety of contemporary MOEAs. Current MOEA theoretical developments are evaluated; specific topics addressed include fitness functions, Pareto ranking, niching, fitness sharing, mating restriction, and secondary populations. Since the development and application of MOEAs is a dynamic and rapidly growing activity, we focus on key analytical insights based upon critical MOEA evaluation of current research and applications. Recommended MOEA designs are presented, along with conclusions and recommendations for future work.
Read moreA genetic local search algorithm for multiobjective time-dependent route planning
The multiobjective time-dependent route planning problem is a hard multiobjective combinatorial optimization problem. Metaheuristics showed success in solving many hard optimization problems and recently many efforts have been directed to hybridize elements from different metaheuristics and search methods. The hybridization of genetic algorithms and local search methods proved to be successful in many domains. In this paper we present a genetic local search algorithm for solving the multiobjective time-dependent route planning problem taking the multiobjective route planning in dynamic multihop ridesharing as an example problem. The behavior of the proposed algorithm is compared, on two problem instances using a set of widely used quality indicators, with the behavior of a genetic algorithm proposed for solving the same problem. Experimentation results indicated that the proposed algorithm outperforms the genetic algorithm regarding all quality indicators.
Read moreUsing Shuffled Frog Leaping Algorithm to View Selection Problem Subject to Dual Constraints
This paper presents a shuffled frog leaping algorithm (SFLA) based solution to solve the View Selection Problem (VSP) subject to dual constraints, which is often used to accelerate data warehouse queries. Since VSP is both discrete and constrained, a greedy-repaired strategy under dual constraints is proposed to handle unfeasible solutions. This proposed solution also profits from a mutation strategy in order to improve the quality of solutions, particularly to avoid being trapped in local optima. Experimental results show that under different constraints combinations, SFLA is able to find a near-optimal feasible solution, with maximum error less than 1%. Comparisons with GA and PSO show that SFLA has better solution quality and faster convergence rate, and also scales with the problem size.
Read moreUsing PlatEMO to Solve Multi-Objective Optimization Problems in Applications: A Case Study on Feature Selection
Many real-world optimization problems are characterized by multiple conflicting objectives, which are known as multi-objective optimization problems (MOPs). In the last two decades, evolutionary algorithms have shown promising performance in solving various MOPs, and a large number of multi-objective evolutionary algorithms (MOEAs) have been proposed. In order to determine the most suitable MOEA for a specific MOP, it is usually necessary to perform experiments to compare the performance of multiple candidate MOEAs. In 2017, an evolutionary multi-objective optimization platform was proposed by us, called PlatEMO, which provides the source codes of many state-of-the-art MOEAs and helps researchers perform batch experiments on these MOEAs. In this work, we illustrate the method of using the newest version of PlatEMO to solve MOPs in applications, by means of a case study on the feature selection problem, which is an important and difficult task in machine learning and data mining. This paper details the method of adding the feature selection problem to PlatEMO, and presents the experimental results of eight MOEAs on nine datasets in feature selection.
Read moreA new multi-objective evolutionary algorithm for service restoration: Non-dominated Sorting Genetic Algorithm-II in subpopulation tables
Distribution system (DS) service restoration is a very complex multi-objective and multi-constraint optimization problem, which requires high-quality Pareto fronts for helping the DS operators' work. This paper proposes a multi-objective evolutionary algorithm (MOEA) that combines the advantages of MOEAs in Tables with the properties of the Non-dominated Sorting Genetic Algorithm-II (NSGA-II) to recover Pareto fronts of the service problem. Its main differentials are its ability of recovering high-quality Pareto fronts, even in large-scale DSs, and of prioritizing switching operations in Remotely Controlled Switches (RCSs), independently of amount of RCSs. In opposite of Manually Controlled Switches (MCSs), RCS can be faster and cheapest operated from operating center. Several tests were carried out to validate and to compare the proposed MOEA to three MOEAs published in literature. Four metrics were used for assessing the quality of Pareto fronts generated and a Welch's t hypothesis test enabled the comparison of MOEAs' performance. Test results indicate the proposed MOEA outperforms those three MOEAs published in literature.
Read moreIterative Approach to Many-Objective Engineering Design: Balancing Conflicting Objectives for Engineered Injection and Extraction
Engineering design problems are often characterized by multiple, conflicting objectives. Multiobjective evolutionary algorithms (MOEAs) can optimize these complex problems using embedded simulation models to calculate objective function values without simplifying assumptions, which are often required by traditional search algorithms. This study contributes a demonstration of how the performance of the MOEA can be improved through iterative reformulations of the problem objectives and constraints. We illustrate this iterative design approach using a case study of engineered injection and extraction (EIE), a strategy developed to enhance contaminant degradation during in situ groundwater remediation. During EIE, clean water is injected or extracted at wells surrounding a contaminated groundwater plume following the one-time injection of a treatment chemical. This sequence of injections and extractions reconfigures the plume and increases its contact with the treatment chemical, which ultimately increases the rate of contaminant degradation. The MOEA is used to optimize the EIE design, which consists of the sequence of pumping rates and well locations that dictate the plume reconfiguration. The optimization problem is a challenging one, characterized by uncertainty from randomness associated with hydrodynamic dispersion of the plume. The objectives and constraints of the remediation are refined iteratively during the optimization process, after considering tradeoffs between objectives for design solutions generated by the MOEA. Tradeoff data not only informs the design of solutions, it provides valuable information about the structure and conflicts of the design problem, which represents crucial support for stakeholders in their decision-making analyses for in situ groundwater remediation project planning with EIE.
Read more