- Book Chapter
7
- 10.1016/b978-0-444-51862-0.50029-0
Luckiness and Regret in Minimum Description Length Inference
- Jan 01, 2011
- Philosophy of Statistics
- Steven De Rooij + 1 more +1
Luckiness and Regret in Minimum Description Length Inference
The advance in the global environment, rapidly changing markets, and information technology has created a new stage for design. In such an environment, one strategy for success is the Collaborative Product Development (CPD). Organizing people effectively is the goal of Collaborative Product Development, and it solves the problem with certain foreseeability. The development group activities are influenced not only by the methods and decisions available, but also by correlation among personnel. Grouping the personnel according to their correlation intensity is defined as collaboration space division (CSD). Upon establishment of a correlation matrix (CM) of personnel and an analysis of the collaboration space, the genetic algorithm (GA) and minimum description length (MDL) principle may be used as tools in optimizing collaboration space. The MDL principle is used in setting up an object function, and the GA is used as a methodology. The algorithm encodes spatial information as a chromosome in binary. After repetitious crossover, mutation, selection and multiplication, a robust chromosome is found, which can be decoded into an optimal collaboration space. This new method can calculate the members in sub-spaces and individual groupings within the staff. Furthermore, the intersection of sub-spaces and public persons belonging to all sub-spaces can be determined simultaneously.
Loading PDF
Luckiness and Regret in Minimum Description Length Inference
Luckiness and Regret in Minimum Description Length Inference
An Introduction to the Minimum Description Length Principle
We give a brief introduction to the minimum description length (MDL) principle. The MDL principle is a mathematical formulation of Occam’s razor. It says ‘simple explanations of a given phenomenon are to be preferred over complex ones.’ This is recognized as one of basic stances of scientists, and plays an important role in statistics and machine learning. In particular, Rissanen proposed MDL criterion for statistical model selection based on information theory in 1978. After that, much literature has been published and the notion of MDL principle was founded in the 1990s. In this article, we review some important results on the MDL principle.KeywordsBayes mixtureLaplace estimatorMDLModel selectionMinimax regretUniversal code
Read moreAn extension on learning Bayesian belief networks based on MDL principle
Bayesian belief network (BBN) is a framework for representation/inference of some knowledge with uncertainty. Since the process of constructing a BBN manually by experts is time-consuming in general, some method supporting the task is needed. We proposed an algorithm for acquiring some BBN automatically from finite examples based on minimum description length (MDL) principle. This paper addresses an improvement which relaxes a constraint that the original scheme held on the representation. In BBNs, attributes and stochastic dependencies between them are expressed as nodes and directed links connecting them, respectively, where each attribute may be a predicate, a numerical data, etc., and each dependence is numerically expressed as the conditional probability of one attribute given other attributes if their dependence exists. Therefore, in general, BBNs are represented in terms of the network structure and the conditional probabilities.
Read moreCluster structure inference for microarray data based on an information theoretic criterion
This paper presents a new method for estimating the number of clusters and cluster structure in microarray data sets, based on minimum description length (MDL) principle and normalized maximum likelihood (NML) model for linear regression. The cluster-structure models for microarray data are compared based on the MDL principle. The performance of the new method is studied using simulated and real microarray data sets.
Read moreConstruction of tree structured classifiers by the MDL principle
An approach to the problem of constructing tree structured classifiers that is based on Rissanen's (1983) minimum description length (MDL) principle is presented. Simple and efficient rules for sequential growing and pruning of the tree are derived using this approach. The two rules are derived from a single MDL-based criterion. These splitting and pruning rules are intuitively pleasing and computationally simple. The computational load of the pruning rule is substantially smaller than the alternative pruning schemes of A. Mabbet et al. (1980) and L. Breiman et al. (1984). The extension of these splitting and pruning criteria to the case of multiple classes is straightforward. Experimental results illustrating the performance of this technique in automatic character recognition are provided.< <ETX xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">></ETX>
Read moreSpatially adaptive wavelet denoising using the minimum description length principle.
This paper presents a new spatially adaptive wavelet denoising method. Based on a doubly stochastic process model of wavelet coefficients, the method gives a new threshold, which varies spatially according to the variances of the coefficients, using the minimum description length (MDL) principle. The new threshold is not only easier to analyze since it is in a closed form, but also provides more facility for future compression than several other methods, almost without deteriorating mean square error risk.
Read moreIntegrative Parameter-Free Clustering of Data with Mixed Type Attributes
Integrative mining of heterogeneous data is one of the major challenges for data mining in the next decade. We address the problem of integrative clustering of data with mixed type attributes. Most existing solutions suffer from one or both of the following drawbacks: Either they require input parameters which are difficult to estimate, or/and they do not adequately support mixed type attributes. Our technique INTEGRATE is a novel clustering approach that truly integrates the information provided by heterogeneous numerical and categorical attributes. Originating from information theory, the Minimum Description Length (MDL) principle allows a unified view on numerical and categorical information and thus naturally balances the influence of both sources of information in clustering. Moreover, supported by the MDL principle, parameter-free clustering can be performed which enhances the usability of INTEGRATE on real world data. Extensive experiments demonstrate the effectiveness of INTEGRATE in exploiting numerical and categorical information for clustering. As an efficient iterative algorithm INTEGRATE is scalable to large data sets.
Read moreQuery-based summarization using MDL principle
Query-based text summarization is aimed at extracting essential information that answers the query from original text. The answer is presented in a minimal, often predefined, number of words. In this paper we introduce a new unsupervised approach for query-based extractive summarization, based on the minimum description length (MDL) principle that employs Krimp compression algorithm (Vreeken et al., 2011). The key idea of our approach is to select frequent word sets related to a given query that compress document sentences better and therefore describe the document better. A summary is extracted by selecting sentences that best cover query-related frequent word sets. The approach is evaluated based on the DUC 2005 and DUC 2006 datasets which are specifically designed for query-based summarization (DUC, 2005 2006). It competes with the best results.
Read moreDetecting Latent Structure Uncertainty with Structural Entropy
This paper proposes a new method for detecting the uncertainty of a latent structure. We consider the case where the latent structure of dataset changes gradually over time, with the goal of selecting the optimal model at any given time. In selecting the optimal model, we use the minimum description length (MDL) principle, specifically the normalized maximum likelihood (NML), which is the optimal code length in the sense of Shtarkov’s minimax regret. To detect the uncertainty of a latent structure, the main idea proposed here is that the uncertainty of the model selection will increase at the initial stage when the model change occurs. Here, we propose a new indicator called structural entropy (SE), which defines model selection uncertainty based on the MDL principle. We use several models for model selection, including the clustering structures of the Gaussian mixture model and Poisson mixture model, and a time-dependent structure model such as the autoregression model. We show the behavior of the proposed indicator (SE) using an artificial dataset and a real marketing dataset.
Read moreDynamic Model of Visual Recognition Predicts Neural Response Properties in the Visual Cortex
The responses of visual cortical neurons during fixation tasks can be significantly modulated by stimuli from beyond the classical receptive field. Modulatory effects in neural responses have also been recently reported in a task where a monkey freely views a natural scene. In this article, we describe a hierarchical network model of visual recognition that explains these experimental observations by using a form of the extended Kalman filter as given by the minimum description length (MDL) principle. The model dynamically combines input-driven bottom-up signals with expectation-driven top-down signals to predict current recognition state. Synaptic weights in the model are adapted in a Hebbian manner according to a learning rule also derived from the MDL principle. The resulting prediction-learning scheme can be viewed as implementing a form of expectation-maximization (EM) algorithm. The architecture of the model posits an active computational role of the reciprocal connections between adjoining visual cortical areas in determining neural response properties. In particular, the model demonstrates the possible role of feedback from higher cortical areas in mediating neurophysiological effects due to stimuli from beyond the classical receptive field. Simulations of the model are provided that help explain the experimental observations regarding neural responses in both free viewing and fixation conditions.
Read moreApplication of the minimum description length principle to object-oriented video image compression
Presents a new formulation of moving object segmentation and motion estimation to quantify the potential gain in using object-oriented motion compensation in image sequence coding. Motivated by Rissanen's (1983) minimum description length (MDL) principle that allows unified treatment of coding model complexity and parameter values, an objective function is constructed in terms of an ideal coding length function for each motion-compensated frame. Based on this new objective function, a procedure to segment and estimate moving objects in a scene is proposed. Each object in the scene is represented by its boundaries, motion parameters, and motion-compensated prediction error (MCPE). A number of experimental comparisons between block-oriented and object oriented coding schemes suggest that significant potential coding gain using the new MDL-based criterion is possible. >
Read moreMDL Based Structure Selection of Union of Ellipse Models for Scaled and Smoothed Histological Images
In this chapter, we investigate refinements to the structure selection method used in our recently developed minimum description length (MDL) method for interpreting clumps of nuclei in histological images. We start from the SNEF method, which fits elliptical shapes to the clump image based on the extracted contours and on the image gradient information. Introducing some variability in the parameters of the algorithm, we obtain a number of competing interpretations and we select the least redundant interpretation based on the MDL principle, where the description codelengths are evaluated by a simple implementable coding scheme. We investigate in this paper two ways for allowing additional variability in the basic SNEF method: first by utilizing a pre-processing stage of smoothing the original image using various degrees of smoothing and second by using re-scaling of the original image at various downsizing scales. Both transformations have the potential to hide artifacts and features of the original image that prevented the proper interpretation of the nuclei shapes, and we show experimentally that the set of candidate segmentations obtained will contain variants with better MDL values than the MDL of the initial SNEF segmentations. We compare the results of the automatic interpretation algorithm against the ground truth defined by annotations of human subjects.
Read moreRecovering High-Level Structure of Software Systems Using a Minimum Description Length Principle
In [12] a system was described for finding good hierarchical decompositions of complex systems represented as collections of nodes and links, using a genetic algorithm, with an information theoretic fitness function (representing complexity) derived from a minimum description length principle. This paper describes the application of this approach to the problem of reverse engineering the high-level structure of software systems.KeywordsGenetic AlgorithmDestination NodeTuring MachineReverse EngineeringDependency GraphThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Read moreFitness Function Comparison for GA-Based Feature Construction
When primitive data representation yields attribute interactions, learning requires feature construction. MFE2/GA, a GA-based feature construction has been shown to learn more accurately than others when there exist several complex attribute interactions. A new fitness function, based on the principle of Minimum Description Length (MDL), is proposed and implemented as part of the MFE3/GA system. Since the individuals of the GA population are collections of new features constructed to change the representation of data, an MDL-based fitness considers not only the part of data left unexplained by the constructed features (errors), but also the complexity of the constructed features as a new representation (theory). An empirical study shows the advantage of the new fitness over other fitness not based on MDL, and both are compared to the performance baselines provided by relevant systems.
Read moreOn the Use of MDL Principle in Gene Expression Prediction
The structure and biological behavior of a cell are determined by the pattern of gene expressions within that cell. The so-called gene prediction problem refers to finding rules, or sets of possible rules, on how certain genes expressions determine the expression level of a given target gene. In this paper, we investigate the gene prediction problem and propose the use of new predictors, selected according to the minimum description length (MDL) principle. We compare the use of Boolean predictors, ternary predictors and perceptron predictors. We resort to MDL as a tool for selecting the proper size of the prediction window. MDL is also well suited for comparing predictors having different complexities. We show that the best description can be achieved by the Boolean and ternary predictors, since they obtain better fitting of the data with a lower complexity of the model. To illustrate the comparison, both synthetic and experimental data are used.
Read more