- Research Article
3
- 10.1016/j.ins.2010.02.008
On the bottleneck tree alignment problems
- Feb 06, 2010
- Information Sciences
- Yen Hung Chen + 1 more +1
On the bottleneck tree alignment problems
Given a set W of k sequences (strings) and a tree structure T with k leaves, each of which is labeled with a unique sequence in W, a tree alignment is to label a sequence to each internal node of T. The weight of an edge of the tree alignment is the distance, such as Hamming distance, Levenshtein (edit) distance or reversal distance, between the two sequences labeled to the two ends of the edge. The bottleneck tree alignment problem is to find a tree alignment such that the weight of the largest edge is minimized. A lifted tree alignment is a tree alignment, where each internal node v is labeled one of the sequences that was labeled to the children of v. The bottleneck lifted tree alignment problem is to find a lifted tree alignment such that the weight of the largest edge is minimized. In this paper, we show that the bottleneck tree alignment problem is NP-complete even when the tree structure is the binary tree and the weight function is metric. For special cases, we present an exact algorithm to solve the bottleneck lifted tree alignment problem in polynomial time. If the weight function is ultrametric, we show that any lifted tree alignment is an optimal bottleneck tree alignment.KeywordsEdit distancebottleneck tree alignmentmetricultrametricNP-complete
On the bottleneck tree alignment problems
On the bottleneck tree alignment problems
Pair hidden Markov models on tree structures.
Computationally identifying non-coding RNA regions on the genome has much scope for investigation and is essentially harder than gene-finding problems for protein-coding regions. Since comparative sequence analysis is effective for non-coding RNA detection, efficient computational methods are expected for structural alignments of RNA sequences. On the other hand, Hidden Markov Models (HMMs) have played important roles for modeling and analysing biological sequences. Especially, the concept of Pair HMMs (PHMMs) have been examined extensively as mathematical models for alignments and gene finding. We propose the pair HMMs on tree structures (PHMMTSs), which is an extension of PHMMs defined on alignments of trees and provides a unifying framework and an automata-theoretic model for alignments of trees, structural alignments and pair stochastic context-free grammars. By structural alignment, we mean a pairwise alignment to align an unfolded RNA sequence into an RNA sequence of known secondary structure. First, we extend the notion of PHMMs defined on alignments of 'linear' sequences to pair stochastic tree automata, called PHMMTSs, defined on alignments of 'trees'. The PHMMTSs provide various types of alignments of trees such as affine-gap alignments of trees and an automata-theoretic model for alignment of trees. Second, based on the observation that a secondary structure of RNA can be represented by a tree, we apply PHMMTSs to the problem of structural alignments of RNAs. We modify PHMMTSs so that it takes as input a pair of a 'linear' sequence and a 'tree' representing a secondary structure of RNA to produce a structural alignment. Further, the PHMMTSs with input of a pair of two linear sequences is mathematically equal to the pair stochastic context-free grammars. We demonstrate some computational experiments to show the effectiveness of our method for structural alignments, and discuss a complexity issue of PHMMTSs.
Read moreAlignment of trees — an alternative to tree edit
Alignment of trees — an alternative to tree edit
The tree alignment problem.
BackgroundThe inference of homologies among DNA sequences, that is, positions in multiple genomes that share a common evolutionary origin, is a crucial, yet difficult task facing biologists. Its computational counterpart is known as the multiple sequence alignment problem. There are various criteria and methods available to perform multiple sequence alignments, and among these, the minimization of the overall cost of the alignment on a phylogenetic tree is known in combinatorial optimization as the Tree Alignment Problem. This problem typically occurs as a subproblem of the Generalized Tree Alignment Problem, which looks for the tree with the lowest alignment cost among all possible trees. This is equivalent to the Maximum Parsimony problem when the input sequences are not aligned, that is, when phylogeny and alignments are simultaneously inferred.ResultsFor large data sets, a popular heuristic is Direct Optimization (DO). DO provides a good tradeoff between speed, scalability, and competitive scores, and is implemented in the computer program POY. All other (competitive) algorithms have greater time complexities compared to DO. Here, we introduce and present experiments a new algorithm Affine-DO to accommodate the indel (alignment gap) models commonly used in phylogenetic analysis of molecular sequence data. Affine-DO has the same time complexity as DO, but is correctly suited for the affine gap edit distance. We demonstrate its performance with more than 330,000 experimental tests. These experiments show that the solutions of Affine-DO are close to the lower bound inferred from a linear programming solution. Moreover, iterating over a solution produced using Affine-DO shows little improvement.ConclusionsOur results show that Affine-DO is likely producing near-optimal solutions, with approximations within 10% for sequences with small divergence, and within 30% for random sequences, for which Affine-DO produced the worst solutions. The Affine-DO algorithm has the necessary scalability and optimality to be a significant improvement in the real-world phylogenetic analysis of sequence data.
Read morePlanar multifacility location problems with tree structure and finite dominating sets
Planar multifacility location problems with tree structure and finite dominating sets
A Theoretical Analysis of Tree Edit Distance Measures
The notion of the tree edit distance provides a unifying framework for measuring distance and finding approximate common patterns between two trees. A diversity of tree edit distance measures have been proposed to deal with tree related problems, such as minor containment, maximum common subtree isomorphism, maximum common embedded subtree, and alignment of trees. These classes of problems are characterized by the conditions of the tree mappings, which specify how to associate the nodes in one tree with the nodes in the other. In this paper, we study the declarative semantics of edit distance measures based on the tree mapping. In prior work, the edit distance measures have been not well-formalized. So the relationship among various algorithms based on the tree edit distance has hardly been studied. Our framework enables us to study the relationship. By using our framework, we reveal the declarative semantics of the alignment of trees, which has remained unknown in prior work.
Read moreImplementing approximate regularities
Implementing approximate regularities
Iterative pass optimization of sequence data
The problem of determining the minimum-cost hypothetical ancestral sequences for a given cladogram is known to be NP-complete. This "tree alignment" problem has motivated the considerable effort placed in multiple sequence alignment procedures. Wheeler in 1996 proposed a heuristic method, direct optimization, to calculate cladogram costs without the intervention of multiple sequence alignment. This method, though more efficient in time and more effective in cladogram length than many alignment-based procedures, greedily optimizes nodes based on descendent information only. In their proposal of an exact multiple alignment solution, Sankoff et al. in 1976 described a heuristic procedure--the iterative improvement method--to create alignments at internal nodes by solving a series of median problems. The combination of a three-sequence direct optimization with iterative improvement and a branch-length-based cladogram cost procedure, provides an algorithm that frequently results in superior (i.e., lower) cladogram costs. This iterative pass optimization is both computation and memory intensive, but economies can be made to reduce this burden. An example in arthropod systematics is discussed.
Read moreSpaces of algebraic measure trees and triangulations of the circle
In this paper we present with algebraic trees a novel notion of (continuum) trees which generalizes countable graph-theoretic trees to (potentially) uncountable structures. For that purpose we focus on the tree structure given by the branch point map which assigns to each triple of points their branch point. We give an axiomatic definition of algebraic trees, define a natural topology, and equip them with a probability measure on the Borel-$\sigma$-field.Under an order-separability condition, algebraic (measure) trees can be considered as tree structure equivalence classes of metric (measure) trees (i.e.\ subtrees of R-trees). Using Gromov-weak convergence (i.e.\ sample distance convergence) of the particular representatives given by the metric arising from the distribution of branch points, we define a metrizable topology on the space of equivalence classes of algebraic measure trees. In many applications, binary trees are of particular interest. We introduce on that subspace with the sample shape and the sample subtree mass convergence two additional, natural topologies. Relying on the connection to triangulations of the circle, we show that all three topologies are actually the same, and the space of binary algebraic measure trees is compact. To this end, we provide a formal definition of triangulations of the circle, and show that the coding map which sends a triangulation to an algebraic measure tree is a continuous surjection onto the subspace of binary algebraic non-atomic measure trees.
Read moreQuantifying similarity in animal vocal sequences: which metric performs best?
SummaryMany animals communicate using sequences of discrete acoustic elements which can be complex, vary in their degree of stereotypy, and are potentially open‐ended. Variation in sequences can provide important ecological, behavioural or evolutionary information about the structure and connectivity of populations, mechanisms for vocal cultural evolution and the underlying drivers responsible for these processes. Various mathematical techniques have been used to form a realistic approximation of sequence similarity for such tasks.Here, we use both simulated and empirical data sets from animal vocal sequences (rock hyrax,Procavia capensis; humpback whale,Megaptera novaeangliae; bottlenose dolphin,Tursiops truncatus; andCarolina chickadee,Poecile carolinensis) to test which of eight sequence analysis metrics are more likely to reconstruct the information encoded in the sequences, and to test the fidelity of estimation of model parameters, when the sequences are assumed to conform to particular statistical models.Results from the simulated data indicated that multiple metrics were equally successful in reconstructing the information encoded in the sequences of simulated individuals (Markov chains,n‐gram models, repeat distribution and edit distance) and data generated by different stochastic processes (entropy rate andn‐grams). However, the string edit (Levenshtein) distance performed consistently and significantly better than all other tested metrics (including entropy,Markov chains,n‐grams, mutual information) for all empirical data sets, despite being less commonly used in the field of animal acoustic communication.TheLevenshtein distance metric provides a robust analytical approach that should be considered in the comparison of animal acoustic sequences in preference to other commonly employed techniques (such asMarkov chains, hiddenMarkov models orShannon entropy). The recent discovery that non‐Markovian vocal sequences may be more common in animal communication than previously thought provides a rich area for future research that requires non‐Markovian‐based analysis techniques to investigate animal grammars and potentially the origin of human language.
Read morePolynomiality for Bin Packing with a Constant Number of Item Types
We consider the bin packing problem with d different item sizes s i and item multiplicities a i , where all numbers are given in binary encoding. This problem formulation is also known as the one-dimensional cutting stock problem . In this work, we provide an algorithm that, for constant d , solves bin packing in polynomial time. This was an open problem for all d\ge 3 . In fact, for constant d our algorithm solves the following problem in polynomial time: Given two d -dimensional polytopes P and Q , find the smallest number of integer points in P whose sum lies in Q . Our approach also applies to high multiplicity scheduling problems in which the number of copies of each job type is given in binary encoding and each type comes with certain parameters such as release dates, processing times, and deadlines. We show that a variety of high multiplicity scheduling problems can be solved in polynomial time if the number of job types is constant.
Read moreThe Complexity of Finding Optimal Subgraphs to Represent Spatial Correlation
Understanding spatial correlation is vital in many fields including epidemiology and social science. Lee, Meeks and Pettersson (Stat. Comput. 2021) recently demonstrated that improved inference for areal unit count data can be achieved by carrying out modifications to a graph representing spatial correlations; specifically, they delete edges of the planar graph derived from border-sharing between geographic regions in order to maximise a specific objective function. In this paper we address the computational complexity of the associated graph optimisation problem. We demonstrate that this problem cannot be solved in polynomial time unless P = NP; we further show intractability for two simpler variants of the problem. We follow these results with two parameterised algorithms that exactly solve the problem in polynomial time in restricted settings. The first of these utilises dynamic programming on a tree decomposition, and runs in polynomial time if both the treewidth and maximum degree are bounded. The second algorithm is restricted to problem instances with maximum degree three, as may arise from triangulations of planar surfaces, but is an FPT algorithm when the maximum number of edges that can be removed is taken as the parameter.
Read moreOn algebraic connectivity as a function of an edge weight
We consider the Laplacian matrix of a weighted graph, and how the algebraic connectivity, α, behaves when considered as a function of a single edge weight. Under suitable differentiability conditions, we bound the first derivative of α from above, show that α is necessarily concave down, and produce a lower bound on the second derivative of α. When α is simple, we discuss the effect of increasing an edge weight on the corresponding Fiedler vector. We also compute the limiting value of α as the edge weight increases to infinity.
Read moreError Tree: A Tree Structure for Hamming and Edit Distances and Wildcards Matching
Approximate pattern matching is a fundamental problem in the bioinformatics and information retrieval applications. The problem involves different matching relations such as Hamming distance, edit distances, and the wildcards matching problem. The input is usually a text of length n over a fixed alphabet of length Σ, a pattern of length m, and an integer k. The output is to find all positions that have ≤ k Hamming distance, edit distance, or wildcards matching with P. Many algorithms and indexes have been proposed to solve the problems more efficiently, but due to the space and time complexities of the problems, most tools adopted heuristics approaches based on, for instance, suffix tree, suffix array, or Burrows Wheeler Transform to reach practical implementations. Error Tree is a novel tree structure that is mainly oriented to solve the approximate pattern matching problems, using less space and faster computation time. The algorithm proposes for Hamming distance and wildcards matching a tree structure that needs [Formula: see text] words and takes [Formula: see text] in the average case) of query time for any online/offline pattern, where occ is the number of outputs. In addition, a tree structure of [Formula: see text] words and [Formula: see text] in the average case) query time for edit distance for any online/offline pattern.
Read moreQuantum Attacks Against Indistinguishablility Obfuscators Proved Secure in the Weak Multilinear Map Model
We present a quantum polynomial time attack against the GMMSSZ branching program obfuscator of Garg et al. (TCC’16), when instantiated with the GGH13 multilinear map of Garg et al. (EUROCRYPT’13). This candidate obfuscator was proved secure in the weak multilinear map model introduced by Miles et al. (CRYPTO’16). Our attack uses the short principal ideal solver of Cramer et al. (EUROCRYPT’16), to recover a secret element of the GGH13 multilinear map in quantum polynomial time. We then use this secret element to mount a (classical) polynomial time mixed-input attack against the GMMSSZ obfuscator. The main result of this article can hence be seen as a classical reduction from the security of the GMMSSZ obfuscator to the short principal ideal problem (the quantum setting is then only used to solve this problem in polynomial time). As an additional contribution, we explain how the same ideas can be adapted to mount a quantum polynomial time attack against the DGGMM obfuscator of Döttling et al. (ePrint 2016), which was also proved secure in the weak multilinear map model.
Read more