• Home
  • Search
  • Maximal Words in Sequence Comparisons Based on Subword Composition
  • Cite Icon21
  • https://doi.org/10.1007/978-3-642-12476-1_2Copy DOI Icon

Maximal Words in Sequence Comparisons Based on Subword Composition

  • Jan 1, 2010
  • Alberto Apostolico
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Measures of sequence similarity and distance based more or less explicitly on subword composition are attracting an increasing interest driven by intensive applications such as massive document classification and genome-wide molecular taxonomy. A uniform character of such measures is in some underlying notion of relative compressibility, whereby two similar sequences are expected to share a larger number of common substrings than two distant ones. This paper reviews some of the approaches to sequence comparison based on subword composition and suggests that their common denominator may ultimately reside in special classes of subwords, the nature of which resonates in interesting ways with the structure of popular subword trees and graphs.

Similar Papers
  • Research Article
  • Citations1

Clustering Molecular Sequences with Their Components.

  • Jan 01, 1997
  • Genome Informatics
  • Matsuda +3
  • Research Article
  • Citations2

Scalable neighbour search and alignment with uvaia.

  • Mar 06, 2024
  • PeerJ
  • Leonardo De Oliveira Martins +2
  • Research Article
  • Citations103

CLEANUP: a fast computer program for removing redundancies from nucleotide sequence databases.

  • Jan 01, 1996
  • Computer applications in the biosciences : CABIOS
  • Giorgio Grillo +3
  • Research Article
  • Citations24

Characterizing the D2 Statistic: Word Matches in Biological Sequences

  • Jan 08, 2009
  • Statistical Applications in Genetics and Molecular Biology
  • Sylvain Forêt +2
  • Research Article
  • Citations3

A novel index which precisely derives protein coding regions from cross-species genome alignments.

  • Jan 01, 2002
  • Genome Informatics
  • Tetsushi Yada +2
  • Research Article
  • Citations13

A Novel Similarity Measure for Sequence Data

  • Sep 30, 2011
  • Journal of Information Processing Systems
  • M Pandi +2
  • Research Article
  • Citations8

Host–parasite relations of bacteria and phages can be unveiled by Oligostickiness, a measure of relaxed sequence similarity

  • Jan 06, 2009
  • Bioinformatics
  • Shamim Ahmed +4
  • Research Article
  • Citations3

Cluster analysis of cancer data using semantic similarity, sequence similarity and biological measures

  • Aug 12, 2014
  • Network Modeling Analysis in Health Informatics and Bioinformatics
  • Sajid Nagi +1
  • Research Article
  • Citations14

A review of alignment based similarity measures for web usage mining

  • May 28, 2019
  • Artificial Intelligence Review
  • Vinh-Trung Luu +5
  • PDF
  • Research Article
  • Citations8

Context-Aware Practice Problem Recommendation Using Learners’ Skill Level Navigation Patterns

  • Jan 01, 2023
  • Intelligent Automation & Soft Computing
  • P N Ramesh +1
  • Conference Article
  • Citations5

Web usage prediction and recommendation using web session clustering

  • Sep 01, 2016
  • Vinh-Trung Luu +4
  • Book Chapter
  • Citations5

Similarity Search for Interval Time Sequences

  • Jan 01, 2004
  • Byoung-Kee Yi +1
  • Research Article
  • Citations18

A New Similarity Metric for Sequential Data

  • Oct 01, 2010
  • International Journal of Data Warehousing and Mining
  • Pradeep Kumar +2
  • Book Chapter
  • Citations9

An Efficient Collaborative Recommender System for Removing Sparsity Problem

  • Jan 01, 2020
  • Avita Fuskele Jain +2
  • Research Article

EMap 2.0: A Web-Based Platform for Identifying electron Transfer Pathways in Proteins and Protein Families.

  • Feb 01, 2026
  • Wiley interdisciplinary reviews. Computational molecular science
  • James R Gayvert +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.