• Home
  • Search
  • Proteome Analyst - Transparent High-throughput Protein Annotation: Function, Localization and Custom Predictors
  • Cite Icon10
  • https://doi.org/10.7939/r37m0415rCopy DOI Icon

Proteome Analyst - Transparent High-throughput Protein Annotation: Function, Localization and Custom Predictors

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Motivation: Modern sequencing technology now permits the sequencing of entire genomes, leading to thousands of new gene sequences in need of detailed annotation. It is too time consuming to predict the properties of each protein sequence manually and to organize the results of many prediction tools by hand. The prediction process must be automated so the predictions can be automatically organized, but the predictions must also be transparent. That is, the rationale for each prediction should be easily examinable by anyone that wishes to use the prediction. Results: Proteome Analyst (PA) is a web-based system for predicting the properties of each protein in a proteome. PA has three interesting features. First, it provides a single web-based tool that allows the user to select a wide range of analytic tools and automatically apply them to each protein in a proteome. In essence, PA provides one-stop automatic high-throughput analysis. Second, PA has the ability to explain its predictions to users. PA is based on established machine learning techniques, but makes every prediction transparent to its users. Third, PA allows users to easily create their own transparent custom predictors without programming. Availability: http://www.cs.ualberta.ca/~bioinfo/PA Supplementary Information: http://www.cs.ualberta.ca/~bioinfo/PA/Walkthrough http://www.cs.ualberta.ca/~bioinfo/PA/Experiments Contact: bioinfo@cs.ualberta.ca INTRODUCTION High-throughput sequencing technology has made it possible for even the smallest laboratory to sequence the genome of organelles, viruses and bacteria. Indeed there are now more than 1200 genome sequences deposited in public databases (http://www.ebi.ac.uk/genomes/) and this number is growing rapidly. Given the size and complexity of these data sets, most researchers are compelled to use automated annotation systems to identify or classify individual genes and proteins in their genomic data. A number of excellent systems have been developed over the past few years that permit automated genome-wide or proteome-wide annotation. These include GeneQuiz (Andrade et al., 1999), GeneAtlas (Kitson et al., 2002), EnsEMBL (Hubbard et al., 2002), PEDANT (Frishman et al., 2001), Genotator (Harris, 1997), Magpie (Gaasterland and Sensen, 1996) and GAIA (Overton et al., 1998). These use web-based tools to identify genes, parse data, translate sequences, search against public databases, identify domains or motifs and perform predictive analyses of protein sequences. Many of these packages provide user-customizable search schemas and graphical, hyperlinked output. The level of interpretation or inference offered by these annotation systems varies widely, with some offering only raw data (lists of homologues, calculated properties, etc.) in a consolidated format and others inferring function or ontology through detailed lexical analysis. Our PA system focuses on the task of prediction or classification. Our results show that classification/prediction can be used for many annotations. The most obvious is general-function, but we are also working on sub-cellular localization, specific function and protein-protein interaction are others we are working on. To use classification, we require an ontology or “dictionary” of class terms. Consequently, a number of controlled vocabularies or ontologies of protein function have started appearing. Among the first was the enzyme classification scheme (E.C. number) developed by the IUBMB and first employed in the ENZYME database (Bairoch, 1993). Later, Riley (1993) developed functional classification schemas for E. coli and other microbes. Variations on these functional vocabularies were added as other genomes and genome annotation systems were

Similar Papers
  • Research Article

Completion of human Chromosome 21, the Human Genome Project, and Steps towards Understanding Ourselves through Comparative Genomics

  • Sep 01, 2000
  • Journal of Genetics and Molecular Biology
  • Todd D Taylor
  • Research Article

Integration of Alignment and Phylogeny in the Whole-Genome Era

  • Jun 18, 2015
  • Open Scholarship Institutional Repository (Washington University in St. Louis)
  • Hongying Sun
  • Supplementary Content

A bioinformatic analysis of Mycobacterium tuberculosis and host genomic data

  • Jan 09, 2018
  • LSHTM Research Online (London School of Hygiene and Tropical Medicine)
  • Jody Phelan
  • Research Article
  • Citations37

Integrating genomic data to predict transcription factor binding.

  • Jan 01, 2005
  • Genome Informatics
  • Dustin T Holloway +2
  • Research Article

Prediction and analysis on secreted proteins of Echinococcus multilocularis by genome-wide bioinformatics approaches

  • Nov 28, 2014
  • Int J Med Parasit Dis
  • Ting Zhang +4
  • Research Article

Molecular interactomes: Network-guided cancer prognosis prediction a multi-way chromatin interaction analysis

  • Nov 12, 2018
  • Research Repository (Delft University of Technology)
  • Amin Allahyar
  • Research Article

پیشبینی ژنهای کاذب جدید در ژنوم مرجع گوسفند

  • Mar 21, 2018
  • SHILAP Revista de lepidopterología
  • محمد رضا بختیاری زاده +2
  • Research Article
  • Citations1

Effects of sample processing modes on high-throughput sequencing of coronavirus whole genome

  • Feb 29, 2020
  • Chinese journal of microbiology and immunology
  • Chen Jing +6
  • Research Article

What do 1000 genomes tell us about biased gene conversion theory

  • Jun 26, 2013
  • F1000Research
  • Arnab Saha Mandal +3
  • Research Article

Detection and analysis on complete genomic sequence of human bocavirus GZ2010-1 strain in Guangzhou

  • Aug 15, 2011
  • International Medicine and Health Guidance News
  • Jia‐Yu Zhong +5
  • Research Article

The application progress in whole-genome and exome sequencings of liver cancer

  • Nov 28, 2017
  • Chinese Journal of Hepatobiliary Surgery
  • Wei Chen
  • Supplementary Content

Characterization of the cytochrome [beta] gene in plant pathogenic, basidiomycetes and consequences for QoI resistance

  • Jan 01, 2005
  • edoc (University of Basel)
  • V Grasso
  • Research Article

Application of whole genome exon sequencing in type 2 diabetes mellitus

  • Mar 15, 2018
  • Chin J Clinicians(Electronic Edition)
  • Xiaowei Zhou
  • Research Article
  • Citations11

The Importance of Marine Genomics to Life

  • Jan 23, 2015
  • Popoola Raimot Titilade +1
  • Research Article
  • Citations6

Population genomics study of Vibrio alginolyticus.

  • Apr 20, 2021
  • Yi chuan = Hereditas
  • Hong Yuan Zheng +10
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.