• Home
  • Search
  • Computational Method for Identifying Malonylation Sites by Using Random Forest Algorithm.
  • Cite Icon1
  • https://doi.org/10.2174/1386207322666181227144318Copy DOI Icon

Computational Method for Identifying Malonylation Sites by Using Random Forest Algorithm.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

As a newly uncovered post-translational modification on the ε-amino group of lysine residue, protein malonylation was found to be involved in metabolic pathways and certain diseases. Apart from experimental approaches, several computational methods based on machine learning algorithms were recently proposed to predict malonylation sites. However, previous methods failed to address imbalanced data sizes between positive and negative samples. In this study, we identified the significant features of malonylation sites in a novel computational method which applied machine learning algorithms and balanced data sizes by applying synthetic minority over-sampling technique. Four types of features, namely, amino acid (AA) composition, position-specific scoring matrix (PSSM), AA factor, and disorder were used to encode residues in protein segments. Then, a two-step feature selection procedure including maximum relevance minimum redundancy and incremental feature selection, together with random forest algorithm, was performed on the constructed hybrid feature vector. An optimal classifier was built from the optimal feature subset, which featured an F1-measure of 0.356. Feature analysis was performed on several selected important features. Results showed that certain types of PSSM and disorder features may be closely associated with malonylation of lysine residues. Our study contributes to the development of computational approaches for predicting malonyllysine and provides insights into molecular mechanism of malonylation.

Similar Papers
  • PDF
  • Research Article
  • Citations19

Predicting A-to-I RNA editing by feature selection and random forest.

  • Oct 22, 2014
  • PLoS ONE
  • Yang Shu +4
  • PDF
  • Research Article
  • Citations18

Incorporating hybrid models into lysine malonylation sites prediction on mammalian and plant proteins

  • Jun 29, 2020
  • Scientific Reports
  • Chia-Ru Chung +6
  • Research Article
  • Citations25

RF-MaloSite and DL-Malosite: Methods based on random forest and deep learning to identify malonylation sites

  • Jan 01, 2020
  • Computational and Structural Biotechnology Journal
  • Hussam Al-Barakati +5
  • Research Article
  • Citations7

Computational Prediction of Protein Epsilon Lysine Acetylation Sites Based on a Feature Selection Method.

  • Oct 23, 2017
  • Combinatorial Chemistry & High Throughput Screening
  • Jianzhao Gao +5
  • Research Article
  • Citations17

Recursive DBPSO for Computationally Efficient Electronic Nose System

  • Jan 01, 2018
  • IEEE Sensors Journal
  • Atiq Ur Rehman +1
  • PDF
  • Research Article
  • Citations17

Prediction of protein motions from amino acid sequence and its application to protein-protein interaction

  • Jan 01, 2010
  • BMC Structural Biology
  • Shuichi Hirose +6
  • Research Article
  • Citations16

Application of Machine Learning Algorithms to Identify Problematic Nuclear Data

  • Jul 18, 2021
  • Nuclear Science and Engineering
  • Pavel A Grechanuk +2
  • Research Article
  • Citations135

StackDPPred: a stacking based prediction of DNA-binding protein from sequence.

  • Jul 19, 2018
  • Bioinformatics
  • Avdesh Mishra +2
  • Research Article
  • Citations28

Relevance of Machine Learning Techniques and Various Protein Features in Protein Fold Classification: A Review

  • Dec 13, 2019
  • Current Bioinformatics
  • Komal Patil +1
  • PDF
  • Research Article
  • Citations28

A Two-Step Feature Selection Method to Predict Cancerlectins by Multiview Features and Synthetic Minority Oversampling Technique.

  • Jan 01, 2018
  • BioMed Research International
  • Runtao Yang +3
  • Conference Article
  • Citations24

Detecting scareware by mining variable length instruction sequences

  • Aug 01, 2011
  • Raja Khurram Shahzad +1
  • PDF
  • Research Article
  • Citations47

Prediction of Protein Cleavage Site with Feature Selection by Random Forest

  • Sep 18, 2012
  • PLoS ONE
  • Bi-Qing Li +3
  • Research Article
  • Citations1

Identification of the malonylation modification in Staphylococcus aureus and insight into the regulators in biofilm formation.

  • Jan 01, 2025
  • Frontiers in microbiology
  • Xiaoyan Yu +6
  • PDF
  • Front Matter
  • Citations19

Computational systems biology methods in molecular biology, chemistry biology, molecular biomedicine, and biopharmacy.

  • Jan 01, 2014
  • BioMed Research International
  • Yudong Cai +3
  • PDF
  • Research Article
  • Citations6

Identifying In Vitro Cultured Human Hepatocytes Markers with Machine Learning Methods Based on Single-Cell RNA-Seq Data.

  • May 30, 2022
  • Frontiers in bioengineering and biotechnology
  • Zhandong Li +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.