- Research Article
39
- 10.1016/j.jocs.2011.06.008
A time series approach for clustering mass spectrometry data
- Jul 12, 2011
- Journal of Computational Science
- Francesco Gullo + 4 more +4
A time series approach for clustering mass spectrometry data
Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.
A time series approach for clustering mass spectrometry data
A time series approach for clustering mass spectrometry data
Interlaboratory Comparison of Untargeted Mass Spectrometry Data Uncovers Underlying Causes for Variability.
Despite the value of mass spectrometry in modern natural products discovery workflows, it remains very difficult to compare data sets between laboratories. In this study we compared mass spectrometry data for the same sample set from two different laboratories (quadrupole time-of-flight and quadrupole-Orbitrap) and evaluated the similarity between these two data sets in terms of both mass spectrometry features and their ability to describe the chemical composition of the sample set. Somewhat surprisingly, the two data sets, collected with appropriate controls and replication, had very low feature overlap (25.7% of Laboratory A features overlapping 21.8% of Laboratory B features). Our data clearly demonstrate that differences in fragmentation, charge state, and adduct formation in the ionization source are a major underlying cause for these differences. Consistent with other recent literature, these findings challenge the conventional wisdom that electrospray ionization mass spectrometry (ESI-MS) yields a simple one-to-one correspondence between analytes in solution and features in the data set. Importantly, despite low overlap in feature lists, principal component analysis (PCA) generated qualitatively similar PCA plots. Overall, our findings demonstrate that comparing untargeted metabolomics data between laboratories is challenging, but that data sets with low feature overlap can yield the same qualitative description of a sample set using PCA.
Read moreIsoCor: isotope correction for high-resolution MS labeling experiments.
Mass spectrometry (MS) is widely used for isotopic studies of metabolism and other (bio)chemical processes. Quantitative applications in systems and synthetic biology require to correct the raw MS data for the contribution of naturally occurring isotopes. Several tools are available to correct low-resolution MS data, and recent developments made substantial improvements by introducing resolution-dependent correction methods, hence opening the way to the correction of high-resolution MS (HRMS) data. Nevertheless, current HRMS correction methods partly fail to determine which isotopic species are resolved from the tracer isotopologues and should thus be corrected. We present an updated version of our isotope correction software (IsoCor) with a novel correction algorithm which ensures to accurately exploit any chemical species with any isotopic tracer, at any MS resolution. IsoCor v2 also includes a novel graphical user interface for intuitive use by end-users and a command-line interface to streamline integration into existing pipelines. IsoCor v2 is implemented in Python 3 and was tested on Windows, Unix and MacOS platforms. The source code and the documentation are freely distributed under GPL3 license at https://github.com/MetaSys-LISBP/IsoCor/ and https://isocor.readthedocs.io/.
Read moreB-274 Determining Novel Biomarker Candidates Associated With Heroin and Cocaine Exposure Using Urine Metabolomics
Background Substance abuse poses an escalating health risk in the United States. While acute drug exposure can be tracked using drugs and their metabolites as biomarkers in clinical settings, there is a notable absence of biomarkers for chronic drug exposure or addiction. The development and availability of such biomarkers could significantly enhance clinicians’ ability to identify patients at high risk of addiction, and more accurately predict disease progression and treatment responses. Methods Our study aimed to discover novel biomarkers associated with substance abuse by analyzing raw mass spectrometry (MS) data from urine comprehensive drug screening (UCDS) conducted at the University of Pittsburgh Medical Center (UPMC). The MS data, acquired through untargeted data collection using LC-qToF-MS at the UPMC Clinical Toxicology Laboratory, was initially preprocessed (including peak identification, alignment, and putative annotation) using MS-DIAL to generate a data matrix. This matrix, consisting of objects (cases) and features (m/z and retention time), was further normalized and analyzed using MetaboAnalyst. A combination of statistical methods, including point biserial correlation analysis, volcano plot, significance analysis of microarrays and metabolites (SAM), empirical Bayesian analysis of microarrays and metabolites (EBAM), partial least squares discriminant analysis (PLS-DA), and random forest (RF), were employed to select features significantly associated with the EIA-results. These features were subsequently evaluated for analyte identification using MS-FINDER. Results Our analysis identified nine significant features associated with 6MAM-EIA for heroin abuse, including 6-monoacetylmorphine itself and norfentanyl. For features associated with cocaine metabolite-EIA, 37 significant features were selected, including multiple cocaine metabolites, nicotine metabolites, and norfentanyl. Some of these selected features are likely products of in-source fragmentation or endogenous metabolites. Conclusions Our study has identified a set of potential biomarker candidates for illicit drug exposure. Notably, norfentanyl was found to be significantly associated with both 6MAM-EIA and cocaine metabolite-EIA, reflecting current trends in illicit drug use. Further chemical identification of these biomarkers is planned for future work.
Read morePlasma biomarkers of abdominal aortic aneurysm
Plasma biomarkers of abdominal aortic aneurysm
Chapter 6 - Centrifugal partition chromatography isolation of glaucolides sesquiterpenes and LC-ESIMS/MS technique for differentiation the mass spectrometry behavior of hirsutinolide and glaucolide skeletons
Chapter 6 - Centrifugal partition chromatography isolation of glaucolides sesquiterpenes and LC-ESIMS/MS technique for differentiation the mass spectrometry behavior of hirsutinolide and glaucolide skeletons
Read moreDifferent software processing affects the peak picking and metabolic pathway recognition of metabolomics data
Different software processing affects the peak picking and metabolic pathway recognition of metabolomics data
Using solid-phase extraction to facilitate a focused tile-based Fisher ratio analysis of comprehensive two-dimensional gas chromatography time-of-flight mass spectrometry data: comparative analysis of aerospace fuel composition.
Tile-based Fisher ratio (F-ratio) analysis of comprehensive two-dimensional gas chromatography time-of-flight mass spectrometry (GC × GC-TOFMS) data is a powerful, supervised discovery methodology for pinpointing sample class-distinguishing analytes between two or more sample classes. Herein, we extend this analytical methodology to focus upon specific chemical groups in kerosene-based aerospace fuel using solid-phase extraction (SPE). Treating samples with SPE removes specific compounds depending on the SPE stationary phase (i.e., silica), creating an altered "pass" sample, identical to the original "neat" sample except for the extracted compounds. Application of F-ratio analysis to the neat samples against the pass samples provides global discovery with a numerically sorted hit list of all analytes affected by the SPE procedure. Sections of GC × GC-TOFMS data from the top analyte hits are reconstructed to form a "stitch" chromatogram to visualize the sample class-distinguishing compounds, revealing excellent agreement with the extract chromatogram. Additionally, utilizing the four-grid tiling scheme developed for tile-based F-ratio analysis, we demonstrate a tile-based pairwise analysis method, referred to as 1v1 analysis, to discover analytes that differ in concentration between two fuel chromatograms. Application of 1v1 analysis is highly efficient since replicates do not necessarily need to be run on the GC × GC-TOFMS instrument, which is beneficial for sample-limited applications. The 1v1 analyses discovered most of the same features as F-ratio analysis, ranging from 69 to 81% of the features discovered by F-ratio analysis while requiring one-sixth the data. Lastly, the overall methodology is applied to three candidate rocket fuels to better understand the compound class-distinguishing differences. The separate hit lists produced for high-concentration bulk hydrocarbon differences and low-concentration level polar compound differences provided valuable insight into these candidate rocket fuels.
Read moreMS-Decipher: a user-friendly proteome database search software with an emphasis on deciphering the spectra of O-linked glycopeptides.
The interpretation of mass spectrometry (MS) data is a crucial step in proteomics analysis, and the identification of post-translational modifications (PTMs) is vital for the understanding of the regulation mechanism of the living system. Among various PTMs, glycosylation is one of the most diverse ones. Though many search engines have been developed to decipher proteomic data, some of them are difficult to operate and have poor performance on glycoproteomic datasets compared to advanced glycoproteomic software. To simplify the analysis of proteomic datasets, especially O-glycoproteomic datasets, here, we present a user-friendly proteomic database search platform, MS-Decipher, for the identification of peptides from MS data. Two scoring schemes can be chosen for peptide-spectra matching. It was found that MS-Decipher had the same sensitivity and confidence in peptide identification compared to traditional database searching software. In addition, a special search mode, O-Search, is integrated into MS-Decipher to identify O-glycopeptides for O-glycoproteomic analysis. Compared with Mascot, MetaMorpheus and MSFragger, MS-Decipher can obtain about 139.9%, 48.8% and 6.9% more O-glycopeptide-spectrum matches. A useful tool is provided in MS-Decipher for the visualization of O-glycopeptide-spectra matches. MS-Decipher has a user-friendly graphical user interface, making it easier to operate. Several file formats are available in the searching and validation steps. MS-Decipher is implemented with Java, and can be used cross-platform. MS-Decipher is freely available at https://github.com/DICP-1809/MS-Decipher for academic use. For detailed implementation steps, please see the user guide. Supplementary data are available at Bioinformatics online.
Read moreAdductHunter: identifying protein-metal complex adducts in mass spectra
Mass spectrometry (MS) is an analytical technique for molecule identification that can be used for investigating protein-metal complex interactions. Once the MS data is collected, the mass spectra are usually interpreted manually to identify the adducts formed as a result of the interactions between proteins and metal-based species. However, with increasing resolution, dataset size, and species complexity, the time required to identify adducts and the error-prone nature of manual assignment have become limiting factors in MS analysis. AdductHunter is a open-source web-based analysis tool that automates the peak identification process using constraint integer optimization to find feasible combinations of protein and fragments, and dynamic time warping to calculate the dissimilarity between the theoretical isotope pattern of a species and its experimental isotope peak distribution. Empirical evaluation on a collection of 22 unique MS datasetsshows fast and accurate identification of protein-metal complex adducts in deconvoluted mass spectra.
Read moreUntargeted high-resolution paired mass distance data mining for retrieving general chemical relationships
Untargeted metabolomics analysis captures chemical reactions among small molecules. Common mass spectrometry-based metabolomics workflows first identify the small molecules significantly associated with the outcome of interest, then begin exploring their biochemical relationships to understand biological fate or impact. We suggest an alternative by which general chemical relationships including abiotic reactions can be directly retrieved through untargeted high-resolution paired mass distance (PMD) analysis without a priori knowledge of the identities of participating compounds. PMDs calculated from the mass spectrometry data are linked to chemical reactions obtained via data mining of small molecule and reaction databases, i.e. ‘PMD-based reactomics’. We demonstrate applications of PMD-based reactomics including PMD network analysis, source appointment of unknown compounds, and biomarker reaction discovery as complements to compound discovery analyses used in traditional untargeted workflows. An R implementation of reactomics analysis and the reaction/PMD databases is available as the pmd package.
Read moreMSMCE: A novel representation module for classification of raw mass spectrometry data.
Mass spectrometry (MS) analysis plays a crucial role in the biomedical field; however, the high dimensionality and complexity of MS data pose significant challenges for feature extraction and classification. Deep learning has become a dominant approach in data analysis, and while some deep learning methods have achieved progress in MS classification, their feature representation capabilities remain limited. Most existing methods rely on single-channel representations, which struggle to effectively capture structural information within MS data. To address these limitations, we propose a Multi-Channel Embedding Representation Module (MSMCE), which focuses on modeling inter-channel dependencies to generate multi-channel representations of raw MS data. Additionally, we implement a feature fusion mechanism by concatenating the initial encoded representation with the multi-channel embeddings along the channel dimension, significantly enhancing the classification performance of subsequent models. Experimental results on four public datasets demonstrate that the proposed MSMCE module not only achieves substantial improvements in classification performance but also enhances computational efficiency and training stability, highlighting its effectiveness in raw MS data classification and its potential for robust application across diverse datasets.
Read moreStructural proteomics of a bacterial mega membrane protein complex: FtsH-HflK-HflC
Structural proteomics of a bacterial mega membrane protein complex: FtsH-HflK-HflC
Abstract 4528: Quantitative mass spectrometry to interrogate proteomic heterogeneity in metastatic lung adenocarcinoma and validate a novel somatic mutation CDK12-G879V
Lung cancer is the leading cause of cancer mortality. Tumor heterogeneity is a major cause of treatment failure. Intra- and inter-metastatic tumor heterogeneity has been demonstrated by next generation sequencing (NGS) studies. However, heterogeneity in the proteome and phosphoproteome has been less studied. Integrated proteogenomics is essential to understanding the intricacies of tumor heterogeneity affecting treatment response. Here, we performed integrated mass spectrometry-based proteogenomics to characterize spatial and temporal heterogeneity of an exceptional responder lung adenocarcinoma patient who survived with metastatic disease for more than 7 years while on combination treatment with HER2-targeted and chemotherapy. We employed Super-SILAC and TMT labeling strategies to quantify the proteome and phosphoproteome of a lung metastatic site and ten different metastatic progressive lymph nodes from our patient collected across a span of seven years, including at autopsy. To further interrogate the mass spectrometry data, patient-specific database was built to incorporate all the somatic variants identified by NGS. An extensive validation pipeline was built for confirmation of variant peptides. CRISPR-Cas9-mediated gene knockout, cell viability assays, and confocal microscopy were used for further validation of novel variants. A total of 6214 and 4061 proteins were identified from Super-SILAC and TMT experiments, respectively. 3648 proteins were identified and quantified in both experiments. More than 2000 proteins had catalytic activity, including kinases, phosphatases and metabolic enzymes. We identified 78 and 23 mutant peptides from Super-SILAC and TMT experiments, respectively. Three somatic variants, CDK12-G879V, FASN-R1439Q and HNRNPF-A105T, were confirmed using our variant peptide detection pipeline. Multiple reaction monitoring in a triple quadrupole mass spectrometer successfully identified and relatively quantified two of the variant tryptic peptides harboring the mutations, CDK12-G879V and FASN-R1439Q from the lung and lymph node metastatic sites, respectively. We investigated the consequences of loss of CDK12 function, as predicted from the novel CDK12-G879V mutant, in chemotherapy sensitivity. A549 lung adenocarcinoma cells, upon knockdown of CDK12, exhibited greater chemotherapy sensitivity that was rescued by wild type CDK12, but not by CDK12-G879V mutant. We demonstrate the importance of integrated proteogenomic analyses to identify variant peptides in mass spectrometry data and studying proteomic heterogeneity affecting treatment response. CDK12-G879V mutation results in a nonfunctional CDK12 kinase and chemotherapy susceptibility in lung metastatic sites, likely explaining the “cure” of lung metastatic sites in this patient. Citation Format: Xu Xhang, Khoa Dang P. Nguyen, Paul Rudnick, Nitin Roper, Emily Kawaler, Tapan K. Maity, Shivangi Awasthi, Shaojian Gao, Romi Biswas, Abhilash Venugopalan, Constance Cultraro, David Fenyo, Udayan Guha. Quantitative mass spectrometry to interrogate proteomic heterogeneity in metastatic lung adenocarcinoma and validate a novel somatic mutation CDK12-G879V [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 4528.
Read moreDevelopment of Stereotactic Mass Spectrometry for Brain Tumor Surgery
Surgery remains the first and most important treatment modality for the majority of solid tumors. Across a range of brain tumor types and grades, postoperative residual tumor has a great impact on prognosis. The principal challenge and objective of neurosurgical intervention is therefore to maximize tumor resection while minimizing the potential for neurological deficit by preserving critical tissue. To introduce the integration of desorption electrospray ionization mass spectrometry into surgery for in vivo molecular tissue characterization and intraoperative definition of tumor boundaries without systemic injection of contrast agents. Using a frameless stereotactic sampling approach and by integrating a 3-dimensional navigation system with an ultrasonic surgical probe, we obtained image-registered surgical specimens. The samples were analyzed with ambient desorption/ionization mass spectrometry and validated against standard histopathology. This new approach will enable neurosurgeons to detect tumor infiltration of the normal brain intraoperatively with mass spectrometry and to obtain spatially resolved molecular tissue characterization without any exogenous agent and with high sensitivity and specificity. Proof of concept is presented in using mass spectrometry intraoperatively for real-time measurement of molecular structure and using that tissue characterization method to detect tumor boundaries. Multiple sampling sites within the tumor mass were defined for a patient with a recurrent left frontal oligodendroglioma, World Health Organization grade II with chromosome 1p/19q codeletion, and mass spectrometry data indicated a correlation between lipid constitution and tumor cell prevalence. The mass spectrometry measurements reflect a complex molecular structure and are integrated with frameless stereotaxy and imaging, providing 3-dimensional molecular imaging without systemic injection of any agents, which can be implemented for surgical margins delineation of any organ and with a rapidity that allows real-time analysis.
Read more