- Research Article
298
- 10.1016/j.tplants.2014.08.004
Machine learning for Big Data analytics in plants.
- Sep 14, 2014
- Trends in Plant Science
- Chuang Ma + 2 more +2
Machine learning for Big Data analytics in plants.
Exponential growth in scientific research data demands novel measures for managing the extremely large datasets. In particular, due to advancements in high-resolution microscopy, the nanoscopy scientific research community is producing datasets up to the range of multiple TeraBytes (TB). Systematically acquired datasets of biological specimens are composed of multiple high-resolution images, in the range of 150-200 TB. The management of these extremely large datasets requires an optimized Generic Client Service (GCS) API with an integration into a data repository system. The novel API proposed in this paper provides an abstract interface that connects various disparate systems. The API is optimized to provide an efficient and automated ingest, download of the data and management of its metadata. The ingest and download processes are based on well-defined workflows stated in this paper. The base metadata model for comprehensive description of the datasets is also stated in the paper. The API is seamlessly integrated with a digital data repository system, namely KIT Data Manager to make it adaptable for a wide range of communities. Finally, a simple and easy to use command line tool is realized based on GCS API to manage large datasets of nanoscopy research community.
Machine learning for Big Data analytics in plants.
Machine learning for Big Data analytics in plants.
RNA CoMPASS: A Dual Approach for Pathogen and Host Transcriptome Analysis of RNA-Seq Datasets
High-throughput RNA sequencing (RNA-seq) has become an instrumental assay for the analysis of multiple aspects of an organism's transcriptome. Further, the analysis of a biological specimen's associated microbiome can also be performed using RNA-seq data and this application is gaining interest in the scientific community. There are many existing bioinformatics tools designed for analysis and visualization of transcriptome data. Despite the availability of an array of next generation sequencing (NGS) analysis tools, the analysis of RNA-seq data sets poses a challenge for many biomedical researchers who are not familiar with command-line tools. Here we present RNA CoMPASS, a comprehensive RNA-seq analysis pipeline for the simultaneous analysis of transcriptomes and metatranscriptomes from diverse biological specimens. RNA CoMPASS leverages existing tools and parallel computing technology to facilitate the analysis of even very large datasets. RNA CoMPASS has a web-based graphical user interface with intrinsic queuing to control a distributed computational pipeline. RNA CoMPASS was evaluated by analyzing RNA-seq data sets from 45 B-cell samples. Twenty-two of these samples were derived from lymphoblastoid cell lines (LCLs) generated by the infection of naïve B-cells with the Epstein Barr virus (EBV), while another 23 samples were derived from Burkitt's lymphomas (BL), some of which arose in part through infection with EBV. Appropriately, RNA CoMPASS identified EBV in all LCLs and in a fraction of the BLs. Cluster analysis of the human transcriptome component of the RNA CoMPASS output clearly separated the BLs (which have a germinal center-like phenotype) from the LCLs (which have a blast-like phenotype) with evidence of activated MYC signaling and lower interferon and NF-kB signaling in the BLs. Together, this analysis illustrates the utility of RNA CoMPASS in the simultaneous analysis of transcriptome and metatranscriptome data. RNA CoMPASS is freely available at http://rnacompass.sourceforge.net/.
Read moreEnhancing student learning in database courses with large data sets
Rapidly increasing storage device capacities at ever decreasing costs have resulted in mushrooming of publicly available large data sets on the Web. In this paper, we describe a novel approach to teaching relational database course by using such data repositories. We demonstrate our approach using the Amazon.com product database, though the approach is generic and is applicable to other data repositories. The Amazon database is supposedly the largest product database ever in existence. We have used the Amazon Web Services API and .NET/C# application to extract a subset of the product database to enhance student learning in a relational database course. This realistic data served various activities of the course and provided a rich backdrop to demonstrate more interesting features of SQL and Oracle cost-based query optimization. Central to the course is a semester-long team project. We discuss the details of data extraction from Amazon.com, conceptual and logical data modeling, logical and physical database design, database creation and data loading, database querying, and database application development.
Read moreUnlocking massively parallel spectral proper orthogonal decompositions in the PySPOD package
Unlocking massively parallel spectral proper orthogonal decompositions in the PySPOD package
Modern Data Containers for Scalable Archiving and Access of Distributed Acoustic Sensing Data
In recent years, Distributed Acoustic Sensing (DAS) has emerged as a powerful technology in seismology, enabling the acquisition of high-resolution seismic data using optical fibers as sensors. As the number of DAS experiments continues to grow and DAS interrogators become able to record along ever longer fiber-optic cables, the volume of generated data is rapidly increasing, creating significant challenges for long-term archiving and efficient data access. These challenges include not only the storage of very large datasets—often on the order of hundreds of terabytes—but also the ability to fast random access and process subsets of the data.Currently, the DAS ecosystem is dominated by proprietary data formats defined by individual vendors. While HDF5 is increasingly adopted as a more open alternative, it presents strong limitations regarding scalability in multi-threaded and multi-process environments. In contrast, modern data container formats such as Zarr and TileDB offer native support for parallel I/O and flexible storage backends, ranging from local file systems to on-premise and cloud-based object storage.In this contribution, we present a comparative evaluation of these modern data formats for DAS applications, focusing on performance, scalability, and usability. We discuss the latest results obtained from the activities of the Geo-INQUIRE* project and assess the feasibility and potential benefits of their adoption for the long-term management and analysis of DAS datasets.* Geo-INQUIRE is funded by the European Union (GA 101058518)
Read moreIs power everything? What can we learn from large data sets
Is power everything? What can we learn from large data sets
Analysis-ready VCF at Biobank scale using Zarr
BackgroundVariant Call Format (VCF) is the standard file format for interchanging genetic variation data and associated quality control metrics. The usual row-wise encoding of the VCF data model (either as text or packed binary) emphasizes efficient retrieval of all data for a given variant, but accessing data on a field or sample basis is inefficient. The Biobank-scale datasets currently available consist of hundreds of thousands of whole genomes and hundreds of terabytes of compressed VCF. Row-wise data storage is fundamentally unsuitable and a more scalable approach is needed.ResultsZarr is a format for storing multidimensional data that is widely used across the sciences, and is ideally suited to massively parallel processing. We present the VCF Zarr specification, an encoding of the VCF data model using Zarr, along with fundamental software infrastructure for efficient and reliable conversion at scale. We show how this format is far more efficient than standard VCF-based approaches, and competitive with specialized methods for storing genotype data in terms of compression ratios and single-threaded calculation performance. We present case studies on subsets of 3 large human datasets (Genomics England: n=78,195; Our Future Health: n=651,050; All of Us: n=245,394) along with whole genome datasets for Norway Spruce (n=1,063) and SARS-CoV-2 (n=4,484,157). We demonstrate the potential for VCF Zarr to enable a new generation of high-performance and cost-effective applications via illustrative examples using cloud computing and GPUs.ConclusionsLarge row-encoded VCF files are a major bottleneck for current research, and storing and processing these files incurs a substantial cost. The VCF Zarr specification, building on widely used, open-source technologies, has the potential to greatly reduce these costs, and may enable a diverse ecosystem of next-generation tools for analysing genetic variation data directly from cloud-based object stores, while maintaining compatibility with existing file-oriented workflows.
Read moreIndexing genomic sequence libraries
Indexing genomic sequence libraries
Annotation as a support to user interaction for content enhancement in digital libraries
This work describes the interface design and interaction of a generic annotation service for Digital Library Management Systems (DLMSs), called Digital Library Annotation Service (DiLAS), that has been designed and is currently undergoing development and user test in the framework of the DELOS European Network of Excellence. The objective of DiLAS is to design and develop an architecture and a framework able to support and evaluate a generic annotation service, i. e. a service that can be easily used into different DLMSs enhancing their User Interfaces (UIs) in order to offer to Digital Library (DL) users a set of uniform, user-tested (under certain required conditions), and recognizable functionalities.
Read more<title>Medical image archive node simulation and architecture</title>
It is a well known fact that managed care and new treatment technologies are revolutionizing the health care provider world. Community Health Information Network and Computer-based Patient Record projects are underway throughout the United States. More and more hospitals are installing digital, `filmless' radiology (and other imagery) systems. They generate a staggering amount of information around the clock. For example, a typical 500-bed hospital might accumulate more than 5 terabytes of image data in a period of 30 years for conventional x-ray images and digital images such as Magnetic Resonance Imaging and Computer Tomography images. With several hospitals contributing to the archive, the storage required will be in the hundreds of terabytes. Systems for reliable, secure, and inexpensive storage and retrieval of digital medical information do not exist today. In this paper, we present a Medical Image Archive and Distribution Service (MIADS) concept. MIADS is a system shared by individual and community hospitals, laboratories, and doctors' offices that need to store and retrieve medical images. Due to the large volume and complexity of the data, as well as the diversified user access requirement, implementation of the MIADS will be a complex procedure. One of the key challenges to implementing a MIADS is to select a cost-effective, scalable system architecture to meet the ingest/retrieval performance requirements. We have performed an in-depth system engineering study, and developed a sophisticated simulation model to address this key challenge. This paper describes the overall system architecture based on our system engineering study and simulation results. In particular, we will emphasize system scalability and upgradability issues. Furthermore, we will discuss our simulation results in detail. The simulations study the ingest/retrieval performance requirements based on different system configurations and architectures for variables such as workload, tape access time, number of drives, number of exams per patient, number of Central Processing Units, patient grouping, and priority impacts. The MIADS, which could be a key component of a broader data repository system, will be able to communicate with and obtain data from existing hospital information systems. We will discuss the external interfaces enabling MIADS to communicate with and obtain data from existing Radiology Information Systems such as the Picture Archiving and Communication System (PACS). Our system design encompasses the broader aspects of the archive node, which could include multimedia data such as image, audio, video, and free text data. This system is designed to be integrated with current hospital PACS through a Digital Imaging and Communications in Medicine interface. However, the system can also be accessed through the Internet using Hypertext Transport Protocol or Simple File Transport Protocol. Our design and simulation work will be key to implementing a successful, scalable medical image archive and distribution system.© (1996) COPYRIGHT SPIE--The International Society for Optical Engineering. Downloading of the abstract is permitted for personal use only.
Read moreEssentials of data repositories
What is a data repository and what is a data archive? We do not have any definitive answers and opinions often vary, but it is useful to note how the terms are used in practice. Certainly in one sense the two terms are synonymous. Each term carries with it certain implied characteristics, and is used within its own contexts and traditions. Social science data archives, as we discussed in Chapter 1, have a tradition going back decades. They can be established at various levels, most commonly national or sub-institutional, such as departmental. The term ‘archive’ to some implies rigorous long-term preservation procedures, and can also be associated with long-term storage that is not publicly accessible. In some cases these are collections intended to become open at some point in the future for legal reasons, also known as ‘dark archives’. To others, ‘archive’ is seen as a simple IT storage service akin to long-term back-up, but without any of the added curation that helps keep data usable and understandable over time and across communities (‘keeping the bits safe’). Of course the word archive is both a noun and a verb and the demand to ‘archive research data’ is a common one; but we suggest that on closer examination there is not enough shared understanding among communities about what it means to ‘archive your data’. For this reason we feel the term is best avoided, except as part of a name of a known institution.
Read moreData Extract: Mining Context from the Web for Dataset Extraction
In this paper we address the problem of dataset extraction from research articles. With the growing digital data repositories and the demand of data centric research in data mining community, finding appropriate dataset for a research problem has become an essential step in scientific research. But given the wide variety of data usage in scientific research it is very difficult to figure out which datasets are most useful for a particular research topic. To alleviate this problem, an automated dataset search engine is a powerful tool. In this work we propose a novel approach to extract dataset names from research articles. We propose a novel way of using web intelligence from academic search engines and online dictionaries to mine dataset names from research articles. We also show a comparison between different sources of web knowledge by comparing different academic search engines such as Google scholar, Microsoft academic search. The performance of this approach is evaluated using standard information retrieval metric such as precision, recall and F-measure. We get an F-measure of 80%. This accuracy is significant for an unsupervised approach.
Read moreGENE2D: A NoSQL Integrated Data Repository of Genetic Disorders Data
There are few sources from which to obtain clinical and genetic data for use in research in Saudi Arabia. Numerous obstacles led to the difficulty of integrating these data from silos and scattered sources to provide standardized access to large data sets for patients with common health conditions. To this end, we sought to contribute to this area and offer a practical and easy-to-implement solution. In this paper, we aim to design and implement a “not only SQL” (NoSQL) based integration framework to generate an Integrated Data Repository of Genetic Disorders Data (GENE2D) to integrate data from various genetic clinics and research centers in Saudi Arabia and provide an easy-to-use query interface for researchers to conduct their studies on large datasets. The major components involved in the GENE2D architecture consists of the data sources, the integrated data repository (IDR) as a central database, and the application interface. The IDR uses a NoSQL document store via MongoDB (an open source document-oriented database program) as a backend database. The application interface called Query Builder provides multiple services for data retrieval from the database using a custom query to answer simple or complex research questions. The GENE2D system demonstrates its potential to help grow and develop a national genetic disorders database in Saudi Arabia.
Read moreFostering scientists’ data sharing behaviors via data repositories, journal supplements, and personal communication methods
Fostering scientists’ data sharing behaviors via data repositories, journal supplements, and personal communication methods
Read moreCommunity science: A typology and its implications for governance of social-ecological systems
Community science: A typology and its implications for governance of social-ecological systems