- Research Article
17
- 10.1016/s0169-023x(96)00045-6
Data schema design as a schema evolution process
- Apr 01, 1997
- Data & Knowledge Engineering
- H.A Proper
Data schema design as a schema evolution process
This research offers a novel AI solution for detecting data consistency checking and schema evolution problems in SQL to migrations in the cloud. The solution utilizes machine learning models to execute data integrity and schema change discovery in migrations. The significant findings report that the proposed model has a 17–22% better discovery than traditional means and that data consistency defects by up to 9%. The model is also effective in reducing migration time and, thus, efficient for large migrations. These findings illustrate the application of AI-powered solutions in optimizing migration efficiency and trustworthiness to enable businesses to migrate legacy SQL databases to the cloud. However, the model’s need for quality tagged data sets and scalability in accommodating large and diverse data sets is to be examined. Future research can focus on deep learning approaches, multi-cloud support, and real-world application testing to enable flexibility and performance. This research provides a foundation for developing cloud database migration technology to achieve more efficient and error-free migrations.
Data schema design as a schema evolution process
Data schema design as a schema evolution process
VC-SLAM—A Handcrafted Data Corpus for the Construction of Semantic Models
Ontology-based data management and knowledge graphs have emerged in recent years as efficient approaches for managing and utilizing diverse and large data sets. In this regard, research on algorithms for automatic semantic labeling and modeling as a prerequisite for both has made steady progress in the form of new approaches. The range of algorithms varies in the type of information used (data schema, values, or metadata), as well as in the underlying methodology (e.g., use of different machine learning methods or external knowledge bases). Approaches that have been established over the years, however, still come with various weaknesses. Most approaches are evaluated on few small data corpora specific to the approach. This reduces comparability and also limits statements for the general applicability and performance of those approaches. Other research areas, such as computer vision or natural language processing solve this problem by providing unified data corpora for the evaluation of specific algorithms and tasks. In this paper, we present and publish VC-SLAM to lay the necessary foundation for future research. This corpus allows the evaluation and comparison of semantic labeling and modeling approaches across different methodologies, and it is the first corpus that additionally allows to leverage textual data documentations for semantic labeling and modeling. Each of the contained 101 data sets consists of labels, data and metadata, as well as corresponding semantic labels and a semantic model that were manually created by human experts using an ontology that was explicitly built for the corpus. We provide statistical information about the corpus as well as a critical discussion of its strengths and shortcomings, and test the corpus with existing methods for labeling and modeling.
Read moreA New Benchmark and Low Computational Cost Localization Method for Cephalometric Analysis
In this study, we present WebCeph2k, an extensive and diverse cephalometric landmark localization dataset that surpasses previous benchmark datasets in terms of number of landmark annotations. This diverse cephalometric landmarks dataset has significant value in medical imaging research. Existing studies predominantly focus on datasets obtained from a single medical center and provider, which offers a limited number of landmarks and a limited diversity of cephalograms, resulting in models that exhibit low robustness and generalization when applied to more diverse datasets. The clinical application of cephalometry is hampered by significant localization errors in landmark localization models, in addition to the inadequacy of existing datasets’ landmarks for clinical cephalometric diagnosis. The limited generalization ability and the occurrence of “overfitting” in deep learning models are mainly caused by the small size in the dataset. In the medical field, the inclusion of large and diverse datasets can greatly improve the generalization and performance of landmark localization models. This paper presents our WebCeph2k dataset from 9 medical centers, covering 9 different imaging devices, which surpasses the only publicly available ISBI2015 dataset in terms of sample size and number of landmarks. In addition, this study employs a low computational cost methodology to achieve optimal landmarks localization: 1) ROI regions of X-ray images are derived by exploiting the prior distribution of the data, 2) the model computational cost is reduced by adopting a spatial-depth transformation strategy, 3) the standard heatmap decoding method is optimized by integrating a compensation strategy. The results show that the proposed method not only achieves competitive localization results to other state-of-the-art approaches, but also offers a reduction of the model computational cost, resulting in faster inference. Consequently, this research offers valuable prospects in the field of general-purpose medical landmark localization methods. We also find that our proposed dataset is more complex and challenging than the ISBI dataset. The dataset and code are available at https://github.com/switch626/WebCeph2k.
Read moreSemLinker: automating big data integration for casual users
A data integration approach combines data from different sources and builds a unified view for the users. Big data integration inherently is a complex task, and the existing approaches are either potentially limited or invariably rely on manual inputs and interposition from experts or skilled users. SemLinker, an ontology-based data integration system, is part of a metadata management framework for personal data lake (PDL), a personal store-everything architecture. PDL is for casual and unskilled users, therefore SemLinker adopts an automated data integration workflow to minimize manual input requirements. To support the flat architecture of a lake, SemLinker builds and maintains a schema metadata level without involving any physical transformation of data during integration, preserving the data in their native formats while, at the same time, allowing them to be queried and analyzed. Scalability, heterogeneity, and schema evolution are big data integration challenges that are addressed by SemLinker. Large and real-world datasets of substantial heterogeneities are used in evaluating SemLinker. The results demonstrate and confirm the integration efficiency and robustness of SemLinker, especially regarding its capability in the automatic handling of data heterogeneities and schema evolutions.
Read moreMulti-deep learning models based analysis for classification of thunderstorms on meteorological dataset over Ranchi region
Deep learning (DL) is now generally acknowledged as the benchmark and evolution in machine learning (ML) fields. Further more, it has steadily suited the most extensively in computational techniques for ML, delivering excellent outcomes in several challenging intellectual works that have equal or still outperformed human ability. DL has the advantage of learning from enormous amounts of data. Here, the DL approach has been applied to classify thunderstorms and non-thunderstorms with a meteorological dataset. Daily observational hourly data sets from 2016 to 2021 have been used to classify the incidence of thunderstorms and non-thunderstorms. Three DL approaches are used: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Long Short-Term Memory (LSTM) to classify thunderstorm incidences. These DL techniques are compared with conventional ML techniques, showing that DL approaches outperformed ML. The CNN method provides the highest performance with 92.01% accuracy among all DL and ML approaches.
Read moreCatMapper: user interface support for large complex categories and semantic data exploration
Scientists and policymakers are increasingly leveraging complex, multi-scale data from diverse, worldwide sources to understand the causes and consequences of economic development, social stratification, climate change, cultural diversity, and violent conflict. This work frequently requires integrating data across diverse datasets by complex, dynamic categories (e.g., ethnicities, languages, religions, subdistricts). However, different datasets encode corresponding categories in disparate formats and at different resolutions (e.g., Guatemala Indigenous vs. Maya vs. K’iche’). These diverse encodings must be translated across datasets before bringing them together for analysis. At global scales across thousands of categories, the combinatorial complexity creates thorny challenges for manual reconciliation and for transparent documentation and sharing of researcher decisions. There is a need to investigate direct and uncomplicated ways to support search and explore the semantics for complex and diverse datasets.We design and deploy such a tool, CatMapper, to support semantic discovery through exploration and manipulation for large, complex and diverse datasets. CatMapper enables exploring contextual information about specific categories, translating new sets of categories from existing datasets and published studies, identify and integrating novel combinations of datasets for researchers’ custom needs, including automatically generated syntax to merge datasets of interest, and publishing and sharing merging templates for public re-use and open science. CatMapper does not store observational data. Rather, it is a dynamic, interactive dictionary of keys to help users integrate observational data from diverse external datasets in disparate formats, thereby complementing and leveraging a fast-growing ecology of datasets storing observational data. We have conducted heuristic evaluation on CatMapper usability. Results shed lights on enriching semantic data discovery.
Read moreCombining Magnification and Measurement for Non-Contact Cardiac Monitoring
Deep learning approaches currently achieve the state-of-the-art results on camera-based vital signs measurement. One of the main challenges with using neural models for these applications is the lack of sufficiently large and diverse datasets. Limited data increases the chances of overfitting models to the available data which in turn can harm generalization. In this paper, we show that the generalizability of imaging photoplethysmography models can be improved by augmenting the training set with "magnified" videos. These augmentations are specifically designed to reveal useful features for recovering the photoplethysmogram. We show that using augmentations of this form is more effective at improving model robustness than other commonly used data augmentation approaches. We show better within-dataset and especially cross-dataset performance with our proposed data augmentation approach on three publicly available datasets.
Read moreAdvances in Neuroimaging and Deep Learning for Emotion Detection: A Systematic Review of Cognitive Neuroscience and Algorithmic Innovations.
Background/Objectives: The following systematic review integrates neuroimaging techniques with deep learning approaches concerning emotion detection. It, therefore, aims to merge cognitive neuroscience insights with advanced algorithmic methods in pursuit of an enhanced understanding and applications of emotion recognition. Methods: The study was conducted following PRISMA guidelines, involving a rigorous selection process that resulted in the inclusion of 64 empirical studies that explore neuroimaging modalities such as fMRI, EEG, and MEG, discussing their capabilities and limitations in emotion recognition. It further evaluates deep learning architectures, including neural networks, CNNs, and GANs, in terms of their roles in classifying emotions from various domains: human-computer interaction, mental health, marketing, and more. Ethical and practical challenges in implementing these systems are also analyzed. Results: The review identifies fMRI as a powerful but resource-intensive modality, while EEG and MEG are more accessible with high temporal resolution but limited by spatial accuracy. Deep learning models, especially CNNs and GANs, have performed well in classifying emotions, though they do not always require large and diverse datasets. Combining neuroimaging data with behavioral and cognitive features improves classification performance. However, ethical challenges, such as data privacy and bias, remain significant concerns. Conclusions: The study has emphasized the efficiencies of neuroimaging and deep learning in emotion detection, while various ethical and technical challenges were also highlighted. Future research should integrate behavioral and cognitive neuroscience advances, establish ethical guidelines, and explore innovative methods to enhance system reliability and applicability.
Read moreComparative evaluation of text classification techniques using a large diverse Arabic dataset
A vast amount of valuable human knowledge is recorded in documents. The rapid growth in the number of machine-readable documents for public or private access necessitates the use of automatic text classification. While a lot of effort has been put into Western languages—mostly English—minimal experimentation has been done with Arabic. This paper presents, first, an up-to-date review of the work done in the field of Arabic text classification and, second, a large and diverse dataset that can be used for benchmarking Arabic text classification algorithms. The different techniques derived from the literature review are illustrated by their application to the proposed dataset. The results of various feature selections, weighting methods, and classification algorithms show, on average, the superiority of support vector machine, followed by the decision tree algorithm (C4.5) and Naïve Bayes. The best classification accuracy was 97 % for the Islamic Topics dataset, and the least accurate was 61 % for the Arabic Poems dataset.
Read moreMeta data management
By meta data management, we mean techniques for manipulating schemas and schema-like objects (such as interface definitions and web site maps) and mappings between them. Work on meta data problems goes back to at least the early 1970s, when data translation was the hot database research topic, even before relational databases caught on. Many popular research problems in the past five years are primarily meta data problems, such as data warehouse tools (e.g., ETL – to extract, transform and load), data integration, the semantic web, generation of XML or object-oriented wrappers for SQL databases, and generation of wrappers for web sites. Other classical meta data problems are information resource management, design tool support and integration, and schema evolution and data migration. Despite its longevity and continued importance, there is no widely-accepted conceptual framework for the meta data field, as there is for many other database topics, such as access methods, query processing, and transaction management. In this seminar, we propose such a conceptual framework. It consists of three layers: applications, design patterns, and basic operators. Applications are the end-user problems to be solved, like those listed in the previous paragraph. Design patterns are generic problems that need to be solved in support of many different applications, such as meta modeling (for all meta data problems), answering queries using views (for data integration and the semantic web), and change propagation (for data translation, schema evolution, and round-trip engineering). Basic operators are procedures that are needed to support multiple design patterns and applications, such as matching schemas to produce a mapping, merging schemas based on a mapping, and composing mappings. We will describe several meta data management problems, and for each, we will explain the design patterns and operators that are needed to solve it. We will summarize the main approaches to each design pattern and operator – the main choices of language, data structures, and algorithms – and will highlight the relevant papers that address it. This seminar is targeted at both practicing engineers and researchers. The former will learn about the latest solutions to important meta data problems and the many difficult unsolved problems that are best to avoid. Database researchers, especially professors, will benefit from considering the conceptual framework that we propose, since no database textbooks treat meta data management as a separate topic as far as we know.
Read moreGlycoDash: automated, visually assisted curation of glycoproteomics datasets for large sample numbers
The challenge of robust and automated glycopeptide quantitation from liquid chromatography-mass spectrometry (LC–MS) data has yet to be adequately addressed by commercial software. Recently, open-source tools like Skyline and LaCyTools have advanced the field of label-free MS1 level quantitation. Yet, important steps late in the data processing workflow remain manual. Because manual data curation is time-consuming and error-prone, it presents a bottleneck, especially in an era of emerging high-throughput methodologies and increasingly complex analyses such as antigen-specific antibody glycosylation. We addressed this gap by developing GlycoDash, an R Shiny-based interactive web application designed to democratize label-free high-throughput glycoproteomics data analysis. The software comes in at a stage where analytes have been identified and quantified, but whole measurement and individual analyte signals of insufficient quality for quantitation remain and reduce the quality of the overall dataset. GlycoDash focuses on these challenges by incorporating several options for measurement and metadata linking, spectral and analyte curation, normalization, and repeatability assessment, and additionally includes glycosylation trait calculation, data visualization, and reporting capabilities that adhere to FAIR principles. The performance and versatility of GlycoDash were demonstrated across antibody glycoproteomics data of increasing complexity, ranging from relatively simple monoclonal antibody glycosylation analysis to a clinical cohort with over a thousand measurements. In a matter of hours, these large, diverse, and complex datasets were curated and explored. High-quality datasets with integrated metadata ready for final analysis and visualization were obtained. Critical aspects of the curation strategy underlying GlycoDash are discussed. GlycoDash effectively automates and streamlines the curation of glycopeptide quantitation data, addressing a critical need for high-throughput glycoproteomics data analysis. Its robust performance across diverse datasets and its comprehensive feature toolbox significantly enhance both research and clinical applications in glycoproteomics.Graphical
Read moreLattice Histograms: a Resilient Synopsis Structure
Despite the surge of interest in data reduction techniques over the past years, no method has been proposed to date that can always achieve approximation quality preferable to that of the optimal plain histogram for a target error metric. In this paper, we introduce the lattice histogram: a novel data reduction method that discovers and exploits any arbitrary hierarchy in the data, and achieves approximation quality provably at least as high as an optimal histogram for any data reduction problem. We formulate LH construction techniques with approximation guarantees for general error metrics. We show that the case of minimizing a maximum-error metric can be solved by a specialized, memory-sparing approach; we exploit this solution to design reduced-space heuristics for the general- error case. We develop a mixed synopsis approach, applicable to the space-efficient high-quality summarization of very large data sets. We experimentally corroborate the superiority of LHs in approximation quality over previous techniques with representative error metrics and diverse data sets.
Read moreA transparent schema-evolution system based on object-oriented view technology
When a database is shared by many users, updates to the database schema are almost always prohibited because there is a risk of making existing application programs obsolete when they run against the modified schema. The paper addresses the problem by integrating schema evolution with view facilities. When new requirements necessitate schema updates for a particular user, then the user specifies schema changes to his personal view, rather than to the shared base schema. Our view schema evolution approach then computes a new view schema that reflects the semantics of the desired schema change, and replaces the old view with the new one. We show that our system provides the means for schema change without affecting other views (and thus without affecting existing application programs). The persistent data is shared by different views of the schema, i.e., both old as well as newly developed applications can continue to interoperate. The paper describes a solution approach of realizing the evolution mechanism as a working system, which as its key feature requires the underlying object oriented view system to support capacity augmenting views. We present algorithms that implement the complete set of typical schema evolution operations as view definitions. Lastly, we describe the transparent schema evolution system (TSE) that we have built on top of GemStone, including our solution for supporting capacity augmenting view mechanisms.
Read moreAutomated multiclass segmentation of liver vessel structures in CT images using deep learning approaches: a liver surgery pre-planning tool.
Accurate liver vessel segmentation is essential for effective liver surgery pre-planning, and reducing surgical risks since it enables the precise localization and extensive assessment of complex vessel structures. Manual liver vessel segmentation is a time-intensive process reliant on operator expertise and skill. The complex, tree-like architecture of hepatic and portal veins, which are interwoven and anatomically variable, further complicates this challenge. This study addresses these challenges by proposing the UNETR (U-Net Transformers) architecture for the multi-class segmentation of portal and hepatic veins in liver CT images. UNETR leverages a transformer-based encoder to effectively capture long-range dependencies, overcoming the limitations of convolutional neural networks (CNNs) in handling complex anatomical structures. The proposed method was evaluated on contrast-enhanced CT images from the IRCAD as well as a locally dataset developed from a hospital. On the local dataset, the UNETR model achieved Dice coefficients of 49.71% for portal veins, 69.39% for hepatic veins, and 76.74% for overall vessel segmentation, while reaching Dice coefficients of 62.54% for vessel segmentation on the IRCAD dataset. These results highlight the method's effectiveness in identifying complex vessel structures across diverse datasets. These findings underscore the critical role of advanced architectures and precise annotations in improving segmentation accuracy. This work provides a foundation for future advancements in automated liver surgery pre-planning, with the potential to enhance clinical outcomes significantly. The implementation code is available on GitHub: https://github.com/saharsarkar/Multiclass-Vessel-Segmentation .
Read moreEnhancing lung abnormalities diagnosis using hybrid DCNN-ViT-GRU model with explainable AI: A deep learning approach
Enhancing lung abnormalities diagnosis using hybrid DCNN-ViT-GRU model with explainable AI: A deep learning approach