- Research Article
13
- 10.1016/j.smhl.2021.100263
A review of harmonization methods for studying dietary patterns.
- Mar 01, 2022
- Smart Health
- Venkata Sukumar Gurugubelli + 8 more +8
A review of harmonization methods for studying dietary patterns.
In the age of big data, it is important for primary research data to follow the FAIR principles of findability, accessibility, interoperability, and reusability. Data harmonization enhances interoperability and reusability by aligning heterogeneous data under standardized representations, benefiting both repository curators responsible for upholding data quality standards and consumers who require unified datasets. However, data harmonization is difficult in practice, requiring significant domain and technical expertise. We present a software framework to facilitate principled and reproducible harmonization protocols. Our framework implements a novel strategy of building harmonization transformations from parameterizable primitive operations, such as the assignment of numerical values to user-specified categories, with automated bookkeeping for executed transformations. We establish our data representation model and harmonization strategy and then report a proof-of-concept application in the context of the RADx Data Hub. Our framework enables data practitioners to execute transparent and reproducible harmonization protocols that align closely with their research goals.
A review of harmonization methods for studying dietary patterns.
A review of harmonization methods for studying dietary patterns.
LLM-based harmonized data ingestion for dataspace: a novel system for automating data ingestion across heterogeneous data sources
• Offers a new vision on automated data harmonization in Dataspace (DS) systems. • Introduces LLM-based methods for scalable DS ingestion of heterogeneous datasources. • Presents a system with Harmonizer, Transformer, Evaluator components for ingestion. • Demonstrated an automated data ingestion prototype using LLM agents. • Validates the system with healthcare use case harmonizing heterogeneous data sources. Dataspaces (DS) enable stakeholders to collaborate on innovative, data-driven services by integrating data across domains. However, the realization and adoption of DS remain challenging due to domain-specific heterogeneity at the system, service, and data levels. While system and service-level heterogeneity can often be addressed through standards, data-level heterogeneity, namely data structures and semantics variations, remains challenging. To effectively ingest data into the DS, two communication endpoints must correctly interpret each other’s data models, therefore, DS ecosystems rely on “harmonization”, the process of generating a unified target data model from heterogeneous sources and transforming incoming data accordingly. Currently, harmonization and transformation are performed manually whenever new data sources are integrated. This is time-consuming, costly, and difficult to scale, posing a critical barrier to the realization and adoption of DS in practice. This study proposes a novel methodology for automated data harmonization during ingestion into DS ecosystems. The approach integrates harmonization, transformation, and human-in-the-loop evaluation within an automated system powered by modern LLM-based AI agents. These agents address data-level heterogeneity and generate harmonized target data models, representing a substantial departure from current manually-handled data harmonization. The system is validated through a healthcare use case, demonstrating its practical feasibility for harmonization during data ingestion into the DS. Overall, this work provides a foundational step toward seamless, efficient, and scalable data integration in DS. By automating data harmonization, it delivers substantial value to industry digital solutions as well as domains where data heterogeneity persists, including IoT or Big Data platforms.
Read moreThe Use of the General Animal-Based Measures Codified Terms in the Scientific Literature on Farm Animal Welfare
Background: The approach to farm animal welfare evaluation has changed and animal-based measures (ABM), defined as the responses of an animal or effects on an animal, were introduced to assess animal welfare. Animal-based measures can be taken directly on the animal or indirectly and include the use of animal records. They can result from a specific event or be the cumulative outcome of many days, weeks, or months. The objective of the current study was to analyze the use of general ABM codified terms in the scientific literature, the presence of their definitions, and the gap mapping of their use across animal species, categories, years of publication, and geographical areas of the corresponding author's institution. The ultimate aim was to propose a common standard terminology to improve communication among stakeholders. In this study, data models were populated by collecting information coming from scientific papers extracted through a transparent and reproducible protocol using Web of ScienceTM and filtering for the general ABM codified terms (or synonyms/equivalents). A total of 199 papers were retained, and their full texts were assessed. The frequency of general codified ABM terms was analyzed according to the classification factors listed in the objectives. These papers were prevalently European (159 documents), and the most represented species was cattle. Fifty percent of the papers did not provide a definition of the general ABM terms, and 54% cited other sources as reference for their definition. The results of the study showed a very low penetration of the general codified ABM term in the literature on farm animal welfare, with only 1.5% of the papers including the term ABM. This does not mean that specific ABM are not studied, but rather that these specific ABM are not defined as such under a common umbrella, and there is no consensus on the use of terminology, not even among scientists. Thus, we cannot expect the stakeholders to use a common language and a standardized terminology. The recognition and the inclusion of ABM in the lists of commonly accepted abbreviations of scientific journals could be a first step to harmonize the terminology in the scientific literature.
Read moreAccess to longitudinal mental health data in Africa: Lessons from a landscape analysis by the INSPIRE network datahub
Background Data from longitudinal mental health research in Africa is critical to understanding the complexities of mental health disorders in the continent’s diverse contexts. To be useful, data need to adhere to the FAIR principles of Findability, Accessibility, Interoperability, and Reusability. Methods A literature search from 1970 to 2022 identified longitudinal studies on depression, anxiety, and psychosis in was done. Using Artificial Intelligence (AI) and Natural Language Processing techniques, the search found data from more studies. The search engaged stakeholders in understanding data sharing practices and barriers and categorizing methods and challenges for sharing data. Results The initial search yielded 18,019 articles, of which 284 were eligible for review, and 226 passed quality assessments. A significant effort to access data directly from researchers yielded positive responses for 100 articles with available data statements, from which datasets were requested through online repositories and direct correspondence. Analysis revealed significant disparities in the distribution of mental health research across countries, with a concentration of studies in specific areas and on certain conditions. The study also highlighted a varied adherence to FAIR principles, with only 17 (17%) datasets adhering to data-sharing practices. Conclusions Despite the challenges encountered in data accessibility and the manual adjustments required, the study’s findings irradiate the path toward a more equitable and effective mental health research ecosystem on the continent. By fostering collaboration and embracing advanced methodologies and technologies, this study advocates for a concerted effort to improve the accessibility, interoperability, and reusability of mental health data. Ultimately, the project aims to contribute to understanding data-sharing dynamics in Africa, paving the way for informed interventions and policies.
Read moreA dataset of cold-water coral distribution records
Species distribution data are key for monitoring present and future biodiversity patterns and informing conservation and management strategies. Large biodiversity information facilities often contain spatial and taxonomic errors that reduce the quality of the provided data. Moreover, datasets are frequently shared in varying formats, inhibiting proper integration and interoperability. Here, we provide a quality-controlled dataset of the diversity and distribution of cold-water corals, which provide key ecosystem services and are considered vulnerable to human activities and climate change effects. We use the common term cold-water corals to refer to species of the orders Alcyonacea, Antipatharia, Pennatulacea, Scleractinia, Zoantharia of the subphylum Anthozoa, and order Anthoathecata of the class Hydrozoa. Distribution records were collated from multiple sources, standardized using the Darwin Core Standard, dereplicated, taxonomically corrected and flagged for potential vertical and geographic distribution errors based on peer-reviewed published literature and expert consulting. This resulted in 817,559 quality-controlled records of 1,170 accepted species of cold-water corals, openly available under the FAIR principle of Findability, Accessibility, Interoperability and Reusability of data. The dataset represents the most updated baseline for the global cold-water coral diversity, and it can be used by the broad scientific community to provide insights into biodiversity patterns and their drivers, identify regions of high biodiversity and endemicity, and project potential redistribution under future climate change. It can also be used by managers and stakeholders to guide biodiversity conservation and prioritization actions against biodiversity loss.
Read moreA Resource of Microbiome Benchmark Datasets with Biological Ground Truth
Little consensus exists about which statistical methods are most suitable for identifying differentially abundant (DA) taxa in microbiome research. A limitation of benchmarking DA methods is the lack of datasets with ground truth. Here, we introduce the MicrobiomeBenchmarkData package/resource of datasets with known biological ground truth and varying ecological complexity. We compiled and standardized three datasets: 1) a high-diversity dataset from the Human Microbiome Project contrasting subgingival and supragingival oral plaques using both 16S and shotgun data, 2) a low-diversity 16S dataset of healthy vaginal and bacterial vaginosis samples, and 3) a 16S spike-in dataset containing fixed amounts of foreign spiked-in bacteria across stool samples of patients who underwent allogeneic stem transplantation. We used the first two datasets to benchmark 17 differential abundance methods broadly categorized as classical, compositional, microbiome-specific, RNA-seq, and scRNA-seq tests. We used the third dataset to assess low variability of spiked-in bacteria. Overall, RNAseq methods performed best in the high-diversity dataset (gingiva) while most methods performed well in the low-diversity dataset (vagina). Notably, compositional methods proposed specifically for microbiome data analysis, performed equivalently to other methods in the high-complexity dataset but produced spurious results in the low-diversity dataset and increased unwanted variability of spike-in bacteria. The MicrobiomeBenchmarkData project aims to be a gold-standard resource of microbiome datasets for benchmarking DA methods based on biologically relevant signatures, following the data FAIRness principles of Findability, Accessibility, Interoperability, and Reusability. Datasets are available as TSV text files through Zenodo (https://doi.org/10.5281/zenodo.6911026), and as TreeSummarizedExperiment class objects through the MicrobiomeBenchmarkData R/Bioconductor package (https://waldronlab.io/MicrobiomeBenchmarkData).
Read moreAccess to longitudinal mental health data in Africa: Lessons from a landscape analysis by the INSPIRE network datahub
Background Data from longitudinal mental health research in Africa is critical to understanding the complexities of mental health disorders in the continent’s diverse contexts. To be useful, data need to adhere to the FAIR principles of Findability, Accessibility, Interoperability, and Reusability. Methods A literature search from 1970 to 2022 identified longitudinal studies on depression, anxiety, and psychosis in was done. Using Artificial Intelligence (AI) and Natural Language Processing techniques, the search found data from more studies. The search engaged stakeholders in understanding data sharing practices and barriers and categorizing methods and challenges for sharing data. Results The initial search yielded 18,019 articles, of which 284 were eligible for review, and 226 passed quality assessments. A significant effort to access data directly from researchers yielded positive responses for 100 articles with available data statements, from which datasets were requested through online repositories and direct correspondence. Analysis revealed significant disparities in the distribution of mental health research across countries, with a concentration of studies in specific areas and on certain conditions. The study also highlighted a varied adherence to FAIR principles, with only 16 datasets adhering to data-sharing practices. Conclusions Despite the challenges encountered in data accessibility and the manual adjustments required, the study’s findings irradiate the path toward a more equitable and effective mental health research ecosystem on the continent. By fostering collaboration and embracing advanced methodologies and technologies, this study advocates for a concerted effort to improve the accessibility, interoperability, and reusability of mental health data. Ultimately, the project aims to contribute to understanding data-sharing dynamics in Africa, paving the way for informed interventions and policies.
Read moreThe three pillars approach and a practicum portfolio for FAIR research in political science
This paper presents for a political science audience the Three Pillars Approach to the FAIR principles of Findability, Accessibility, Interoperability, and Re-usabilty for data and metadata. A portfolio of illustrative practical activities is offered that scholarly communities can take up to make their research more FAIR at disciplinary and subdisciplinary levels.
Read moreMaking Linked Data accessible for One Health Surveillance with the "One Health Linked Data Toolbox"
In times of emerging diseases, data sharing and data integration are of particular relevance for One Health Surveillance (OHS) and decision support. Furthermore, there is an increasing demand to provide governmental data in compliance to the FAIR (Findable, Accessible, Interoperable, Reusable) data principles. Semantic web technologies are key facilitators for providing data interoperability, as they allow explicit annotation of data with their meaning, enabling reuse without loss of the data collection context. Among these, we highlight ontologies as a tool for modeling knowledge in a field, which simplify the interpretation and mapping of datasets in a computer readable medium; and the Resource Description Format (RDF), which allows data to be shared among human and computer agents following this knowledge model. Despite their potential for enabling cross-sectoral interoperability and data linkage, the use and application of these technologies is often hindered by their complexity and the lack of easy-to-use software applications. To overcome these challenges the OHEJP Project ORION developed the Health Surveillance Ontology (HSO). This knowledge model forms a foundation for semantic interoperability in the domain of One Health Surveillance. It provides a solution to add data from the target sectors (public health, animal health and food safety) in compliance with the FAIR principles of findability, accessibility, interoperability, and reusability, supporting interdisciplinary data exchange and usage. To provide use cases and facilitate the accessibility to HSO, we developed the One Health Linked Data Toolbox (OHLDT), which consists of three new and custom-developed web applications with specific functionalities. The first web application allows users to convert surveillance data available in Excel files online into HSO-RDF and vice versa. The web application demonstrates that data provided in well-established data formats can be automatically translated in the linked data format HSO-RDF. The second application is a demonstrator of the usage of HSO-RDF in a HSO triplestore database. In the user interface of this application, the user can select HSO concepts based on which to search and filter among surveillance datasets stored in a HSO triplestore database. The service then provides automatically generated dashboards based on the context of the data. The third web application demonstrates the use of data interoperability in the OHS context by using HSO-RDF to annotate meta-data, and in this way link datasets across sectors. The web application provides a dashboard to compare public data on zoonosis surveillance provided by EFSA and ECDC. The first solution enables linked data production, while the second and third provide examples of linked data consumption, and their value in enabling data interoperability across sectors. All described solutions are based on the open-source software KNIME and are deployed as web service via a KNIME Server hosted at the German Federal Institute for Risk Assessment. The semantic web extension of KNIME, which is based on the Apache Jena Framework, allowed a rapid an easy development within the project. The underlying open source KNIME workflows are freely available and can be easily customized by interested end users. With our applications, we demonstrate that the use of linked data has a great potential strengthening the use of FAIR data in OHS and interdisciplinary data exchange.
Read moreReal-world evidence in achondroplasia: considerations for a standardized data set
BackgroundCollection of real-world evidence (RWE) is important in achondroplasia. Development of a prospective, shared, international resource that follows the principles of findability, accessibility, interoperability, and reuse of digital assets, and that captures long-term, high-quality data, would improve understanding of the natural history of achondroplasia, quality of life, and related outcomes.MethodsThe Europe, Middle East, and Africa (EMEA) Achondroplasia Steering Committee comprises a multidisciplinary team of 17 clinical experts and 3 advocacy organization representatives. The committee undertook an exercise to identify essential data elements for a standardized prospective registry to study the natural history of achondroplasia and related outcomes.ResultsA range of RWE on achondroplasia is being collected at EMEA centres. Whereas commonalities exist, the data elements, methods used to collect and store them, and frequency of collection vary. The topics considered most important for collection were auxological measures, sleep studies, quality of life, and neurological manifestations. Data considered essential for a prospective registry were grouped into six categories: demographics; diagnosis and patient measurements; medical issues; investigations and surgical events; medications; and outcomes possibly associated with achondroplasia treatments.ConclusionsLong-term, high-quality data are needed for this rare, multifaceted condition. Establishing registries that collect predefined data elements across age spans will provide contemporaneous prospective and longitudinal information and will be useful to improve clinical decision-making and management. It should be feasible to collect a minimum dataset with the flexibility to include country-specific criteria and pool data across countries to examine clinical outcomes associated with achondroplasia and different therapeutic approaches.
Read moreMaking digital objects FAIR in high energy physics: An implementation for Universal FeynRules Output (UFO) models
Research in the data-intensive discipline of high energy physics (HEP) often relies on domain-specific digital contents. Reproducibility of research relies on proper preservation of these digital objects. This paper reflects on the interpretation of principles of Findability, Accessibility, Interoperability, and Reusability (FAIR) in such context and demonstrates its implementation by describing the development of an end-to-end support infrastructure for preserving and accessing Universal FeynRules Output (UFO) models guided by the FAIR principles. UFO models are custom-made python libraries used by the HEP community for Monte Carlo simulation of collider physics events. Our framework provides simple but robust tools to preserve and access the UFO models and corresponding metadata in accordance with the FAIR principles.
Read moreBig Data Focus on Physicians: Start-up Profiles Physicians’ Pharma Friendliness
Big Data Focus on Physicians: Start-up Profiles Physicians’ Pharma Friendliness
Blockchain adoption in agri-food supply chain: a DEMATEL-based analysis of interdependent barriers and strategic imperatives
Purpose This study scrutinizes the challenges hindering blockchain adoption in agro-food supply chains, despite its potential to enhance transparency and efficiency. It employs the Decision-Making Trial and Evaluation Laboratory (DEMATEL) method to systematically identify these barriers and analyze their causal relationships. Design/methodology/approach The study adopted a rigorous three-phase methodology. First, a bibliometric analysis using RStudio mapped the existing research landscape and identified preliminary challenges. Second, a diverse panel of 15 domain experts utilized DEMATEL approach to assess the direct influence among these identified barriers. In the third phase, the quantitative results were interpreted to generate a cause–effect diagram from the calculated C + R and C–R values, visualizing influential relationships and classifying barriers as either causes or effects. Findings The analysis revealed that “High Implementation Costs/Financial Constraints (C1), Evolving and Ambiguous Policy Frameworks (C5), and Trust issues among various partners (C7)” are critical key drivers. These foundational bottlenecks strongly influence other challenges, demanding immediate attention. Dependent variables, such as “Data Security and Privacy Concerns (C10), Lack of Digital and Physical Infrastructure (C8), and Limited Digital Literacy and Technical Expertise (C6),” were identified as primary effects, indicating that their mitigation depends on addressing the key drivers. Less influential factors like “Interoperability Challenges (C9)” and “Energy Consumption for Sustainability (C2)” become more critical as adoption matures. Research limitations/implications This study offers significant insights into barriers to blockchain adoption in India's agri-food supply chain; however, the reliance on domain experts, though carefully selected, introduces an inherent degree of subjectivity to the qualitative assessments. While the DEMATEL method excels at identifying causal relationships, it does not explicitly quantify the precise impact strength of each barrier on the overall adoption rate. Originality/value The findings provide strategic insights for policymakers and managers. Priority should be given to developing robust financial incentives, investing substantially in rural digital and physical infrastructure and implementing comprehensive skill development programs, along with establishing clear and adaptive regulatory frameworks. Addressing these core causal barriers can unlock blockchain's transformative potential, enhancing transparency, fostering trust and empowering farmers towards a more resilient and efficient agri-food supply chain, thereby aligning with modernization and sustainable development goals.
Read moreStakeholder access to lands for agriculture and land governance: A systematic review
Equitable access to agricultural land remains a central challenge for sustainable and inclusive rural development in Africa. Despite the recognized importance of this issue, a comprehensive and systematic synthesis of land conflict types, key governance actors, and prevailing tenure systems is scarcely documented across Africa. This review fills this gap by applying a transparent and reproducible systematic review protocol – including database searching, screening, and thematic coding – which resulted in the selection of 55 studies from Web of Science, ScienceDirect, and African Journals Online. The findings reveal that inheritance (63.64%) is the most prevalent land tenure regime, while borrowing is the least common (23.67%), indicating a shift away from communal arrangements. Land conflicts are predominantly centred on ownership (28.57%) and usage rights (25%), often exacerbated by fraudulent allocations. The review underscores the central, yet often conflicting, roles of state and customary institutions, with 80% of studies highlighting corruption as a major impediment to effective land law enforcement. Women are identified as the most vulnerable actors in land access and security. The study concludes that transparent, inclusive land governance reforms, which harmonize formal and customary systems, are urgently needed to mitigate conflicts and empower marginalized stakeholders.
Read morePolicy Specialisation: A Meta-Analysis of Abil Hasan's Thoughts Al-Mawardi in the Book of "Ahkam us-Shulthaniyah"
Policy specialisation is a concept developed in the application of knowledge and technical skills in a particular field to design and implement effective policies. This research has an urgency on the need to understand and apply the principles of policy specialisation in modern times, as well as to explore the relevance of classical thought in contemporary policy theory. The purpose of the research is to analyse the concept of policy specialisation and map the relevance of these principles with the thoughts of Abil Hasan Al-Mawardi in his book "Ahkam us-Shulthaniyah". This research uses a meta-analysis method with a literature study approach (Library Research) by reviewing various literature including classic works and contemporary research, in data collection is done by synthesising data from various secondary sources, such as books, journal articles, and policy reports, as well as primary data from historical documents. The results showed that 1) indicators of the principles of policy specialisation including technical expertise, specific issues, and division of tasks, in line with Abil Hasan Al-Mawardi's thoughts in "Ahkam us-Shulthaniyah" are still in line with many of the principles covered in the concept of Policy Specialisation. Al-Mawardi strongly emphasises the technical expertise of stakeholders in state administration, both a clear division of roles in government and consistency in the application of laws and policies; 2) Meta-analysis of Al-Mawardi's thought shows that although his work was written in the 11th century many of the principles of state administration put forward in the book "Ahkam us-Shulthaniyah" can be applied in contemporary policy design. Al-Mawardi's contribution provides the foundation (Policy Specialisation) of policy specialisation in public administration, especially in governance that requires high coordination and consistency. This research confirms that the integration between classical theory and modern policy can enrich the understanding and implementation of public policy and provide a solid basis for the development of contemporary policy theory and practice.
Read more