- Book Chapter
1
- 10.1007/978-3-031-36678-9_4
Data Integration and Harmonisation
- Jan 01, 2023
- Maxim Moinat + 2 more +2
Publications from 2021 to 2026
Showing 4 of 4 papers
Data Integration and Harmonisation
MatchMiner: an open-source platform for cancer precision medicine
Widespread, comprehensive sequencing of patient tumors has facilitated the usage of precision medicine (PM) drugs to target specific genomic alterations. Therapeutic clinical trials are necessary to test new PM drugs to advance precision medicine, however, the abundance of patient sequencing data coupled with complex clinical trial eligibility has made it challenging to match patients to PM trials. To facilitate enrollment onto PM trials, we developed MatchMiner, an open-source platform to computationally match genomically profiled cancer patients to PM trials. Here, we describe MatchMiner’s capabilities, outline its deployment at Dana-Farber Cancer Institute (DFCI), and characterize its impact on PM trial enrollment. MatchMiner’s primary goals are to facilitate PM trial options for all patients and accelerate trial enrollment onto PM trials. MatchMiner can help clinicians find trial options for an individual patient or provide trial teams with candidate patients matching their trial’s eligibility criteria. From March 2016 through March 2021, we curated 354 PM trials containing a broad range of genomic and clinical eligibility criteria and MatchMiner facilitated 166 trial consents (MatchMiner consents, MMC) for 159 patients. To quantify MatchMiner’s impact on trial consent, we measured time from genomic sequencing report date to trial consent date for the 166 MMC compared to trial consents not facilitated by MatchMiner (non-MMC). We found MMC consented to trials 55 days (22%) earlier than non-MMC. MatchMiner has enabled our clinicians to match patients to PM trials and accelerated the trial enrollment process.
Read moreUsing the Data Quality Dashboard to Improve the EHDEN Network
Federated networks of observational health databases have the potential to be a rich resource to inform clinical practice and regulatory decision making. However, the lack of standard data quality processes makes it difficult to know if these data are research ready. The EHDEN COVID-19 Rapid Collaboration Call presented the opportunity to assess how the newly developed open-source tool Data Quality Dashboard (DQD) informs the quality of data in a federated network. Fifteen Data Partners (DPs) from 10 different countries worked with the EHDEN taskforce to map their data to the OMOP CDM. Throughout the process at least two DQD results were collected and compared for each DP. All DPs showed an improvement in their data quality between the first and last run of the DQD. The DQD excelled at helping DPs identify and fix conformance issues but showed less of an impact on completeness and plausibility checks. This is the first study to apply the DQD on multiple, disparate databases across a network. While study-specific checks should still be run, we recommend that all data holders converting their data to the OMOP CDM use the DQD as it ensures conformance to the model specifications and that a database meets a baseline level of completeness and plausibility for use in research.
Read moreThe introduction of the FAIR –Findable, Accessible, Interoperable, Reusable– principles has caused quite an uproar within the scientific community. Principles which, if everyone adheres to them, could result in new, revolutionary ways of performing research and fulfill the promise of open science. Furthermore, it allows for concepts such as personalized medicine and personal health monitoring to -finally- become implemented in daily practice. However, to bring about these changes, data users need to rethink the way they treat scientific data. Just passing a dataset along, without extensive metadata will not suffice anymore. Such new ways of executing research require a significantly different approach from the entire scientific community or, for that matter, anyone who wants to reap the benefits from going FAIR. Yet, how do you initiate behavioral change? One important solution is by changing the software scientists use and requiring data owners, or data stewards, to FAIRify their dataset. Data catalogs are a great starting point for FAIRifying data as the software already intends to make data Findable and Accessible, while the metadata is Interoperable and relying on users to provide sufficient metadata to ensure Reusability. In this paper we analyse how well the FAIR principles are implemented in several data catalogs. To determine how FAIR a catalog is, the FAIR metrics were created by the GO-FAIR initiative. These metrics help determine to what extend data can be considered FAIR. However, the metrics were only recently developed, being first released at the end of 2017. At the moment software does not come standard with a FAIR metrics review. Still, this insight is highly desired by the scientific community. How else can they be sure that (public) money is spend in a FAIR way? The Hyve has tested/evaluated three popular open source data catalogs based on the FAIR metrics: CKAN, Dataverse, and Invenio. Most data stewards will be familiar with at least one of these. Within this white paper we provide answers to the following questions: Which of the three data catalogs performs best in making data FAIR? Which data catalog utilizes FAIR datasets the most? Which one creates the most FAIR metadata? Which catalog has the highest potential to increase its FAIRness, and how? Which data catalog facilitates the FAIRifying process the best?
Read more