- Research Article
2
- 10.1016/j.inffus.2025.104092
VLDBench Evaluating multimodal disinformation with regulatory alignment
- Jun 01, 2026
- Information Fusion
- Shaina Raza + 10 more +10
Publications from 2021 to 2026
Showing 10 of 283 papers
VLDBench Evaluating multimodal disinformation with regulatory alignment
From conventional screening to self-driving discovery: Organ-on-Chip platforms as engines for AI-guided nanomedicine.
Nanoparticles have become an essential platform for next-generation drug delivery and therapeutic development, yet clinical translation remains limited by an incomplete understanding of their interactions within human biological systems. Organ-on-a-chip technology offers a powerful approach to evaluate efficacy and safety of nanomedicine under physiologically relevant conditions in human cells by recreating fluid flow, mechanical stress, barrier function, immune interactions and inter-organ communications. These advanced in vitro systems allow quantitative assessments of nanoparticle transport, distribution, and safety with improved human relevance compared to conventional cell culture and animal models. The integration of sensors within organ-on-a-chip platforms enables real-time monitoring of tissue responses and nanoparticle kinetics. Advances in automation and robotic liquid handling support scalable and reproducible testing across multiple tissue models. Artificial intelligence and active learning tools facilitate automated data analysis and experimental optimization, paving the way towards self-driving nanomedicine evaluation platforms that accelerate discovery and clinical translation.
Read moreGeneralist biological artificial intelligence in modeling the language of life.
Generalist biological artificial intelligence (GBAI) represents a transformative approach to modeling the 'language of life'-the flow of information from DNA to cellular function. This Review synthesizes rapid advances in biological AI to interpret and generate DNA, RNA, proteins and cellular systems. We chart a course toward comprehensive systems that can concurrently process and predict across these domains, performing several critical biological tasks simultaneously. Substantial opportunities lie in synergizing language and structural AI, leveraging specialized models and improving AI agents for autonomous discovery. After addressing challenges in data, biological complexity, scalability and experimental validation, GBAI has the potential to deepen our understanding of disease pathways and biomarkers, advance automated therapeutic design and evaluation, and integrate within virtual cells to meaningfully simulate biological activity.
Read moreLUMI-lab: A foundation model-driven autonomous platform enabling discovery of ionizable lipid designs for mRNA delivery.
BarcodeBERT: transformers for biodiversity analyses
MotivationIn the global effort to characterize biodiversity, short species-specific genomic sequences known as DNA barcodes enable fine-grained comparisons among organisms within the same kingdom of life. Although machine learning algorithms specifically designed for the analysis of DNA barcodes are becoming more popular, most existing methodologies rely on generic supervised training algorithms.ResultsWe introduce BarcodeBERT, a family of models tailored to biodiversity analysis and trained exclusively on data from a reference library of 1.5 M invertebrate DNA barcodes. We evaluate BarcodeBERT on taxonomic identification tasks against a spectrum of machine learning approaches, including supervised training of classical neural architectures and fine-tuning of general DNA foundation models. Our self-supervised pretraining strategies on domain-specific data outperform fine-tuned foundation models, especially in identification tasks involving lower taxa such as genera and species. Compared with BLAST, a widely used sequence-search tool, BarcodeBERT achieves comparable species-level classification accuracy while being 55× faster. Our analysis of masking and tokenization strategies also provides practical guidance for building customized DNA language models, emphasizing the importance of aligning model training strategies with dataset characteristics and domain knowledge.Availability and implementationThe code repository is available at https://github.com/bioscan-ml/BarcodeBERT.
Read moreSampling of natural speech for the assessment of psychopathology: data collection procedure and inter-rater reliability.
Learning enhanced ensemble filters
Insights into tribal-level adaptive evolution and phylogeny in Soricinae from morphology and mitogenome of the Chinese endemic Sorex cansulus
Sorex cansulus is a critically endangered shrew species endemic to China. To elucidate adaptive evolution and evolutionary relationships of five tribes of the subfamily Soricinae. This study presents the morphology and complete mitogenome of S. cansulus. Codon usage bias analyses indicated that natural selection was the predominant force shaping its mitochondrial protein-coding genes, with exception for the atp8 gene. We performed a comparative mitogenomic analysis of 43 species across 13 genera and 5 tribes. The evolutionary rates (Ka/Ks) of 13 protein-coding genes across all tribes were significantly less than 1, suggesting strong purifying selection and evolutionary conservation. Notably, tribes with more similar Ka/Ks values exhibited closer evolutionary relationships. Phylogenetic trees yielded consistent topologies and strongly supported the monophyly of Soricinae (PP=1.00; BS=100%). Our results supported key taxonomic revisions: the reclassification of Episoriculus fumidus to the genus Pseudosoriculus; the elevation of Blarinella griselda to the genus Parablarinella, and the recognition of S. cansulus and Sorex sinalis as distinct sister species. Within the inferred phylogeny, Blarinellini was placed at the base as the earliest-diverging lineage, while Nectogalini was at the top as the latest-diverging lineage. However, some internal nodes within tribes showed relatively low support, which may be attributed to incomplete taxon sampling, rapid radiative evolution. All five tribes retained the typical mammalian mitogenomic organization, with the notable exception of Sorex daphaenodon, Sorex tundrensis, and Sorex araneus, which possessed an extra trnW gene (23 tRNA genes total). These three species shared this unique feature and formed a distinct clade, whose implications for their mitochondrial function and adaptation warrant further research.
Read moreDeep Learning Improves Photometric Redshifts in All Regions of Color Space
Abstract Photometric redshifts (photo- z ’s) are crucial for the cosmology, galaxy evolution, and transient science drivers of next-generation imaging facilities like the Euclid Mission, the Vera C. Rubin Observatory, and the Nancy Grace Roman Space Telescope. Previous work has shown that image-based deep learning photo- z methods produce smaller scatter than photometry-based classical machine learning (ML) methods on the Sloan Digital Sky Survey (SDSS. Main Galaxy Sample, a test bed photo- z dataset. However, global assessments can obscure local trends. To explore this possibility, we used a self-organizing map (SOM) to cluster SDSS galaxies based on their ugriz colors. Deep learning methods achieve lower photo- z scatter than classical ML methods for all SOM cells. The fractional reduction in scatter is roughly constant across most of color space with the exception of the most bulge-dominated and reddest cells where it is smaller in magnitude. Interestingly, classical ML photo- z ’s suffer from a significant color-dependent attenuation bias, where photo- z ’s for galaxies within an SOM cell are systematically biased towards the cell’s mean spectroscopic redshift and away from extreme values, which is not readily apparent when all objects are considered. In contrast, deep learning photo- z ’s suffer from very little color-dependent attenuation bias. The increased attenuation bias for classical ML photo- z methods is the primary reason why they exhibit larger scatter than deep learning methods. This difference can be explained by the deep learning methods weighting redshift information from the individual pixels of a galaxy image more optimally than integrated photometry.
Read moreDiscovery of predictive biomarkers for cancer therapy through computational approaches.
Precision oncology involves the use of predictive biomarkers to personalize treatment. However, for most cancer therapeutics or combination regimens, effective biomarkers have been elusive. This challenge has fuelled efforts to interrogate increasingly diverse and complex clinical and molecular determinants of treatment response. Some molecular predictors have been identified (for example, based on analysis of transcriptomic or imaging data), although the limited reproducibility and robustness of many of these candidate biomarkers make them difficult to apply in clinical practice. Moreover, different types of predictor must often be combined to optimize treatment selection (for example, gene signatures plus patient characteristics). Computational methods, including machine learning and artificial intelligence approaches, provide opportunities to identify predictive patterns in both clinical data and preclinical datasets and to predict treatment response for individual patients. Such approaches also offer opportunities to predict the efficacy or synergy of drug combinations, for example, via extrapolation from correlations of monotherapy responses or by linking the cellular responses observed in preclinical drug screens with molecular and clinical data from patients. In this Review, we describe the application of computational methods to predictive biomarker discovery, including current progress, key challenges facing this field, and future opportunities.
Read more