- Research Article
- 10.1016/j.jtct.2025.04.010
Heterogeneity of Survival Benefit Conferred By Letermovir.
- Jul 01, 2025
- Transplantation and cellular therapy
- Yu Akahoshi + 22 more +22
Heterogeneity of Survival Benefit Conferred By Letermovir.
Modern machine learning (ML) models are being used heavily in business domains to build effective decision support systems. As a primary requirement, supervised ML models need large labeled datasets. However, obtaining a high volume of labeled training data is both expensive and time-consuming. Researchers have proposed several labeling approaches to avoid manual labeling efforts. Active learning (AL) and Data Programming (DP) are two state-of-the-art techniques used to label datasets. Nevertheless, both approaches have their strengths and weaknesses. For example, AL is computationally expensive to apply on large industrial datasets; and labels generated by DP are often inaccurate and difficult to interpret. To address these challenges, in this paper, we propose a novel hybrid method that integrates the scalability of DP with the user engagement and accuracy of AL. The proposed approach aims at optimizing the labeling process by applying DP to generate initial noisy training data and then use AL to query the user to label only those points that maximize the accuracy of the final labels with a minimum annotation cost. To evaluate the proposed approach, we have used five open source datasets and a real-world business dataset of 1.5 million records. We use traditional active learning and data programming techniques as baselines to compare the performance and annotation cost of our proposed approach. The results show that the proposed method can achieve higher labeling accuracy than data programming. It also can minimize the labeling cost in real-world business scenarios, while delivering a comparable level of performance (accuracy) with active learning.
Heterogeneity of Survival Benefit Conferred By Letermovir.
Heterogeneity of Survival Benefit Conferred By Letermovir.
Comparative Studyof Machine Learning and System Identificationfor Process Systems Engineering Dynamics
This study provides a comprehensive benchmarking of traditionalsystem identification and modern machine learning (ML) models forthe data-driven modeling of dynamical systems, with a focus on processsystems engineering (PSE) applications. To achieve this, we deployAutoSID, an automated end-to-end framework inspired by Machine LearningOperations (MLOps) principles. While AutoSID facilitates model selection,training, validation, and evaluation, its purpose here is to serveas a platform to investigate how well MLOps-inspired tools can beadapted for system identification tasks in PSE. Our investigationincludes a comparison of 12 diverse model architectures from the systemidentification, machine learning, and deep learning literature, evaluatedacross 11 PSE case studies under varying data regimes. We employ fourmodel search or hyperparameter optimization algorithms and three modelselection criteria to ensure a thorough assessment. Our findings highlightthe importance of model selection as the crucial step in system identification.Specifically, our results demonstrate the effectiveness of Bayesianoptimization with tree-structured parzen estimators (TPE) for balancedmodel selection, while k-fold cross-validation proves to be a robustmetric for performance evaluation during the selection process. Inlarge-scale data scenarios, where performance differences between k-fold cross-validation and information criteria are small,information criteria emerge as a computationally efficient alternative.Once the “best” model structure is decided, in termsof model performance, we find that ML models with balanced complexity,such as tree ensemble models, consistently achieve superior predictiveaccuracy and computational efficiency, outperforming both simplisticand overly complex models. These findings provide actionable insightsinto model selection and performance evaluation for PSE practitionersand demonstrate the potential of incorporating MLOps-inspired workflowsinto the system identification process.
Read moreRethinking Zero-Loss in AI Data Centers: Towards Bounded-Loss Reliability in Deep Learning Systems
As modern machine learning (ML) workloads grow in size and complexity, communication has emerged as a critical systems bottleneck—not only during training, but across the full lifecycle of ML computation, including inference, pretraining, and fine-tuning: Distributed execution frameworks routinely rely on synchronized communication for exchanging gradients, intermediate activations, token embeddings, routing metadata, and attention context, all of which are vulnerable to tail latency, especially in large-scale environments. Delays in any one of these exchanges can stall iteration progress, degrade throughput, and underutilize expensive compute (e.g., GPU) resources. This dissertation rethinks communication in ML systems through the lens of domain-specific loss tolerance, and proposes two complementary systems that leverage the statistical resilience of ML workloads to partial information, offering a new approach to reducing latency and improving scalability. The first, OptiReduce, is a software-based solution targeting collective communication during distributed data parallel training, the most commonly used distributed ML technique. It introduces a best-effort, time-bounded AllReduce mechanism that tolerates bounded gradient loss to bypass stragglers and reduce iteration stalls. OptiReduce contributes three key techniques: (1) a Transpose AllReduce (TAR) algorithm that establishes direct peerto- peer communication among nodes, avoiding the loss propagation inherent in ring-based collectives where errors accumulate through intermediate aggregations; (2) an Unreliable Bounded Transport (UBT) with adaptive timeouts that caps communication latency regardless of network variability, using dynamic incast control and early timeout strategies to maximize gradient delivery within each time window; and (3) Hadamard Transform encoding that disperses the impact of dropped gradients across entire tensors, converting localized packet loss into uniformly distributed noise that ML models can absorb. Across public cloud environments including CloudLab, OptiReduce achieves up to 70% and 30% faster timeto- accuracy compared to Gloo and NCCL collective communication libraries, respectively, while limiting gradient loss to under 0.1% and maintaining model convergence accuracy even under lossy or congested network conditions. The second, OptiNIC, generalizes these principles to the RDMA transport layer in hardware. To extend beyond AllReduce to all collective patterns, support both training and inference workloads, and meet the 100–400 Gbps line-rate requirements of modern ML clusters, OptiNIC offloads the core ideas of bounded, best-effort communication directly into the NIC. It eliminates retransmissions and in-order delivery from the NIC datapath, enabling a best-effort, out-of-order transport model for RDMA. Unlike traditional transports that signal completion only after full data delivery, OptiNIC introduces bounded completion semantics—where operations complete within application-specified timeouts, triggering forward progress even when data is missing. By removing reorder buffers, retransmission queues, and per-packet sequencing logic, OptiNIC cuts NIC BRAM usage by 2.7× and nearly doubles mean-time-between-failure. Loss recovery shifts to lightweight software mechanisms such as the Hadamard Transform, while standard congestion control (e.g., DCQCN, EQDS, Swift) remains intact. Across public cloud deployments, OptiNIC improves time-to-accuracy by 2× for training, increases inference throughput by 1.6×, and reduces 99th-percentile latency by 3.5×—all while preserving model accuracy. Together, these systems demonstrate that domain-specific collective and transport designs, tailored to ML’s unique statistical properties, can unlock new levels of efficiency, resilience, and scalability in distributed machine learning.
Read moreMulti-Layer AI Defense Models Against Real-Time Phishing and Deepfake Financial Fraud
The rise of AI-powered cyberattacks-especially deepfake-driven financial scams and sophisticated phishing schemes-demands an equally advanced defense paradigm.This research proposes a multi-layer AI defense model that integrates real-time threat detection, behavioral analysis, adversarial learning, and biometric authentication to counter phishing and deepfake financial fraud.Drawing from both classical and modern machine learning models, the study explores hybrid architectures combining Natural Language Processing (NLP), Convolutional Neural Networks (CNNs), and Graph Neural Networks (GNNs).The system is tested against synthetic fraud scenarios, showing promising precision, adaptability, and low false positives.This paper also presents a comprehensive literature review, architecture design, sequence diagrams, and comparative analysis to reinforce the effectiveness of multi-layered defense in real-world financial systems.
Read moreWhen is memorization of irrelevant training data necessary for high-accuracy learning?
Modern machine learning models are complex and frequently encode surprising amounts of information about individual inputs. In extreme cases, complex models appear to memorize entire input examples, including seemingly irrelevant information (social security numbers from text, for example). In this paper, we aim to understand whether this sort of memorization is necessary for accurate learning. We describe natural prediction problems in which every sufficiently accurate training algorithm must encode, in the prediction model, essentially all the information about a large subset of its training examples. This remains true even when the examples are high-dimensional and have entropy much higher than the sample size, and even when most of that information is ultimately irrelevant to the task at hand. Further, our results do not depend on the training algorithm or the class of models used for learning. Our problems are simple and fairly natural variants of the next-symbol prediction and the cluster labeling tasks. These tasks can be seen as abstractions of text- and image-related prediction problems. To establish our results, we reduce from a family of one-way communication problems for which we prove new information complexity lower bounds. Additionally, we present synthetic-data experiments demonstrating successful attacks on logistic regression and neural network classifiers.
Read moreModelling of stock market security price Dynamics Using market microstructure Data
In modern electronic stock exchanges there is an opportunity to analyze event driven market microstructure data. This data is highly informative and describes physical price formation which makes it possible to find complex patterns in price dynamics. It is very time consuming and hard to find this kind of patterns by handcrafted rules. However, modern machine learning models are able to solve such issues automatically by learning price behavior which is always changing. The present study presents profitable trading system based on a machine learning model and market microstructure data. Data for the research was collected from Moscow stock exchange MICEX and represents a limit order book change log and all market trades of a liquid security for a certain period. Logistic regression model was used and compared to neural network models with different configuration. According to the study results logistic regression model has almost the same prediction quality as neural network models have but also has a high speed of response which is very important for stock market trading. The developed trading system has medium frequency of deals submission that lets it to avoid expensive infrastructure which is usually needed in high-frequency trading systems. At the same time, the system uses the potential of high quality market microstructure data to the full extent. This paper describes the entire process of trading system development including feature engineering, models behavior comparison and creation of trading strategy with testing on historical data.
Read moreDynamic GPU Energy Optimization for Machine Learning Training Workloads
GPUs are widely used to accelerate the training of machine learning workloads. As modern machine learning models become increasingly larger, they require a longer time to train, leading to higher GPU energy consumption. This paper presents GPOEO, an online GPU energy optimization framework for machine learning training workloads. GPOEO dynamically determines the optimal energy configuration by employing novel techniques for online measurement, multi-objective prediction modeling, and search optimization. To characterize the target workload behavior, GPOEO utilizes GPU performance counters. To reduce the performance counter profiling overhead, it uses an analytical model to detect the training iteration change and only collects performance counter data when an iteration shift is detected. GPOEO employs multi-objective models based on gradient boosting and a local search algorithm to find a trade-off between execution time and energy consumption. We evaluate the GPOEO by applying it to 71 machine learning workloads from two AI benchmark suites running on an NVIDIA RTX3080Ti GPU. Compared with the NVIDIA default scheduling strategy, GPOEO delivers a mean energy saving of 16.2% with a modest average execution time increase of 5.1%.
Read moreOn the Adversarial Robustness of Quantized Neural Networks Against Common Adversaries in Time-Series Forecasting
Real-world edge applications now use modern machine learning models which require both resource efficiency and robustness against adversarial threats. Deep neural networks which include time series forecasting models still face risks from adversarial perturbations while quantization techniques used for memory and compute efficiency create unpredictable robustness challenges. This project investigates the adversarial resistance of Long Short-Term Memory (LSTM) models after applying post-training quantization at three different precision levels: 16-bit floating point (FP16), 8-bit integer (INT8) and custom 4-bit quantization. The Jena Climate dataset serves as our main benchmark for training a fullprecision LSTM model followed by multiple quantization strategies which include TensorFlow Lite converters and a custom 4-bit quantization method based on ACIQ techniques that use clipping and scaling with bias correction. The evaluation of all models occurs under clean and adversarial conditions through Fast Gradient Sign Method (FGSM) and Basic Iterative Method (BIM) attacks with various perturbation strengths. The results indicate that quantized models show minor accuracy losses in clean conditions yet 4-bit models demonstrate better resistance to adversarial attacks. We create non-retraining defense strategies to enhance robustness which include input clipping and temporal median filtering and Gaussian smoothing and feature squeezing. These lightweight defenses effectively restore performance after an attack by reducing adversarial degradation without requiring modifications to model parameters. The complete pipeline is validated on the Jena Climate dataset and further tested on two additional multivariate datasets—Beijing PM2.5 and Appliances Energy Prediction—demonstrating consistent robustness patterns across data types. Our research demonstrates essential trade-offs between compression and efficiency and robustness which provides useful guidance for creating secure time-series models suitable for edge AI deployment
Read moreCost-aware active learning for named entity recognition in clinical text.
Active Learning (AL) attempts to reduce annotation cost (ie, time) by selecting the most informative examples for annotation. Most approaches tacitly (and unrealistically) assume that the cost for annotating each sample is identical. This study introduces a cost-aware AL method, which simultaneously models both the annotation cost and the informativeness of the samples and evaluates both via simulation and user studies. We designed a novel, cost-aware AL algorithm (Cost-CAUSE) for annotating clinical named entities; we first utilized lexical and syntactic features to estimate annotation cost, then we incorporated this cost measure into an existing AL algorithm. Using the 2010 i2b2/VA data set, we then conducted a simulation study comparing Cost-CAUSE with noncost-aware AL methods, and a user study comparing Cost-CAUSE with passive learning. Our cost model fit empirical annotation data well, and Cost-CAUSE increased the simulation area under the learning curve (ALC) scores by up to 5.6% and 4.9%, compared with random sampling and alternate AL methods. Moreover, in a user annotation task, Cost-CAUSE outperformed passive learning on the ALC score and reduced annotation time by 20.5%-30.2%. Although AL has proven effective in simulations, our user study shows that a real-world environment is far more complex. Other factors have a noticeable effect on the AL method, such as the annotation accuracy of users, the tiredness of users, and even the physical and mental condition of users. Cost-CAUSE saves significant annotation cost compared to random sampling.
Read moreVirtual Adversarial Active Learning
In traditional active learning, one of the most well-known strategies is to select the most uncertain data for annotation. By doing that, we acquire as most as we can obtain from the labeling oracle so that the training in the next run can be much more effective than the one from this run once the informative labeled data are added to training. The strategy however, may not be suitable when deep learning become one of the dominant modeling techniques. Deep learning is notorious for its failure to achieve a certain degree of effectiveness under the adversarial environment. Often we see the sparsity in deep learning training space which gives us a result with low confidence. Moreover, to have some adversarial inputs to fool the deep learners, we should have an active learning strategy that can deal with the aforementioned difficulties. We propose a novel Active Learning strategy based on Virtual Adversarial Training (VAT) and the computation of local distributional roughness (LDR). Instead of selecting the data that are closest to the decision boundaries, we select the data that are located in a place with rough enough surface if measured by the posterior probability. The proposed strategy called Virtual Adversarial Active Learning (VAAL) can help us to find the data with rough surface, reshape the model with smooth posterior distribution output thanks to the active learning framework. Moreover, we shall prefer the labeling data that own enough confidence once they are annotated from an oracle. In VAAL, we have the VAT that can not only be used as a regularization term but also helps us effectively and actively choose the valuable samples for active learning labeling. Experiment results show that the proposed VAAL strategy can guide the convolutional networks model converging efficiently on several well-known datasets.
Read moreComparative Analysis of Traditional Machine Learning and Active Learning in Science Experiments
In recent years, machine learning has been vastly used in many scientific experiments to analyze and predict the data [1]. Since there are many types of machine learning and each of them has additional further branches, it's necessary to understand which one can predict or summarize the data with more accuracy. This paper mainly focuses on traditional machine learning and active learning. To compare the effects of these two types of machine learning, the paper uses a dataset about using metal oxide semiconductor (MOX) to measure the concentration of carbon monoxide (CO). In the first part, the paper tests the accuracy of predicting CO concentration with 3 types of traditional machine learning: Decision Tree, Random Forest and K-nearest neighbors. In the second part, the paper chooses the traditional machine learning which has the best performance in the first part and compares the accuracy of it with the active learning on predicting CO concentration. The result is that the classification accuracy of Random Forest is 0.625, which is the highest among them. And active learning is generally better than traditional machine learning when training samples are small, and they have a similar accuracy when the training sample is enough.
Read moreLearning how to Active Learn: A Deep Reinforcement Learning Approach
Active learning aims to select a small subset of data for annotation such that a classifier learned on the data is highly accurate. This is usually done using heuristic selection methods, however the effectiveness of such methods is limited and moreover, the performance of heuristics varies between datasets. To address these shortcomings, we introduce a novel formulation by reframing the active learning as a reinforcement learning problem and explicitly learning a data selection policy, where the policy takes the role of the active learning heuristic. Importantly, our method allows the selection policy learned using simulation on one language to be transferred to other languages. We demonstrate our method using cross-lingual named entity recognition, observing uniform improvements over traditional active learning.
Read moreIntelligent Part Comparison in Computer Aided Design
Increasing product complexity and customisation lead to a growing number of parts in computer-aided design, which poses a challenge for part management. This paper presents an approach for automated part comparison using part geometry representation. The solution involves converting CAD models, generating shape distribution histograms and using bounding box dimensions for part comparisons. The method, tested on an open source dataset and a large industrial dataset, shows significant improvements in the efficiency of identifying similar parts, making it highly applicable for industrial use.
Read moreStructural Modeling and Measuring Impact of Active Learning Methods in Engineering Education
<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Contribution:</i> This article presents a novel approach that demonstrates how students’ learning can be improved by increasing classroom adherence to active learning (AL) application, in a typical engineering education (EE) environment. It does that by using classroom observation protocol data and student assessment grades analyzed by a statistical tool, representing seven engineering programs. <p xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <i>Background:</i> AL applications in EE have been growing in recent years and there have been relevant discussions about their effect on students learning. This research is grounded on a 2.5-year-long data collection process through objective measures, following a strict experimental design project. It differs from other studies where surveys are used for data collection. <p xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <i>Research Question:</i> Can differences in students’ learning between traditional teaching and AL methods be distinguished by means of classroom observation protocol measures coupled with partial least-square structural equation modeling? <p xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <i>Methodology:</i> An experimental research design was conducted in an Engineering Higher Education Institution, by taking independent measures of latent constructs’ indicators, such as student grades and AL classroom adherence levels. The data were subject to a partial least-squares structural equation modeling (PLS-SEM) approach for modeling, analysis, and validation. <p xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <i>Findings:</i> The results suggested a nonlinear positive cause-and-effect relationship between AL adherence and learning, validated by several performance indexes, such as a learning prediction relevance <inline-formula> <tex-math notation="LaTeX">$Q^{2}~\mathrm{predict}$</tex-math> </inline-formula> <inline-formula> <tex-math notation="LaTeX">$=$</tex-math> </inline-formula> 0.453 and goodness of fit <inline-formula> <tex-math notation="LaTeX">$=$</tex-math> </inline-formula> 0.588. The model demonstrated a learning score improvement from 45.89 to 74.90 as an effect of the adherence to AL score increase from 0 to 35.97.
Read moreUsing Deep Active Learning to Save Sensing Cost When Estimating Overall Air Quality
Air quality is widely concerned by the governments and people. To save cost, air quality monitoring stations are deployed at only a few locations, and the stations are actuated at partial time. Therefore, it is necessary to study how to actively collect a subset of air quality data to maximize the estimation accuracy of air quality at other locations and time. In order to solve this challenge, we propose the active variational adversarial model (AVAM) that selects the most valuable unlabeled samples through two iterative phases of active learning. In the first phase of our model, a candidate set with unlabeled samples is selected through traditional active learning. In the second phase, variational auto-encoder (VAE) is used to obtain the compressed representation of the candidate set and the training set with labeled samples, then a discriminator based on three-layer neural network is trained from the compressed representation. Finally the discriminator can output the most valuable unlabeled samples from the candidate set. The experimental results show that the AVAM proposed in this paper is superior to active learning models with the first or second phase only.
Read more