Data-Driven Prediction and Uncertainty Quantification on Chemical Concentration in Electroless Plating Process
Electroless plating is a chemical process commonly used in semiconductor manufacturing that deposits a uniform metal coating onto a substrate through an auto-catalytic reduction reaction without using electricity. High quality deposition hinges on the precise control of bath composition and operating conditions, which in turn requires accurate and accessible monitoring of the bath chemical concentration. While chemical analyzers present an effective monitoring technique, they are generally limited by high costs, complexity, and maintenance. Time-delayed measurements also impose challenges on real-time operations, and any reduction to the monitoring lag would be highly beneficial. We introduce a data-driven machine learning (ML) approach able to achieve fast concentration predictions directly from operating conditions, and able to compute the prediction uncertainty that is valuable for subsequent robust control and optimization and for improving transparency and trustworthiness of ML tools in manufacturing settings. Notably, our ML procedure overcomes challenges of asynchronous time-series training data with missing values, and where available data are often noisy and sparse. Our ML approach begins with a data preprocessing step to handle asynchronous measurements and missing values by engineering features rooted in the underlying physical processes. Notably, a systematic feature selection is performed to down select features that are most correlated with the prediction target, thereby reduce model overfitting. We then compare a number of regression model architectures for capturing the sequential data relationships, including linear models, random forest, extreme gradient boosting, and fully connected, long short-term memory, and transformer neural networks. We quantify the model uncertainty by training them in a Bayesian manner using scalable Stein variational gradient descent, to compute the posterior probability distributions conditioned on the training data that reflect the uncertainty in the models induced by the quality and quantity of the available observations. We demonstrate the overall ML approach on a real-world dataset from a semiconductor manufacturing plant that exhibits the aforementioned data challenges. The best performing models exhibited 1, 5, 12% accuracy (defined as achieving within a 3% margin) improvements on predicting the metal ion, alkali, reductant concentrations, respectively, over a naive feature extraction technique. It also achieved 13, 10, 32% accuracy improvements over an autoregressive integrated moving average (ARIMA) baseline. These results show the approach’s effectiveness in providing reliable predictions with quantified uncertainty.
Read more