- Research Article
2
- 10.1002/wea.1982
Challenges for data assimilation – from convective‐scale to climate
- Sep 26, 2012
- Weather
- Richard Marriott
Challenges for data assimilation – from convective‐scale to climate
Reconstructing spatiotemporal systems from sparse observations remains a long-standing challenge in several domains, including geoscience, air pollution and fluid dynamics. While various data assimilation (DA) and machine learning (ML) methods have shown potential, they still face significant challenges (see [1]): The computational burden of conventional DA algorithms (including error covariance specification), particularly for multivariate, high-dimensional systems. Sparse and movable sensor placement, which makes conventional ML models (typically requiring fixed and regularly distributed input data) cumbersome. The ill-defined nature of the sparse reconstruction problem, which poses significant risks of overfitting. We will present our recent works aimed at addressing these challenges. More specifically, we developed latent DA algorithms [2] to reduce the computational burden of variational DA methods. These algorithms demonstrate great potential in efficiently assimilating sparse observations within a reduced-order latent space constructed by neural networks, thanks to the TorchDA library [3]. The latter enables GPU implementation of mainstream data assimilation methods and supports non-explicit state-observation transformation functions, provided they can be learned by a neural network.We have also employed advanced deep learning techniques, including Voronoi-tessellation CNNs [4] and Vision Transformer-based autoencoders [5], to learn mappings from sparse observations to the complete physical space. These approaches effectively address challenges such as movable sensor placements and varying sensor numbers. Their integration with DA algorithms has also been evaluated.Finally, our recent work explores [6] the utility of generative AI techniques, particularly denoising diffusion models, for field reconstruction from sparse observations. Generative AI methods offer two main advantages: first, they produce a sample from a probability distribution rather than predicting the mean as a fixed output, which can help mitigate overfitting caused by the illy-defined problem. Second, they inherently function as ensemble predictors by generating several samples, facilitating uncertainty quantification, which is essential in data assimilation. The numerical results tested on cases ranging from fluid dynamics benchmarks to semi-operational air pollution simulations will also be discussed.[1] Cheng, S., Quilodrán-Casas, C., Ouala, S., Farchi, A., Liu, C., Tandeo, P., Fablet, R., Lucor, D., Iooss, B., Brajard, J., Xiao, D., Janjic, T., Ding, W., Guo, Y., Carrassi, A., Bocquet, M. and Arcucci, R, 2023. Machine learning with data assimilation and uncertainty quantification for dynamical systems: a review. IEEE/CAA Journal of Automatica Sinica[2] Cheng, S., Chen, J., Anastasiou, C., Angeli, P., Matar, O.K., Guo, Y.K., Pain, C.C. and Arcucci, R., 2023. Generalised latent assimilation in heterogeneous reduced spaces with machine learning surrogate models. Journal of Scientific Computing[3] Cheng, S., Min, J., Liu, C. and Arcucci, R., 2025. TorchDA: A Python package for performing data assimilation with deep learning forward and transformation functions. Computer Physics Communications[4] Cheng, S., Liu, C., Guo, Y. and Arcucci, R., 2024. Efficient deep data assimilation with sparse observations and time-varying sensors. Journal of Computational Physics[5] Fan, H., Cheng, S., de Nazelle, A.J. and Arcucci, R., 2024. ViTAE-SL: a vision transformer-based autoencoder and spatial interpolation learner for field reconstruction. Computer Physics Communications [6] Zhuang, Y., Cheng, S. and Duraisamy, K., 2025. Spatially-aware diffusion models with cross-attention for global field reconstruction with sparse observations. Computer Methods in Applied Mechanics and Engineering
Challenges for data assimilation – from convective‐scale to climate
Challenges for data assimilation – from convective‐scale to climate
Discharge estimation under uncertainty using variational methods with application to the full Saint‐Venant hydraulic network model
SummaryEstimating river discharge from in situ and/or remote sensing data is a key issue for evaluation of water balance at local and global scales and for water management. Variational data assimilation (DA) is a powerful approach used in operational weather and ocean forecasting, which can also be used in this context. A distinctive feature of the river discharge estimation problem is the likely presence of significant uncertainty in principal parameters of a hydraulic model, such as bathymetry and friction, which have to be included into the control vector alongside the discharge. However, the conventional variational DA method being used for solving such extended problems often fails. This happens because the control vector iterates (i.e., approximations arising in the course of minimization) result into hydraulic states not supported by the model. In this paper, we suggest a novel version of the variational DA method specially designed for solving estimation‐under‐uncertainty problems, which is based on the ideas of iterative regularization.The method is implemented with SIC2, which is a full Saint‐Venant based 1D‐network model. The SIC2 software is widely used by research, consultant and industrial communities for modeling river, irrigation canal, and drainage network behavior. The adjoint model required for variational DA is obtained by means of automatic differentiation. This is likely to be the first stable consistent adjoint of the 1D‐network model of a commercial status in existence.The DA problems considered in this paper are offtake/tributary estimation under uncertainty in the cross‐device parameters and inflow discharge estimation under uncertainty in the bathymetry defining parameters and the friction coefficient. Numerical tests have been designed to understand identifiability of discharge given uncertainty in bathymetry and friction. The developed methodology, and software seems useful in the context of the future Surface Water and Ocean Topography satellite mission. Copyright © 2016 John Wiley & Sons, Ltd.
Read moreSpatially and Temporally Complete Satellite Soil Moisture Data Based on a Data Assimilation Method
Multiple soil moisture products have been generated from data acquired by satellite. However, these satellite soil moisture products are not spatially or temporally complete, primarily due to track changes, radio-frequency interference, dense vegetation, and frozen soil. These deficiencies limit the application of soil moisture in land surface process simulation, climatic modeling, and global change research. To fill the gaps and generate spatially and temporally complete soil moisture data, a data assimilation algorithm is proposed in this study. A soil moisture model is used to simulate soil moisture over time, and the shuffled complex evolution optimization method, developed at the University of Arizona, is used to estimate the control variables of the soil moisture model from good-quality satellite soil moisture data covering one year, so that the temporal behavior of the modeled soil moisture reaches the best agreement with the good-quality satellite soil moisture data. Soil moisture time series were then reconstructed by the soil moisture model according to the optimal values of the control variables. To analyze its performance, the data assimilation algorithm was applied to a daily soil moisture product derived from the Advanced Microwave Scanning Radiometer for the Earth Observing System (AMSR-E), the Microwave Radiometer Imager (MWRI), and the Advanced Microwave Scanning Radiometer 2 (AMSR2). Preliminary analysis using soil moisture data simulated by the Global Land Data Assimilation System (GLDAS) Noah model and soil moisture measurements at a multi-scale Soil Moisture and Temperature Monitoring Network on the central Tibetan Plateau (CTP-SMTMN) was performed to validate this method. The results show that the data assimilation algorithm can efficiently reconstruct spatially and temporally complete soil moisture time series. The reconstructed soil moisture data are consistent with the spatial precipitation distribution and have strong positive correlations with the values simulated by the GLDAS Noah model over large areas of the region. Compared to the soil moisture measurements at the medium and large networks, the reconstructed soil moisture data have almost the same accuracy as the soil moisture product derived from AMSR-E/MWRI/AMSR2 for ascending and descending orbits.
Read moreCyclic Steam Injection Modeling and Optimization for Candidate Selection, Steam Volume Optimization, and SOR Minimization, Powered by Unique, Fast, Modeling and Data Assimilation Algorithms
The application of a novel modeling and data assimilation approach is presented, demonstrating the impact of quantitatively modeled and optimized cyclic steam candidate and steam volume selection in a mature heavy oil field. Results are reviewed for a cyclic steam operation in the San Joaquin Basin in California, including steam savings, production increases and SOR reduction in excess of 20 percent. The new approach is based on a novel modeling and data assimilation method. This paper focuses on the data assimilation aspect of the workflow, which is fast, quantifies uncertainty and enables simultaneous assimilation of multiple data sources. The approach is a combination of a modified Ensemble Kalman Filter (EnKF) with quadratic programming that enables assimilation of data from thousands of wells and different data sources with a relatively small ensemble. Additionally, since the EnKF is known to underestimate uncertainty, statistical techniques are used to correct the uncertainty estimates to conform with empirical estimates. The predictive capacity of the calibrated models is demonstrated through a statistical back-test process, wherein, the wells are fitted to historical data up to before the last steam job, and the response of the last steam job is predicted and compared to actual production. Once a calibrated model is validated, the models can then be used to predict the performance of future jobs, and thereby identify the best wells to steam, and also optimize steam volumes that minimize SOR and maximize incremental oil production due to the steam job. The above described modeling and optimization workflow has been applied to multiple fields in the San Joaquin valley. For the field chosen for this case study, production comes from the unconsolidated sands of the Pliocene Chanac and Kern River formations with porosity averaging 30%, permeability averaging 1,500-2,000 mD and net thicknesses typically between 50 and 300 feet. Dip is generally monoclinal across the field at approximately 5 – 15 degrees. The reservoir is shallow, with depths ranging from 750 – 2,250 feet. Oil gravity averages 14 – 16° API. Reservoir pressure is well below bubble point and averages 25 – 100 PSI. Data sources assimilated included production and injection rates, wellhead pressures, steam and producer temperatures, temperature profiles, steam ID logs and various petro-physical logs. Results to date indicate that application of this workflow has increased incremental oil production from the steam jobs by over 44% compared to a prior control period. And most interestingly, such an increase in production is achieved while cutting overall steam injection, as the steam volume optimization identifies inefficient use of steam. The most significant new finding is that cyclic steam injection can be quantitatively modeled and optimized rapidly to maximize profit and minimize SOR. This is important, particularly for mature heavy oil fields, in order to maximize recovery amidst varying commodity prices. The novelty of the new model is its combination of speed of data integration (less than a week) and runtime (minutes) with long-term predictive accuracy (years or decades). This is due to the unique modeling and data assimilation methodology.
Read moreDeep Learning Augmented Data Assimilation: Reconstructing Missing Information with Convolutional Autoencoders
Remote sensing data play a critical role in improving numerical weather prediction (NWP). However, the physical principles of radiation dictate that data voids frequently exist in physical space (e.g., subcloud area for satellite infrared radiance or no-precipitation region for radar reflectivity). Such data gaps impair the accuracy of initial conditions derived from data assimilation (DA), which has a negative impact on NWP. We use the barotropic vorticity equation to demonstrate the potential of deep learning augmented data assimilation (DDA), which involves reconstructing spatially complete pseudo-observation fields from incomplete observations and using them for DA. By training a convolutional autoencoder (CAE) with a long simulation at a coarse “forecast” resolution (T63), we obtained a deep learning approximation of the “reconstruction operator,” which maps spatially incomplete observations to a model state with full spatial coverage and resolution. The CAE was applied to an incomplete streamfunction observation (∼30% missing) from a high-resolution benchmark simulation and demonstrated satisfactory reconstruction performance, even when only very sparse (1/16 of T63 grid density) observations were used as input. When only spatially incomplete observations are used, the analysis fields obtained from ensemble square root filter (EnSRF) assimilation exhibit significant error. However, in DDA, when EnSRF takes in the combined data from the incomplete observations and CAE reconstruction, analysis error reduces significantly. Such gains are more pronounced with sparse observation and small ensemble size because the DDA analysis is much less sensitive to observation density and ensemble size than the conventional DA analysis, which is based solely on incomplete observations. Significance Statement Data assimilation plays a critical role in improving the skills of modern numerical weather prediction by establishing accurate initial conditions. However, unobservable regions are common in observation data, particularly those derived from remote sensing. The nonlinear relationship between data from observable regions and the physical state of unobservable regions may impede DA efficiency. As a result, we propose that deep learning be used to improve data assimilation in such cases by reconstructing a spatially complete first guess of the physical state with deep learning and then applying data assimilation to the reconstructed field. Such deep learning augmentation is found effective in improving the accuracy of data assimilation, especially for sparse observation and small ensemble size.
Read moreComparison of Hybrid-3DEnVar against 3DVar for the assimilation of surface observations over the Alpine terrain in AROME-Austria
Surface observations can provide crucial information for NWP models. If not assimilated carefully, however, they can degrade forecast accuracy, especially in complex terrains like the Alps. The horizontal and vertical covariances of climatological background error covariances used in the three-dimensional variational (3DVar) data assimilation (DA) method can produce unrealistic increments over sloped terrain. For instance, an observation from a valley station can still generate increments at the mountaintop, even though the valley observation may not accurately represent the mountaintop's weather conditions. We used a hybrid three-dimensional ensemble variational (Hybrid-3DEnVar) DA method to address this issue, incorporating a 50-member convection-permitting ensemble. This method was recently tested in Geosphere Austria's convective scale limited-area NWP model AROME at a 2.5 km horizontal resolution. We assimilated 2-meter temperature, 2-meter relative humidity, geopotential, and 10-meter wind components from 680 surface stations, including the Austrian TAWES network and SYNOP observations from neighbouring countries. 400 stations were actively assimilated from the observation dataset, and the rest were used to verify the analysis.  Our results present the effectiveness of this newly tested Hybrid-3DEnVar against GeoSphere Austria's operational 3DVar in assimilating surface observations over complex Alpine terrain.
Read moreA Python interface to the Fortran-based Parallel Data Assimilation Framework: pyPDAF v1.0.2
Abstract. Data assimilation (DA) is an essential component of numerical weather and climate prediction. Efficient implementation of DA algorithms benefits both research and operational prediction. Currently, a variety of DA software programs are available. One of the notable DA libraries is the Parallel Data Assimilation Framework (PDAF) designed for ensemble data assimilation. The DA framework is widely used with complex high-dimensional climate models, and is applied for research on atmosphere, ocean, sea ice and marine ecosystem modelling, as well as operational ocean forecasting. Meanwhile, there are increasing demands for flexible and efficient DA implementations using Python due to the increasing amount of intermediate complexity models as well as machine learning based models coded in Python. To accommodate for such demands, we introduce a Python interface to PDAF, pyPDAF. pyPDAF allows for flexible DA system development while retaining the efficient implementation of the core DA algorithms in the Fortran-based PDAF. The ideal use-case of pyPDAF is a DA system where the model integration is independent from the DA program, which reads the model forecast ensemble, produces an analysis, and updates the restart files of the model, or a DA system where the model can be used in Python. With implementations of both PDAF and pyPDAF, this study demonstrates the use of pyPDAF and PDAF in a coupled data assimilation (CDA) setup in a coupled atmosphere-ocean model, the Modular Arbitrary-Order Ocean-Atmosphere Model (MAOOAM). This study demonstrates that pyPDAF allows for PDAF functionalities from Python where users can utilise Python functions to handle case-specific information from observations and numerical model. The study also shows that pyPDAF can be used with high-dimensional systems with little slow-down per analysis step of only up to 13 % for the localized ensemble Kalman filter LETKF in the example used in this study. The study also shows that, compared to PDAF, the overhead of pyPDAF is comparatively smaller when computationally intensive components dominate the DA system. This can be the case for systems with high-dimensional state vectors.
Read moreVariational Autoencoder-Enhanced Variational Methods for Data Assimilation
Data assimilation (DA) is a statistical approach used to estimate the states of physical systems by integrating prior model predictions (background states xb​) with observational data (y). This integration produces an accurate estimate, called analysis states (​xa), by sampling or maximizing the posterior likelihood p(xxb, y). In weather forecasting, background states are generated by imperfect models, and the likelihood p(xxb) is often unknown. Observations, sourced from diverse instruments, are mapped to model space using observation operators (H). Effective DA algorithms must accurately estimate p(xxb) while accommodating various observation operators, including those involving sparse, noisy or irregular data.Traditional DA methods, such as variational assimilation, assume that the background error (​x - xb) follows a Gaussian distribution independent of xb​. This allows explicit computation of p(xxb, y) and optimization via techniques like gradient descent. While robust to various observation operators, these methods depend heavily on expert knowledge to construct error correlations and are limited by their Gaussian assumptions.Generative neural networks, particularly diffusion models, have emerged as alternatives for modeling p(xxb). Notable examples include SDA and DiffDA, which use diffusion models to learn background distributions. SDA incorporates observations via diffusion posterior sampling, while DiffDA employs the repaint technique. These approaches improve on traditional methods by capturing more complex distributions but often struggle with sparse, irregular observations. For instance, DiffDA assumes grid-aligned data, while SDA relies on assumptions that can reduce accuracy in real-world scenarios.In this research, we aim to develop a neural network-based data assimilation algorithm that not only captures the non-Gaussian characteristics of the conditional background distribution for enhanced accuracy but also effectively assimilates data under real-world observations (sparse, noisy and outside of the grid). We introduce VAE-Var, a novel data assimilation algorithm in which a variational autoencoder is first employed to learn the conditional background distribution and then the decoder component is utilized to construct a variational cost function, which, when optimized, yields the analysis states.Key advantages of VAE-Var include:This algorithm inherits the framework of traditional variational assimilation by explicitly modeling the posterior probability function p(xxb, y) and maximizing it to derive the analysis states. As a result, compared to other neural network data assimilation methods such as SDA and DiffDA, VAE-Var can better handle different types of observation operators, particularly irregular observations that do not fall on the grid points of the physical field. Unlike traditional variational assimilation algorithms, VAE-Var alleviates the dependence on expert knowledge for constructing the conditional background distribution, enabling the model to effectively capture non-Gaussian structures. This makes VAE-Var perform better in sparse observational settings. Experiments with the FengWu weather forecasting system at 0.25° resolution show that VAE-Var achieves higher accuracy than DiffDA and traditional algorithms (interpolation and 3DVar) in sparse observational settings. When integrated with FengWu, VAE-Var reliably assimilates real-world GDAS prepbufr observations over a one-year period.
Read moreA unified framework for the analysis of accuracy and stability of a class of approximate Gaussian filters for the Navier–Stokes equations
Bayesian state estimation of a dynamical system utilising a stream of noisy measurements is important in many geophysical and engineering applications. In these cases, nonlinearities, high (or infinite) dimensionality of the state space, and sparse observations pose key challenges for deriving efficient and accurate data assimilation (DA) techniques. A number of DA algorithms used commonly in practice, such as the Ensemble Kalman Filter (EnKF) or Ensemble Square Root Kalman Filter (EnSRKF), suffer from serious drawbacks such as catastrophic filter divergence and filter instability. The analysis of stability and accuracy of these DA schemes has thus far focused either on finite-dimensional dynamics, or on complete observations in case of infinite-dimensional dissipative systems. We develop a unified framework for the analysis of several well-known and empirically efficient DA techniques derived from various Gaussian approximations of the Bayesian filtering schemes for geophysical-type dissipative dynamics with quadratic nonlinearities. We establish rigorous results on (time-asymptotic) accuracy and stability of these algorithms with general covariance and observation operators. The accuracy and stability results for EnKF and EnSRKF for dissipative PDEs are, to the best of our knowledge, completely new in this general setting. It turns out that a hitherto unexploited cancellation property involving the ensemble covariance and observation operators and the concept of covariance localization in conjunction with covariance inflation play a pivotal role in the accuracy and stability for EnKF and EnSRKF. Our approach also elucidates the links, via determining functionals, between the approximate-Bayesian and control-theoretic approaches to DA. We consider the ‘model’ dynamics governed by the two-dimensional incompressible Navier–Stokes equations (NSEs) and observations given by noisy measurements of averaged volume elements or spectral/modal observations of the velocity field. In this setup, several continuous-time DA techniques, namely the so-called 3DVar, EnKF and EnSRKF reduce to a stochastically forced NSEs. For the first time, we derive conditions for accuracy and stability of EnKF and EnSRKF. The derived bounds are given for the limit supremum of the expected value of the L 2 norm and of the H 1 Sobolev norm of the difference between the approximating solution and the actual solution as the time tends to infinity. Moreover, our analysis reveals an interplay between the resolution of the observations (roughly, the ‘richness’ of the observation space) associated with the observation operator underlying the DA algorithms and covariance inflation and localization which are employed in practice for improved filter performance.
Read moreImproving Water Distribution Network Models with Data Assimilation: Unleashing the Power of Real-Time Improvements
Ensuring a reliable water supply in the face of changing conditions and growing demand is a critical global challenge. Water distribution networks (WDNs) are essential infrastructure, but traditional modeling methods based on historical data often struggle to adapt in real-time and integrate new information. In order to lower model errors in WDN models, this study explores the use of a Data Assimilation (DA) method that makes use of the Ensemble Kalman Filter. The study explores the effectiveness of the DA method, and a novel Greedy Algorithm (GA) for optimizing the location of sensors. The study shows that the DA method improves the WDN model and the GA is able to successfully determine optimal sensor locations. It was observed that increasing the number of sensors in the WDN increased the effectiveness of the DA method. This study highlights the potential of data assimilation to improve WDN modeling. Water utilities can gain from more precise predictions, greater system performance, and efficient decision-making for the management and maintenance of water distribution networks by allowing models to adapt and improve dynamically
Read moreReservoir Inverse Modeling by Ensemble Smoother with Multiple Data Assimilation for Seismic and Production Data
Time-lapse seismic data has been widely used for the detailed reservoir characterization, and data assimilation algorithms are commonly used for petroleum reservoir history matching of production data. However, hardly any seismic data has been integrated into the reservoir inverse modeling workflow, due to the large data size, finer gridding, and especially the scarcity in time of the seismic data. Popular ensemble-based reservoir inverse modeling methods such as ensemble Kalman filter (EnKF) also face the problems of high computational cost due to the storage of the intermediate variables, restarting the reservoir simulating process and the inconsistency of the full-step and step-wise simulations. The intrinsic sequential data assimilations characteristics of the EnKF set obstacles to assimilate 4D seismic data. The alternative method can be the ensemble smoother (ES), which is an ensemble based method for data assimilation. The ES is based on a Bayesian updating scheme of the reservoir model to match the production history and improve the production forecast. To improve the algorithm convergence, a multiple data assimilation (MDA) method was proposed by Emerick and Reynolds (2012a). The algorithm we used is called ensemble smoother with multiple data assimilation (ESMDA), we modified the ESMDA to integrate the geophysical data, and created the new workflow to history match both the production and geophysical data. The available production data is well measurements, including oil production rate, well water cut and bottom-hole pressure, as well as time-lapse geophysical data, including P-wave impedance. In this paper, we first formulated the mathematical descriptions of ESMDA, we then down-scaled the seismic data, and modified the data assimilation algorithms, illustrated the history matching workflow for both production and geophysical data, finally, we showed how it works with a case study on a water flooding operation in a synthetic reservoir. In comparison, we also showed the history matching only on the production data, which yields an inferior results than the one matches both production and seismic data.
Read moreReply on RC3
Quantifying continental-scale river discharge is essential to understanding the terrestrial water cycle but is susceptible to errors caused by a lack of observations and the limitations of hydrodynamic modeling. Data assimilation (DA) methods are increasingly used to estimate river discharge in combination with emerging river-related remote sensing products (e.g., water surface elevation [WSE], water surface slope, river width, and flood extent). However, directly comparing simulated WSE to satellite altimetry data remains challenging (e.g., because of large biases between simulations and observations or uncertainties in parameters), and large errors can be introduced when satellite observations are assimilated into hydrodynamic models. In this study we performed direct, anomaly, and normalized value assimilation experiments to investigate the capacity of DA to improve river discharge within the current limitations of hydrodynamic modeling. We performed hydrological DA using a physically-based empirical localization method applied to the Amazon Basin. We used satellite altimetry data from ENVISAT, Jason 1, and Jason 2. Direct DA was the baseline assimilation method and was subject to errors due to biases in the simulated WSE. To overcome these errors, we used anomaly DA as an alternative to direct DA. We found that the modeled and observed WSE distributions differed considerably (e.g., differences in amplitude, seasonal flow variation, and a skewed distribution due to limitations of the hydrodynamic models). Therefore, normalized value DA was performed to improve discharge estimation. River discharge estimates were improved at 24 %, 38 %, and 62 % of stream gauges in the direct, anomaly, and normalized value assimilations relative to simulations without DA. Normalized value assimilation performed best for estimating river discharge given the current limitations of hydrodynamic models. Most gauges within the river reaches covered by satellite observations accurately estimated river discharge, with Nash-Sutcliffe efficiency (NSE) > 0.6. The amplitudes of WSE variation were improved in the normalized DA experiment. Furthermore, in the Amazon Basin, normalized assimilation (median NSE = 0.50) improved river discharge estimation compared to open-loop simulation with the global hydrodynamic model (median NSE = 0.42). River discharge estimation using direct DA methods was improved by 7 % with calibration of river bathymetry based on NSE. The direct DA approach outperformed the other DA approaches when runoff was considerably biased, but anomaly DA performed best when the river bathymetry was erroneous. The uncertainties in hydrodynamic modeling (e.g., river bottom elevation, river width, simplified floodplain dynamics, and the rectangular cross-section assumption) should be improved to fully realize the advantages of river discharge DA through the assimilation of satellite altimetry. This study contributes to the development of a global river discharge reanalysis product that is consistent spatially and temporally.
Read moreComparative Study on Assimilating Remote Sensing High Frequency Radar Surface Currents at an Atlantic Marine Renewable Energy Test Site
A variety of data assimilation approaches have been applied to enhance modelling capability and accuracy using observations from different sources. The algorithms have varying degrees of complexity of implementation, and they improve model results with varying degrees of success. Very little work has been carried out on comparing the implementation of different data assimilation algorithms using High Frequency radar (HFR) data into models of complex inshore waters strongly influenced by both tides and wind dynamics, such as Galway Bay. This research entailed implementing four different data assimilation algorithms: Direct Insertion (DI), Optimal Interpolation (OI), Nudging and indirect data assimilation via correcting model forcing into a three-dimensional hydrodynamic model and carrying out detailed comparisons of model performances. This work will allow researchers to directly compare four of the most common data assimilation algorithms being used in operational coastal hydrodynamics. The suitability of practical data assimilation algorithms for hindcasting and forecasting in shallow coastal waters subjected to alternate wetting and drying using data collected from radars was assessed. Results indicated that a forecasting system of surface currents based on the three-dimensional model EFDC (Environmental Fluid Dynamics Code) and the HFR data using a Nudging or DI algorithm was considered the most appropriate for Galway Bay. The largest averaged Data Assimilation Skill Score (DASS) over the ≥6 h forecasting period from the best model NDA attained 26% and 31% for east–west and north–south surface velocity components respectively. Because of its ease of implementation and its accuracy, this data assimilation system can provide timely and useful information for various practical coastal hindcast and forecast operations.
Read moreA Mini‐Batch Stochastic Optimization‐Based Adaptive Localization Scheme and Its Implementation in NLS‐i4DVar
This paper proposes a mini‐batch stochastic optimization‐based adaptive localization scheme for computing the “optimal” localization radius in data assimilation (DA) applications. After constructing a cost function of the localization radius by estimating forecast and observation error statistics, a mini‐batch stochastic gradient descent method with a novel sampling strategy is proposed to minimize the cost function. The proposed stochastic optimization algorithm is further incorporated into the DA method NLS‐i4DVar (the nonlinear least squares integral correcting four‐dimensional variational DA method), which was developed by the authors in Tian et al. (2021, https://doi.org/10.1029/2021EA001767). It is utilized to compute the “optimal” covariance localization radii adaptively and flow‐dependently inside NLS‐i4DVar. The computational cost of NLS‐i4DVar with the proposed adaptive localization scheme only increases slightly due to the use of the mini‐batch stochastic optimization algorithm. Numerical experimental results using the shallow‐water equations demonstrate that NLS‐i4DVar with the proposed adaptive localization scheme shows substantial performance improvement over the standard NLS‐i4DVar method.
Read moreData assimilation using support vector machines and ensemble Kalman filter for multi-layer soil moisture prediction
Data assimilation using support vector machines and ensemble Kalman filter for multi-layer soil moisture prediction