- Research Article
- 10.1016/j.atech.2026.101996
Developing a robust yield prediction model for bean cultivars (Fabaceae) in the delmarva region using multi-faceted and multi-year data
- Aug 01, 2026
- Smart Agricultural Technology
- Alfadhl Y Alkhaled + 2 more +2
• Multi-year ML framework improves yield prediction under high climatic variability. • Integrating agronomic and environmental data outperforms single-factor models. • Random Forest and XGBoost capture nonlinear yield responses across bean cultivars. • Heat stress and reproductive traits emerge as key yield-driving predictors. • Framework supports climate-smart, cultivar-specific decision making for growers. Accurate yield prediction remains a major challenge for legume production in environmentally variable regions such as the USA Delmarva Peninsula, where sandy soils and fluctuating climatic conditions intensify crop sensitivity to heat and moisture stresses. To address this challenge, this study developed a multi-year machine learning framework to predict grain yield for four bean species, mung bean, pigeon pea, kidney bean, and cowpea, represented by a total of 11 cultivars evaluated across eight growing seasons. The framework integrated four categories of predictors: Genotype (G) (cultivars), agronomic (A) traits (plant height, pods and seeds metrics, plant biomass, and yield), management (M) practices (trial structure and replications), and environmental (E) factors (temperature, heat stress index: HSI, growing degree days: GDD, and precipitation). Two nonlinear algorithms, Random Forest (RF) and eXtreme Gradient Boosting (XGBoost), were evaluated under single-input, two-input, three-input, and full multi-component configurations to assess the relative predictive contributions of G, A, M, and E. Mean annual grain yield varied substantially across seasons, ranging from 93.5 kg/ha in 2014 to 6,279.9 kg/ha in 2023, reflecting pronounced interannual climatic variability and environmental heterogeneity in the Delmarva region. Both models achieved their highest predictive accuracy in years with more consistent E patterns and narrower yield distributions (coefficient of determination: R² = 0.98–0.87 in 2017–2019), whereas predictive performance declined in seasons characterized by greater climatic variability and more dispersed yield values. Among single-input models, A traits showed the strongest predictive ability ( R² ≈ 0.53), followed by E variables ( R² ≤ 0.49), while G-only and M-only inputs contributed minimally. Across multi-input configurations, the E+A combination delivered the highest overall two-input performance (RF R² = 0.69; XGBoost R² = 0.71), and the three-input models that combined A and E with G or M also showed strong predictive ability ( R² ≈ 0.70–0.73), while the full G+E+A+M model achieved the strongest combined-year accuracy (RF R² = 0.73; XGBoost R² = 0.76). These results demonstrate that integrating agronomic structure with environmental variability substantially improves multi-year yield prediction, and support cultivar-specific recommendations and climate-informed management strategies for Delmarva and other coastal agroecosystems.
Read more