• Home
  • Search
  • Using machine learning to select variables in data envelopment analysis: Simulations and application using electricity distribution data
  • Cite Icon27
  • https://doi.org/10.1016/j.eneco.2023.106621Copy DOI Icon

Using machine learning to select variables in data envelopment analysis: Simulations and application using electricity distribution data

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Agencies that regulate electricity providers often apply nonparametric data envelopment analysis (DEA) to assess the relative efficiency of each firm. The reliability and validity of DEA are contingent upon selecting relevant input variables. In the era of big (wide) data, the assumptions of traditional variable selection techniques are often violated due to challenges related to high-dimensional data and their standard empirical properties. Currently, regulators have access to a large number of potential input variables. Therefore, our aim is to introduce new machine learning methods for regulators of the energy market. We also propose a new two-step analytical approach where, in the first step, the machine learning-based adaptive least absolute shrinkage and selection operator (ALASSO) is used to select variables and, in the second step, selected variables are used in a DEA model. In contrast to previous research, we find, by using a more realistic data-generating process common for production functions (i.e., Cobb–Douglas and Translog), that the performance of different machine learning techniques differs substantially in different empirically relevant situations. Simulations also reveal that the ALASSO is superior to other machine learning and regression-based methods when the collinearity is low or moderate. However, in situations of multicollinearity, the LASSO approach exhibits the best performance. We also use real data from the Swedish electricity distribution market to illustrate the empirical relevance of selecting the most appropriate variable selection method.

Similar Papers
  • Research Article
  • Citations1

Ranking of Banks’ Risk Reporting Using Data Envelopment Analysis

  • Oct 01, 2021
  • Advances in Mathematical Finance and Applications
  • Azar Moslemi +2
  • Research Article
  • Citations201

Stepwise selection of variables in data envelopment analysis: Procedures and managerial perspectives

  • Jul 01, 2007
  • European Journal of Operational Research
  • Janet M Wagner +1
  • Book Chapter
  • Citations19

Oracle Efficient Estimation and Forecasting With the Adaptive Lasso and the Adaptive Group Lasso in Vector Autoregressions

  • Jun 26, 2014
  • Laurent A.F Callot +1
  • Research Article

Development and validation of a clinical risk-prediction model for immune checkpoint inhibitor-related pneumonitis in patients with gastrointestinal cancer based on four machine learning algorithms.

  • Dec 30, 2025
  • Chinese journal of cancer research = Chung-kuo yen cheng yen chiu
  • Yixuan Wang +4
  • Research Article
  • Citations55

The appreciative democratic voice of DEA: A case of faculty academic performance evaluation

  • Sep 18, 2013
  • Socio-Economic Planning Sciences
  • Muhittin Oral +3
  • PDF
  • Research Article
  • Citations32

Comparative Study on Variable Selection Approaches in Establishment of Remote Sensing Model for Forest Biomass Estimation

  • Jun 17, 2019
  • Remote Sensing
  • Xiaohui Yu +5
  • Research Article
  • Citations11

Meso-level Comparison of Mental Health Service Availability and Use in Chile and Spain

  • Apr 01, 2008
  • Psychiatric Services
  • L Salvador-Carulla +6
  • Research Article

Data driven approach for weight restricted data envelopment analysis models with single output

  • Dec 29, 2023
  • Journal of Turkish Operations Management
  • Şenol Kurt +2
  • PDF
  • Research Article
  • Citations3

Evaluating the Financial Efficiency of the Healthcare System: A Three-Stage DEA Model Analysis

  • Jan 01, 2023
  • R-Economy
  • Aida S Omir +2
  • Conference Article

Three-Dimensional Visualization for Multidimensional Analysis and Performance Management of Socio-Economic Systems

  • Jul 01, 2019
  • Alexander Afanasiev +3
  • Research Article
  • Citations4

Variable Selection via Regression Trees in the Presence of Irrelevant Variables

  • Jan 02, 2013
  • Communications in Statistics - Simulation and Computation
  • Youngjae Chang
  • Research Article
  • Citations26

Classifying 2-year recurrence in patients with dlbcl using clinical variables with imbalanced data and machine learning methods

  • Jun 09, 2020
  • Computer Methods and Programs in Biomedicine
  • Lei Wang +7
  • Conference Article

A slacks-based measure of super-efficiency in DEA with interval data

  • Jul 01, 2013
  • Xu Xiao-Ning +2
  • Research Article
  • Citations44

A survey on measuring efficiency through the determination of the least distance in data envelopment analysis

  • Dec 02, 2016
  • Journal of Centrum Cathedra
  • Juan Aparicio
  • Research Article
  • Citations103

Variable selection for multiply‐imputed data with application to dioxin exposure study

  • Mar 25, 2013
  • Statistics in Medicine
  • Qixuan Chen +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.