- Research Article
2
- 10.2139/ssrn.3236837
Machine Learning and Transparency: A Scoping Exercise
- Aug 31, 2018
- SSRN Electronic Journal
- Vidushi Marda
Machine Learning and Transparency: A Scoping Exercise
The fact that machine learning is growing more and more entrenched in almost every aspect of society, combined with the opacity of various of its algorithms has induced the relatively young research area of transparent machine learning. The aim of this domain is to provide explanations for automated decisions to increase public trust. In my thesis, I am going to consider certain problems that arise from this research agenda. Particularly, I so far considered, the dilemma of conflicting explanations and the issue of privacy concerns arising from transparency.
Machine Learning and Transparency: A Scoping Exercise
Machine Learning and Transparency: A Scoping Exercise
Explainability and transparency of machine learning in ADM systems
Artificially Intelligent (AI) systems have become part of society. They support decision-making processes and take decisions autonomously. Due to their ability to analyze very large data sets, they can be superior to humans in certain aspects and they will have a major impact on people's lives. Tensions with societal, legal norms or constitutional values have already arisen. They can be observed in different scenarios, ranging from credit scoring and autonomous driving to social profiling and predictive policing. The lack of transparency of AI systems, which were considered as “black-boxes,” has already been identified as a problem. These so-called “black boxes” have been a weak spot for those who develop or use an AI system. This paper demystifies AI by discussing the role of transparency and explainability of machine learning. The paper describes different aspects of transparency and explains methods to increase our understanding of the behavior of AI systems.
Read moreIn Defense of Sociotechnical Pragmatism
The current discourse on fairness, accountability, and transparency in machine learning is driven by two competing narratives: sociotechnical dogmatism, which holds that society is full of inefficiencies and imperfections that can only be solved by better algorithms; and sociotechnical skepticism, which opposes many instances of automation on principle. Both perspectives, we argue, are reductive and unhelpful. In this chapter, we review a large, diverse body of literature in an attempt to move beyond this restrictive duality, toward a pragmatic synthesis that emphasizes the central role of context and agency in evaluating new and emerging technologies. We show how epistemological and ethical considerations are inextricably intertwined in contemporary debates on algorithmic bias and explainability. We trace the dialectical interplay between dogmatic and skeptical narratives across disciplines, merging insights from social theory and philosophy. We review a number of theories of explanation, ultimately endorsing a sociotechnical pragmatism that combines elements of Floridi’s levelism and Mayo’s reliabilism to place a special emphasis on notions of agency and trust. We conclude that this hybrid does more to promote fairness, accountability, and transparency in machine learning than dogmatic or skeptical alternatives.
Read moreInterpretable Machine Learning: A Case Study on Predicting Fuel Consumption in VLGC Ship Propulsion
The integration of machine learning (ML) in marine engineering has been increasingly subjected to stringent regulatory scrutiny. While environmental regulations aim to reduce harmful emissions and energy consumption, there is also a growing demand for the interpretability of ML models to ensure their reliability and adherence to safety standards. This research highlights the need to develop models that are both transparent and comprehensible to domain experts and regulatory bodies. This paper underscores the importance of transparency in machine learning through a use case involving a VLGC ship two-stroke propulsion engine. By adhering to the CRISP-DM standard, we fostered close collaboration between marine engineers and machine learning experts to circumvent the common pitfalls of automated ML. The methodology included comprehensive data exploration, cleaning, and verification, followed by feature selection and training of linear regression and decision tree models that are not only transparent but also highly interpretable. The linear model achieved an RMSE of 23.16 and an MRAE of 14.7%, while the accuracy of decision trees ranged between 96.4% and 97.69%. This study demonstrates that machine learning models for predicting propulsion engine fuel consumption can be interpretable, adhering to regulatory requirements, while still achieving adequate predictive performance.
Read moreEnhancing Transparency in Healthcare Machine Learning Models Using Shap and Deeplift a Methodological Approach
This paper intends to provide a better understanding of how these models produce predictions, particularly in complex medical diagnoses, and at the same time bridging the gap between technical model outputs and clinical applications. This study addresses the critical problem of transparency of machine learning (ML) models in health care where interpretability is an essential aspect for ethical decision making and trust building.The goal of this paper is to present a clearer understanding of how these models generate predictions, especially in important fields like complex medical diagnoses, thereby bridging the gap between technical model outputs and clinical applications. The crucial issue of transparency in machine learning (ML) models within healthcare is addressed in this study where interpretability plays a vital role in ethical decision-making and fostering trust. The focus of the research is enhancing model transparency by using SHapley Additive exPlanations (SHAP) and Deep Learning Important FeaTures (DeepLIFT), two crucial methods that are designed to elucidate the decision-making processes of ML models.This mode of approach helps to have more distributed comprehension of the decision pathways by models thus aiding in knowing how each feature contributed to the last prediction. It is this method that has been employed to showcase the efficiency of predicting melanoma and also diabetic retinopathy which are two vital medical diagnostic areas. In healthcare, SHAP along with DeepLIFT has improved the models’ explainability and trustworthiness significantly and hence making them easy for those in the field. The advanced interpretability methods presented in this document enhances ML model transparency especially when dealing with health issues. As a result, interpretability becomes an even bigger issue and they are supposed to be able to use these tools for reliable and open decisions when it comes to medical specialists.
Read moreRegulation of AI and AI Crimes
Regulation of AI and AI Crimes
ARTIFICIAL INTELLIGENCE AND THE ETHICS OF REPRESENTING NIGERIAN CULTURE IN THE DIGITAL AGE
The convergence of AI and cultural representation raises critical ethical concerns in the digital age, particularly in culturally diverse and complex societies like Nigeria. This paper takes a conceptual methodological stance to explore the intersection of artificial intelligence (AI) technologies, such as language models and translation systems, in relation to Nigerian cultural and linguistic realities by highlighting the power dynamics of digital language preservation and the risk of reinforcing stereotypes or erasing nuance. Using a bipartite framework involving Floridi’s Information Ethics and the Fairness, Accountability, and Transparency in Machine Learning (FAT/ML) model, this study examines how artificial intelligence (AI) systems, including ChatGPT and Google Translate, may misrepresent or distort Nigerian languages and cultures, such as Yoruba, Hausa, Igbo, and Pidgin. The study advocates for the development of AI that is culturally aware and ethically grounded to promote inclusive frameworks that prioritize linguistic equity and cultural sustainability in the digital representation of Nigerian identities. This research contributes to the development of responsible AI practices that respect and promote cultural diversity in the digital age by exploring the global relevance of these issues.
Read moreA comprehensive review on detection of plant disease using machine learning and deep learning approaches
A comprehensive review on detection of plant disease using machine learning and deep learning approaches
Overview of deep learning in medical imaging.
The use of machine learning (ML) has been increasing rapidly in the medical imaging field, including computer-aided diagnosis (CAD), radiomics, and medical image analysis. Recently, an ML area called deep learning emerged in the computer vision field and became very popular in many fields. It started from an event in late 2012, when a deep-learning approach based on a convolutional neural network (CNN) won an overwhelming victory in the best-known worldwide computer vision competition, ImageNet Classification. Since then, researchers in virtually all fields, including medical imaging, have started actively participating in the explosively growing field of deep learning. In this paper, the area of deep learning in medical imaging is overviewed, including (1) what was changed in machine learning before and after the introduction of deep learning, (2) what is the source of the power of deep learning, (3) two major deep-learning models: a massive-training artificial neural network (MTANN) and a convolutional neural network (CNN), (4) similarities and differences between the two models, and (5) their applications to medical imaging. This review shows that ML with feature input (or feature-based ML) was dominant before the introduction of deep learning, and that the major and essential difference between ML before and after deep learning is the learning of image data directly without object segmentation or feature extraction; thus, it is the source of the power of deep learning, although the depth of the model is an important attribute. The class of ML with image input (or image-based ML) including deep learning has a long history, but recently gained popularity due to the use of the new terminology, deep learning. There are two major models in this class of ML in medical imaging, MTANN and CNN, which have similarities as well as several differences. In our experience, MTANNs were substantially more efficient in their development, had a higher performance, and required a lesser number of training cases than did CNNs. "Deep learning", or ML with image input, in medical imaging is an explosively growing, promising field. It is expected that ML with image input will be the mainstream area in the field of medical imaging in the next few decades.
Read moreA Next-Generation of Biomonitoring to Detect Global Ecosystem Change
There is growing interest in the potential for combining eDNA and artificial intelligence (machine learning) to detect and evaluate in real time changes in ecosystems at the global scale, in a more sensitive and cost-effective way than current biomonitoring methods. Machine learning might make better use of the eDNA census data that can currently be collected to evaluate the network of ecological interactions that are at the base of the services that ecosystems supply and that we wish to protect. To date, eDNA and machine learning developments have effectively progressed in parallel and in isolation in various spheres of ecosystem monitoring (disease, invasion, conservation, etc. in aerial, terrestrial, and aquatic systems). The goal of this Research Topic is to explore the range of ongoing activities to build the next generation of biomonitoring tools and in doing so to make researchers in the different spheres aware of the breadth of work being undertaken, and to set a unifying research agenda (the key questions) for the development of global biomonitoring using eDNA and machine learning. The scope of this Research Topic will be to explore: 1. eDNA approaches currently being used in case study systems from all spheres of monitoring; 2. Theoretical underpinnings of machine learning for biomonitoring; 3. What type of networks do we need to reconstruct for effective monitoring (co-occurrence, trophic, etc); 4. Examples of learning large scale, replicated networks from eDNA in the different spheres; 5. Statistical and analytical approaches to analysing large-scale, highly replicated networks; 6. Technological developments necessary to build a next-generation biomonitoring framework at the global scale; 7. A research agenda paper that develops “10 key questions for eDNA and machine learning in biomonitoring”. Details for Authors: The Research Topic “A next-generation of global biomonitoring to detect ecosystem change” will publish conceptual, data, case study, technological and synthetic papers on eDNA and machine learning approaches for developing a unified next-generation biomonitoring framework. Paper length conforms to the guidelines of the journal Frontiers in Ecology and Evolution.
Read moreHRD in SMEs: A research agenda whose time has come
As can be seen from its website, and reiterated in numerous editorials (e.g., Anderson, 2017; Nimon, 2017; Reio & Werner, 2017), Human Resource Development Quarterly (HRDQ) provides a central focus on human resource development (HRD) issues as well as the means for disseminating empirical research across the breadth of the discipline. Furthermore, the listing of keywords on its website indicates the importance HRDQ places on knowing more about learning in workplace settings as it includes words and phrases such as workplace issues, workplace learning, organizational studies, and workplace performance. This is in line with general increased interest in organizational learning in recent years (Higgins & Aspinall, 2011). Therefore, it is concerning that HRDQ seldom reports on an area of workplace learning in a sector that, in many countries throughout the world, encompasses approximately 99% of all businesses, provides over 50% of employment, and can generate around 50% of national turnover (Chartered Institute for Personnel & Development [CIPD], 2015; Coetzer & Perry, 2008; European Commission, 2016; Federation of Small Businesses, 2015; Hamburg, Engert, Anke, & Marin, 2008; Matlay, 2014; Mellett & O'Brien, 2014; U.K. Parliament, 2014; U.S. Census Bureau, 2012). If you have not yet guessed, this area of learning, which is vital to economies across the globe, occurs in small and medium‐sized enterprises (SMEs). Consequently, in this editorial, we seek to explore the extent of this omission, not only in HRDQ but also in other journals, and then investigate possible reasons for this. We hope that by emphasizing both the importance of and the lack of reported research into HRD in SMEs, we will encourage further dialogue and submissions related to this important topic.
Read moreChallenges in AutoML and Declarative Studies Using Systematic Literature Review
Machine Learning (ML) technologies have become essential tools, transforming industries and unlocking incredible potential in various fields. ML is now widely used for data-driven decision-making and predictive analytics across fields like healthcare, finance, transportation, and more. However, building and implementing ML models can be complex and time-consuming, often requiring programming proficiency and data science skills. Despite significant progress in ML, non-experts often struggle with selecting algorithms, optimizing models, and deploying ML solutions. This paper conducts a systematic literature review to explore challenges in the area of machine learning based on multiple categories involving features engineering and data extraction, learning model structure and activities, learning-based analysis and visualization, analysis algorithms in data-based systems, machine learning algorithms and systems development, and declarative ML-based prediction. Addressing these challenges underlines the importance of following AutoML and Declarative ML strategies in simplifying the ML process.
Read moreAdvanced Machine Learning Based Malware Detection Systems
In the area of machine learning (ML) training data optimization through the construction of compact data, the focus of this paper is presented. The concept of compact data design, aimed at creating an optimized dataset that maximizes benefits without the need to manage a vast amount of complex data, is introduced. Improvements in the methods for optimizing ML training have been incorporated into the development of artificial intelligence (AI) systems. The introduction of understanding ML training datasets as a facet of Explainable AI (XAI), comprehensible to humans, has been made. Among the methods of XAI, the evaluation of input feature importance stands out as a way to enhance the accuracy of complex ML models. The innovative method of compact data design for optimizing ML training through dataset reduction is proposed. The performance of an ML-based malware detection system, along with its variant utilizing compact data, has been assessed, demonstrating the maintenance of 99% accuracy. By applying a 76% reduced input dataset, the speed ofMLtraining with the novel compact data design could be maximized, suggesting that anMLsystem trained in this manner could achieve statistically equivalent accuracy with only 57% of the original data sample size.
Read moreTransfer learning for AiTR: comparing deep learning to other machine learning approaches
Aided target recognition (AiTR), the problem of classifying objects from sensor data, is an important problem with applications across industry and defense. While classification algorithms continue to improve, they often require more training data than is available or they do not transfer well to settings not represented in the training set. These problems are mitigated by transfer learning (TL), where knowledge gained in a well-understood source domain is transferred to a target domain of interest. In this context, the target domain could represents a poorly-labeled dataset, a different sensor, or an altogether new set of classes to identify. While TL for classification has been an active area of machine learning (ML) research for decades, transfer learning within a deep learning framework remains a relatively new area of research. Although deep learning (DL) provides exceptional modeling flexibility and accuracy on recent real world problems, open questions remain regarding how much transfer benefit is gained by using DL versus other ML architectures. Our goal is to address this shortcoming by comparing transfer learning within a DL framework to other ML approaches across transfer tasks and datasets. Our main contributions are: 1) an empirical analysis of DL and ML algorithms on several transfer tasks and domains including gene expressions and satellite imagery, and 2) a discussion of the limitations and assumptions of TL for aided target recognition -- both for DL and ML in general. We close with a discussion of future directions for DL transfer.
Read moreNew Real Time Clinical Decision System Using Machine Learning – Disease Prediction
Deep learning can be known as an area of machine learning which is useful to reorganize patient care. It is not yet part of standard care, especially when it comes to individual patient care. It is not clear to what extent various data-driven techniques are being used to support clinical decision making systems (CDS). Therefore, there has not been a review of ways in which research in machine learning and other types of data-driven techniques can contribute effectively to clinical care. There are many different parts in clinical decision making. Few of them are disease predictions, online patient tracking; artifact detection, state estimation etc. There are many diseases present now-a-days. Every disease has its own symptoms. Our project is to find the disease based on the symptoms. The disease prediction is made using machine learning techniques.
Read more