- Research Article
97
- 10.1016/j.knosys.2020.106598
Multi-class imbalanced big data classification on Spark
- Nov 07, 2020
- Knowledge-Based Systems
- William C Sleeman Iv + 1 more +1
Multi-class imbalanced big data classification on Spark
Currently, there are many problems in imbalanced big data classification based on rough set with virtual reality technology in cloud computing. For example, redundant big data cleaning is not clear, the effect is poor for big data denoising and feature extraction, and the precision of classification is low. In this paper, an imbalanced big data classification is proposed based on Hubness and K nearest neighbor to address such problems. First, the SNM algorithm is used in order to efficient cleaning of redundant big data. Then, wavelet threshold denoising algorithm is used to denoise the big data to improve the denoising effect. Meantime, feature of big data is extracted based on Lyapunov theorem. Moreover, the Hubness and K-nearest neighbor algorithms are used to achieve high precision of imbalanced big data classification. Experiments verify that the proposed method effectively strengthens current cleaning and denoising methods of redundant imbalanced big data, as well as improves accuracy of extraction and classification of big data.
Multi-class imbalanced big data classification on Spark
Multi-class imbalanced big data classification on Spark
Cloud computing and big data: Technologies and applications
Cloud computing and big data: Technologies and applications
Meta Learning for Imbalanced Big Data Analysis by using Generative Adversarial Networks
Imbalanced big data means big data where the ratio of a certain class is relatively small compared to other classes. When the machine learning model is trained by using imbalanced big data, the problem with performance drops for the minority class occurs. For this reason, various oversampling methodologies have been proposed, but simple oversampling leads to problem of the overfitting. In this paper, we propose a meta learning methodology for efficient analysis of imbalanced big data. The proposed meta learning methodology uses the meta information of the data generated by the generative model based on Generative Adversarial Networks. It prevents the generative model from becoming too similar to the real data in minority class. Compared to the simple oversampling methodology for analyzing imbalanced big data, it is less likely to cause overfitting. Experimental results show that the proposed method can efficiently analyze imbalanced big data.
Read moreBig data with cloud computing: Discussions and challenges
With the recent advancements in computer technologies, the amount of data available is increasing day by day. However, excessive amounts of data create great challenges for users. Meanwhile, cloud computing services provide a powerful environment to store large volumes of data. They eliminate various requirements, such as dedicated space and maintenance of expensive computer hardware and software. Handling big data is a time-consuming task that requires large computational clusters to ensure successful data storage and processing. In this work, the definition, classification, and characteristics of big data are discussed, along with various cloud services, such as Microsoft Azure, Google Cloud, Amazon Web Services, International Business Machine cloud, Hortonworks, and MapR. A comparative analysis of various cloud-based big data frameworks is also performed. Various research challenges are defined in terms of distributed database storage, data security, heterogeneity, and data visualization.
Read moreLegal Governance of Brain Data Derived from Artificial Intelligence
Photo by Josh Riemer on Unsplash
 Introduction
 With the rapid advancements in neurotechnological machinery and improved analytical insights from machine learning in neuroscience, the availability of big brain data has increased tremendously. Neurological health research is done using digitized brain data.[1] There must be adequate data governance to secure the privacy of subjects participating in brain research and treatments. If not properly regulated, the research methods could lead to significant breaches of the subject’s autonomy and privacy. This paper will address the necessity for neuroprotection laws, which effectively govern the use of big brain data to ensure respect for patient privacy and autonomy.
 Background
 Artificial intelligence and machine learning can be integrated with neuroscience big brain data to drive research studies. This integrative technology allows patterns of electrical activity in neurons to be studied in detail.[2]Specifically, it uses a robotic system which can reason, plan, and exhibit biologically intelligent behavior. Machine learning is a method of computer programming where the code can adapt its behavior based on big brain data.[3] The big brain data is the collection of large amounts of information for the purpose of deciphering patterns through computer analysis using machine learning.[4] The information that these technologies provide is extensive enough to allow a researcher to read a patient’s mind. AI and machine learning technologies work by finding the underlying structure of brain data, which is then described by patterns known as latent factors, eventually resulting in an understanding of the brain’s temporal dynamics.[5]
 Through these technologies, researchers are able to decipher how the human brain computes its performances and thoughts. However, due to the extensive and complex nature of the data processed through AI and machine learning, researchers may gain access to personal information a patient may not wish to reveal. From a bioethical lens, tensions arise in the realm of patient autonomy. Patients are not able to control the transmission of data from their brains that is analyzed by researchers. Governing brain data through laws may enhance the extent of patient privacy in the case where brain data is being used through AI technologies.[6] A responsible approach to governing brain data would require a sophisticated legal structure.
 Analysis
 Impact on Patient Autonomy and Privacy 
 In research pertaining to big brain data, the consent forms do not fully cover the vast amounts of information that is collected. According to research, personal data has become the most sought out commodity to provide content to corporations and the web-based service industry. Unfortunately, data leaks that release private information frequently occur.[7] The storage of an individual’s data on technologies accessible on the internet during research studies makes it vulnerable to leaks, jeopardizing an individual’s privacy. These data leaks may cause the patient to be identified easily, as the degree of information provided by AI technologies are personalized and may be decoded through brain fingerprinting methods.[8]
 There has been an extensive growth in the development and use of AI. It is efficient in providing information to radiologists who diagnose various diseases including brain cancer and psychiatric disease, and AI assists in the delivery of telemedicine.[9] However, the ethical pitfall of reduced patient autonomy must be addressed by analyzing current AI technologies and creating more options for patient preference in how the data may be used. For instance, facial recognition technology[10] commonly used in health care produces more information than listed in common consent forms, threatening to undermine informed consent. Facial recognition software collects extensive data and may disclose more information than a person would prefer to provide despite being a useful tool for diagnosing medical and genetic conditions.[11] In addition, people may not be aware that their images are being used to generate more clinical data for other purposes. It is difficult to guarantee the data is anonymized. Consent requirements must include informing people about the complexity of the potential uses of the data; software developers should maximize patient privacy.[12] Furthermore, there is a “human element” in the use of AI technologies as medical providers control the use and the extent to which data is captured or accessed through the AI technologies.[13] People must understand the scope of the technology and have clear communication with the physician or health care provider about how the medical information will be used. 
 Existing Laws for Brain Data Governance 
 A strict system of defined legal responsibilities of medical providers will ensure a higher degree of patient privacy and autonomy when AI technologies and data from machine learning are used. Governing specific algorithmic data is crucial in safeguarding a patient’s privacy and developing a gold standard treatment protocol following the procurement of the information.[14] Certain AI technologies provide more data than others, and legal boundaries should be established to ensure strong performance, quality control, and scope for patient privacy and autonomy. For instance, currently AI technologies are being used in the realm of intensive neurological care. However, there is a significant level of patient uncertainty about how much control patients have over the data’s uses.[15] Calibrated legal and ethical standards will allow important brain data to be securely governed and monitored.
 Once brain signals are recorded and processed from one individual, the data may be merged with other data in Brain Computer Interface Technology (BCI).[16] To ensure a right and ability to retrieve personal data or pull it from the collection, specific regulations for varying types of data are needed.[17] The importance of consent and patient privacy must be considered through giving patients a transparent view of how brain data is governed.[18] The legal system must address discriminatory issues and risks to patients whose data is used in studies. Laws like the General Data Protection Regulation (GDPR) and the California Consumer Privacy Protection Act (CCPA) can serve as effective models to protect aggregated data. These laws govern consumer information and ensure the compliance when personal data is collected.[19] California voters recently approved expansion of the CCPA to health data. The Washington Privacy Act, which would have provided rights to access, change, and withdraw personal data, failed to pass. Other states should improve privacy as well,[20] although a federal bill would be preferable. Scientists at the Heidelberg Academy of Sciences argue for data security to be governed in a manner that balances patient privacy and autonomy with the commercial interests of researchers.[21] The balance could be achieved through privacy protections like those in the Washington Privacy Act. Although the Health Insurance Portability and Accountability Act (HIPAA) provides an overall framework to deter the likelihood of dangers to patient protection and privacy, more thorough laws are warranted to combat pervasive data transfer and analysis that technology has brought to the health care industry.[22] Breaches of patient privacy under current HIPAA regulations include releasing patient information to a reporter without their consent and sending HIV data to a patient’s employer without consent.[23] HIPAA does not cover information being shared with outside contractors who do not have an agreement with technology companies to keep patient data confidential. HIPAA regulations also do not always address blatant breaches on patient data confidentiality.[24] Patients must be provided with methods to monitor the data being analyzed to be able to view the extent of private information being generated via AI technologies. In health research, the medical purposes of better diagnosis, earlier detection of diseases, or prevention are ethical justifications for the use of the data if it was collected with permission, the person understood and approved the uses of the data, and the data was deidentified.
 A standard governance framework is required in providing the fairest system of care to patients who allow their brain data to be examined. Informed consent in the neuroscience field could reaffirm the privacy and autonomy of patients by ensuring that they understand the type of information collected. Laws also could protect data after a patient’s death. Malpractice in the scope of brain data could give people a cause of action critical in safeguarding patient’s rights. Data breach lawsuits will become common but generally do not cover deidentified data that becomes part of big data collection. A more synchronized approach to the collection and consent process will encourage an understanding of how big data is used to diagnose and treat patients. Some altruistic people may even be more likely to consent if they know the largescale data collection is helpful to treat and diagnose people. Others should have the ability to opt out of sharing neurological data, especially when there is not certainty surrounding deidentification.[25]
 Conclusion
 Artificial intelligence and machine learning technologies have the potential to aid in the diagnosis and treatment of people globally by extracting and aggregating brain data specific to individuals. However, the secure use of the data is necessary to build trust between care providers and patients, as well as in balancing the bioethical principles of beneficence and patient autonomy. We must ensure the highest quality of care to patients, while protecting their privacy, informed consent, and clinical trust. More sophis
Read moreGrey Wolf Shuffled Shepherd Optimization Algorithm-Based Hybrid Deep Learning Classifier for Big Data Classification
In recent days, big data is a vital role in information knowledge analysis, predicting, and manipulating process. Moreover, big data is well-known for organized extraction and analysis of large or difficult databases. Furthermore, it is widely useful in data management as compared with the conventional data processing approach. The development in big data is highly increasing gradually, such that traditional software tools faced various issues during big data handling. However, data imbalance in huge databases is a main limitation in the research area. In this paper, the Grey wolf Shuffled Shepherd Optimization Algorithm (GWSSOA)-based Deep Recurrent Neural Network (DRNN) algorithm is devised to classify the big data. In this technique, for classifying the big data a hybrid classifier, termed as Holoentropy driven Correlative Naive Bayes classifier (HCNB) and DRNN classifier is introduced. In addition, the developed hybrid classification model utilizes the MapReduce structure to solve big data issues. Here, the training process of the DRNN classifier is employed using GWSSOA. However, the developed GWSSOA is devised by integrating Shuffled Shepherd Optimization Algorithm (SSOA) and Grey Wolf Optimizer (GWO) algorithms. The developed GWSSOA-based DRNN model outperforms other big data classification techniques with regards to accuracy, specificity, and sensitivity of 0.966, 0.964, 0.870, and 209837ms.
Read moreMulti-window based ensemble learning for classification of imbalanced streaming data
Imbalanced streaming data is commonly encountered in real-world data mining and machine learning applications, and has attracted much attention in recent years. Both imbalanced data and streaming data in practice are normally encountered together; however, little research work has been studied on the two types of data together. In this paper, we propose a multi-window based ensemble learning method for the classification of imbalanced streaming data. Three types of windows are defined to store the current batch of instances, the latest minority instances, and the ensemble classifier. The ensemble classifier consists of a set of latest sub-classifiers, and the instances employed to train each sub-classifier. All sub-classifiers are weighted prior to predicting the class labels of newly arriving instances, and new sub-classifiers are trained only when the precision is below a predefined threshold. Extensive experiments on synthetic datasets and real-world datasets demonstrate that the new approach can efficiently and effectively classify imbalanced streaming data, and generally outperforms existing approaches.
Read moreSelecting a suitable Cloud Computing technology deployment model for an academic institute
Purpose– Cloud Computing (CC) technology is getting implemented rapidly in the educational sector to improve learning, research and other administrative process. As evident from the literature review, most of these implementations are happening in the western countries such as USA, UK, while the level of implementation of CC in developing countries such as India is rare. Moreover, implementing CC technology in the educational sector require various decisions to be made by the managers of the Information Technology (IT) department such as selecting suitable deployment model, vendor providing cloud service, etc. in their respective university or institute. The purpose of this paper is to attempt to address one such decision. Since, different types of CC deployment are available; selecting a suitable one plays a key role, as it might have an impact on the requirements of various stakeholders such as students, teachers, administrative staff (especially the staff members in the IT department), etc. apart from affecting the overall performance of the facilities such as a laboratory. Naturally, a proper decision by analysing multiple perspectives has to be made while carrying out such strategic initiatives by any educational institute.Design/methodology/approach– A case study methodology has been chosen as the research methodology to discuss and demonstrate the above decision problem that was faced in real time by one of the educational institutes in India, offering high-quality management education. The IT managers of this institute were planning to switch over to CC technology for the computer laboratory and they have to make a decision of choosing suitable alternative CC deployment models such as private cloud (PRC), public cloud (PUC), community cloud (COC), hybrid cloud (HYC), etc. by analysing and comparing them based on various factors and perspectives such as elasticity, availability, scalability, etc. Since, multiple factors are involved in making such a strategic decision, the most commonly used Multi-Criteria Decision Making (MCDM) model – namely, the Analytic Hierarchy Process (AHP) is used as a decision support during the decision making process.Findings– The team of decision makers, who were planning to implement CC in the case institute, found that PRC is best as they believed that it would provide adequate cost savings, apart from providing necessary security to maintain confidential information such as student's detail, grades, etc.Research limitations/implications– The results obtained are based on a single case study. Hence, they cannot be generalized for institutions across educational sector. However, the decision making situation and understanding its impact on the stakeholders of the educational institute can be common across various educational institute.Practical implications– Using a real-life case study of an educational institute, this paper presented a strategic decision making situation, which needs to be considered by the IT managers of the educational institutes when they decide to switch over to CC technology. Various criteria to be considered during the decision making process was identified from the literature review were identified and enumerated. These factors would useful for the IT managers of the different educational institute and they can suitably add or delete these decision criteria as per their requirements and situation at hand. Moreover, the algorithm of AHP, which was used as a decision support, was presented in a step-by-step manner, which should be beneficial for the practitioners to apply the same for similar decision making situations.Originality/value– It is believed that this paper would be the first to report on a strategic decision of choosing the deployment model for CC technology especially in the educational sector. Similarly, this paper would also contribute to the field of CC, as it lists out the decision criteria that are to be considered for making the above decision, which has not got adequate importance. Lastly, this paper is also unique in the realm of AHP because application for a decision problem in the field of CC especially in the educational sector is least reported.
Read moreToward big data analysis to improve enterprise information security
In recent years, big data and cloud computing are considered key trends of modern computer technology. Extracting valuable information is the key purpose of analyzing big data that needs to be secured in order to avoid any potential risks. Most cloud systems applications contain sensitive data, such as; financial, legal and private information. Therefore, threats on such data may put cloud systems holding this data at high risk. The demand on securing cloud systems applications has been increasing rapidly; however, big data protection is still a challenge. This paper proposes a new methodology to protect big data during analysis by classifying data before any action such as moving, copying or processing take place. Big data files are classified according to the criticality level of their contents into three categories from the most to the least sensitive: restricted, confidential and public. Based on big data classification, the encryption algorithm AES 128 is applied on confidential big data, while the encryption algorithm AES 256 is applied on the restricted big data files. The experimental results show that our method enhances the performance of big data analysis systems and outperforms other approaches in the literature.
Read moreThe Effective Application of Cloud Computing Technology in Large Data User Behavior Engine Design
With the constant development of socialist modernization construction in our country, our country's computer information technology has made effective progress, and makes the world into an information age. People's production and life style experience a series of change. However, with the emergence of information diversity and multi-user mode, traditional computer information technology already can't satisfy people's needs, and the development and application of cloud computing technology come to the stage. This article focuses on analyzing the big data user behavior engine design under cloud computing technology, covering big data flow management, multi-user system design, and other aspects to explore the effect of the system test. In recent years, China's mobile Internet technology has made full development, which makes the Internet operators in our country face a new development opportunity and began from traffic management to flow management. The law of user behavior is analyzed to explore the real demand of market and the users. In order to fully meet the changing needs of users, operators must constantly develop and launch new products and strengthen the function of computer technology. Cloud computing technology is such kind of engine system which can satisfy the analysis and processing of huge amounts of data. RESEARCH ON CLOUD COMPUTING SYSTEM OVERALL DESIGN The overall architecture of cloud computing system This study mainly uses cloud computing technology’s huge amounts of data computing to set up complete mobile Internet data mining analysis system. Realize the analysis for Internet user behavior engine, and according to the user's preferences online habits and behavior, provide users with targeted personalized service and form a unified organic whole of data collection, analysis and service type and marketing strategy to improve enterprise's marketing efficiency. In addition, cloud computing system mainly realizes data collection with the help of FTP server, and makes distributed computing and data batch processing at the system interface. Large data shall be deposited in Hbase database. The system can not only realize mass data storage, but those unstructured data storage. Then through Hive integration layer and summary EIL processing, use MapReduce data analysis model to transmit the processing result into the database. The system’s overall structure is shown in Figure 1: 3rd International Conference on Materials Engineering, Manufacturing Technology and Control (ICMEMTC 2016) © 2016. The authors Published by Atlantis Press 1351 Figure 1 The overall architecture of cloud computing system System topology and function distribution System topology mainly refers that a server is taken as Hapdoop platform’s master node server, and the others are Hapdoop platform’s node servers. In general, node server can be dynamically extended according to the actual need, and the master node server not only assigns tasks and flow to node server, but also monitors the work execution of the node server. More than one node server participates to complete the task, and it can improve data processing efficiency. The master node server’s software running status is shown in Figure 2: Figure 2 Master node server function structure
Read moreResearch on Big Data Resource Management and Optimization Based on Cloud Computing Technology
This paper points out the problems of irrational, complex and low utilization rate of massive data storage in most industries at present, analyzes the meaning and characteristics of big data, thus determining the necessity of distribution and management of big data, providing reference ideas for the management of Chinese big data resources, and further promoting the big data resources to play its due role. To provide support for the transformation of science and technology management decision-making, to achieve social and technological progress. On this basis, from different aspects of the mass data management methods and strategies and the significance and efficiency of these methods to achieve, improve the utilization rate of cluster resources, increase the ability of data sharing, a variety of computing frameworks can also share a distributed storage data, finally through the experiment demonstrated the effectiveness and security of the above strategies.
Read moreAssessing big data analytics and characteristics in tourism: Agodi Gardens, Ibadan, Nigeria
There is a plethora of both organized and haphazard data in tourism destinations. Analyzing this data appropriately is crucial for optimal engagement. This study focuses on the connection between big data analytics and big data's characteristics in Agodi Gardens, Ibadan, Nigeria. Specific objectives were to examine the characteristics of big data; as well as to examine Descriptive and predictive data analytics. Respondents were chosen purposively. Survey instrument (questionnaire) was used to elicit data. Data was collected using structured questionnaire. The collected data were analyzed descriptively and inferentially. The study revealed that significant relationship exists between the prescriptive/descriptive big data analytics and the characteristics of big data. Precisely, there is a significant relationship between prescriptive data analytics and velocity, veracity, volume as well as value. Similarly, there is a significant relationship between descriptive data analytics and volume, variety, value as well as veracity. Likewise, variety and veracity of big data could influence big data analytics. The study therefore recommends that the management of Agodi Gardens should engage thorough big data analytics, so that data elicited by customers can be appropriately analysed and topical inference could be drawn from the analysis.
Read moreModeling of class imbalance handling with optimal deep learning enabled big data classification model
Big data is the amount of data that surpasses the ability to process the data of a system concerning memory usage and computation time. It is commonly applied in several domains like healthcare, education, social networks, e-commerce, etc., as they have progressively obtained a massive quantity of input data. A major research problem is big data analytics, which can be carried out using expert systems and deep structured architectures. Besides, data wrangling and class imbalance data handling are challenging issues that need to be resolved in big data analytics. Class imbalance data degrade the performance of the classification model, which remains a challenging process due to the heterogeneous and complex structure of the comparatively huge datasets. Thus, the research focused on presenting a Class Imbalance Handling with Optimal Deep Learning Enabled Big Data Classification (CIHODL-BDC) framework. The core perception of the CIHODL-BDC framework helps to classify the big data in the Hadoop MapReduce framework. To accomplish this, the presented CIHODL-BDC model initially performs a data wrangling process is performed to alter the unrefined data into a useful layout. Next, the CIHODL-BDC model handles the class imbalance problem using a grey wolf optimizer (GWO) with Synthetic Minority Oversampling (SMOTE) technique. Besides, the Adam optimizer procedure with the Bidirectional Long Short Term Memory (BiLSTM) approach is performed to categorize the big data. The result analysis of the proposed CIHODL-BDC model is evaluated by two standard datasets. The simulation outcomes revealed the elevated performance of the CIHODL-BDC approach over existing methods.
Read moreVideo Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues
The proliferation of multimedia devices over the Internet of Things (IoT) generates an unprecedented amount of data. Consequently, the world has stepped into the era of big data. Recently, on the rise of distributed computing technologies, video big data analytics in the cloud has attracted the attention of researchers and practitioners. The current technology and market trends demand an efficient framework for video big data analytics. However, the current work is too limited to provide a complete survey of recent research work on video big data analytics in the cloud, including the management and analysis of a large amount of video data, the challenges, opportunities, and promising research directions. To serve this purpose, we present this study, which conducts a broad overview of the state-of-the-art literature on video big data analytics in the cloud. It also aims to bridge the gap among large-scale video analytics challenges, big data solutions, and cloud computing. In this study, we clarify the basic nomenclatures that govern the video analytics domain and the characteristics of video big data while establishing its relationship with cloud computing. We propose a service-oriented layered reference architecture for intelligent video big data analytics in the cloud. Then, a comprehensive and keen review has been conducted to examine cutting-edge research trends in video big data analytics. Finally, we identify and articulate several open research issues and challenges, which have been raised by the deployment of big data technologies in the cloud for video big data analytics. To the best of our knowledge, this is the first study that presents the generalized view of the video big data analytics in the cloud. This paper provides the research studies and technologies advancing video analyses in the era of big data and cloud computing.
Read moreAllocation of Resources after Disaster Based on Big Data from SNS and Spatial Scan
After a disaster such as earthquakes, debris flows, forest fires, or landslides, etc., a lot of people have to be away from their home and gather in a shelter. In addition, the refugees suffer from the shortage of necessary resources due to impaired life infrastructure, such as damaged roads and communication networks. The degree of reducing damage depends on the amount of food, water, daily necessities and communication resources required by each shelter. How to effectively and efficiently allocate resources according to grasp the exact need of a disaster situation will be an important issue. We estimate the degree of the disaster by collecting and analyzing big data from the SNS, and building a platform for the communication resources to be efficiently and effectively allocated. In order to achieve this goal, we are challenging the following issues A) Understanding situations (user requirements) after disaster occur The SNS streams large scale semantic information about real time situation in society, especially during and after disaster. It is both domain-specific and computational challenge in processing the heterogeneous large data set to extract the exact situational content with reduced semantic uncertainty. The machine learning (ML) and natural language processing (NPL) tool kits are useful in semantic analysis, but still needs domain-specific implementation and computational improvement for the situation understanding from the SNS big data. B) Understanding distribution patterns of situations/users' requirements The disaster related situation is spatiotemporally correlated, and varies dynamically in space and time. It is also domain-specific and computational challenge in estimating the spatiotemporal distribution patterns of the disaster affect based on the spatial big data from SNS. The scan statistics such as the spatial scan have provided well tested mathematical tools and software for spatial data mining. However, new methodologies are necessary since the assumptions have to be different when it meets the spatial big data in SNS. And the computational complexity in spatial big data is also a bottleneck for real-time processing. C) Solving uncertainty of big crowd data One of the major features in big crowd data, e.g., SNS data, is uncertainty behind the data. Especially in a disaster scenario, the collecting time period cannot be long enough to smooth the data automatically. How to efficiently solve uncertainty problem in the big crowd data in a disaster scenario becomes a new and big challenge for disaster management.
Read more