- Research Article
7
- 10.1002/cpe.4517
Cloud computing and big data: Technologies and applications
- May 20, 2018
- Concurrency and Computation: Practice and Experience
- Mostapha Zbakh + 3 more +3
Cloud computing and big data: Technologies and applications
With the recent advancements in computer technologies, the amount of data available is increasing day by day. However, excessive amounts of data create great challenges for users. Meanwhile, cloud computing services provide a powerful environment to store large volumes of data. They eliminate various requirements, such as dedicated space and maintenance of expensive computer hardware and software. Handling big data is a time-consuming task that requires large computational clusters to ensure successful data storage and processing. In this work, the definition, classification, and characteristics of big data are discussed, along with various cloud services, such as Microsoft Azure, Google Cloud, Amazon Web Services, International Business Machine cloud, Hortonworks, and MapR. A comparative analysis of various cloud-based big data frameworks is also performed. Various research challenges are defined in terms of distributed database storage, data security, heterogeneity, and data visualization.
Loading PDF
Cloud computing and big data: Technologies and applications
Cloud computing and big data: Technologies and applications
Imbalanced big data classification based on virtual reality in cloud computing
Currently, there are many problems in imbalanced big data classification based on rough set with virtual reality technology in cloud computing. For example, redundant big data cleaning is not clear, the effect is poor for big data denoising and feature extraction, and the precision of classification is low. In this paper, an imbalanced big data classification is proposed based on Hubness and K nearest neighbor to address such problems. First, the SNM algorithm is used in order to efficient cleaning of redundant big data. Then, wavelet threshold denoising algorithm is used to denoise the big data to improve the denoising effect. Meantime, feature of big data is extracted based on Lyapunov theorem. Moreover, the Hubness and K-nearest neighbor algorithms are used to achieve high precision of imbalanced big data classification. Experiments verify that the proposed method effectively strengthens current cleaning and denoising methods of redundant imbalanced big data, as well as improves accuracy of extraction and classification of big data.
Read moreSeveral Typical Paradigms of Industrial Big data Application
Industrial big data is an important part of big data family, which has important application value for industrial production scheduling, risk perception, state identification, safety monitoring and quality control, etc. Due to the particularity of the industrial field, some concepts in the existing big data research field are unable to reflect accurately the characteristics of industrial big data, such as what is industrial big data, how to measure industrial big data, how to apply industrial big data, and so on. In order to overcome the limitation that the existing definition of big data is not suitable for industrial big data, this paper intuitively proposes the concept of big data cloud and the 3M (Multi-source, Multi-dimension, Multi-span in time) definition of cloud-based big data. Based on big data cloud and 3M definition, three typical paradigms of industrial big data applications are built, including the fusion calculation paradigm, the model correction paradigm and the information compensation paradigm. These results are helpful for grasping systematically the methods and approaches of industrial big data applications.
Read moreExtreme learning with projection relational algebraic secured data transmission for big cloud data
SummaryCloud Computing (CC) and big data are growing technology in the business. Big data is demonstrated in terms of volume, variety, and velocity. CC is employed for storing, processing, and accessing data. Many cryptographic techniques have been developed to enhance big data security in cloud computing. However, security and privacy are the primary concerns in protecting data, as it is highly sensitive. Yet, it faces the major problems of inefficient performance, increased time consumption, and lack of data confidentiality and integrity. To address this issue, proposed Extreme Learning with Projection Relational Algebraic Secured Data Transmission (ELPRA‐SDT) is introduced to secure data transactions from cloud users to cloud servers with enhanced data confidentiality and reduced time consumption for big cloud data. The proposed ELPRA‐SDT consists of two major processes namely registration and key generation. At first, the user's IP address is registered employing a transitive advanced set relation theory graph model in a cloud server (CS) for retrieving the numerous services. The CS generates private and public keys for each registered user's IP address using the Transitive Operational and Time Synchronized Random Winternitz Key generation model. After, the user sends a request to the CS for acquiring data. The CS validates the requested user based on security policy attributes. Second, the Projection Relational Algebraic Signcryption and Unsigncryption algorithm performs signature verification to ensure secure data access for protecting the data. Results of experiments carried out by using Coburg Intrusion Detection Data Sets‐001 dataset in Java. ELPRA‐SDT method is more efficient and more suitable for providing security and privacy to network traces in the Cloud. The result shows maximum performance with data confidentiality by 10% and data integrity by 13%. In addition, delay is reduced by 32%, and data delivery time and communication complexity is decreased by 28% and 24% to other existing methods.
Read moreCloud Computing for Big Data Analytics: Scalable Solutions for Data-Intensive Applications
The explosion of data in the digital era has posed major challenges handling, computing and analyzing enormous and complex datasets. Cloud computing has arisen as a revolutionary solution providing scalable and elastic infrastructure necessary to deal with incoming big data workloads. This study employs an empirical approach to evaluate the performance, cost efficiency, and scalability of the three dominant cloud service models, Infrastructure-as-a-Service (IaaS), Platform-as-a-Service (PaaS) and Function-as-a-Service (FaaS) on Amazon Web Services, Microsoft Azure and Google Cloud Platform. Standard big data analytics workloads running real-time stream processing and machine learning activities were then implemented in Apache Spark, Hadoop, and Kafka on harmonized cloud environments. Key performance metrics such as execution time, CPU utilization, memory, cost per task and throughput were taken, analyzed statistically using ANOVA and Tukey’s post hoc tests. Results show that FaaS configurations are always faster in execution speed, memory efficiency and cost compared to IaaS, while IaaS delivers better CPU usage for continual workloads. AWS and GCP platform performed relatively balanced when compared to Azure. It is concluded that serverless architecture is, in fact, optimal for modular and burst-oriented analytics, and hybrid models might be more appropriate for complex pipelines. These results can offer cloud architects practical directions towards scalable and cost-effective big data solutions.
Read moreAPPLICATION AND PLATFORM DESIGN OF GEOSPATIAL BIG DATA
Abstract. With the wide application of Big Data, Artificial Intelligence and Internet of Things in geographic information technology and industry, geospatial big data arises at the historic moment. In addition to the traditional "5V" characteristics of big data, which are Volume, Velocity, Variety, Veracity and Valuable, geospatial big data also has the characteristics of "Location Attribute". At present, the study of geospatial big data are mainly concentrated in: knowledge mining and discovery of geospatial data, Spatiotemporal big data mining, the impact of geospatial big data on visualization, social perception and smart city, geospatial big data services for government decision-making support four aspects. Based on the connotation and extension of geospatial big data, this paper comprehensively defines geospatial big data comprehensively. The application of geospatial big data in location visualization, industrial thematic geographic information comprehensive service and geographic data science and knowledge service is introduced in detail. Furthermore, the key technologies and design indicators of the National Geospatial Big Data Platform are elaborated from the perspectives of infrastructure, functional requirements and non-functional requirements, and the design and application of the National Geospatial Public Service Big Data Platform are illustrated. The challenges and opportunities of geospatial big data are discussed from the perspectives of open resource sharing, management decision support and data security. Finally, the development trend and direction of geospatial big data are summarized and prospected, so as to build a high-quality geospatial big data platform and play a greater role in social public application services and administrative management decision-making.
Read moreDesign and Implement Network Load Balancer with Multi Availability Zone using Amazon Web Server
Cloud computing is the on-demand delivery of IT resources over the Internet with pay-as-you-go pricing. Instead of buying, owning, and maintaining physical data centers and servers, you can access technology services, such as computing power, storage, and databases, on an as-needed basis from a cloud provider like Amazon Web Services (AWS), Microsoft Azure, Google cloud. Cloud computing is one of the most growing technologies. The fundamental idea behind cloud computing is to distribute an array of computing services by unifying and scheduling a pool of computing resources, thereby minimizing the burden on the users and helping them focus on their core businesses. These computing resources are hosted on virtual hosts and distributed on-demand to the users by cloud service providers. For efficient resource utilization, systematic load balancing of incoming user traffic across virtual hosts is imperative. Aim of this paper is to design and implement Network load balancer with cross zone feature which balance incoming user requests and avoid the problem of server down in real scenario. Elastic Load Balancer automatically equally distribute incoming traffic to the available EC2, targets, like Amazon EC2 instances, IP addresses containersand Lambda functions. Load balancer effectively distribute load equally on available server and AZ to improve fault tolerance and availability. It can manage the differing load of your application traffic in one Availability Zone or over different Availability Zones.
Read moreSurvey on Big Data Analytics in Health Care
Massive amount of data in different forms need to be handled in any healthcare applications. Type of data, size of data, data security and other features has more significance in handling the data. The term big data refers to data with certain characteristics, volume, velocity, value, veracity and variability. Such big data need to be stored, processed, and analyzed for required results. Medical data has more complexity in predicting the results from it, which will have more significance in patient's treatment. Because of its significance, there is need of developing efficient and better performing algorithms, techniques and tools to analyze medical big data. Whereas, the traditional algorithms are not capable for analyzing such complex data. Machine learning algorithms well fit for these kinds of data and analytics. In this Keywords: Big data, Health care, disease prediction, SVM, CNN survey paper, we discussed about characteristic of big data, features of big data, how to represent big data, different types of machine learning algorithms used in big data analytics. We discussed about big data analytics in major healthcare areas like EHR maintenance, disease diagnose, prediction of emergency condition of patients, etc.,. Also stated different machine algorithms usage in disease diagnose and patient's data analysis and discussed about importance of various machine learning algorithms. Here, we have highlighted the areas where big data analytics have been applied in healthcare sectors. It describes the characteristics and features of big data, importance of big data analytics in healthcare sectors, various machine learning algorithms used in big data analytics and their efficiency.
Read moreArtificial Intelligence joining forces with cloud computing: Pros and pitfalls
Artificial intelligence (AI) joining hands with cloud computing is shaking things up in tech changing the game for companies and changing the way they do business. AI is super good at handling big piles of data making tricky jobs run on their own, and figuring out what could happen next. It's getting a helping hand from cloud computing, which brings the kind of big-time infrastructure and brainpower needed for AI's heavy lifting. When you put them together, it's like a turbo boost for coming up with new ideas making work smoother, and helping businesses find all sorts of extra value. AI's not just something that chills in the cloud these days—it's shaping how cloud structures grow and change. Big shots like Amazon Web Services (AWS), Microsoft Azure, and Google Cloud, they're slipping AI tricks into their setups. Think stuff like smart scaling that predicts what you need, computers managing their own resources, and tight security 'cause of AI looking out for threats Plus, there's this new gear made just for AI—like GPUs and TPUs—that's switching up how much it costs to do the cloud thing by making AI tasks zip along faster
Read moreComparison of the IoT Platform Vendors, Microsoft Azure, Amazon Web Services, and Google Cloud, from Users’ Perspectives
The largest Internet of Things (IoT) cloud platform vendors are Microsoft Azure, Amazon Web Services, and Google Cloud. These companies are known as the big three, and have all agreed to join the IoT domain and concentrate on improving the services on their IoT platforms. While these platform descriptions are extensive, users are constantly experiencing difficulties in making the right choice of which platform to use, between the three. This paper presents a comparison of the big three, using the constraints of hubs, analytics, and security. The study also provides some recommendations as to which IoT cloud platform vendor is more ideal, notwithstanding the limitations of the study. In view of these results, users will be able to more confidently select vendors, based on their demands and goals.
Read moreLegal Governance of Brain Data Derived from Artificial Intelligence
Photo by Josh Riemer on Unsplash
 Introduction
 With the rapid advancements in neurotechnological machinery and improved analytical insights from machine learning in neuroscience, the availability of big brain data has increased tremendously. Neurological health research is done using digitized brain data.[1] There must be adequate data governance to secure the privacy of subjects participating in brain research and treatments. If not properly regulated, the research methods could lead to significant breaches of the subject’s autonomy and privacy. This paper will address the necessity for neuroprotection laws, which effectively govern the use of big brain data to ensure respect for patient privacy and autonomy.
 Background
 Artificial intelligence and machine learning can be integrated with neuroscience big brain data to drive research studies. This integrative technology allows patterns of electrical activity in neurons to be studied in detail.[2]Specifically, it uses a robotic system which can reason, plan, and exhibit biologically intelligent behavior. Machine learning is a method of computer programming where the code can adapt its behavior based on big brain data.[3] The big brain data is the collection of large amounts of information for the purpose of deciphering patterns through computer analysis using machine learning.[4] The information that these technologies provide is extensive enough to allow a researcher to read a patient’s mind. AI and machine learning technologies work by finding the underlying structure of brain data, which is then described by patterns known as latent factors, eventually resulting in an understanding of the brain’s temporal dynamics.[5]
 Through these technologies, researchers are able to decipher how the human brain computes its performances and thoughts. However, due to the extensive and complex nature of the data processed through AI and machine learning, researchers may gain access to personal information a patient may not wish to reveal. From a bioethical lens, tensions arise in the realm of patient autonomy. Patients are not able to control the transmission of data from their brains that is analyzed by researchers. Governing brain data through laws may enhance the extent of patient privacy in the case where brain data is being used through AI technologies.[6] A responsible approach to governing brain data would require a sophisticated legal structure.
 Analysis
 Impact on Patient Autonomy and Privacy 
 In research pertaining to big brain data, the consent forms do not fully cover the vast amounts of information that is collected. According to research, personal data has become the most sought out commodity to provide content to corporations and the web-based service industry. Unfortunately, data leaks that release private information frequently occur.[7] The storage of an individual’s data on technologies accessible on the internet during research studies makes it vulnerable to leaks, jeopardizing an individual’s privacy. These data leaks may cause the patient to be identified easily, as the degree of information provided by AI technologies are personalized and may be decoded through brain fingerprinting methods.[8]
 There has been an extensive growth in the development and use of AI. It is efficient in providing information to radiologists who diagnose various diseases including brain cancer and psychiatric disease, and AI assists in the delivery of telemedicine.[9] However, the ethical pitfall of reduced patient autonomy must be addressed by analyzing current AI technologies and creating more options for patient preference in how the data may be used. For instance, facial recognition technology[10] commonly used in health care produces more information than listed in common consent forms, threatening to undermine informed consent. Facial recognition software collects extensive data and may disclose more information than a person would prefer to provide despite being a useful tool for diagnosing medical and genetic conditions.[11] In addition, people may not be aware that their images are being used to generate more clinical data for other purposes. It is difficult to guarantee the data is anonymized. Consent requirements must include informing people about the complexity of the potential uses of the data; software developers should maximize patient privacy.[12] Furthermore, there is a “human element” in the use of AI technologies as medical providers control the use and the extent to which data is captured or accessed through the AI technologies.[13] People must understand the scope of the technology and have clear communication with the physician or health care provider about how the medical information will be used. 
 Existing Laws for Brain Data Governance 
 A strict system of defined legal responsibilities of medical providers will ensure a higher degree of patient privacy and autonomy when AI technologies and data from machine learning are used. Governing specific algorithmic data is crucial in safeguarding a patient’s privacy and developing a gold standard treatment protocol following the procurement of the information.[14] Certain AI technologies provide more data than others, and legal boundaries should be established to ensure strong performance, quality control, and scope for patient privacy and autonomy. For instance, currently AI technologies are being used in the realm of intensive neurological care. However, there is a significant level of patient uncertainty about how much control patients have over the data’s uses.[15] Calibrated legal and ethical standards will allow important brain data to be securely governed and monitored.
 Once brain signals are recorded and processed from one individual, the data may be merged with other data in Brain Computer Interface Technology (BCI).[16] To ensure a right and ability to retrieve personal data or pull it from the collection, specific regulations for varying types of data are needed.[17] The importance of consent and patient privacy must be considered through giving patients a transparent view of how brain data is governed.[18] The legal system must address discriminatory issues and risks to patients whose data is used in studies. Laws like the General Data Protection Regulation (GDPR) and the California Consumer Privacy Protection Act (CCPA) can serve as effective models to protect aggregated data. These laws govern consumer information and ensure the compliance when personal data is collected.[19] California voters recently approved expansion of the CCPA to health data. The Washington Privacy Act, which would have provided rights to access, change, and withdraw personal data, failed to pass. Other states should improve privacy as well,[20] although a federal bill would be preferable. Scientists at the Heidelberg Academy of Sciences argue for data security to be governed in a manner that balances patient privacy and autonomy with the commercial interests of researchers.[21] The balance could be achieved through privacy protections like those in the Washington Privacy Act. Although the Health Insurance Portability and Accountability Act (HIPAA) provides an overall framework to deter the likelihood of dangers to patient protection and privacy, more thorough laws are warranted to combat pervasive data transfer and analysis that technology has brought to the health care industry.[22] Breaches of patient privacy under current HIPAA regulations include releasing patient information to a reporter without their consent and sending HIV data to a patient’s employer without consent.[23] HIPAA does not cover information being shared with outside contractors who do not have an agreement with technology companies to keep patient data confidential. HIPAA regulations also do not always address blatant breaches on patient data confidentiality.[24] Patients must be provided with methods to monitor the data being analyzed to be able to view the extent of private information being generated via AI technologies. In health research, the medical purposes of better diagnosis, earlier detection of diseases, or prevention are ethical justifications for the use of the data if it was collected with permission, the person understood and approved the uses of the data, and the data was deidentified.
 A standard governance framework is required in providing the fairest system of care to patients who allow their brain data to be examined. Informed consent in the neuroscience field could reaffirm the privacy and autonomy of patients by ensuring that they understand the type of information collected. Laws also could protect data after a patient’s death. Malpractice in the scope of brain data could give people a cause of action critical in safeguarding patient’s rights. Data breach lawsuits will become common but generally do not cover deidentified data that becomes part of big data collection. A more synchronized approach to the collection and consent process will encourage an understanding of how big data is used to diagnose and treat patients. Some altruistic people may even be more likely to consent if they know the largescale data collection is helpful to treat and diagnose people. Others should have the ability to opt out of sharing neurological data, especially when there is not certainty surrounding deidentification.[25]
 Conclusion
 Artificial intelligence and machine learning technologies have the potential to aid in the diagnosis and treatment of people globally by extracting and aggregating brain data specific to individuals. However, the secure use of the data is necessary to build trust between care providers and patients, as well as in balancing the bioethical principles of beneficence and patient autonomy. We must ensure the highest quality of care to patients, while protecting their privacy, informed consent, and clinical trust. More sophis
Read moreWhat is your definition of Big Data? Researchers’ understanding of the phenomenon of the decade
The term Big Data is commonly used to describe a range of different concepts: from the collection and aggregation of vast amounts of data, to a plethora of advanced digital techniques designed to reveal patterns related to human behavior. In spite of its widespread use, the term is still loaded with conceptual vagueness. The aim of this study is to examine the understanding of the meaning of Big Data from the perspectives of researchers in the fields of psychology and sociology in order to examine whether researchers consider currently existing definitions to be adequate and investigate if a standard discipline centric definition is possible.MethodsThirty-nine interviews were performed with Swiss and American researchers involved in Big Data research in relevant fields. The interviews were analyzed using thematic coding.ResultsNo univocal definition of Big Data was found among the respondents and many participants admitted uncertainty towards giving a definition of Big Data. A few participants described Big Data with the traditional “Vs” definition—although they could not agree on the number of Vs. However, most of the researchers preferred a more practical definition, linking it to processes such as data collection and data processing.ConclusionThe study identified an overall uncertainty or uneasiness among researchers towards the use of the term Big Data which might derive from the tendency to recognize Big Data as a shifting and evolving cultural phenomenon. Moreover, the currently enacted use of the term as a hyped-up buzzword might further aggravate the conceptual vagueness of Big Data.
Read moreVideo Big Data Analytics in the Cloud: A Reference Architecture, Survey, Opportunities, and Open Research Issues
The proliferation of multimedia devices over the Internet of Things (IoT) generates an unprecedented amount of data. Consequently, the world has stepped into the era of big data. Recently, on the rise of distributed computing technologies, video big data analytics in the cloud has attracted the attention of researchers and practitioners. The current technology and market trends demand an efficient framework for video big data analytics. However, the current work is too limited to provide a complete survey of recent research work on video big data analytics in the cloud, including the management and analysis of a large amount of video data, the challenges, opportunities, and promising research directions. To serve this purpose, we present this study, which conducts a broad overview of the state-of-the-art literature on video big data analytics in the cloud. It also aims to bridge the gap among large-scale video analytics challenges, big data solutions, and cloud computing. In this study, we clarify the basic nomenclatures that govern the video analytics domain and the characteristics of video big data while establishing its relationship with cloud computing. We propose a service-oriented layered reference architecture for intelligent video big data analytics in the cloud. Then, a comprehensive and keen review has been conducted to examine cutting-edge research trends in video big data analytics. Finally, we identify and articulate several open research issues and challenges, which have been raised by the deployment of big data technologies in the cloud for video big data analytics. To the best of our knowledge, this is the first study that presents the generalized view of the video big data analytics in the cloud. This paper provides the research studies and technologies advancing video analyses in the era of big data and cloud computing.
Read moreAssessing big data analytics and characteristics in tourism: Agodi Gardens, Ibadan, Nigeria
There is a plethora of both organized and haphazard data in tourism destinations. Analyzing this data appropriately is crucial for optimal engagement. This study focuses on the connection between big data analytics and big data's characteristics in Agodi Gardens, Ibadan, Nigeria. Specific objectives were to examine the characteristics of big data; as well as to examine Descriptive and predictive data analytics. Respondents were chosen purposively. Survey instrument (questionnaire) was used to elicit data. Data was collected using structured questionnaire. The collected data were analyzed descriptively and inferentially. The study revealed that significant relationship exists between the prescriptive/descriptive big data analytics and the characteristics of big data. Precisely, there is a significant relationship between prescriptive data analytics and velocity, veracity, volume as well as value. Similarly, there is a significant relationship between descriptive data analytics and volume, variety, value as well as veracity. Likewise, variety and veracity of big data could influence big data analytics. The study therefore recommends that the management of Agodi Gardens should engage thorough big data analytics, so that data elicited by customers can be appropriately analysed and topical inference could be drawn from the analysis.
Read moreAWS Cloud Computing Solutions: Optimizing Implementation for Businesses
This study delves into the optimization strategies of Amazon Web Services (AWS) cloud computing solutions tailored specifically for businesses. It provides an in-depth exploration of the transformative impact that cloud computing has had on modern business operations, emphasizing the pivotal role played by AWS in delivering scalable, flexible, and cost-effective resources. With projections indicating the global public cloud services market's growth to $332.3 billion in 2021, it becomes evident that cloud solutions are increasingly relied upon, with AWS commanding a significant 32% market share. This paper highlights the importance of optimizing AWS solutions, considering its extensive range of services designed to meet diverse business needs. However, transitioning to cloud environments brings challenges related to data security, compliance, integration, and migration complexities, as evidenced by pertinent studies and literature reviews. This research synthesizes insights gleaned from various studies, emphasizing the adaptability and cost-effectiveness of AWS while stressing the critical role of robust security measures in its implementation across industries and geographic regions. Understanding prevalent business requisites and challenges in cloud adoption, including scalability, cost efficiency, reliability, availability, and compliance, is essential for organizations seeking to harness cloud technology effectively. Furthermore, this paper provides an overview of AWS's services encompassing computing, storage, databases, machine learning, and AI, showcasing how these empower businesses to streamline operations, foster innovation, and scale dynamically. It elucidates successful instances where AWS was strategically implemented in Netflix, Airbnb, Jollibee Group, and Capital One, demonstrating its impact on scalability, innovation, and service enhancement. The paper outlines anticipated trends in AWS and cloud computing, focusing on user-friendly tools, advancements in machine learning, infrastructure improvements, increased automation, and deeper integration of AI and advanced analytics. Ultimately, this research emphasizes business's need to optimize AWS solutions, navigate complexities, foster innovation, and achieve operational excellence in an increasingly competitive market.
Read more