- Book Chapter
8
- 10.1016/b978-155860916-7/50015-4
Chapter 14 - Knowledge Discovery and Data Mining
- Jan 01, 2003
- Business Intelligence
- David Loshin
Chapter 14 - Knowledge Discovery and Data Mining
In recent years, the exponential growth of both structured and unstructured data has highlighted the importance of efficient knowledge discovery techniques. Among these, image mining and data mining play pivotal roles in extracting meaningful patterns and insights from different forms of information. At the same time, data mining focuses on structured datasets such as databases and transaction records, while image mining deals with unstructured visual content, requiring advanced feature extraction, pattern recognition, and machine learning methods. This paper presents a comparative analysis of image mining and data mining techniques, emphasising their methodologies, applications, and challenges. The comparison explores key dimensions, including data preprocessing, feature selection, algorithmic approaches, and application domains such as healthcare, security, business intelligence, and multimedia. Experimental results show clustering accuracy of 87% for data mining and image classification accuracy of 92% for image mining, highlighting the effectiveness of specialised approaches. The study highlights both the similarities and differences in knowledge discovery processes, demonstrating how integrating image mining and data mining can enhance decision-making in diverse fields. The findings provide a comprehensive understanding of their complementary roles, offering valuable insights for researchers and practitioners aiming to develop hybrid approaches to knowledge discovery.
Chapter 14 - Knowledge Discovery and Data Mining
Chapter 14 - Knowledge Discovery and Data Mining
Profiling internet banking users: A knowledge discovery in data mining process model based approach
Analysing datasets using data mining techniques can enhance decision making in organizations. However, to ensure that the full potential of these techniques is realised it is important that decision makers understand there are Knowledge Discovery and Data Mining (KDDM) processes that are mature enough to be adopted. This paper demonstrates the benefits of using a KDDM process to evaluate survey data for internet banking users in Jamaica which includes demographic as well as attitudinal and behavioral variables. The major benefits of following this process include the selection of a set of models, rather than a single model, which are more relevant to the business/research objectives and use of a more targeted knowledge discovery process as the data mining analyst is now directed to consider the effects the decisions in each phase will have on subsequent phases. This leads to more relevant knowledge being extracted from the data mining process.
Read moreTowards MKDA: A Knowledge Discovery Assistant for Researches in Medicine
Nowadays doctors are generating a huge amount of raw data. These data, analyzed with data mining techniques, could be sources of new knowledge. Unluckily such tasks need skilled data analysts, and not so much researchers in Medicine are also data mining experts. In this paper we present a web based system for knowledge discovery assistance in Medicine able to advice a medical researcher in this kind of tasks. The user must define only the experiment specifications in a formal language we have defined. The system GUI helps users in their composition. Then the system plans a Knowledge Discovery Process (KDP) on the basis of rules in a knowledge base. Finally the system executes the KDP and produces a model as result. The system works through the co-operation of different web services specialized in different tasks. The system is still under development.
Read moreRapidMiner
Powerful, Flexible Tools for a Data-Driven WorldAs the data deluge continues in todays world, the need to master data mining, predictive analytics, and business analytics has never been greater. These techniques and tools provide unprecedented insights into data, enabling better decision making and forecasting, and ultimately the solution of increasingly complex problems. Learn from the Creators of the RapidMiner Software Written by leaders in the data mining community, including the developers of the RapidMiner software, RapidMiner: Data Mining Use Cases and Business Analytics Applications provides an in-depth introduction to the application of data mining and business analytics techniques and tools in scientific research, medicine, industry, commerce, and diverse other sectors. It presents the most powerful and flexible open source software solutions: RapidMiner and RapidAnalytics. The software and their extensions can be freely downloaded at www.RapidMiner.com. Understand Each Stage of the Data Mining ProcessThe book and software tools cover all relevant steps of the data mining process, from data loading, transformation, integration, aggregation, and visualization to automated feature selection, automated parameter and process optimization, and integration with other tools, such as R packages or your IT infrastructure via web services. The book and software also extensively discuss the analysis of unstructured data, including text and image mining. Easily Implement Analytics Approaches Using RapidMiner and RapidAnalytics Each chapter describes an application, how to approach it with data mining methods, and how to implement it with RapidMiner and RapidAnalytics. These application-oriented chapters give you not only the necessary analytics to solve problems and tasks, but also reproducible, step-by-step descriptions of using RapidMiner and RapidAnalytics. The case studies serve as blueprints for your own data mining applications, enabling you to effectively solve similar problems.
Read moreBusiness Intelligence Through Big Data Analytics, Data Mining and Machine Learning
There is a huge amount of data creating during the fourth industry revaluation and the data are generating explosively by various fields of the Internet of Things (IoT). The organizations are producing and storing the huge amount of data into the data servers every moment. This data comes from social media, sensors, tracking, website, and online news articles. The Google, Facebook, Walmart, and Taobao are the most remarkable organizations are generating most of the data in the web servers. Data comes into three forms as structured (text/numeric), semi structured (audio, video, and image) and unstructured (XML and RSS feeds). A business makes revenue from the analysis of 20% of such data, which is a structured form while 80% of data is unstructured. Therefore, unstructured data contains valuable information that can help the organization to improve the business productive, better decision-making, extract the insights, new products and services and understand the market conditions in various fields such as shopping, finance, education, manufacturing, and healthcare. The unstructured data are needed to be analyzed and distribute in a structured manner, that is required information’s are to be gathered through the data mining techniques are used to mining the data. In this paper, expose the importance of data analytics and data management for beneficial usage of business intelligence, big data, data mining and machine and data management. In addition, the different techniques that can be used to discover the knowledge and useful information from such data been analyzed. This can be beneficial for numerous users concern on text mining and convert complex data into meaningful information for researchers, analyst, data scientist, and business decision makers as well.
Read moreChallenges from Clustering Analysis to Knowledge Discovery in Molecular Biomechanics
Throughout endless experimental work, short records of dynamic molecular data are generated from time to time. Biomechanics data mining and knowledge discovery have become an important study area to turn the abundance of generated raw data into pieces of information. In data mining, researchers often encounter challenging issues and constraints, ranging from nature of the collected microarray data and developed clustering algorithms to informative discovery for rhythmic data decision-making processes. This article presents the review of the commonly practiced clustering techniques in molecular biomechanical systems towards better applications in bioengineering research. It highlights the constraints and challenges encountered in temporal molecular bioengineering mechanisms. The findings revealed that the molecular data are commonly analyzed based on data mining computation and mathematical applications to link both developmental stages interfaces and the mechanical principles of living organisms. In this area, mathematical analyses are extensively carried out to investigate dynamic microarray using clustering techniques. The main goal is to extract informative knowledge. Therefore, in order to derive collective patterns and reliable information from microarray, there is a need to consider effects from the nature of data, clustering algorithms and knowledge discovery processes which require substantial understanding on biological systems. Keywords: Clustering, data mining, gene expression, information, knowledge discovery, molecular biomechanics, molecular data, bioengineering mechanisms, mathematical applications, microarray data, Clustering Algorithms.
Read moreP125. Development of a novel ensemble machine learning algorithm for prediction of complications and readmission after anterior cervical spinal fusion
P125. Development of a novel ensemble machine learning algorithm for prediction of complications and readmission after anterior cervical spinal fusion
Read morePattern mining algorithms for data streams using itemset
Pattern mining algorithms for data streams using itemset
Early Prediction of Hydrocephalus Using Data Mining Techniques
Hydrocephalus is a condition which is characterized by head enlargement in infants due to enlargement of brain ventricles. An excess of fluid secretion and collection of fluid within the brain cavities are treated as Hydrocephalus. The extra fluid puts stress on the brain and can damage the brain. Hence, the increase in the fluid level in the brain's cavities may increase intracranial pressure and lead to brain damage. It is most usual in infants on children and rarely in the adult age group. Children often have a full life span if hydrocephalus is early detected and treated. This paper presents various data mining techniques and used to find out the disease in an early manner. Magnetic resonance imaging is one of the detection tools which is used to predict the disease properly. This approach includes the basic four data mining processes, namely preprocessing, segmentation, feature extraction, and classification as stage-by-stage manner using MRI dataset. Along with this process, the tree augmented Naïve Bayes nearest neighbor (TANNN) algorithm is also implemented to improve the accuracy in detecting the disease and also gave the best detection rate. The TANNN algorithm may provide the best results in diplomatic, uniqueness, perfection, and overall running time. The first stage in the data mining technique is preprocessing, which converts the original data into a useful format. The second stage is a key technique, and it groups the original data into possible divisions according to its category. The third stage is feature extraction, which is used to extract the needed data from the source. The fourth stage is the classification that appoints data in a collection to destination categories or data groups. This paper also concentrates on image mining, which includes experiments in image elements such as texture, shape, and size. Image classification is an important task in the field of medicine and technology. This helps the radiologist in the process of diagnosing hydrocephalus.KeywordsHydrocephalusData miningPreprocessingSegmentationFeature extractionClassification
Read moreImproving healthcare services using source anonymous scheme with privacy preserving distributed healthcare data collection and mining
The trends of data mining on healthcare data for improving medical services have increased because of the electronic healthcare record(EHR) system, which collects a massive amount of data on a daily basis. In the current scenario, hospital maintains its EHR system and stores the detailed information of patients. Data mining for healthcare improvement requires the data from all the EHR systems located at a different location to be stored at the central data mining server. Collection of healthcare data at some untrusted central data mining server raises privacy threats. Healthcare data contains patients’ private information and sharing this information for data mining creates privacy issues. Most of the previous research either focused on k-anonymity technique which causes information loss and decreases data mining accuracy or privacy preserving data mining which is focused on only specific data mining technique. We adopt source anonymous technique as privacy preserving scheme and present a novel scheme for healthcare data collection and mining in this paper. Our scheme collects data from all EHR systems without any information loss and stores at a single central data mining server, also ensuring privacy is preserved. Central data mining server helps to analyze the collected data with different data mining techniques (Association rule mining, Classification, Clustering, etc.) without the involvement of EHR systems. Our scheme is collusion resilient against central data mining server and EHR systems. Theoretical and experimental analysis show the efficiency of our scheme in terms of computation and communication cost. The experimental results using Heart disease dataset show the advantage to EHR systems using the proposed approach in terms of disease prediction accuracy.
Read moreRetinal fundus diseases diagnosis using image mining
Nowadays in world diseases are increasing day by day and get affected in various form which lead to damage of some or other body part. Glaucoma and Diabetic Retinopathy [DR] are one of the factors of Eye diseases, which include vision loss. Image will undergo a standard method of applying image processing and mining technique to give exact identification of retinal category of Normal, Glaucoma and DR. Identifying the diseases with respect to existing the proposed system will give better accuracy and efficiency.
Read moreFuture Trends of Data Mining in Predicting the Various Diseases in Medical Healthcare System
The thriving medical applications of data mining in the fields of medicine and public health has led to the popularity of its use in knowledge discovery in databases (KDD). Data mining has revealed novel biomedical and healthcare acquaintances for clinical decision making that has great potential to improve the treatment quality of hospitals and increase the survival rate of patients. Disease diagnosis is one of the applications where data mining tools are establishing the successful results. Data mining intends to endow with a systematic survey of current techniques of knowledge discovery in databases using data mining techniques that are in use in today’s medical research. Discussion is made to enable the disease diagnosis and the breakthrough of hidden healthcare patterns from related databases is offered. Also, the use of data mining to discover such relationships as those between health conditions and a disease is presented. It further discusses about the tools that can be used for the processing and classification of data. This paper summarizes various technical articles on medical diagnosis and prognosis. It has also been focused on current research being carried out using the data mining techniques to enhance the disease(s) forecasting process. This research paper provides future trends of current techniques of KDD, using data mining tools for healthcare. It also confers significant issues and challenges associated with data mining and healthcare in general. The research found a growing number of data mining applications, including analysis of health care centers for better health policy-making, detection of disease outbreaks and preventable hospital deaths. The root causes of all diseases get closer towards drugs i.e. the foremost risk factor of all hilarious diseases. Drug addiction using WEKA has been used that brings into light concerning majority of drug abusers started abusing drugs at age below 20yrs.It is to make aware the druggist about the various diseases that are caused with heavy or long term intake of drugs in their life. So, to make an expert system that will awake the youth about precarious use of drugs and also alert the affected person.
Read moreRevisiting Interestingness Measures for Knowledge Discovery in Databases
The voluminous amount of data stored in databases contains hidden knowledge which could be valuable to improve decision making process of any organization. As it is not humanely possible to analyze large databases, it has become essential to apply advanced data mining algorithms for extracting patterns (models) from data to support decision making. A number of data mining algorithms produce information of a statistical nature that allows the user to assess how accurate and reliable the discovered knowledge is? However, in many cases this is not enough for the users. Even if the discovered knowledge is highly accurate from a statistical point of view, it might not be interesting to the user. Therefore the process of knowledge discovery in databases (KDD) aims at discovering knowledge that is interesting and useful to the user. Most of the data mining algorithms so far have paid lot of attention to discovery of accurate and comprehensible knowledge. Though, the question of interestingness has been addressed time to time, it is being increasingly realized by data mining community that this subject needs a renewed focus. This paper is an attempt to review the measures of interestingness used in the data mining literature. The main contribution of the paper is to improve the understanding of interestingness measures for discovery of knowledge and identify the unresolved problems to set the directions for the future research in this area.
Read moreKnowledge Discovery in Aerodynamic Design Space for Flyback-Booster Wing Using Data Mining
The data mining has been performed for the aerodynamic design optimization result of two-stage-to-orbit reusable launch vehicle flyback booster wing. Three data mining techniques were used such as self-organizing map, functional analysis of variance, and rough set theory. The optimization problem had four aerodynamic objective functions and 71 design variables regarding wing shape. The optimization obtained the result as the hypothetical design database with 302 all solutions including the 102 non-dominated solutions. Consequently, the knowledge in the design space was acquired regarding the correlation between objective functions, and the influence of the design variables to the objective function, for non-dominated and all evaluated solutions, respectively. The features of three data mining techniques were revealed. Although the combination among three techniques discovered detailed design knowledge, self-organizing map was especially a key technique for knowledge discovery. Moreover, design knowledge from all solutions conserved the information from non-dominated solutions. Data mining was essential to solve multi-objective optimization problem.
Read moreRole of Data Mining and Knowledge Discovery in Managing Telecommunication Systems
This chapter is interested in discussing how to use data mining techniques to assist in achieving an acceptable level of quality of service of telecommunication systems. The quality of service is defined as the metrics which are predicated by using the data mining techniques, decision tree, association rules and neural networks. Routing algorithms can use this metric for optimal path selection which in turn will affect positively on the system performance. Also, in this chapter management axis using data mining techniques were handled, i.e., check the status of the telecommunication networks, role of data mining in obtaining optimal configuration, how to use data mining technique to assure high level of security for the telecommunication. The popularity of data mining in the telecommunications industry can be viewed as an extension of the use of expert systems in the telecommunications industry. These systems were developed to address the complexity associated with maintaining a huge network infrastructure and the need to maximize network reliability while minimizing labor costs (Liebowitz, J. 1988). The problem with these expert systems is that they are expensive to develop because it is both difficult and time consuming to elicit the requisite domain knowledge from experts.
Read more