- Book Chapter
8
- 10.1016/b978-155860916-7/50015-4
Chapter 14 - Knowledge Discovery and Data Mining
- Jan 01, 2003
- Business Intelligence
- David Loshin
Chapter 14 - Knowledge Discovery and Data Mining
The amount of data getting generated in any sector at present is enormous. The information flow in the pharma industry is huge. Pharma firms are progressing into increased technology-enabled products and services. Data mining, which is knowledge discovery from large sets of data, helps pharma firms to discover patterns in improving the quality of drug discovery and delivery methods. The paper aims to present how data mining is useful in the pharma industry, how its techniques can yield good results in pharma sector, and to show how data mining can really enhance in making decisions using pharmaceutical data. This conceptual paper is written based on secondary study, research and observations from magazines, reports and notes. The author has listed the types of patterns that can be discovered using data mining in pharma data. The paper shows how data mining is useful in the pharma industry and how its techniques can yield good results in pharma sector. Although much work can be produced for discovering knowledge in pharma data using data mining, the paper is limited to conceptualizing the ideas and view points at this stage; future work may include applying data mining techniques to pharma data based on primary research using the available, famous significant data mining tools. Research papers and conceptual papers related to data mining in Pharma industry are rare; this is the motivation for the paper.
Chapter 14 - Knowledge Discovery and Data Mining
Chapter 14 - Knowledge Discovery and Data Mining
Active Storage with Analytics Capabilities and I/O Runtime System for Petascale Systems
Computational scientists must understand results from experimental, observational and computational simulation generated data to gain insights and perform knowledge discovery. As systems approach the petascale range, problems that were unimaginable a few years ago are within reach. With the increasing volume and complexity of data produced by ultra-scale simulations and high-throughput experiments, understanding the science is largely hampered by the lack of comprehensive I/O, storage, acceleration of data manipulation, analysis, and mining tools. Scientists require techniques, tools and infrastructure to facilitate better understanding of their data, in particular the ability to effectively perform complex data analysis, statistical analysis and knowledge discovery. The goal of this work is to enable more effective analysis of scientific datasets through the integration of enhancements in the I/O stack, from active storage support at the file system layer to MPI-IO and high-level I/O library layers. We propose to provide software components to accelerate data analytics, mining, I/O, and knowledge discovery for large-scale scientific applications, thereby increasing productivity of both scientists and the systems. Our approaches include 1) design the interfaces in high-level I/O libraries, such as parallel netCDF, for applications to activate data mining operations at the lower I/O layers; 2) Enhance MPI-IO runtime systems to incorporate the functionality developed as a part of the runtime system design; 3) Develop parallel data mining programs as part of runtime library for server-side file system in PVFS file system; and 4) Prototype an active storage cluster, which will utilize multicore CPUs, GPUs, and FPGAs to carry out the data mining workload.
Read moreResearch Issues on Datamining
Data Mining refers to a set of methods applicable to large and complex databases to eliminate the randomness and discover the hidden pattern. Datamining (DM), also known as knowledge discovery from databases (KDD), is the extraction of new knowledge from huge databases. Data mining involves the use of sophisticated data analysis tools to discover previously unknown, valid patterns and relationships in large data sets. Data mining tools can forecast the future trends and activities to support the decision of people. The scope of datamining is associated with Uncovering trends and patterns are a great power for the businesses of all sectors and industries. Modern intrusion detection applications are confronted with a variety of issues. These applications must be reliable, extensible, manageable, and minimal in maintenance costs. Data mining-based intrusion detection systems (IDSs) have shown high accuracy, good generalisation to novel types of intrusion, and stable behaviour in a changing environment in recent years. The number of hidden layers in various neural network topologies is evaluated in order to discover the best neural network. The technique of attempting to discover instances of network attacks by comparing current behaviour to the expected actions of an intruder is known as misuse detection. Artificial neural networks have the ability to detect and classify network activity using data that is limited, incomplete, and nonlinear. The main purpose of this work is to identify privacy and security concerns among cloud computing participants and consumers in a distributed environment. Techniques like Machine Learning, Natural Language Processing (NLP), and Data Mining are combined to automatically identify and uncover patterns from many sorts of materials. Predictive analytics is capable of dealing with both continuous and discontinuous changes. Classification, prediction, and to some extent, affinity analysis constitute the analytical methods employed in predictive analytics. The contemporary study in text or document mining is focusing on syntactic components and the semantic environment. In order to accomplish this, and with the motivation gained from our previous research contributions, we investigated a mining model to classify documents based on the Order of Context, Concept, and Semantic Relations (OCCSR). The use of data mining techniques based on Cloud computing will enable users to retrieve meaningful information from virtually integrated data warehouses, lowering infrastructure and storage costs. Data mining can extract useful and potentially useful information from the cloud. Big Data is typically defined by three characteristics known as the 3Vs (Volume, Velocity and Variety). The surveys approaches, environments, and technologies in key areas for Big Data analytics capabilities and discusses how they aid in the development of analytics solutions for Clouds. The clustering technique belongs to an unsupervised learning and it is used to discover a new set of categories. Grid-based clustering has the shortest processing time, which is typically determined by the size of the grid rather than the data. We compare the performance of three clustering algorithms: hierarchical clustering, density-based clustering, and K Means clustering. The majority of current approaches to detecting misuse involve the use of rule-based expert systems to identify indicators of known attacks. We provide a brief overview of the use of various Artificial Intelligence techniques and their advancements in the design, development, and application of Intrusion Detection Systems (IDS) for protecting computer and communication networks from intruders. The goal of Knowledge Discovery in Data (KDD) is to extract information that is not obvious by using careful and detailed analysis and interpretation. To drive decisions and actions, analytics employs KDD, data mining, text mining, statistical and quantitative analysis, explanatory and predictive models, and advanced and interactive visualisation techniques.
Read moreDeployment of Partitioning Around Medoids Clustering Algorithm on a Set of Objects Derived from Analytical CRM Data
The aim of this study is to highlight the importance of the unsupervised learning in Data mining and CRM fields. Data mining commonly known by its acronym KDD: knowledge discovery in data base, it refers to all methods and algorithms used for data exploration or prediction in large data bases volumes, Data mining is very important in various fields such as science, business and other areas deal with a large data set. CRM: Customer Relationship Management is an integrated information system that is used to plan, schedule and control the pre-sales and post-sales activities in an organization, both CRM and data mining techniques helps organizations maximize the value of every customer interaction and drive superior corporate performance. Clustering is one of the favoured used methods in data mining: The objective of this study is to implement the clustering algorithm K-Medoids via a shell script applied on a set of Analytical CRM data stored in Teradata environment.
Read moreReview of Data mining (Knowledge discovery) in the Future
Data mining (sometimes called data or knowledge discovery) is the process of analyzing data from different perspectives and summarizing it into useful information - information that can be used to increase revenue, cuts costs, or both. Data mining software is one of a number of analytical tools for analyzing data. It allows users to analyze data from many different dimensions or angles, categorize it, and summarize the relationships identified. Technically, data mining is the process of finding correlations or patterns among dozens of fields in large relational databases. The term data mining is often used to apply to the two separate processes of knowledge discovery and prediction. Knowledge discovery provides explicit information that has a readable form and can be understood by a user (e.g., association rule mining). Forecasting, or predictive modeling provides predictions of future events and may be transparent and readable in some approaches (e.g., rule-based systems) and opaque in others such as neural networks. Moreover, some data-mining systems such as neural networks are inherently geared towards prediction and pattern recognition, rather than knowledge discovery. In Future different Scope of data mining are. 1. Developing a unifying theory of data mining. 2. Scaling up for high dimensional data and high speed data streams. 3. Data mining in a network setting. 4. Data mining for biological and environmental problems. 5. Security, privacy and data integrity. 6. Dealing with non-static, unbalanced and cost-sensitive data
Read moreAdvancing Knowledge Discovery and Data Mining
Knowledge discovery and data mining have become areas of growing significance because of the recent increasing demand for KDD techniques, including those used in machine learning, databases, statistics, knowledge acquisition, data visualization, and high performance computing. Knowledge discovery and data mining can be extremely beneficial for the field of Artificial Intelligence in many areas, such as industry, commerce, government, education and so on. The relation between Knowledge and Data Mining, and Knowledge Discovery in Database (KDD) process are presented in the paper. Data mining theory, Data mining tasks, Data Mining technology and Data Mining challenges are also proposed. This is an belief abstract for an invited talk at the workshop.
Read moreMethods and problems in data mining
Knowledge discovery in databases and data mining aim at semiautomatic tools for the analysis of large data sets. We consider some methods used in data mining, concentrating on levelwise search for all frequently occurring patterns. We show how this technique can be used in various applications. We also discuss possibilities for compiling data mining queries into algorithms, and look at the use of sampling in data mining. We conclude by listing several open research problems in data mining and knowledge discovery.
Read moreNew Frontiers in Applied Data Mining
Five high-quality workshops were held at the 13th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD 2009) in Bangkok, Thailand during April 27-30, 2009. There were 17, 6, 9, 4 and 5 accepted papers to be presented at the Pacific Asia Workshop on Intelligence and Security Informatics (PAISI 2009), the workshop on Advances and Issues in Biomedical Data Mining (AIBDM 2009), the workshop on Data Mining with Imbalanced Classes and Error Cost (ICEC 2009), the workshop on Open Source in Data Mining (OSDM 2009), and the workshop on Quality Issues, Measures of Interestingness and Evaluation of Data Mining Models (QIMIE 2009). One competition, PAKDD 2009 Data Mining Competition, and one local workshop, Thai Track Session, were arranged. From these workshops (except PAISI which published its works in separate LNCS proceedings), we selected two or three best papers for this LNCS publication. PAKDD is a major international conference in the areas of data mining (DM) and knowledge discovery in database (KDD). It provides an international forum for researchers and industry practitioners to share their new ideas, original research results and practical development experiences from all KDD-related areas including data mining, data warehousing, machine learning, databases, statistics, knowledge acquisition and automatic scientific discovery,data visualization, causal induction and knowledge-based systems.
Read moreHard Data Analytics Problems Make for Better Data Analysis Algorithms: Bioinformatics as an Example.
Data mining and knowledge discovery techniques have greatly progressed in the last decade. They are now able to handle larger and larger datasets, process heterogeneous information, integrate complex metadata, and extract and visualize new knowledge. Often these advances were driven by new challenges arising from real-world domains, with biology and biotechnology a prime source of diverse and hard (e.g., high volume, high throughput, high variety, and high noise) data analytics problems. The aim of this article is to show the broad spectrum of data mining tasks and challenges present in biological data, and how these challenges have driven us over the years to design new data mining and knowledge discovery procedures for biodata. This is illustrated with the help of two kinds of case studies. The first kind is focused on the field of protein structure prediction, where we have contributed in several areas: by designing, through regression, functions that can distinguish between good and bad models of a protein's predicted structure; by creating new measures to characterize aspects of a protein's structure associated with individual positions in a protein's sequence, measures containing information that might be useful for protein structure prediction; and by creating accurate estimators of these structural aspects. The second kind of case study is focused on omics data analytics, a class of biological data characterized for having extremely high dimensionalities. Our methods were able not only to generate very accurate classification models, but also to discover new biological knowledge that was later ratified by experimentalists. Finally, we describe several strategies to tightly integrate knowledge extraction and data mining in order to create a new class of biodata mining algorithms that can natively embrace the complexity of biological data, efficiently generate accurate information in the form of classification/regression models, and extract valuable new knowledge. Thus, a complete data-to-information-to-knowledge pipeline is presented.
Read moreInteractive Knowledge Discovery for Temporal Lobe Epilepsy
Medical data mining and knowledge discovery can benefit from the experience and knowledge of clinicians, however, the implementation of this data mining system is challenging. Unlike traditional data mining methods, in this class of applications we process data with some posterior knowledge and the target function is more complex and even may include the opinion of user. Despite the success of the classical reasoning algorithms in many common data mining applications, they failed to address medical record processing where we need to extract information from incomplete, small samples along with an external rulebase to generate ‘meaningful’ interpretation of biological phenomenon. Swarm intelligence is an alternative class of flexible approaches that is promising in data mining. With full control over the rule extraction target function, particle swarm optimization (PSO) is a suitable approach for data mining subject to a rulebase which defines the quality of rules and constancy with previous observations. In this chapter we describe a complex clinical problem that has been addressed using PSO data mining. A large group of temporal lobe epilepsy patients are studied to find the best surgery candidates. Since there are many parameters involved in the decision process, the problem is not tractable from traditional data mining point of view, while the new approach that uses the field knowledge could extract valuable information. The proposed method allows expert to interaction with data mining process by offering manual manipulation of generated rules. The algorithm adjusts the rule set with regard to manipulations. Each rule has a reasoning which is based on the provided rulebase and similar observed cases. Support vector machine (SVM) classifier and swarm data miner are integrated to handle joint processing of raw data and rules. This approach is used to establish the limits of observations and build decision boundaries based on these critical observations.
Read moreUsing Grids for Distributed Knowledge Discovery
Knowledge discovery is a compute and data intensive process that allows for finding patterns, trends, and models in large datasets. The Grid can be effectively exploited for deploying knowledge discovery applications because of the high-performance it can offer and its distributed infrastructure. For effective use of Grids in knowledge discovery, the development of middleware is critical to support data management, data transfer, data mining and knowledge representation. To such purpose, we designed the Knowledge Grid, a high-level environment providing for Grid-based knowledge discovery tools and services. Such services allow users to create and manage complex knowledge discovery applications, composed as workflows that integrate data sources and data mining tools provided as distributed Grid services. This chapter describes the Knowledge Grid architecture and describes how its components can be used to design and implement distributed knowledge discovery applications. Then, the chapter describes how the Knowledge Grid services can be made accessible using the Open Grid Services Architecture (OGSA) model.
Read moreStatistical Data Mining and Knowledge Discovery
The Role of Bayesian and Frequentist Multivariate Modeling in Statistical Data Mining, S. James Press Intelligent Statistical Data Mining with Information Complexity and Genetic Algorithms, Hamparsum Bozdogan Econometric and Statistical Data Mining, Prediction and Policy-Making, Arnold Zellner Data Mining Strategies for the Detection of Chemical Warfare Agents, Jeffrey. L. Solka, Edward J. Wegman, and David J. Marchette Disclosure Limitation Methods Based on Bounds for Large Contingency Tables with Applications to Disability, Adrian Dobra, Elena A. Erosheva and Stephen E. Fienberg Partial Membership Models with Application to Disability Survey Data, Elena A. Erosheva Automated Scoring of Polygraph Data, Aleksandra B. Slavkovic Missing Value Algorithms in Decision Trees, Hyunjoong Kim and Sumer Yates Unsupervised Learning from Incomplete Data Using a Mixture Model Approach, Lynette Hunt and Murray Jorgensen Improving the Performance of Radial Basis Function (RBF) Classification Using Information Criteria, Zhenqiu Liu and Hamparsum Bozdogan Use of Kernel Based Techniques for Sensor Validation in Nuclear Power Plants, Andrei V. Gribok, Aleksey M. Urmanov, J. Wesley Hines, Robert E. Uhrig Data Mining and Traditional Regression, Christopher M. Hill, Linda C. Malone, and Linda Trocine An Extended Sliced Inverse Regression, Masahiro Mizuta Hokkaido University, Sapporo, Japan Using Genetic Programming to Improve the Group Method of Data Handling in Time Series Prediction, M. Hiassat, M.F. Abbod, and N. Mort Data Mining for Monitoring Plant Devices Using GMDH and Pattern Classification, B.R. Upadhyaya and B. Lu Statistical Modeling and Data Mining to Identify Consumer Preferences, Francois Boussu and Jean Jacques Denimal Testing for Structural Change Over Time of Brand Attribute Perceptions in Market Segments, Sara Dolnicar and Friedrich Leisch Kernel PCA for Feature Extraction with Information Complexity, Zhenqiu Liu and Hamparsum Bozdogan Global Principal Component Analysis for Dimensionality Reduction in Distributed Data Mining, Hairong Qi, Tsei-Wei Wang, J. Douglas Birdwell A New Metric for Categorical Data, S. H. Al-Harbi, G. P. McKeown and V. J. Rayward-Smith Ordinal Logistic Modeling Using ICOMP as a Goodness-of-Fit Criterion J. Michael Lanning and Hamparsum Bozdogan Comparing Latent Class Factor Analysis with the Traditional Approach in Data Mining, Jay Magidson and Jeroen Vermunt On Cluster Effects in Mining Complex Econometric Data, M. Ishaq Bhatti Neural Networks Based Data Mining Techniques For Steel Making, Ravindra K. Sarma, Amar Gupta, and Sanjeev Vadhavkar Solving Data Clustering Problem as a String Search Problem, V. Olman, D. Xu, and Y. Xu Behavior-Based Recommender Systems as Value-Added Services for Scientific Libraries, Andreas Geyer-Schulz, Michael Hahsler, Andreas Neumann, and Anke Thede GTP (General Text Parser) Software for Text Mining, Justin T. Giles, Ling Wo, Michael W. Berry Implication Intensity: From the Basic Statistical Definition to the Entropic Version Julien Blanchard, Pascale Kuntz, Fabrice Guillet, Regis Gras Use of a Secondary Splitting Criterion in Classification Forest Construction, Chang-Yung Yu and Heping Zhang A Method Integrating Self-Organizing Maps to Predict the Probability of Barrier Removal, Zhicheng Zhang, and Frederic Vanderhaegen Cluster Analysis of Imputed Financial Data Using an Augmentation-Based Algorithm, H. Bensmail, R. P. DeGennaro Data Mining in Federal Agencies, David L. Banks and Robert T. Olszewski STING: Evaluation of Scientific & Technological Innovation and Progress, S. Sirmakessis, K. Markello, P. Markellou, G. Mayritsakis, K. Perdikouri, Tsakalidis, and Georgia Panagopoulou The Semantic Conference Organizer, Kevin Heinrich, Michael W. Berry, Jack J. Dongarra, Sathish Vadhiyar
Read moreInter-transactional association rules for multi-dimensional contexts for prediction and their application to studying meteorological data
Inter-transactional association rules for multi-dimensional contexts for prediction and their application to studying meteorological data
Read moreGrid-Based Data Mining and Knowledge Discovery
The increasing use of computers in all the areas of human activities is resulting in huge collections of digital data. Databases are common everywhere and are used as repositories of every kind of data. Knowledge discovery techniques and tools are used today to analyze those very large data sets to identify interesting patterns and trends in them. When data is maintained over geographically distributed sites the computational power of distributed and parallel systems can be exploited for knowledge discovery in databases. In this scenario the Grid can provide an effective computational support for distributed knowledge discovery on large data sets. To this purpose we designed a system called Knowledge Grid This chapter describes the Knowledge Grid architecture and discusses some related systems and models recently proposed for knowledge discovery on Grids. The chapter presents also how to design and implement distributed data mining applications by using the Knowledge Grid tools starting from searching Grid resources, composing software and data elements, and executing the resulting application on a Grid.
Read moreDatabase Issues in Knowledge Discovery and Data Mining
In recent years both the number and the size of organisational databases have increased rapidly. However, although available processing power has also grown, the increase in stored data has not necessarily led to a corresponding increase in useful information and knowledge. This has led to a growing interest in the development of tools capable of harnessing the increased processing power available to better utilise the potential of stored data. The terms Knowledge Discovery in Databases and Mining have been adopted for a field of research dealing with the automatic discovery of knowledge implicit within databases. Data mining is useful in situations where the volume of data is either too large or too complicated for manual processing or, to a lesser extent, where human experts are unavailable to provide knowledge. The success already attained by a wide range of data mining applications has continued to prompt further investigation into alternative data mining techniques and the extension of data mining to new domains. This paper surveys, from the standpoint of the database systems community, current issues in data mining research by examining the architectural and process models adopted by knowledge discovery systems, the different types of discovered knowledge, the way knowledge discovery systems operate on different data types, various techniques for knowledge discovery and the ways in which discovered knowledge is used.
Read more