The role of domain knowledge in data mining

The ideal situation for a Data Mining or Knowledge Discovery system would be for the user to be able to pose a query of the form “Give me something interesting that could be useful” and for the system to discover some useful knowledge for the user. But such a system would be unrealistic as databases in the real world are very large and so it would be too inefficient to be workable. So the role of the human within the discovery process is essential. Moreover, the measure of what is meant by “interesting to the user” is dependent on the user as well as the domain within which the Data Mining system is being used. In this paper we discuss the use of domain knowledge within Data Mining. We define three classes of domain knowledge: Hierarchical Generalization Trees ( HG-Trees), Attribute Relationship Rules (AR-rules) and EnvironmentBased Constraints (EBC). We discuss how each one of these types of domain knowledge is incorporated into the discovery process within the EDM (Evidential Data Mining) framework for Data Mining proposed earlier by the authors [ANAN94], and in particular within the STRIP (Strong Rule Induction in Parallel) algorithm [ANAN95] implemented within the EDM framework. We highlight the advantages of using domain knowledge within the discovery process by providing results from the application of the STRIP algorithm in the actuarial domain.

[1]  G. G. Stokes "J." , 1890, The New Yale Book of Quotations.

[2]  M. Handzic 5 , 1824, The Banality of Heidegger.

[3]  William Frawley,et al.  Knowledge Discovery in Databases , 1991 .

[4]  D. Bell,et al.  Evidence Theory and Its Applications , 1991 .

[5]  Gregory Piatetsky-Shapiro,et al.  Discovery, Analysis, and Presentation of Strong Rules , 1991, Knowledge Discovery in Databases.

[6]  David A. Bell,et al.  Discounting and Combination Operations in Evidential Reasoning , 1993, UAI.

[7]  Tomasz Imielinski,et al.  Database Mining: A Performance Perspective , 1993, IEEE Trans. Knowl. Data Eng..

[8]  David A. Bell,et al.  From Data Properties to Evidence , 1993, IEEE Trans. Knowl. Data Eng..

[9]  Kenneth H. Fasman,et al.  The GDB human genome data base anno 1993 , 1993, Nucleic Acids Res..

[10]  Jiawei Han,et al.  Dynamic Generation and Refinement of Concept Hierarchies for Knowledge Discovery in Databases , 1994, KDD Workshop.

[11]  Jiawei Han,et al.  Towards Efficient Induction Mechanisms in Database Systems , 1994, Theor. Comput. Sci..

[12]  K. H. Fasman,et al.  The GDB Human Genome Data Base anno 1994. , 1994, Nucleic acids research.

[13]  David A. Bell,et al.  Evidential techniques in parallel database mining , 1995, HPCN Europe.

[14]  D. Bell,et al.  Database mining in the Northern Ireland Housing Executive , 1995 .

[15]  J. Mallen,et al.  Utilising domain knowledge in inductive knowledge discovery , 1995 .