Jasni, Mohamad Zain and Rahmat Widia, Sembiring (2011) The Design of Pre-Processing Multidimensional Data Based on Component Analysis. Computer and Information Science, 4 (3). pp. 106-115. ISSN 1913-8989 (Print); 1913-8997 (Online). (Published)
|
PDF
The_Design_of_Pre-Processing_Multidimensional_Data_Based_on_Component_Analysis-Journal-.pdf Download (597kB) |
Abstract
Increased implementation of new databases related to multidimensional data involving techniques to support efficient query process, create opportunities for more extensive research. Pre-processing is required because of lack of data attribute values, noisy data, errors, inconsistencies or outliers and differences in coding. Several types of pre-processing based on component analysis will be carried out for cleaning, data integration and transformation, as well as to reduce the dimensions. Component analysis can be done by statistical methods, with the aim to separate the various sources of data into a statistical pattern independent. This paper aims to improve the quality of pre-processed data based on component analysis. RapidMiner is used for data pre-processing using FastICA algorithm. Kernel K-mean is used to cluster the pre-processed data and Expectation Maximization (EM) is used to model. The model was tested using wisconsin breast cancer datasets, lung cancer datasets and prostate cancer datasets. The result shows that the performance of the cluster vector value is higher and the processing time is shorter.
Item Type: | Article |
---|---|
Uncontrolled Keywords: | Pre-processing data, Data cleansing, Data noisy, FastICA |
Subjects: | Q Science > QA Mathematics > QA75 Electronic computers. Computer science |
Faculty/Division: | Faculty of Computer System And Software Engineering |
Depositing User: | Mr. Zairi Ibrahim |
Date Deposited: | 28 Dec 2011 02:17 |
Last Modified: | 21 May 2018 05:27 |
URI: | http://umpir.ump.edu.my/id/eprint/2067 |
Download Statistic: | View Download Statistics |
Actions (login required)
View Item |