A comparison of different procedures for principal component analysis in the presence of outliers


Alkan B. B., ATAKAN C., Alkan N.

JOURNAL OF APPLIED STATISTICS, cilt.42, sa.8, ss.1716-1722, 2015 (SCI-Expanded) identifier identifier

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 42 Sayı: 8
  • Basım Tarihi: 2015
  • Doi Numarası: 10.1080/02664763.2015.1005063
  • Dergi Adı: JOURNAL OF APPLIED STATISTICS
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus
  • Sayfa Sayıları: ss.1716-1722
  • Anahtar Kelimeler: multiple imputation, outliers, missing value, expectation-maximization, principal component analysis, MISSING DATA, ROBUST
  • Akdeniz Üniversitesi Adresli: Hayır

Özet

Principal component analysis (PCA) is a popular technique that is useful for dimensionality reduction but it is affected by the presence of outliers. The outlier sensitivity of classical PCA (CPCA) has caused the development of new approaches. Effects of using estimates obtained by expectation-maximization - EM and multiple imputation - MI instead of outliers were examined on the artificial and a real data set. Furthermore, robust PCA based on minimum covariance determinant (MCD), PCA based on estimates obtained by EM instead of outliers and PCA based on estimates obtained by MI instead of outliers were compared with the results of CPCA. In this study, we tried to show the effects of using estimates obtained by MI and EM instead of outliers, depending on the ratio of outliers in data set. Finally, when the ratio of outliers exceeds 20%, we suggest the use of estimates obtained by MI and EM instead of outliers as an alternative approach.