kda, pda, compare             package:ks             R Documentation

_K_e_r_n_e_l _a_n_d _p_a_r_a_m_e_t_r_i_c _d_i_s_c_r_i_m_i_n_a_n_t _a_n_a_l_y_s_i_s

_D_e_s_c_r_i_p_t_i_o_n:

     Kernel and parametric discriminant analysis.

_U_s_a_g_e:

     kda(x, x.group, Hs, y, prior.prob=NULL)
     pda(x, x.group, y, prior.prob=NULL, type="quad")
     compare(x.group, est.group)

_A_r_g_u_m_e_n_t_s:

       x: matrix of training data values

 x.group: vector of group labels for training data

est.group: vector of estimated group labels

       y: matrix of test data

      Hs: (stacked) matrix of bandwidth matrices

prior.prob: vector of prior probabilities

    type: '"line"' = linear discriminant, '"quad"' = quadratic
          discriminant

_D_e_t_a_i_l_s:

     If you have prior probabilities then set 'prior.prob' to these.
     Otherwise the default is to use the sample proportions as 
     estimates of the prior probabilities.

     The parametric discriminant analysers use the code from the 'MASS'
     library namely 'lda' and 'qda' for linear and quadratic
     discriminants.

_V_a_l_u_e:

     The discriminant analysers are 'kda' and 'pda' and these return a
     vector of group labels assigned via discriminant analysis.  If the
     test data 'y' are given then these are classified. Otherwise the
     training data 'x' are classified.

     The function 'compare' creates a comparison between the true group
     labels 'x.group' and the estimated ones 'est.group'. It returns a
     list with fields 

   cross: cross-classification table with the rows indicating the true
          group and the columns the estimated group

   error: misclassification rate (MR) where 

 MR = (number of points wrongly classified) / (total number of points)

     Note that this MR is only suitable when we have test data. If  we
     don't have test data, then the cross validated estimate is more
     appropriate.  See Silverman (1986).

_R_e_f_e_r_e_n_c_e_s:

     Silverman, B. W. (1986) _Data Analysis for Statistics and Data
     Analysis_. Chapman & Hall. London.

     Simonoff, J. S. (1996) _Smoothing Methods in Statistics_.
     Springer-Verlag. New York

     Venables, W.N. & Ripley, B.D. (1997) _Modern Applied Statistics
     with S-PLUS_. Springer-Verlag. New York.

_S_e_e _A_l_s_o:

     'kda.kde', 'pda.pde'

_E_x_a_m_p_l_e_s:

     ### bivariate example - restricted iris dataset  
     library(MASS)
     data(iris)
     iris.mat <- rbind(iris[,,1], iris[,,2], iris[,,3])
     ir <- iris.mat[,c(1,2)]
     ir.gr <- iris.mat[,5]

     H <- Hkda(ir, ir.gr, bw="plugin", pre="scale")
     kda.gr <- kda(ir, ir.gr, H, ir)
     lda.gr <- pda(ir, ir.gr, ir, type="line")
     qda.gr <- pda(ir, ir.gr, ir, type="quad")

     compare(kda.gr, ir.gr)
     compare(qda.gr, ir.gr)
     compare(lda.gr, ir.gr)

     ### multivariate example - full iris dataset
     ir <- iris[,1:4]
     ir.gr <- iris[,5]

     H <- Hkda(ir, ir.gr, bw="plugin", pre="scale")
     kda.gr <- kda(ir, ir.gr, H, ir)
     lda.gr <- pda(ir, ir.gr, ir, type="line")
     qda.gr <- pda(ir, ir.gr, ir, type="quad")

     compare(kda.gr, ir.gr)
     compare(qda.gr, ir.gr)
     compare(lda.gr, ir.gr)

