rm.coef              package:subselect              R Documentation

_C_o_m_p_u_t_e_s _t_h_e _R_M _c_o_e_f_f_i_c_i_e_n_t _f_o_r _v_a_r_i_a_b_l_e _s_u_b_s_e_t _s_e_l_e_c_t_i_o_n

_D_e_s_c_r_i_p_t_i_o_n:

     Computes the RM coefficient, measuring the similarity of the
     spectral decompositions of a p-variable data matrix, and of the
     matrix which results from regressing all the variables on a subset
     of only k variables.

_U_s_a_g_e:

     rm.coef(mat, indices)

_A_r_g_u_m_e_n_t_s:

     mat: the full data set's covariance (or correlation) matrix

 indices: a numerical vector, matrix or 3-d array of integers giving
          the indices of the variables in the subset. If a matrix is
          specified, each row is taken to represent a different
          _k_-variable subset. If a 3-d array is given, it is assumed
          that the third dimension corresponds to different
          cardinalities.

_D_e_t_a_i_l_s:

     Computes the RM coefficient that measures the similarity of the
     spectral decompositions of a p-variable data matrix, and of the
     matrix which results from regressing those variables on a subset
     (given by "indices") of the variables.  Input data is expected in
     the form of a (co)variance or correlation matrix. If a non-square
     matrix is given, it is assumed to be a data matrix, and its
     (co)variance matrix is used as input.

     The definition of the RM coefficient is as follows:

        RM = sqrt{frac{mathrm{tr}(X^t P_v X)}{mathrm{X^t X}}}

     {RM = sqrt(tr(X' Pv X)/tr(X'X))} where X is the full 
     (column-centered) data matrix and Pv is the matrix of  orthogonal
     projections on the subspace spanned by a k-variable subset.

     This definition is equivalent to:

           RM = sqrt(sum_i(lambda_i r_i^2)/sum_j(lambda_j))

     where lambda_i stands for the $i$-th largest eigenvalue of the
     covariance matrix defined by X and r stands for the multiple
     correlation between the 'i'-th Principal Component and the
     k-variable subset.

     These definitions are also equivalent to the expression used in
     the code, which only requires the covariance (or correlation)
     matrix of the data under consideration.

     The fact that 'indices' can be a matrix or 3-d array allows for
     the computation of the RM values of subsets produced by the search
     functions 'anneal', 'genetic' and 'improve' (whose output option
     '$subsets' are matrices or 3-d arrays), using a different
     criterion (see the example below).

_V_a_l_u_e:

     The value of the RM coefficient.

_R_e_f_e_r_e_n_c_e_s:

     Cadima, J. and Jolliffe, I.T. (2001), "Variable Selection and the
     Interpretation of Principal Subspaces", _Journal of Agricultural,
     Biological and Environmental Statistics_, Vol. 6, 62-79.

     McCabe, G.P. (1986) "Prediction of Principal Components by
     Variable Subsets", _Technical Report 86-19, Department of
     Statistics, Purdue University_.

     Ramsay, J.O., ten Berge, J. and Styan, G.P.H. (1984), "Matrix
     Correlation", _Psychometrika_, 49, 403-423.

_E_x_a_m_p_l_e_s:

     ## An example with a very small data set.

     data(iris3) 
     x<-iris3[,,1]
     rm.coef(var(x),c(1,3))
     ## [1] 0.8724422

     ## An example computing the RMs of three subsets produced when the
     ## anneal function attempted to optimize the RV criterion (using an
     ## absurdly small number of iterations).

     data(swiss)
     rvresults<-anneal(cor(swiss),2,nsol=4,niter=5,criterion="Rv")
     rm.coef(cor(swiss),rvresults$subsets)

     ##              Card.2
     ##Solution 1 0.7982296
     ##Solution 2 0.7945390
     ##Solution 3 0.7649296
     ##Solution 4 0.7623326

