gcd.coef              package:subselect              R Documentation

_C_o_m_p_u_t_e_s _Y_a_n_a_i'_s _G_C_D _i_n _t_h_e _c_o_n_t_e_x_t _o_f _t_h_e _v_a_r_i_a_b_l_e-_s_u_b_s_e_t _s_e_l_e_c_t_i_o_n _p_r_o_b_l_e_m

_D_e_s_c_r_i_p_t_i_o_n:

     Computes Yanai's Generalized Coefficient of Determination for the
     similarity of the subspaces spanned by a subset of variables and a
     subset of the full data set's Principal Components.

_U_s_a_g_e:

     gcd.coef(mat, indices, pcindices =  NULL)

_A_r_g_u_m_e_n_t_s:

     mat: the full data set's covariance (or correlation) matrix.

 indices: a numerical vector, matrix or 3-d array of integers giving
          the indices of the variables in the subset. If a matrix is
          specified, each row is taken to represent a different
          _k_-variable subset. If a 3-d array is given, it is assumed
          that the third dimension corresponds to different
          cardinalities.

pcindices: a numerical vector of indices of Principal Components. By
          default, the first _k_ PCs are chosen, where _k_ is the
          cardinality of the subset of variables whose criterion value
          is being computed. *If a vector of PCs is specified by the
          user, those PCs will be used for all cardinalities that were
          requested*.

_D_e_t_a_i_l_s:

     Computes Yanai's Generalized Coefficient of Determination for the
     similarity of the subspaces spanned by a subset of variables
     (specified by 'indices') and a subset of the full-data set's
     Principal Components (specified by 'pcindices'). Input data is
     expected in the form of a (co)variance or correlation matrix. If a
     non-square matrix is given, it is assumed to be a data matrix, and
     its (co)variance matrix is used as input. The number of variables
     (k) and of PCs (q) does not have to be the same.

     Yanai's GCD is defined as:

          GCD = frac{mathrm{tr}(P_vcdot P_c)}{sqrt{kcdot q}}

     {GCD = tr(PvPc)/sqrt(k q)} where Pv and Pc are the matrices of
     orthogonal projections on the subspaces spanned by the k-variable
     subset and by the q-Principal Component subset, respectively.

     This definition is equivalent to:

                   GCD = sum_i (r_i^2) / sqrt(k q)

     where r_i stands for the multiple correlation between the 'i'-th
     Principal Component and the k-variable subset, and the sum is
     carried out over the q PCs (i=1,...,q) selected.

     These definitions are also equivalent to the expression used in
     the code, which only requires the covariance (or correlation)
     matrix of the data under consideration.

     The fact that 'indices' can be a matrix or 3-d array allows for
     the computation of the GCD values of subsets produced by the
     search functions 'anneal', 'genetic' and 'improve' (whose output
     option '$subsets' are matrices or 3-d arrays), using a different
     criterion (see the example below).

_V_a_l_u_e:

     The value of the GCD coefficient.

_R_e_f_e_r_e_n_c_e_s:

     Cadima, J. and Jolliffe, I.T. (2001), "Variable Selection and the
     Interpretation of Principal Subspaces", _Journal of Agricultural,
     Biological and Environmental Statistics_, Vol. 6, 62-79.

     Ramsay, J.O., ten Berge, J. and Styan, G.P.H. (1984), "Matrix
     Correlation", _Psychometrika_, 49, 403-423.

_E_x_a_m_p_l_e_s:

     ## An example with a very small data set.

     data(iris3) 
     x<-iris3[,,1]
     gcd.coef(cor(x),c(1,3))
     ## [1] 0.7666286
     gcd.coef(cor(x),c(1,3),pcindices=c(1,3))
     ## [1] 0.584452
     gcd.coef(cor(x),c(1,3),pcindices=1)
     ## [1] 0.6035127

     ## An example computing the GCDs of three subsets produced when the
     ## anneal function attempted to optimize the RV criterion (using an
     ## absurdly small number of iterations).

     data(swiss)
     rvresults<-anneal(cor(swiss),2,nsol=4,niter=5,criterion="Rv")
     gcd.coef(cor(swiss),rvresults$subsets)

     ##              Card.2
     ##Solution 1 0.4962297
     ##Solution 2 0.7092591
     ##Solution 3 0.4748525
     ##Solution 4 0.4649259

