agreement                package:clue                R Documentation

_A_g_r_e_e_m_e_n_t _B_e_t_w_e_e_n _P_a_r_t_i_t_i_o_n_s _o_r _H_i_e_r_a_r_c_h_i_e_s

_D_e_s_c_r_i_p_t_i_o_n:

     Compute the agreement between (ensembles) of partitions or
     hierarchies.

_U_s_a_g_e:

     cl_agreement(x, y = NULL, method = "euclidean")

_A_r_g_u_m_e_n_t_s:

       x: an ensemble of partitions or hierarchies, or something
          coercible to that (see 'cl_ensemble').

       y: 'NULL' (default), or as for 'x'.

  method: a character string specifying one of the built-in methods for
          computing agreement, or a function to be taken as a
          user-defined method.  If a character string, its lower-cased
          version is matched against the lower-cased names of the
          available built-in methods using 'pmatch'.  See *Details* for
          available built-in methods.

_D_e_t_a_i_l_s:

     If 'y' is given, its components must be of the same kind as those
     of 'x' (i.e., components must either all be partitions, or all be
     hierarchies).

     If all components are partitions, the following built-in methods
     for measuring agreement between two partitions with respective
     membership matrices u and v (brought to a common number of
     columns) are available:

     '"_e_u_c_l_i_d_e_a_n"' 1 - d, where d is the euclidean dissimilarity of the
          memberships, i.e., the minimal sum of the squared differences
          of u and all column permutations of v.  See Dimitriadou,
          Weingessel and Hornik (2002).

     '"_R_a_n_d"' The Rand index (the rate of distinct pairs of objects
          both in the same class or both in different classes in both
          partitions), see Rand (1971) or Gordon (1999), page 198. For
          soft partitions, (currently) the Rand index of the
          corresponding "nerest" hard partitions is used.

     '"_c_R_a_n_d"' The Rand index corrected for agreement by chance, see
          Hubert and Arabie (1985) or Gordon (1999), page 198. Can only
          be used for hard partitions.

     '"_N_M_I"' Normalized Mutual Information, see Strehl and Ghosh
          (2002).  For soft partitions, (currently) the NMI of the
          corresponding "nerest" hard partitions is used. 

     '"_K_P"' The Katz-Powell index, i.e., the product-moment correlation
          coefficient between the elements of the co-membership
          matrices C(u) = u u' and C(v), respectively, see Katz and
          Powell (1953).  For soft partitions, (currently) the
          Katz-Powell index of the corresponding "nerest" hard
          partitions is used.  (Note that for hard partitions, the
          (i,j) entry of C(u) is one iff objects i and j are in the
          same class.)

     '"_a_n_g_l_e"' The maximal cosine of the angle between the elements of
          u and all column permutations of v.

     '"_d_i_a_g"' The maximal co-classification rate, i.e., the maximal
          rate of objects with the same class ids in both partitions
          after arbitrarily permuting the ids.

     If all components are hierarchies, available built-in methods for
     measuring agreement between two hierarchies with respective
     ultrametrics u and v are as follows.

     '"_e_u_c_l_i_d_e_a_n"' 1 / (1 + d), where d is the euclidean dissimilarity
          of the ultrametrics (i.e., the sum of the squared differences
          of u and v).

     '"_c_o_p_h_e_n_e_t_i_c"' The cophenetic correlation coefficient. (I.e., the
          product-moment correlation of the ultrametrics.

     '"_a_n_g_l_e"' The cosine of the angle between the ultrametrics.

     '"_g_a_m_m_a"' 1 - d, where d is the rate of inversions between the
          associated ultrametrics (i.e., the rate of pairs (i,j) and
          (k,l) for which u_{ij} < u_{kl} and v_{ij} > v_{kl}).  (This
          agreement measure is a linear transformation of Kruskal's
          gamma.)

     If a user-defined agreement method is to be employed, it must be a
     function taking two clusterings as its arguments.

     Symmetric agreement objects of class '"cl_agreement"' are
     implemented as symmetric proximity objects with self-proximities
     identical to one, and inherit from class '"cl_proximity"'.  They
     can be coerced to dense square matrices using 'as.matrix'.  It is
     possible to use 2-index matrix-style subscripting for such
     objects; unless this uses identical row and column indices, this
     results in a (non-symmetric agreement) object of class
     '"cl_cross_agreement"'.

_V_a_l_u_e:

     If 'y' is 'NULL', an object of class '"cl_agreement"' containing
     the agreements between the all pairs of components of 'x'. 
     Otherwise, an object of class '"cl_cross_agreement"' with the
     agreements between the components of 'x' and the components of
     'y'.

_R_e_f_e_r_e_n_c_e_s:

     E. Dimitriadou and A. Weingessel and K. Hornik (2002). A
     combination scheme for fuzzy clustering. _International Journal of
     Pattern Recognition and Artificial Intelligence_, *16*, 901-912.

     A. D. Gordon (1999). _Classification_ (2nd edition). Boca Raton,
     FL: Chapman & Hall/CRC.

     L. Hubert and P. Arabie (1985). Comparing partitions. _Journal of
     Classification_, *2*, 193-218.

     W. M. Rand (1971). Objective criteria for the evaluation of
     clustering methods. _Journal of the American Statistical
     Association_, *66*, 846-850.

     L. Katz and J. H. Powell (1953). A proposed index of the
     conformity of one sociometric measurement to another.
     _Psychometrika_, *18*, 249-256.

     A. Strehl and J. Ghosh (2002). Cluster ensembles - A knowledge
     reuse framework for combining multiple partitions. _Journal on
     Machine Learning Research_, *3*, 583-617.

_S_e_e _A_l_s_o:

     'cl_dissimilarity'; 'classAgreement' in package 'e1071'.

_E_x_a_m_p_l_e_s:

     ## An ensemble of partitions.
     data("CKME")
     pens <- CKME[1 : 20]            # for saving precious time ...
     summary(c(cl_agreement(pens)))
     summary(c(cl_agreement(pens, method = "Rand")))
     summary(c(cl_agreement(pens, method = "diag")))
     cl_agreement(pens[1:5], pens[6:7], method = "NMI")
     ## Equivalently, using subscripting.
     cl_agreement(pens, method = "NMI")[1:5, 6:7]

     ## An ensemble of hierarchies.
     d <- dist(USArrests)
     hclust_methods <- c("ward", "single", "complete", "average",
                         "mcquitty", "median", "centroid")
     hclust_results <- lapply(hclust_methods, function(m) hclust(d, m))
     hens <- cl_ensemble(list = hclust_results)
     names(hens) <- hclust_methods 
     summary(c(cl_agreement(hens)))
     summary(c(cl_agreement(hens, method = "cophenetic")))
     cl_agreement(hens[1:3], hens[4:5], method = "gamma")

