selProbPlot             package:varSelRF             R Documentation

_S_e_l_e_c_t_i_o_n _p_r_o_b_a_b_i_l_i_t_y _p_l_o_t _f_o_r _v_a_r_i_a_b_l_e _i_m_p_o_r_t_a_n_c_e _f_r_o_m _r_a_n_d_o_m _f_o_r_e_s_t_s

_D_e_s_c_r_i_p_t_i_o_n:

     Plot, for the top ranked k variables from the original sample, the
     probability that each of these variables is included among the top
     ranked k genes from the bootstrap samples.

_U_s_a_g_e:

     selProbPlot(object, k = c(20, 100),
                 color = TRUE,
                 legend = FALSE,
                 xlegend = 68,
                 ylegend = 0.93,
                 cexlegend = 1.4,
                 main = NULL,
                 xlab = "Rank of gene",
                 ylab = "Selection probability",
                 pch = 19, ...)

_A_r_g_u_m_e_n_t_s:

  object: An object of class varSelRFBoot such as returned by the
          'varSelRFBoot' function.

       k: A two-component vector with the k-th upper variables for
          which you want the plots. 

   color: If TRUE a color plot; if FALSE, black and white.

  legend: If TRUE, show a legend.

 xlegend: The x-coordinate for the legend.

 ylegend: The y-coordinate for the legend.

cexlegend: The 'cex' argument for the legend.

    main: 'main' for the plot.

    xlab: 'xlab' for the plot.

    ylab: 'ylab' for the plot.

     pch: 'pch' for the plot.

     ...: Additional arguments to plot.

_D_e_t_a_i_l_s:

     Pepe et al., 2003 suggested the use of selection probability plots
     to evaluate the stability and confidence on our selection of
     "relevant genes." This paper also presents several more
     sophisticated ideas not implemented here.

_V_a_l_u_e:

     Used for its side effects of producing a plot. In a single plot
     show the "selection probability plot" for the upper (largest
     variable importance) 'kt'th variables. By default, show the upper
     20 and the upper 100 colored blue and red respectively.

_N_o_t_e:

     This function is in very rudimentary shape and could be used for
     more general types of data. I wrote specifically to produce Fig. 4
     of the paper.

_A_u_t_h_o_r(_s):

     Ramon Diaz-Uriarte rdiaz@ligarto.org

_R_e_f_e_r_e_n_c_e_s:

     Breiman, L. (2001) Random forests. _Machine Learning_, *45*, 5-32.

     Diaz-Uriarte, R. , Alvarez de Andres, S. (2005) Variable selection
     from random forests: application to gene expression data. Tech.
     report. <URL:
     http://ligarto.org/rdiaz/Papers/rfVS/randomForestVarSel.html>

     Pepe, M. S., Longton, G., Anderson, G. L. & Schummer, M. (2003)
     Selecting differentially expressed genes from microarray
     experiments. _Biometrics_, *59*, 133-142. 

     Svetnik, V., Liaw, A. , Tong, C & Wang, T. (2004) Application of
     Breiman's random forest to modeling structure-activity
     relationships of pharmaceutical molecules.  Pp. 334-343 in _F.
     Roli, J. Kittler, and T. Windeatt_ (eds.). _Multiple Classier
     Systems, Fifth International Workshop_, MCS 2004, Proceedings,
     9-11 June 2004, Cagliari, Italy. Lecture Notes in Computer
     Science, vol. 3077.  Berlin: Springer.

_S_e_e _A_l_s_o:

     'randomForest', 'varSelRF', 'varSelRFBoot', 'randomVarImpsRFplot',
     'randomVarImpsRF'

_E_x_a_m_p_l_e_s:

     ## This is a small example, but can take some time.

     x <- matrix(rnorm(25 * 30), ncol = 30)
     x[1:10, 1:2] <- x[1:10, 1:2] + 2
     cl <- factor(c(rep("A", 10), rep("B", 15)))  

     rf.vs1 <- varSelRF(x, cl, ntree = 200, ntreeIterat = 100,
                        vars.drop.frac = 0.2)
     rf.vsb <- varSelRFBoot(x, cl,
                            bootnumber = 10,
                            usingCluster = FALSE,
                            srf = rf.vs1)
     selProbPlot(rf.vsb, k = c(5, 10), legend = TRUE,
                 xlegend = 8, ylegend = 0.8)

