varSelRFBoot            package:varSelRF            R Documentation

_B_o_o_t_s_t_r_a_p _t_h_e _v_a_r_i_a_b_l_e _s_e_l_e_c_t_i_o_n _p_r_o_c_e_d_u_r_e _i_n _v_a_r_S_e_l_R_F

_D_e_s_c_r_i_p_t_i_o_n:

     Use the bootstrap to estimate the prediction error rate (wuth the
     .632+ rule) and the stability of the variable selection procedure
     implemented in 'varSelRF'.

_U_s_a_g_e:

     varSelRFBoot(xdata, Class, c.sd = 1,
                  mtryFactor = 1, ntree = 5000, ntreeIterat = 2000,
                  vars.drop.frac = 0.2, bootnumber = 200,
                  whole.range = TRUE,
                  recompute.var.imp = FALSE,
                  usingCluster = TRUE,
                  TheCluster = NULL, srf = NULL, ...)

_A_r_g_u_m_e_n_t_s:

     Most arguments are the same as for 'varSelRFBoot'. 

   xdata: A data frame or matrix, with subjects/cases in rows and
          variables in columns. NAs not allowed.

   Class: The dependent variable; must be a factor.

    c.sd: The factor that multiplies the sd. to decide on stopping the
          tierations or choosing the final solution. See reference for
          details.

mtryFactor: The multiplication factor of sqrt{number.of.variables} for
          the number of variables to use for the ntry argument of
          randomForest.

   ntree: The number of trees to use for the first forest; same as
          ntree for randomForest.

ntreeIterat: The number of trees to use (ntree of randomForest) for all
          additional forests.

vars.drop.frac: The fraction of variables, from those in the previous
          forest, to exclude at each iteration.

whole.range: If TRUE continue dropping variables until a forest with
          only two variables is built, and choose the best model from
          the complete series of models. If FALSE, stop the iterations
          if the current OOB error becomes larger than the initial OOB
          error (plus c.sd*OOB standard error) or if the current OOB
          error becoems larger than the previous OOB error (plus
          c.sd*OOB standard error).

recompute.var.imp: If TRUE recompute variable importances at each new
          iteration.

bootnumber: The number of bootstrap samples to draw.

usingCluster: If TRUE use a cluster to parallelize the calculations.

TheCluster: The name of the cluster, if one is used.

     srf: An object of class varSelRF. If used, the ntree and
          mtryFactor parameters are taken from this object, not from
          the arguments to this function. If used, it allows to skip
          carrying out a first iteration to build the random forest to
          the complete, original data set.

     ...: Not used.

_D_e_t_a_i_l_s:

     If a cluster is used for the calculations, it will be used for the
     embarrisingly parallelizable task of building as many random
     forests as bootstrap samples.

_V_a_l_u_e:

     An object of class varSelRFBoot, which is a list with components: 

number.of.bootsamples: The number of bootstrap replicates.

bootstrap.pred.error: The .632+ estimate of the prediction error.

leave.one.out.bootstrap: The leave-one-out estimate of the error rate
          (used when computing the .632+ estimate).

all.data.randomForest: A random forest built from all the data, but
          after the variable selection. Thus, beware because the OOB
          error rate is severely biased down.

all.data.vars: The variables selected in the run with all the data.

all.data.run: An object of class varSelRF; the one obtained from a run
          of varSelRF on the original, complete, data set. See
          'varSelRF'.

class.predictions: The out-of-bag predictions from the bootstrap, of
          type "response".See 'predict.randomForest'. This is an array,
          with dimensions number of cases by number of bootstrap
          replicates. 

prob.predictions: The out-of-bag predictions from the bootstrap, of
          type "class probability". See 'predict.randomForest'. This is
          a 3-way array; the last dimension is the bootstrap
          replication; for each bootstrap replication, the 2D array has
          dimensions case by number of classes, and each value is the
          probability of belonging to that class.

number.of.vars: A vector with the number of variables selected for each
          bootstrap sample.

 overlap: The "overlap" between the variables selected from the run in
          original sample and the variables returned from a bootstrap
          sample.  Overlap between the sets of variables A and B is
          defined as

 frac{|variables.in.A cap variables.in.B|}{sqrt{|variables.in.A| |variables.in.B|}}

          or {size (cardinality) of intersection between the two sets /
          sqrt(product of size of each set)}.

all.vars.in.solutions: A vector with all the genes selected in the runs
          on all the bootstrap samples. If the same gene is selected in
          several bootstrap runs, it appears multiple times in this
          vector.

   Class: The original class argument.

allBootRuns: A list of length 'number.of.bootsamples'. Each component
          of this list is an element of class 'varSelRF' and stores the
          results from the runs on each bootstrap sample.

_N_o_t_e:

     The out-of-bag predictions stored in 'class.predictions' and
     'prob.predictions' are NOT the OOB votes from random forest itself
     for a given run. These are predictions from the out-of-bag samples
     for each  bootstrap replication. Thus, these are samples that have
     not been used at all in any of the variable selection procedures
     in the given bootstrap replication.

_A_u_t_h_o_r(_s):

     Ramon Diaz-Uriarte  rdiaz@ligarto.org

_R_e_f_e_r_e_n_c_e_s:

     Breiman, L. (2001) Random forests. _Machine Learning_, *45*, 5-32.

     Diaz-Uriarte, R. and Alvarez de Andres, S. (2005) Variable
     selection from random forests: application to gene expression
     data. Tech. report. <URL:
     http://ligarto.org/rdiaz/Papers/rfVS/randomForestVarSel.html>

     Efron, B. & Tibshirani, R. J. (1997) Improvements on
     cross-validation: the .632+ bootstrap method. _J. American
     Statistical Association_, *92*, 548-560.  

     Svetnik, V., Liaw, A. , Tong, C & Wang, T. (2004) Application of
     Breiman's random forest to modeling structure-activity
     relationships of pharmaceutical molecules.  Pp. 334-343 in _F.
     Roli, J. Kittler, and T. Windeatt_ (eds.). _Multiple Classier
     Systems, Fifth International Workshop_, MCS 2004, Proceedings,
     9-11 June 2004, Cagliari, Italy. Lecture Notes in Computer
     Science, vol. 3077.  Berlin: Springer.

_S_e_e _A_l_s_o:

     'randomForest', 'varSelRF', 'summary.varSelRFBoot',
     'plot.varSelRFBoot',

_E_x_a_m_p_l_e_s:

     ## Not run: 
     ## This is a small example, but can take some time.

     x <- matrix(rnorm(25 * 30), ncol = 30)
     x[1:10, 1:2] <- x[1:10, 1:2] + 2
     cl <- factor(c(rep("A", 10), rep("B", 15)))  

     rf.vs1 <- varSelRF(x, cl, ntree = 200, ntreeIterat = 100,
                        vars.drop.frac = 0.2)
     rf.vsb <- varSelRFBoot(x, cl,
                            bootnumber = 10,
                            usingCluster = FALSE,
                            srf = rf.vs1)
     rf.vsb
     summary(rf.vsb)
     plot(rf.vsb)
     ## End(Not run)

