mfp                   package:mfp                   R Documentation

_F_i_t _a _M_u_l_t_i_p_l_e _F_r_a_c_t_i_o_n_a_l _P_o_l_y_n_o_m_i_a_l _M_o_d_e_l

_D_e_s_c_r_i_p_t_i_o_n:

     Selects the multiple fractional polynomial (MFP) model which best
     predicts  the outcome. The model may be a generalized linear model
     or a proportional hazards (Cox) model.

_U_s_a_g_e:

             mfp(formula, data, family = gaussian, subset, na.action, init, alpha=0.05,
                     select=1, verbose=FALSE, x = TRUE, y = TRUE)

_A_r_g_u_m_e_n_t_s:

 formula: a formula object, with the response of the left of a ~
          operator, and the terms, separated by + operators, on the
          right. Fractional polynomial terms are indicated by fp.  If a
          Cox PH model is required then the outcome should be specified
          using the Surv() notation used by coxph.

    data: a data frame containing the variables occurring in the
          formula. If this is missing, the variables should be on the
          search list.

  family: a family object - a list of functions and expressions for
          defining the link and variance functions, initialization and
          iterative weights. Families supported are gaussian, binomial,
          poisson, Gamma, inverse.gaussian and quasi.  Additionally Cox
          models are specified using "cox".

  subset: expression saying which subset of the rows of the data should
          be used in the fit.  All observations are included by
          default.

na.action: function to filter missing data. This is applied to the
          model.frame after any subset argument has been used. The
          default (with na.fail) is to create an error if any missing
          values are found.

    init: vector of initial values of the iteration (in Cox models
          only).

   alpha: sets the FP selection level for all predictors.  Values for
          individual predictors may be changed via the fp function in
          the formula.

  select: sets the variable selection level for all predictors.  Values
          for individual predictors may be changed via the fp function
          in the formula.

 verbose: run in verbose mode.

       x: return the design matrix in the model object?

       y: return the response in the model object?

_D_e_t_a_i_l_s:

     The estimation algorithm processes the predictors in turn.
     Initially,  mfp silently arranges the predictors in order of
     increasing P-value (i.e. of decreasing statistical significance)
     for omitting each predictor from the model comprising all the
     predictors with each term linear. The aim is to model relatively
     important variables before unimportant ones.

     At the initial cycle, the best-fitting FP function for the first
     predictor is determined, with all the other variables assumed
     linear. The FP selection  procedure is described below. The
     functional form (but NOT the estimated regression coefficients)
     for this predictor is kept, and the  process is repeated for the
     other predictors in turn. The first iteration concludes when all
     the variables have been processed in this way. The next cycle is
     similar, except that the functional forms from the initial cycle
     are retained for all variables excepting the one currently being
     processed.

     A variable whose functional form is prespecified to be linear
     (i.e. to  have 1 df) is tested only for exclusion within the above
     procedure when  its nominal P-value (selection level) according to
     select() is less than 1.

     Updating of FP functions and candidate variables continues until
     the functions and variables included in the overall model do not
     change (convergence). Convergence is usually achieved within 1-4
     cycles.

     _Model Selection_

     mfp uses a form of backward elimination. It start from a most
     complex  permitted FP model and attempt to simplify it by reducing
     the df. The  selection algorithm is inspired by the so-called
     "closed test procedure", a sequence of tests in each of which the
     "familywise error rate" or P-value is maintained at a prespecified
     nominal value such as 0.05. 

     The "closed test" algorithm for choosing an FP model with maximum 
     permitted degree m=2 (4 df) for a single continuous predictor, x,
     is as follows:

     1. Inclusion: test the FP in x for possible omission of x (4 df
     test, significance level determined by select). If x is
     significant, continue, otherwise drop x from the model.

     2. Non-linearity: test the FP in x against a straight line in x (3
     df test, significance level determined by alpha). If significant,
     continue, otherwise the chosen model is a straight line.

     3. Simplification: test the FP with m=2 (4 df) against the best FP
     with m=1 (2 df) (2 df test at alpha level). If significant, choose
     m=2, otherwise choose m=1.

     All significance tests are carried out using an approximate
     P-value calculation based on a difference in deviances (-2 x log
     likelihood)  having a chi-squared or F distribution, depending on
     the regression in  use. Therefore, each of the tests in the
     procedure maintains a  significance level only approximately equal
     to select. The algorithm is  thus not truly a closed procedure.
     However, for a given significance level it does provide some
     protection against over-fitting, that is against choosing
     over-complex MFP models.

_V_a_l_u_e:

     an object of class 'mfp' is returned which either inherits from
     both glm and lm or coxph.

_S_i_d_e _E_f_f_e_c_t_s:

     details are produced on the screen regarding the progress of the 
     backfitting routine. At completion of the algorithm a table is
     displayed showing the final powers selected for each variable
     along with other details.

_K_n_o_w_n _B_u_g_s:

     the program can have problems with factor variables. It appears to
     be best to put factors after fp terms in the formula object.

     glm models should not be specified without an intercept term as
     the software does not yet allow for that possibility.

_A_u_t_h_o_r(_s):

     Gareth Ambler and Axel Benner

_R_e_f_e_r_e_n_c_e_s:

     Royston, P. and Altman, D. (1994) Regression using fractional
     polynomials of continuous covariates. _Appl Stat._ 3: 429-467.

_S_e_e _A_l_s_o:

     mfp.object, fp, plot.fp, glm

_E_x_a_m_p_l_e_s:

             data(GBSG)
             mfp(Surv(rfst, cens) ~ fp(age, df = 4, select = 0.05)
                      + fp(prm, df = 4, select = 0.05), family = cox, data = GBSG)

