ltm                   package:ltm                   R Documentation

_L_a_t_e_n_t _T_r_a_i_t _M_o_d_e_l - _L_a_t_e_n_t _V_a_r_i_a_b_l_e _M_o_d_e_l _f_o_r _B_i_n_a_r_y _D_a_t_a

_D_e_s_c_r_i_p_t_i_o_n:

     Fit a latent trait model under the Item Response Theory (IRT)
     approach.

_U_s_a_g_e:

     ltm(formula, start.val, constraint=NULL, na.action = NULL, 
             control = list())

_A_r_g_u_m_e_n_t_s:

 formula: a two-sided formula providing the responses data matrix and
          describing the latent  structure. In the left side of
          'formula' either a data.frame (that will be converted to  a
          numeric matrix using 'data.matrix()') or a numeric matrix of
          manifest variables must be  supplied. The binary responses
          must be in 0/1 format. In the right side of 'formula' only 
          two latent variables are allowed with codenames 'z1', 'z2'.
          Interaction and quadratic  terms can also be used (see
          *Details* and *Examples* sections for more info).

start.val: a numeric matrix of starting values for the algorithm (rows
          denote items and columns latent structure, with first column
          the intercepts).  If it is not supplied randomly chosen
          starting values are used instead.

constraint: a numeric vector with 3 elements, denoting the item, the
          factor and the value of the constraint. Look at the Details
          section for more info.

na.action: the 'na.action' to be used on the data frame in the
          left-side of 'formula'. In case of missing data, if
          'na.action=NULL' the model uses the available cases, i.e., it
          takes  into account the observed part of sample units with
          missing values (valid under MAR mechanisms if the model is
          correctly specified). If you want to apply a complete case
          analysis then use  'na.action=na.exclude'.

 control: a list of control values,

          _i_t_e_r._e_m the number of EM iterations. Default 40.

          _i_t_e_r._q_N the number of quasi-Newton iterations. Default 150.

          _G_H_k the number of Gauss-Hermite quadrature points. Default
               15.

          _m_e_t_h_o_d the optimization method to be used in 'optim'. Default
               "BFGS".

          _v_e_r_b_o_s_e logical; if 'TRUE' info about the optimization
               procedure are printed.

_D_e_t_a_i_l_s:

     The latent trait model is the analogous of the factor analysis
     model for binary observed data.  The model assumes that the
     dependencies between the observed response variables (known as
     items)  can be interpreted by a small number of latent variables.
     The model formulation is under the IRT  approach; in particular, 

         logit(pi_i)=beta_{i0} + beta_{i1}z_1 + beta_{i2}z_2,

     where pi_i is the the probability of a positive response in the
     ith item, beta_{i0} is the difficulty parameter, beta_{ij} 
     (j=1,2) are the discrimination parameters and z_1, z_2 denote the
     two latent variables. 

     The usual form of the latent trait model assumes linear latent
     variable effects (Bartholomew and  Knott, 1999; Moustaki and
     Knott, 2000). 'ltm' fits the linear one- and two-factor models but
      also provides extensions described by Rizopoulos and Moustaki
     (2005) to include nonlinear latent  variable effects. These are
     incorporated in the linear predictor of the model, i.e., 

 logit(pi_i)=beta_{i0} + beta_{i1}z_1 + beta_{i2}z_2 + beta_{nl}^tf(z_1, z_2),

     where f(z_1, z_2) is a function of z_1 and z_2 (e.g., f(z_1,
     z_2)=z_1z_2, f(z_1, z_2)=z_1^2, etc.) and  beta_{nl} is a matrix
     of nonlinear terms parameters (look also at the Examples).

     The linear one- and two-factor model are unidentified under
     orthogonal rotations on the factors'  space. To achieve
     identifiability you can fix the value of one loading using the
     'constraint'  argument. 

     The parameters are estimated by maximizing the approximate
     marginal log-likelihood under the conditional  independence
     assumption, i.e., conditionally on the latent structure the items
     are independent Bernoulli  variates under the logit link. The
     required integrals are approximated using the Gauss-Hermite rule.
     The  optimization procedure used is a hybrid algorithm. The
     procedure initially uses a moderate number of EM  iterations (see
     'control' argument 'iter.em') and then switches to quasi-Newton
     (see 'control'  arguments 'method' and 'iter.qN') iterations until
     convergence. In case of large number of items,  the `Conjugate
     Gradient' optimization method might be preferable. Due to the use
     of Gauss-Hermite integration  combined with the relatively large
     data sets encountered in multivariate analysis the use of
     optimized BLAS  (e.g., Goto's BLAS) is highly recommended.

_V_a_l_u_e:

     An object of class 'ltm' with components, 

coefficients: the parameter values at convergence.

 log.Lik: the log-likelihood value at convergence.

convergence: the convergence identifier returned by 'optim'.

 hessian: the approximate Hessian matrix at convergence returned by
          'optim'.

patterns: a list with two components: (i) 'mat' a numeric matrix  that
          contains the observed response patterns. (ii) 'dat' a
          data.frame that contains the observed and expected 
          frequencies for each observed response pattern.

      GH: a list with two components used in the Gauss-Hermite rule:
          (i) 'Z' a numeric matrix that contains  the quadrature
          points. (ii) 'GHw' a numeric vector that contains the
          corresponding  weights.

  max.sc: the maximum absolute value of the score vector at
          convergence.

       X: the responses data matrix.

ltn.struct: a character vector describing the latent structure.

 control: the values used in the 'control' argument.

    call: the matched call.

_W_a_r_n_i_n_g:

     In case the Hessian matrix at convergence is not positive
     definite, try to re-fit the model. 'ltm' will use new random
     starting values.

     The inclusion of nonlinear latent variable effects produces more
     complex likelihood surfaces which might possess a number of local
     maxima. To ensure that the maximum likelihood value has been
     reached re-fit the model a number of times (simulations showed
     that usually 10 times are adequate to ensure global convergence).

_A_u_t_h_o_r(_s):

     Dimitris Rizopoulos dimitris.rizopoulos@med.kuleuven.be

_R_e_f_e_r_e_n_c_e_s:

     Baker, F. and Kim, S-H. (2004) _Item Response Theory_, 2nd ed. 
     New York: Marcel Dekker.

     Bartholomew, D. and Knott, M. (1999) _Latent Variable Models and
     Factor Analysis_, 2nd ed. London: Arnold.

     Bartholomew, D., Steel, F., Moustaki, I. and Galbraith, J. (2002)
     _The Analysis and Interpretation of Multivariate Data for Social
     Scientists_. London: Chapman and Hall.

     Moustaki, I. and Knott, M. (2000) Generalized latent trait models.
     _Psychometrika_, *65*, 391-411.

     Rizopoulos, D. and Moustaki, I. (2005) Generalized latent variable
     models with nonlinear terms. _submitted for publication_.

_S_e_e _A_l_s_o:

     'coef.ltm', 'summary.ltm', 'anova.ltm', 'plot.ltm', 'margins',
     'factor.scores'

_E_x_a_m_p_l_e_s:

     ## The two-parameter logistic model for the WIRS data
     ltm(Wirs~z1)

     ## One-factor and a quadratic term
     ## using the Mobility data
     ltm(Mobility~z1+I(z1^2))

     ## Two-factor model with an interaction term
     ## using the WIRS data
     ltm(Wirs~z1*z2)

     ## A two-factor model with the constraint that the loading
     ## for the 3rd item and the 2nd factor equals 1
     ltm(Wirs~z1+z2, constraint = c(3,2,1))

     ## The two-parameter logistic model for the Abortion data 
     ## with 20 quadrature points and 20 EM iterations
     ltm(Abortion~z1, control = list(GHk=20, iter.em=20))

