hapassoc              package:hapassoc              R Documentation

_E_M _a_l_g_o_r_i_t_h_m _t_o _f_i_t _m_a_x_i_m_u_m _l_i_k_e_l_i_h_o_o_d _e_s_t_i_m_a_t_e_s _o_f _t_r_a_i_t _a_s_s_o_c_i_a_t_i_o_n_s _w_i_t_h _S_N_P _h_a_p_l_o_t_y_p_e_s

_D_e_s_c_r_i_p_t_i_o_n:

     This function takes a dataset of haplotypes in which rows for
     individuals of uncertain phase have been augmented by
     "pseudo-individuals" who carry the possible multilocus genotypes
     consistent with the single-locus phenotypes.  The EM algorithm is
     used to find MLE's for trait associations with covariates in
     generalized linear models.

_U_s_a_g_e:

     hapassoc(form,haplos.list,baseline = "missing" ,family = binomial(),
     freq = FALSE, maxit = 50, tol = 0.001, ...)

_A_r_g_u_m_e_n_t_s:

    form: model equation in usual R format

haplos.list: list of haplotype data from 'pre.hapassoc'

baseline: optional, haplotype to be used for baseline coding. Default
          is the most frequent haplotype.

  family: binomial, poisson, gaussian or freq are supported, 
          default=binomial

    freq: initial estimates of haplotype frequencies, default values
          are  calculated in 'pre.hapassoc' using standard
          haplotype-counting  (i.e. EM algorithm without adjustment for
          non-haplotype covariates)

   maxit: maximum number of iterations of the EM algorithm; default=50

     tol: convergence tolerance in terms of the maximum difference in 
          parameter estimates between interations; default=0.001

     ...: additional arguments to be passed to the glm function such 
          as starting values for parameter estimates in the risk model

_V_a_l_u_e:

      it: number of iterations of the EM algorithm

    beta: estimated regression coefficients

    freq: estimated haplotype frequencies

    fits: fitted values of the trait

     wts: final weights calculated in last iteration of the EM
          algorithm. These are estimates of the conditional
          probabilities of each multilocus genotype given the observed 
          single-locus genotypes.

     var: joint variance-covariance matrix of the estimated regression
          coefficients and the estimated haplotype frequencies

dispersionML: maximum likelihood estimate of dispersion parameter  (to
          get the moment estimate, use 'summary.hapassoc')

  family: family of the generalized linear model (e.g. binomial, 
          gaussian, etc.)

response: trait value

converged: TRUE/FALSE indicator of convergence. If the algorithm  fails
          to converge, only the 'converged' indicator is returned.

_N_o_t_e:

     When fitting logistic regression models (i.e.
     'family=binomial()'), you will see warning messages:

     'non-integer #successes in a binomial glm! in: eval(expr, envir,
     enclos)'

     even when the response variable includes only counts. These
     warnings result from fitting a weighted logistic regression at
     each iteration of the EM algorithm and can be safely ignored.

_R_e_f_e_r_e_n_c_e_s:

     Burkett K, McNeney B, Graham J (2004). A note on inference of
     trait associations with SNP haplotypes and other attributes in
     generalized linear models. Human Heredity, *57*:200-206

_S_e_e _A_l_s_o:

     'pre.hapassoc','summary.hapassoc','glm','family'.

_E_x_a_m_p_l_e_s:

     data(hypoDat)
     example.pre.hapassoc<-pre.hapassoc(hypoDat, 3)

     example.pre.hapassoc$initFreq # look at initial haplotype frequencies
     #      h000       h001       h010       h011       h100       h101       h110 
     #0.25179111 0.26050418 0.23606001 0.09164470 0.10133627 0.02636844 0.01081260 
     #      h111 
     #0.02148268 

     names(example.pre.hapassoc$haploDM)
     # "h000"   "h001"   "h010"   "h011"   "h100"   "pooled"

     # Columns of the matrix haploDM score the number of copies of each haplotype 
     # for each pseudo-individual.

     # Logistic regression for a multiplicative odds model having as the baseline 
     # group homozygotes '001/001' for the most common haplotype

     example.regr <- hapassoc(affected ~ attr + h000+ h010 + h011 + h100 + pooled,
                       example.pre.hapassoc, family=binomial())

     # Logistic regression with separate effects for 000 homozygotes, 001 homozygotes 
     # and 000/001 heterozygotes

     example2.regr <- hapassoc(affected ~ attr + I(h000==2) + I(h001==2) +
                        I(h000==1 & h001==1), example.pre.hapassoc, family=binomial())

