Income                package:arules                R Documentation

_I_n_c_o_m_e _D_a_t_a _S_e_t

_D_e_s_c_r_i_p_t_i_o_n:

     The 'Income' data set originates from an example in the book 'The
     Elements of Statistical Learning' (see Section source).  The
     dataset is an extract from this survey.  It consists of 8993
     instances (obtained from the original dataset with 9409 instances,
     by removing those observations with the annual income missing)
     with 14 demographic attributes.  The data set is a good mixture of
     categorical and continuous variables with a lot of missing data. 
     This is characteristic of data mining applications.

_U_s_a_g_e:

     data("Income")
     data("Income_orig")

_F_o_r_m_a_t:

     A data frame with 8993 observations on the following 14 variables.

     _i_n_c_o_m_e a factor with levels '$0-$40,000', '$40,000+' 

     _s_e_x a factor with levels 'male', 'female' 

     _m_a_r_i_t_a_l _s_t_a_t_u_s a factor with levels 'married', 'cohabitation',
          'divorced', 'widowed', 'single'

     _a_g_e a factor with levels '14-34', '35+'

     _e_d_u_c_a_t_i_o_n a factor with levels 'no college graduate', 'college
          graduate'

     _o_c_c_u_p_a_t_i_o_n a factor with levels 'professional/managerial',
          'sales', 'laborer', 'clerical/service', 'homemaker',
          'student', 'military', 'retired', 'unemployed'

     _y_e_a_r_s _i_n _b_a_y _a_r_e_a a factor with levels '1-9', '10+' 

     _d_u_a_l _i_n_c_o_m_e_s a factor with levels 'not married', 'yes', 'no'

     _n_u_m_b_e_r _i_n _h_o_u_s_e_h_o_l_d a factor with levels '1', '2+' 

     _n_u_m_b_e_r _o_f _c_h_i_l_d_r_e_n a factor with levels '0', '1+'

     _h_o_u_s_e_h_o_l_d_e_r _s_t_a_t_u_s a factor with levels 'own', 'rent', 'live with
          parents/family'

     _t_y_p_e _o_f _h_o_m_e a factor with levels 'house', 'condominium',
          'apartment', 'mobile Home', 'other'

     _e_t_h_n_i_c _c_l_a_s_s_i_f_i_c_a_t_i_o_n a factor with levels 'american indian',
          'asian', 'black', 'east indian', 'hispanic', 'pacific
          islander', 'white', 'other' 

     _l_a_n_g_u_a_g_e _i_n _h_o_m_e a factor with levels 'english', 'spanish',
          'other'

_D_e_t_a_i_l_s:

     The original data frame is available as data set
     'Income_original'.  For the 'Income' data set we preprocessed the
     data as described in 'The Elements of Statistical Learning' by
     cutting each ordinal variable (age, education, income, years in
     bay area, number in houshold, and number of children) at its
     median into two values.

_S_o_u_r_c_e:

     Impact Resources, Inc., Columbus, OH (1987).

     Obtained from the web site of the book: Hastie, T., Tibshirani, R.
     & Friedman, J. (2001). _The Elements of Statistical Learning_.
     Springer-Verlag. (<URL:
     http://www-stat.stanford.edu/~tibs/ElemStatLearn/>; called
     'Marketing')

_E_x_a_m_p_l_e_s:

     data("Income")

     Income_transactions <- as(Income, "transactions")
     summary(Income_transactions)

