The hpaXml function family
(hpaXmlProtClass(), hpaXmlTissueExprSum(),
hpaXmlAntibody(), hpaXmlTissueExpr()) each
extract one specific, hand-picked piece of information from an imported
HPA xml file. That covers the most common use cases, but the xml files –
built from the schema at https://www.proteinatlas.org/download/proteinatlas.xsd –
contain much more: protein structure predictions, RNA expression per
cell line and tissue, single-cell type expression, western blot lanes,
and more, with new sections added to the schema on almost every HPA
release.
hpaXmlParse() fills that gap: it generically parses
everything available in an imported xml document, based on the
schema, into a plain R list of tibbles – without needing new package
code every time HPA adds something to the schema.
hpaXmlParse() takes the same kind of input as the other
hpaXml functions: the
"xml_document"/"xml_node" object returned by
hpaXmlGet() (or, for offline/reproducible work,
xml2::read_xml() on a locally saved copy – see the “Working with HPA xml files
offline” vignette).
hpaXmlParse()CCNB1 <- hpaXmlParse(CCNB1xml)
length(CCNB1)
#> [1] 37
names(CCNB1)
#> [1] "antibody" "antibodyTargetWeights"
#> [3] "antibodyTargetWeights_weight" "assayImage"
#> [5] "blotLanes" "cellExpression"
#> [7] "data" "data_level"
#> [9] "data_location" "entry"
#> [11] "entry_synonym" "identifier"
#> [13] "identifier_xref" "image"
#> [15] "image_channel" "image_imageUrl"
#> [17] "lane" "lane_weight"
#> [19] "patient" "patient_level"
#> [21] "patient_location" "proteinClasses"
#> [23] "proteinClasses_proteinClass" "proteinEvidence"
#> [25] "proteinEvidence_evidence" "rnaExpression"
#> [27] "sample" "snomedParameters"
#> [29] "snomedParameters_snomed" "subAssay"
#> [31] "tissueCell" "tissueCell_cellType"
#> [33] "tissueCell_level" "tissueCell_location"
#> [35] "tissueExpression" "tissueExpression_validation"
#> [37] "westernBlot"Every element of the returned list is a tibble – some with a single
row (entry, identifier,
proteinClasses), some with over a thousand
(snomedParameters_snomed). The exact set of tables present
depends on what that gene’s xml file actually contains; a gene with no
western blot data, for example, simply won’t have a
westernBlot entry.
Unlike a straight translation of the xml tree into nested R lists
(which, for this schema, would mean digging 6-7 levels deep to reach a
single value), hpaXmlParse() normalizes everything into a
flat, single-level list of tibbles – similar to a small
relational database extracted from the xml file:
entry row,
carries a foreign key column "<parent>_id" pointing
back at the row (in another tibble) it belongs to.entry, data,
patient, image, antibody, …)
additionally have a surrogate key column "<name>_id",
which is what their own children point at. Tibbles for leaf/property
elements (tissueCell_level, patient_location,
…) are never referenced by anything, so they carry the parent’s foreign
key only.value, and its xml attributes keep
their own names as columns.CCNB1$entry has a name column rather than a
separate entry_name table) and any attributes in columns
named "<tag>_<attribute>".For example, CCNB1$tissueCell rows carry a
data_id that matches the data_id of their
parent row in CCNB1$data, and
CCNB1$tissueCell_level rows carry a
tissueCell_id pointing back at
CCNB1$tissueCell:
CCNB1$data %>% select(data_id, tissueExpression_id, tissue) %>% head(3)
#> # A tibble: 3 x 3
#> data_id tissueExpression_id tissue
#> <int> <int> <chr>
#> 1 1 1 adrenal gland
#> 2 2 1 appendix
#> 3 3 1 bone marrow
CCNB1$tissueCell %>% head(3)
#> # A tibble: 3 x 3
#> data_id tissueCell_id quantity
#> <int> <int> <chr>
#> 1 1 1 <NA>
#> 2 2 2 <NA>
#> 3 2 3 <NA>
CCNB1$tissueCell_level %>% head(3)
#> # A tibble: 3 x 4
#> type value tissueCell_id count
#> <chr> <chr> <int> <chr>
#> 1 expression not detected 1 <NA>
#> 2 expression medium 2 <NA>
#> 3 expression low 3 <NA>Some xml tags mean different things depending on where they appear –
level is a staining intensity under tissueCell
but an RNA abundance under data; location
appears under both tissueCell and patient. To
keep tables unambiguous, this kind of leaf/property-style element is
named "<parent>_<tag>"
(tissueCell_level, patient_location,
tissueCell_cellType, …). Bigger nested entities that
represent “the same kind of record” wherever they occur –
image, data, patient,
antibody, and so on – keep their own plain tag name and are
disambiguated by their foreign key columns instead (e.g. rows in
CCNB1$image that came from an antibody’s western blot carry
a westernBlot_id, while rows from a tissue assay carry a
tissueExpression_id).
To get every IHC staining level recorded for liver samples, join
data (which tissue), tissueCell (which cell
type), tissueCell_cellType, and
tissueCell_level (the actual staining/intensity values) on
their id columns:
cellType <- CCNB1$tissueCell_cellType %>% rename(cellType = value)
level <- CCNB1$tissueCell_level %>% rename(level = value)
CCNB1$data %>%
filter(tissue == "liver", !is.na(tissueExpression_id)) %>%
inner_join(CCNB1$tissueCell, by = "data_id") %>%
inner_join(cellType, by = "tissueCell_id") %>%
inner_join(level, by = "tissueCell_id") %>%
select(tissueExpression_id, tissue, cellType, level_type = type, level)
#> # A tibble: 6 x 5
#> tissueExpression_id tissue cellType level_type level
#> <int> <chr> <chr> <chr> <chr>
#> 1 1 liver bile duct cells expression not detected
#> 2 1 liver hepatocytes expression not detected
#> 3 2 liver bile duct cells staining not detected
#> 4 2 liver bile duct cells intensity Negative
#> 5 2 liver hepatocytes staining not detected
#> 6 2 liver hepatocytes intensity NegativetissueExpression_id 1 is the gene-level tissue assay
(expression, from HPA’s “annotated protein expression”);
tissueExpression_id 2 is the antibody-level assay
(staining/intensity, one row per IHC score
type). Join in CCNB1$tissueExpression on
tissueExpression_id to see which antibody and technology
each row came from.
hpaXmlParse() decides how to shape each part of the tree
purely from the xml schema itself (is this tag’s type allowed to contain
child elements; is this tag allowed, per the schema, to occur more than
once), never from how many times something happens to occur in one
particular gene’s file. That means:
<identifier> with no <xref>
children still gets its own identifier tibble, exactly like
a gene whose identifier does have several; a
<tissueCell> with one <level>
still gets a tissueCell_level tibble, just like one with
four.proteinstructure or cellTypeExpression
branch) is picked up automatically, as its own new tibble(s), with no
changes needed to this package.if (!is.null(CCNB1$westernBlot)) ... or
length(CCNB1$westernBlot) > 0.Anh Tran, 2018-2026
Please cite: Tran, A.N., Dussaq, A.M., Kennell, T. et al. HPAanalyze: an R package that facilitates the retrieval and analysis of the Human Protein Atlas data. BMC Bioinformatics 20, 463 (2019) https://doi.org/10.1186/s12859-019-3059-z