如何使用XML2包解析XML文件的子路径

时间:2018-07-03 03:31:01

标签: r xml tidyverse

我有以下xml page,看起来像这样,我需要使用xml2进行解析

enter image description here

但是,使用此代码,我无法在subcellularLocation xpath下获得列表:

library(xml2)
xmlfile <- "https://www.uniprot.org/uniprot/P09429.xml"

doc <- xmlfile %>%
  xml2::read_xml()

xml_name(doc)
xml_children(doc)
x <- xml_find_all(doc, "//subcellularLocation")
xml_path(x)
# character(0)

正确的方法是什么?


更新

The desired output is a vector:

[1] "Nucleus"                                                   
[2] "Chromosome"                                                
[3] "Cytoplasm"                                                 
[4] "Secreted"                                                  
[5] "Cell membrane"
[6] "Peripheral membrane protein" 
[7] "Extracellular side"
[8] "Endosome"                                                  
[9] "Endoplasmic reticulum-Golgi intermediate compartment"  

2 个答案:

答案 0 :(得分:2)

如果您不介意,可以使用rvest软件包:

library(rvest)
a=read_html(xmlfile)%>%
   html_nodes("subcellularlocation")

a%>%html_children()%>%html_text()

[1] "Nucleus"                                              "Chromosome"                                          
[3] "Cytoplasm"                                            "Secreted"                                            
[5] "Cell membrane"                                        "Peripheral membrane protein"                         
[7] "Extracellular side"                                   "Endosome"                                            
[9] "Endoplasmic reticulum-Golgi intermediate compartment"

答案 1 :(得分:2)

使用x <- xml_find_all(doc, "//d1:subcellularLocation")

每当遇到麻烦的问题时,首先检查文档,请使用?xml_find_all,您会在页面的结尾看到此内容

# Namespaces ---------------------------------------------------------------
# If the document uses namespaces, you'll need use xml_ns to form
# a unique mapping between full namespace url and a short prefix
x <- read_xml('
 <root xmlns:f = "http://foo.com" xmlns:g = "http://bar.com">
   <f:doc><g:baz /></f:doc>
   <f:doc><g:baz /></f:doc>
 </root>
')
xml_find_all(x, ".//f:doc")
xml_find_all(x, ".//f:doc", xml_ns(x))

因此,您然后去检查xml_ns(doc)并找到

d1  <-> http://uniprot.org/uniprot
xsi <-> http://www.w3.org/2001/XMLSchema-instance

更新

xml_find_all(doc, "//d1:subcellularLocation")
   %>% xml_children()
   %>% xml_text()

## [1] "Nucleus"                                             
## [2] "Chromosome"                                          
## [3] "Cytoplasm"                                           
## [4] "Secreted"                                            
## [5] "Cell membrane"                                       
## [6] "Peripheral membrane protein"                         
## [7] "Extracellular side"                                  
## [8] "Endosome"                                            
## [9] "Endoplasmic reticulum-Golgi intermediate compartment"ent"