我想使用R。
从.kml文件中提取描述的值这是文件:
<?xml version="1.0" encoding="UTF-8"?>
<kml xmlns="http://www.opengis.net/kml/2.2"
xmlns:gx="http://www.google.com/kml/ext/2.2"
xmlns:atom="http://www.w3.org/2005/Atom">
<Document>
<open>1</open>
<visibility>1</visibility>
<name><![CDATA[2013-07-06 4:18pm]]></name>
...
<Placemark>
<name><![CDATA[2013-07-06 4:18pm (Start)]]></name>
<description><![CDATA[]]></description>
<TimeStamp><when>2013-07-06T20:18:56.000Z</when></TimeStamp>
<styleUrl>#start</styleUrl>
<Point>
<coordinates>-78.353348,45.020615,340.29998779296875</coordinates>
</Point>
</Placemark>
<Placemark id="tour">
<name><![CDATA[2013-07-06 4:18pm]]></name>
<description><![CDATA[]]></description>
...
<gx:Track>
<when>2013-07-06T20:18:56.000Z</when>
<gx:coord>-78.353348 45.020615 340.29998779296875</gx:coord>
<when>2013-07-06T20:19:12.000Z</when>
<gx:coord>-78.353315 45.020644 340.29998779296875</gx:coord>
<when>2013-07-06T22:12:23.000Z</when>
<gx:coord>-78.353108 45.020736 342.29998779296875</gx:coord>
<ExtendedData>
...
<Placemark>
<name><![CDATA[2013-07-06 4:18pm (End)]]></name>
<description><![CDATA[Created by Google My Tracks on Android.
Name: 2013-07-06 4:18pm
Activity type: cycling
Description: -
Total distance: 49.62 km (30.8 mi)
Total time: 1:53:28
Moving time: 1:50:17
Average speed: 26.24 km/h (16.3 mi/h)
Average moving speed: 27.00 km/h (16.8 mi/h)
Max speed: 61.20 km/h (38.0 mi/h)
Average pace: 2.29 min/km (3.7 min/mi)
Average moving pace: 2.22 min/km (3.6 min/mi)
Fastest pace: 0.98 min/km (1.6 min/mi)
Max elevation: 406 m (1333 ft)
Min elevation: 265 m (868 ft)
Elevation gain: 690 m (2263 ft)
Max grade: 12 %
Min grade: -11 %
Recorded: 2013-07-06 4:18pm
]]></description>
...
</Placemark>
</Document>
</kml>
以下是我想提取的内容,
中包含的文字 <description><![CDATA[Created by Google My Tracks on Android.: ]]></description>
即:
Name: 2013-07-06 4:18pm
Activity type: cycling
Description: -
Total distance: 49.62 km (30.8 mi)
Total time: 1:53:28
Moving time: 1:50:17
Average speed: 26.24 km/h (16.3 mi/h)
Average moving speed: 27.00 km/h (16.8 mi/h)
Max speed: 61.20 km/h (38.0 mi/h)
Average pace: 2.29 min/km (3.7 min/mi)
Average moving pace: 2.22 min/km (3.6 min/mi)
Fastest pace: 0.98 min/km (1.6 min/mi)
Max elevation: 406 m (1333 ft)
Min elevation: 265 m (868 ft)
Elevation gain: 690 m (2263 ft)
Max grade: 12 %
Min grade: -11 %
Recorded: 2013-07-06 4:18p
xmlToList给了我,我认为是NULL,因为CDATA标记意味着解析器不处理后面的内容:
xml <- xmlTreeParse("test1.kml", useInternalNodes=TRUE)
xmllist <- xmlToList(xml)
xmllist$Document$Placemark$description
[[1]]
NULL
我认为这就是this的含义“CDATA用于不应由XML解析器解析的文本数据......解析器会忽略CDATA部分内的所有内容.CDATA部分启动用“”“
以下对我来说也不起作用,也许是出于与CDATA相关的原因:
z1 <- xpathApply(xml, "//description", xmlValue)
z1
list()
任何人都可以帮我提取文件中的文字吗?
以下是该文件的链接:https://docs.google.com/file/d/0B__iOdFGJbXYOHJGbWJVNW0tS3M/edit?usp=sharing
答案 0 :(得分:3)
doc <- xmlTreeParse("test1.kml", useInternalNodes = TRUE)
root <-xmlRoot(doc)
xmlValue(root[["Document"]][["name"]])
R> xmlValue(root[["Document"]][["name"]])
[1] "2013-07-06 4:18pm"
此外,xmlToDataFrame(root)
和xmlToDataFrame(doc)
会在名称列中返回该值。在root或doc上使用xmlToList
会返回NULL
以获取任何CData的值。我正在查看名称节点,因为复制和粘贴您的示例不会xmlParse
。从我自己的小测试看起来这应该适用于任何CData。
答案 1 :(得分:1)
Jake Burkhead在评论中回答了这个问题。他的解决方案做到了。我非常感激。以下是从.kml文件中提取文本的方式:
> xml1 <- xmlTreeParse("2013-07-06 4-18pm.kml", useInternalNodes=TRUE)
> root <-xmlRoot(xml1)
> names(root[["Document"]])
open visibility name author Style Style Style Style
"open" "visibility" "name" "author" "Style" "Style" "Style" "Style"
Style Schema Placemark Placemark Placemark
"Style" "Schema" "Placemark" "Placemark" "Placemark"
> # note that I want the text in the third "Placemark" which is in position [13] so:
> xmlValue(root[["Document"]][[13]][["description"]])
[1] "Created by Google My Tracks on Android.\n\nName: 2013-07-06 4:18pm\nActivity type: cycling\nDescription: -\nTotal distance: 49.62 km (30.8 mi)\nTotal time: 1:53:28\nMoving time: 1:50:17\nAverage speed: 26.24 km/h (16.3 mi/h)\nAverage moving speed: 27.00 km/h (16.8 mi/h)\nMax speed: 61.20 km/h (38.0 mi/h)\nAverage pace: 2.29 min/km (3.7 min/mi)\nAverage moving pace: 2.22 min/km (3.6 min/mi)\nFastest pace: 0.98 min/km (1.6 min/mi)\nMax elevation: 406 m (1333 ft)\nMin elevation: 265 m (868 ft)\nElevation gain: 690 m (2263 ft)\nMax grade: 12 %\nMin grade: -11 %\nRecorded: 2013-07-06 4:18pm\n"
我接受了答案,但我认为我把完整的解决方案放在这里,以防其他人帮忙。
非常感谢你的坚持杰克。还要感谢里卡多和agstudy。
答案 2 :(得分:0)
解决此问题的一个好方法是使用xml2
包读取数据。
# Instead of xmlTreeParse
read_xml("test1.kml", options = "NOCDATA")
然后,您只需使用xml_text()
检索CDATA。
# Instead of xmllist$Document$Placemark$description
read_xml("test1.kml", options = "NOCDATA") %>%
xml_nodes("Placemark") %>%
xml_nodes("description") %>%
xml_text()