我有一个xml文件,如下所示:
<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<feed xml:base="http://data.treasury.gov:8001/Feed.svc/" xmlns:d="http://schemas.microsoft.com/ado/2007/08/dataservices" xmlns:m="http://schemas.microsoft.com/ado/2007/08/dataservices/metadata" xmlns="http://www.w3.org/2005/Atom">
<title type="text">DailyTreasuryYieldCurveRateData</title>
<id>http://data.treasury.gov:8001/feed.svc/DailyTreasuryYieldCurveRateData</id>
<updated>2015-08-30T15:17:09Z</updated>
<link rel="self" title="DailyTreasuryYieldCurveRateData" href="DailyTreasuryYieldCurveRateData" />
<entry>
<id>http://data.treasury.gov:8001/Feed.svc/DailyTreasuryYieldCurveRateData(6404)</id>
<title type="text"></title>
<updated>2015-08-30T15:17:09Z</updated>
<author>
<name />
</author>
<link rel="edit" title="DailyTreasuryYieldCurveRateDatum" href="DailyTreasuryYieldCurveRateData(6404)" />
<category term="TreasuryDataWarehouseModel.DailyTreasuryYieldCurveRateDatum" scheme="http://schemas.microsoft.com/ado/2007/08/dataservices/scheme" />
<content type="application/xml">
<m:properties>
<d:Id m:type="Edm.Int32">6404</d:Id>
<d:NEW_DATE m:type="Edm.DateTime">2015-08-03T00:00:00</d:NEW_DATE>
<d:BC_1MONTH m:type="Edm.Double">0.03</d:BC_1MONTH>
<d:BC_3MONTH m:type="Edm.Double">0.08</d:BC_3MONTH>
<d:BC_6MONTH m:type="Edm.Double">0.17</d:BC_6MONTH>
<d:BC_1YEAR m:type="Edm.Double">0.33</d:BC_1YEAR>
<d:BC_2YEAR m:type="Edm.Double">0.68</d:BC_2YEAR>
<d:BC_3YEAR m:type="Edm.Double">0.99</d:BC_3YEAR>
<d:BC_5YEAR m:type="Edm.Double">1.52</d:BC_5YEAR>
<d:BC_7YEAR m:type="Edm.Double">1.89</d:BC_7YEAR>
<d:BC_10YEAR m:type="Edm.Double">2.16</d:BC_10YEAR>
<d:BC_20YEAR m:type="Edm.Double">2.55</d:BC_20YEAR>
<d:BC_30YEAR m:type="Edm.Double">2.86</d:BC_30YEAR>
<d:BC_30YEARDISPLAY m:type="Edm.Double">2.86</d:BC_30YEARDISPLAY>
</m:properties>
</content>
</entry>
<entry>
<id>http://data.treasury.gov:8001/Feed.svc/DailyTreasuryYieldCurveRateData(6405)</id>
<title type="text"></title>
<updated>2015-08-30T15:17:09Z</updated>
<author>
<name />
</author>
<link rel="edit" title="DailyTreasuryYieldCurveRateDatum" href="DailyTreasuryYieldCurveRateData(6405)" />
<category term="TreasuryDataWarehouseModel.DailyTreasuryYieldCurveRateDatum" scheme="http://schemas.microsoft.com/ado/2007/08/dataservices/scheme" />
<content type="application/xml">
<m:properties>
<d:Id m:type="Edm.Int32">6405</d:Id>
<d:NEW_DATE m:type="Edm.DateTime">2015-08-04T00:00:00</d:NEW_DATE>
<d:BC_1MONTH m:type="Edm.Double">0.05</d:BC_1MONTH>
<d:BC_3MONTH m:type="Edm.Double">0.08</d:BC_3MONTH>
<d:BC_6MONTH m:type="Edm.Double">0.18</d:BC_6MONTH>
<d:BC_1YEAR m:type="Edm.Double">0.37</d:BC_1YEAR>
<d:BC_2YEAR m:type="Edm.Double">0.74</d:BC_2YEAR>
<d:BC_3YEAR m:type="Edm.Double">1.08</d:BC_3YEAR>
<d:BC_5YEAR m:type="Edm.Double">1.6</d:BC_5YEAR>
<d:BC_7YEAR m:type="Edm.Double">1.97</d:BC_7YEAR>
<d:BC_10YEAR m:type="Edm.Double">2.23</d:BC_10YEAR>
<d:BC_20YEAR m:type="Edm.Double">2.59</d:BC_20YEAR>
<d:BC_30YEAR m:type="Edm.Double">2.9</d:BC_30YEAR>
<d:BC_30YEARDISPLAY m:type="Edm.Double">2.9</d:BC_30YEARDISPLAY>
</m:properties>
</content>
</entry>
</feed>
我如何解析&#39; 2.16&#39;对于&#39; BC_10YEAR&#39;?我一直在使用ElementTree和lxml查看其他示例,我似乎无法将这些示例中的xml格式与我的文件格式相匹配。
我尝试的最后一件事是:
from lxml import etree
doc = etree.parse(yield_xml)
memoryElem = doc.find('content')
print memoryElem.text # element text
print memoryElem.get('type') # attribute
我收到错误:AttributeError:&#39; NoneType&#39;对象没有属性&#39; text&#39;
有一种简单的方法吗?
答案 0 :(得分:1)
您可以尝试内置拆分方法:
>>>[data.split('>')[1].split('<')[0] for data in str(xml_file).split('<d:') if 'BC_10YEAR' in data][0]
'2.16'
答案 1 :(得分:0)
我建议使用lxml
的{{1}}方法提供更好的XPath表达式支持:
xpath()
输出
from lxml import etree
doc = etree.parse(yield_xml)
#register prefixes to be used in xpath
ns = {"foo": "http://www.w3.org/2005/Atom",
"d": "http://schemas.microsoft.com/ado/2007/08/dataservices",
"m": "http://schemas.microsoft.com/ado/2007/08/dataservices/metadata"}
#select element <d:BC_10YEAR>, and convert the value to number
result = doc.xpath("number(//foo:content/m:properties/d:BC_10YEAR)", namespaces=ns)
#print the result
print(result)
print(type(result))
如果您想知道上面的xpath表达式中为什么2.16
<type 'float'>
而不仅仅是foo:content
,那是因为foo
隐式地从根元素继承默认命名空间。默认名称空间uri在上面的代码中映射到前缀content
;相关问题:parsing xml containing default namespace to get an element value using lxml