python:使用for循环从elementTree中定义的元素中获取子元素

时间:2017-01-24 11:56:29

标签: python xml

如果我有以下xml:

3

我想从节点<uima.cas.FSArray _id="7429" size="2"> <i>7409</i> <i>7419</i> </uima.cas.FSArray> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7342" codingScheme="SNOMEDCT" code="269435009" oid="269435009#SNOMEDCT" score="0.0" disambiguated="false" cui="C0879626" tui="T046" preferredText="Adverse effects"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7322" codingScheme="SNOMEDCT" code="157754004" oid="157754004#SNOMEDCT" score="0.0" disambiguated="false" cui="C0879626" tui="T046" preferredText="Adverse effects"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7352" codingScheme="SNOMEDCT" code="269432007" oid="269432007#SNOMEDCT" score="0.0" disambiguated="false" cui="C0879626" tui="T046" preferredText="Adverse effects"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7332" codingScheme="SNOMEDCT" code="213029005" oid="213029005#SNOMEDCT" score="0.0" disambiguated="false" cui="C0879626" tui="T046" preferredText="Adverse effects"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7362" codingScheme="SNOMEDCT" code="157762007" oid="157762007#SNOMEDCT" score="0.0" disambiguated="false" cui="C0879626" tui="T046" preferredText="Adverse effects"/> <uima.cas.FSArray _id="7372" size="5"> <i>7362</i> <i>7332</i> <i>7352</i> <i>7322</i> <i>7342</i> </uima.cas.FSArray> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7235" codingScheme="SNOMEDCT" code="274241003" oid="274241003#SNOMEDCT" score="0.0" disambiguated="false" cui="C0004134" tui="T184" preferredText="Ataxia"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7265" codingScheme="SNOMEDCT" code="39384006" oid="39384006#SNOMEDCT" score="0.0" disambiguated="false" cui="C0004134" tui="T184" preferredText="Ataxia"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7255" codingScheme="SNOMEDCT" code="206825002" oid="206825002#SNOMEDCT" score="0.0" disambiguated="false" cui="C0004134" tui="T184" preferredText="Ataxia"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7275" codingScheme="SNOMEDCT" code="20262006" oid="20262006#SNOMEDCT" score="0.0" disambiguated="false" cui="C0004134" tui="T184" preferredText="Ataxia"/> <org.apache.ctakes.typesystem.type.refsem.UmlsConcept _id="7245" codingScheme="SNOMEDCT" code="158202006" oid="158202006#SNOMEDCT" score="0.0" disambiguated="false" cui="C0004134" tui="T184" preferredText="Ataxia"/> <uima.cas.FSArray _id="7285" size="5"> <i>7245</i> <i>7275</i> <i>7255</i> <i>7265</i> <i>7235</i>

获取_id<i>子元素

即。看着第一个节点(前三行),我想找回像

这样的东西
uima.cas.FSArray

和以下_id i 7429 7409 7429 7419 个节点类似。

我意识到同一个节点(没有出现属性)所以我只对uima.cas.FSArray元素的节点感兴趣。

这是我的尝试:

_id

但我明白了:

#!/usr/bin/env python

import sys
import os
import xml.etree.cElementTree as ET

tree = ET.ElementTree(file=sys.argv[-1])

UMLSarr = {}
for x in tree.iterfind('uima.cas.FSArray'):
    UMLSarr[x] = x.attrib
    subArr[x] = SubElement(UMLSarr[x],"subArr",attrib='i')

我已尝试过此代码的各种其他迭代,但我遇到了越来越多的错误,并希望有人能帮助我。

感谢。

1 个答案:

答案 0 :(得分:1)

from lxml import etree

et = etree.fromstring(xml)
for array in et.xpath('//uima.cas.FSArray[@_id]'):
    print(array.xpath('@_id'), array.xpath('./i/text()'))

出:

['7429'] ['7409', '7419']
['7372'] ['7362', '7332', '7352', '7322', '7342']
['7285'] ['7245', '7275', '7255', '7265', '7235']