如何使用Python搜索和替换XML文件中的文本?

时间:2016-06-16 20:34:47

标签: python xml search replace

如何在整个 xml 文件中搜索特定的文本模式,然后在Python 3.5中用新的文本模式替换该文本的每个匹配项?

其他所有内容(格式,属性,注释等)都需要保留原始xml文件中的内容。

我在Windows上运行Python 3.5.1(win32)。

具体来说,我想替换每次出现的" FEATURE NAME"与#34;这工作"并替换每次出现的" FEATURE NUMBER"用" 12345"。

我一直在尝试学习Python和xml.etree.ElementTree,但无法弄清楚这一点。我已经看过"搜索并替换Python中的.xml文件中的一行","用Python搜索并替换文件中的一行"和"如何搜索并使用Python替换文件中的文本?"以及本网站上现有的其他Q / A,但无法解决这个问题 - 我不是一位经验丰富的程序员,所以如果需要更多输入,请告诉我。非常感谢您的帮助!!!

这是我在记事本中打开它时xml代码的样子的副本(除了我添加了空格来缩进每一行并在我将它粘贴到这个问题时点击返回某些行):

<description-topic>
    <access-info>
        <index-term-set>
            <index-term>
                <primary>FID FEATURE NUMBER</primary>
            </index-term>
            <index-term>
                <primary>FEATURE NAME</primary>
            </index-term>
            <index-term>
                <primary>Common features</primary>
                <secondary>FID FEATURE NUMBER</secondary>
            </index-term>
        </index-term-set>
    </access-info>
    <title>FEATURE NUMBER - FEATURE NAME</title>
    <block>
        <label>Platform</label>
        <comment>REVIEWERS: I guessed at the FEATURE NAME</comment>
        <para>
            This feature applies to the following platforms: FEATURE NAME<!--Check the values--></para>
    </block>
    <block branch="no">
        <label>Feature Benefits</label>
        <para>
            <comment>REVIEWERS: What do we put here? See template (link given in review email) for more information.</comment>
        </para>
    </block>
    <block branch="no">
        <label>Dependencies</label>
        <para/>
        <subblock>
            <label>Features</label>
            <comment>What FEATURE NAME do we put here?</comment>
        </subblock>
        <subblock>
            <label>Hardware</label>
            <comment>What FEATURE NAME do we put here?</comment>
            <para>This feature applies to the following: FEATURE NUMBER and text.</para><?Pub Caret -1?>
        </subblock>
        <subblock>
            <label>Dependencies outside the eNodeB</label>
            <comment>What FEATURE NAME do we put here?</comment>
        </subblock>
    </block>
    <block branch="no">
        <label>Impacts</label>
        <comment>REVIEWERS: What FEATURE NUMBER do we put here?</comment>
        <para>
            <comment/>
        </para>
    </block>
</description-topic>

以下是我想要开始工作的最新代码:

from xml.etree import ElementTree as et
tree = et.parse('Atemplate2.xml')
tree.find('description-topic/access-info/index-term-set/index-term/primary/').text = '12345'
tree.write('Atemplate2.xml')

我收到以下错误: Traceback(最近一次调用最后一次):     文件&#34; ajktest18.py&#34;,第15行,in         tree.find(&#39; description-topic / access-info / index-term-set / index-term / primary /&#39;)。text =&#39; 12345&#39;

AttributeError:&#39; NoneType&#39;对象没有属性&#39; text&#39;

我希望能够搜索和修改整个文件中的任何匹配项,但我无法弄清楚如何找到我正在搜索的文本的一个特定位置。

以下是我尝试用来查找路径的代码:

import xml.etree.ElementTree as ET
tree = ET.parse('Atemplate.xml')
root = tree.getroot()

print(root.tag, root.attrib, root.text)

for child in root:
    print(child.tag, child.attrib, child.text)
for label in root.iter('label'):
    print(label.tag, label.attrib, label.text)
for title in root.iter('title'):
    print(title.attrib)

我也尝试了以下代码:

with open('Atemplate2.xml') as f:
    tree = ET.parse(f)
    root = tree.getroot()

for elem in root.getiterator():
    try:
        elem.text = elem.text.replace('FEATURE NAME', 'THIS WORKED')
        elem.text = elem.text.replace('FEATURE NUMBER', '12345')
    except AttributeError:
        pass

tree.write('output.xml')

但是会出现以下错误:

File "<pyshell#40>", line 2, in <module>
    tree = ET.parse(f)
File "C:\MyPath\Python35-32\lib\xml\etree\ElementTree.py", line 1182, in parse
    tree.parse(source, parser)
File "C:\ MyPath \Python35-32\lib\xml\etree\ElementTree.py", line 594, in parse
    self._root = parser._parse_whole(source)
File "C:\ MyPath \Python35-32\lib\encodings\cp1252.py", line 23, in decode
    return codecs.charmap_decode(input,self.errors,decoding_table)[0]

UnicodeDecodeError:&#39; charmap&#39;编解码器不能解码位置1119中的字节0x9d:字符映射到

# #

最终更新 - 这是最终对我有用的代码(谢谢你,Jarad!):

import lxml.etree as ET
#using lxml instead of xml preserved the comments

#adding the encoding when the file is opened and written is needed to avoid a charmap error
with open('filename.xml', encoding="utf8") as f:
  tree = ET.parse(f)
  root = tree.getroot()


  for elem in root.getiterator():
    try:
      elem.text = elem.text.replace('FEATURE NAME', 'THIS WORKED')
      elem.text = elem.text.replace('FEATURE NUMBER', '123456')
    except AttributeError:
      pass

#tree.write('output.xml', encoding="utf8")
# Adding the xml_declaration and method helped keep the header info at the top of the file.
tree.write('output.xml', xml_declaration=True, method='xml', encoding="utf8")

1 个答案:

答案 0 :(得分:4)

注意事项:

  • 我从未使用xml.etree.ElementTree
  • 我从未使用它,因为我从未发现自己在操纵XML
  • 与知道图书馆进出的人相比,我不知道这是否是“最佳”方式
  • 评论家似乎一直在评判你而不是帮助你

这是this excellent answer的修改。问题是,你需要读取XML文件并解析它。

import xml.etree.ElementTree as ET

with open('xmlfile.xml', encoding='latin-1') as f:
  tree = ET.parse(f)
  root = tree.getroot()

  for elem in root.getiterator():
    try:
      elem.text = elem.text.replace('FEATURE NAME', 'THIS WORKED')
      elem.text = elem.text.replace('FEATURE NUMBER', '123456')
    except AttributeError:
      pass

tree.write('output.xml', encoding='latin-1')

请注意,您可以将encoding参数更改为其他内容,例如:utf-8cp1252ISO-8859-1等。真正取决于您的系统和文件。< / p>