从文本文件中提取带标题的段落

时间:2017-09-19 15:19:00

标签: python regex text data-extraction

我有一个像这样的文本文件

  Summary 

    Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop publishing software like Aldus PageMaker including versions of Lorem Ipsum.

Sampler

 is a long established fact that a reader will be distracted by the readable content of a page when looking at its layout. The point of using Lorem Ipsum is that it has a more-or-less normal distribution of letters, as opposed to using 'Content here, content here', making it look like readable English. Many desktop publishing packages and web page editors now use Lorem Ipsum as their default model text, and a search for 'lorem ipsum' will uncover many web sites still in their infancy. Various versions have evolved over the years, sometimes by accident, sometimes on purpose (injected humour and the like) 

我使用的代码可以提取带有摘要标题的段落,但可以使用标题获取。 我使用下面的代码,我需要输出作为文本文件,包括摘要标题。

import os

with open("se.txt", encoding='latin-1')as infile,open("fgh.txt",'w', encoding='latin-1')as outfile:
    copy = False
    for line in infile:
        if line.strip() == "Summary":
            copy = True
        elif line.strip() == "Sampler":
            copy = False
        elif copy:
            outfile.write(line)

我输出为

 Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop publishing software like Aldus PageMaker including versions of Lorem Ipsum.

我需要输出

Summary 

 Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop publishing software like Aldus PageMaker including versions of Lorem Ipsum.

1 个答案:

答案 0 :(得分:0)

将上一个elif更改为简单if

    if line.strip() == "Summary":
        copy = True
    elif line.strip() == "Sampler":
        copy = False
    if copy:
        outfile.write(line)