Question

我想编写一个小脚本（将在无头Linux服务器上运行），它读取PDF，突出显示与我传递的字符串数组中的任何内容匹配的文本，然后保存修改后的PDF。我想我最终会使用python bindings to poppler这样的东西，但不幸的是，文档旁边没有任何文档，而且我在python中几乎没有经验。

如果有人能指点我的教程，示例或一些有用的文档让我入门，我们将不胜感激！

Answer 1

您是否尝试过查看PDFMiner？这听起来像是你想要的。

Answer 2

PDFlib具有Python绑定并支持这些操作。如果要打开PDF，您将需要PDI。 http://www.pdflib.com/products/pdflib-family/pdflib-pdi/和TET。

不幸的是，它是一种商业产品。我过去在生产中使用过这个库，效果很好。绑定非常实用，而不是Python。我已经看到一些尝试使它们更像Pythonic：https://github.com/alexhayes/pythonic-pdflib你将要使用：open_pdi_document（）。

听起来你会想要进行某种搜索突出显示：

http://www.pdflib.com/tet-cookbook/tet-and-pdflib/highlight-search-terms/

Answer 3

是的，可以结合使用pdfminer（pip install pdfminer.six）和PyPDF2。

首先，找到坐标（例如this）。然后突出显示它：

#!/usr/bin/env python

"""Create sample highlight in a PDF file."""

from PyPDF2 import PdfFileWriter, PdfFileReader

from PyPDF2.generic import (
    DictionaryObject,
    NumberObject,
    FloatObject,
    NameObject,
    TextStringObject,
    ArrayObject
)


def create_highlight(x1, y1, x2, y2, meta, color=[0, 1, 0]):
    """
    Create a highlight for a PDF.

    Parameters
    ----------
    x1, y1 : float
        bottom left corner
    x2, y2 : float
        top right corner
    meta : dict
        keys are "author" and "contents"
    color : iterable
        Three elements, (r,g,b)
    """
    new_highlight = DictionaryObject()

    new_highlight.update({
        NameObject("/F"): NumberObject(4),
        NameObject("/Type"): NameObject("/Annot"),
        NameObject("/Subtype"): NameObject("/Highlight"),

        NameObject("/T"): TextStringObject(meta["author"]),
        NameObject("/Contents"): TextStringObject(meta["contents"]),

        NameObject("/C"): ArrayObject([FloatObject(c) for c in color]),
        NameObject("/Rect"): ArrayObject([
            FloatObject(x1),
            FloatObject(y1),
            FloatObject(x2),
            FloatObject(y2)
        ]),
        NameObject("/QuadPoints"): ArrayObject([
            FloatObject(x1),
            FloatObject(y2),
            FloatObject(x2),
            FloatObject(y2),
            FloatObject(x1),
            FloatObject(y1),
            FloatObject(x2),
            FloatObject(y1)
        ]),
    })

    return new_highlight


def add_highlight_to_page(highlight, page, output):
    """
    Add a highlight to a PDF page.

    Parameters
    ----------
    highlight : Highlight object
    page : PDF page object
    output : PdfFileWriter object
    """
    highlight_ref = output._addObject(highlight)

    if "/Annots" in page:
        page[NameObject("/Annots")].append(highlight_ref)
    else:
        page[NameObject("/Annots")] = ArrayObject([highlight_ref])


def main():
    pdf_input = PdfFileReader(open("samples/test3.pdf", "rb"))
    pdf_output = PdfFileWriter()

    page1 = pdf_input.getPage(0)

    highlight = create_highlight(89.9206, 573.1283, 376.849, 591.3563, {
        "author": "John Doe",
        "contents": "Lorem ipsum"
    })

    add_highlight_to_page(highlight, page1, pdf_output)

    pdf_output.addPage(page1)

    output_stream = open("output.pdf", "wb")
    pdf_output.write(output_stream)


if __name__ == '__main__':
    main()

读取，突出显示，以编程方式保存PDF

3 个答案: