Im new on Python. I am using this code to extract text. Is it possible extract all pages and have an output in a file?
import PyPDF2
pdf_file = open('sample.pdf','rb')
read_pdf = PyPDF2.PdfFileReader(pdf_file)
number_of_pages = read_pdf.getNumPages()
page = read_pdf.getPage(10)
page_content = page.extractText()
print (page_content)
答案 0 :(得分:2)
使用循环提取每个页面的文本,并将每个页面的文本写入单个文件。
import PyPDF2
with open('sample.pdf','rb') as pdf_file, open('sample.txt', 'w') as text_file:
read_pdf = PyPDF2.PdfFileReader(pdf_file)
number_of_pages = read_pdf.getNumPages()
for page_number in range(number_of_pages): # use xrange in Py2
page = read_pdf.getPage(page_number)
page_content = page.extractText()
text_file.write(page_content)
答案 1 :(得分:0)
我使用以下代码将多个pdf文件转换为txt
p
df_dir = "D:/search/pdf"
txt_dir = "D:/pdf_to_text"
corpus = (f for f in os.listdir(pdf_dir) if not f.startswith('.') and isfile(join(pdf_dir, f)))
pdfWriter = PyPDF2.PdfFileWriter()
for filename in corpus:
pdf = open(join(pdf_dir, filename),'rb')
pdfReader = PyPDF2.PdfFileReader(pdf)
for page in range(1, pdfReader.numPages):
pageObj = pdfReader.getPage(page)
pdfWriter.addPage(pageObj)
text = pageObj.extractText()
page_name = "{}-page{}.txt".format(filename[:4], page + 1)
with open(join(txt_dir, page_name), mode="w", encoding='UTF-8') as o:
o.write(text)
此代码正常工作,但是对于每个文件,我都有多个页面,当我运行上述代码时,它给我的数据为file1-page1.txt,file1-page2.txt,file1-page3.txt。但我希望file.txt包含所有页面的信息。我该怎么做。
答案 2 :(得分:0)
def getPptContent(path, text):
pdfWriter = PyPDF2.PdfFileWriter()
pdf = open(join(pdf_dir, filename),'rb')
pdfReader = PyPDF2.PdfFileReader(pdf)
for page in range(1, pdfReader.numPages):
pageObj = pdfReader.getPage(page)
pdfWriter.addPage(pageObj)
text = pageObj.extractText()
return text
pdf_dir = "pdf_directory name"
corpus = [str(f) for f in os.listdir(pdf_dir) if not f.startswith('.') and
isfile(join(pdf_dir, f))]
for filename in corpus:
Path = pdf_dir + "/" +filename
print(Path)
file_content = getPptContent(Path)
f = open(pdf_dir + "/output/" + filename.split(".")[0] +".txt" ,"w+",
encoding="utf-8")
f.write(str(file_content))
f.close()
上面的代码对我有用。