Python 操作pdf文件(pdfplumber读取PDF写入Excel)

2023-09-05 16:03| 来源: 网络整理| 查看: 265

Python 操作pdf(pdfplumber读取PDF写入Excel)

文章目录

1. Python 操作pdf(pdfplumber读取PDF写入Excel)

1.1 安装pdfplumber模块库:

1.2 常用操作

1.2.1 Python读取pdf文件案例

1.2.2 Python读取pdf文件代码

1.2.3 Python读取pdf文件存入Excel代码

1. Python 操作pdf(pdfplumber读取PDF写入Excel)

1.1 安装pdfplumber模块库:安装pdfplumber: pip install pdfplumber

pdfplumber.PDF类

pdfplumber.PDF类表示单个PDF ,并具有两个主要属性:

属性说明pdf.metadata从PDF的Info中获取元数据键/值对字典。通常包括"CreationDate，“ModDater"，"Producer"等pdf.pages返回一个包含pdfplumber. Page实例的列表,每一一个实例代表PDF每一页的信息

pdfplumber.Page类

pdfplumber.Page类常用属性

常用方法

1.2 常用操作

PDF是Portable Document Format的缩写，这类文件通常使用.pdf作为其扩展名。在日常开发工作中，最容易遇到的就是从PDF中读取文本内容以及用已有的内容生成PDF文档这两个任务。

1.读取pdf文档信息 2.输出总页数 3.读取第一页宽度、高度等信息 4.读取文本第一页加载pdf pdfplumber.open( "路径/文件名. pdf".pas sword="test "laparams={ "line_ _overlap'”0.7 }) password : 要加载受密码保护的PDF ,请传递password关键字参数 laparams :要将布局分析参数设置为pdfminer. six的布局引擎,请传递laparams关键字参数1.2.1 Python读取pdf文件案例

pdf文件如下

微信图片_20221012221249.png

1.2.2 Python读取pdf文件代码import pdfplumber # 加载pdf path = "C:/Users/Administrator/Desktop/test08/test11 - 多页.pdf" with pdfplumber.open(path) as pdf: print(pdf) print(type(pdf)) # 读取pdf文档信息 print("pdf文档信息:", pdf.metadata) # 输出总页数 print("pdf文档总页数:", len(pdf.pages)) # 1.读取第一页宽度、高度等信息 first_page = pdf.pages[0] # pdfplumber.Page对象第一页 # 查看页码 print('pdf页码:', first_page.page_number) # 查看页宽 print('pdf页宽:', first_page.width) # 查看页高 print('pdf页高:', first_page.height) # 2.读取文本第一页 first_page = pdf.pages[0] # pdfplumber.Page对象第一页 text = first_page.extract_text() print(text) 执行结果： "D:\Program Files1\Python\python.exe" D:/Pycharm-work/pythonTest/打卡/0811读取pdf.py pdf文档信息: {'Author': '', 'Comments': '', 'Company': '', 'CreationDate': "D:20220812102327+02'23'", 'Creator': 'WPS 表格', 'Keywords': '', 'ModDate': "D:20220812102327+02'23'", 'Producer': '', 'SourceModified': "D:20220812102327+02'23'", 'Subject': '', 'Title': '', 'Trapped': 'False'} pdf文档总页数: 2 pdf页码: 1 pdf页宽: 595.25 pdf页高: 841.85 姓名年龄性别地址学习技能张三 20 女北京 python 李四 25 男深圳 java 赵五 28 男上海 C++ 孙六 23 女广州 python 钱七 27 男珠海 python 张101 20 女北京 python ....... ....... 张150 27 男珠海 python 张151 20 女北京 python 张152 25 男深圳 java Process finished with exit code 0 1.2.3 Python读取pdf文件存入Excel代码import pdfplumber import xlwt # 加载pdf path = "C:/Users/Administrator/Desktop/test08/test11 - 多页.pdf" with pdfplumber.open(path) as pdf: page_1 = pdf.pages[0] # pdf第一页 table_1 = page_1.extract_table() # 读取表格数据 print(table_1) # 1.创建Excel对象 workbook = xlwt.Workbook(encoding='utf8') # 2.新建sheet表 worksheet = workbook.add_sheet('Sheet1') # 3.自定义列名 clo1 = table_1[0] # 4.将列表元组clo1写入sheet表单中的第一行 for i in range(0, len(clo1)): worksheet.write(0, i, clo1[i]) # 5.将数据写进sheet表单中 for i in range(0, len(table_1[1:])): data = table_1[1:][i] for j in range(0, len(clo1)): worksheet.write(i + 1, j, data[j]) # 保存Excel文件分两种 workbook.save('test88.xls')

执行结果：

微信图片_20221012221307.png

【本文地址】

公司简介

联系我们