我正在尝试这样做,
import glob
interesting_files = glob.glob("/home/tcs/PYTHONMAP/test1/*.csv")
header_saved = False
with open('/home/tcs/PYTHONMAP/output.csv','wb') as fout:
for filename in interesting_files:
with open(filename) as fin:
header = next(fin)
if not header_saved:
fout.write(header)
header_saved = True
for line in fin:
fout.write(line)
并获得
File "/home/tcs/.config/spyder-py3/temp.py", line 11, in <module>
fout.write(header)
TypeError: a bytes-like object is required, not 'str'
我不太了解python请帮忙 另外,我想知道如何将1个大csv拆分为具有相同标题的多个csv。
答案 0 :(得分:1)
使用pandas:
import pandas as pd
interesting_files = glob.glob("/home/tcs/PYTHONMAP/test1/*.csv")
df = pd.concat((pd.read_csv(f, header = 0) for f in interesting_files))
df.to_csv("output.csv")
还要删除重复的行:
import pandas as pd
interesting_files = glob.glob("/home/tcs/PYTHONMAP/test1/*.csv")
df = pd.concat((pd.read_csv(f, header = 0) for f in interesting_files))
df_deduplicated = df.drop_duplicates()
df_deduplicated.to_csv("output.csv")
在创建数据帧时,这不会消除重复,但之后。因此,通过连接所有文件来创建数据帧。然后它被重复数据删除。然后可以将最终的数据帧保存到csv。
答案 1 :(得分:1)
import glob
import csv
interesting_files = glob.glob("/home/tcs/PYTHONMAP/test1/*.csv")
header_saved = False
with open('/home/tcs/PYTHONMAP/output.csv', 'w') as fout:
writer = csv.writer(fout)
for filename in interesting_files:
with open(filename) as fin:
header = next(fin)
if not header_saved:
writer.writerows(header) # you may need to work here. The writerows require an iterable.
header_saved = True
writer.writerows(fin.readlines())