python 读写文件包含多种编码格式的解决方式

更新时间：2019年12月20日 15:41:01 作者：hm11290219

今天小编就为大家分享一篇python 读写文件包含多种编码格式的解决方式，具有很好的参考价值，希望对大家有所帮助。一起跟随小编过来看看吧

今天写一个脚本文件，需要将多个文件中的内容汇总到一个txt文件中，由于多个文件有三种不同的编码方式，读写出现错误，先将解决方法记录如下：

# -*- coding: utf-8 -*-

import wave

import pylab as pl

import numpy as np

import pandas as pd

import os

import time

import datetime

import arrow

import chardet

import sys

reload(sys)

sys.setdefaultencoding('utf8')

os.chdir("F:/new_srt")

#get words of srt file

###########################################

def get_word():

path = "F:/new_srt"

filelist = os.listdir(path)

for files in filelist:

print files

encoding = chardet.detect(open(files,'r').read())['encoding']

if encoding == 'utf-8':

data=pd.read_csv(files,encoding="utf-8",sep='\r',header=None)

elif encoding == 'GB2312':

try:

data=pd.read_csv(files,encoding="gbk",sep='\r',header=None)

except UnicodeDecodeError:

data=pd.read_csv(files,encoding="utf-8",sep='\r',header=None)

elif encoding == 'UTF-8-SIG':

data=pd.read_csv(files,encoding="UTF-8-SIG",sep='\r',header=None)

else:

print 'this is an error about %s' % files

data_new=pd.DataFrame(np.reshape(data.values, (-1,3)))

data_new.columns=['index','timecut','content']

filename = os.path.splitext(files)[0] #filetype = os.path.splitext(files)[1]

with open('F:/result.txt', 'a') as file:

file.write(str(filename)+' ' )

for item in data_new['content']:

file.write(item.decode("utf-8") +' ') #s=s.decode("utf-8")

file.write('\n')

if __name__ == '__main__':

get_word()

以上这篇python 读写文件包含多种编码格式的解决方式就是小编分享给大家的全部内容了，希望能给大家一个参考，也希望大家多多支持脚本之家。

本站仅提供存储服务，所有内容均由用户发布，如发现有害或侵权内容，请点击举报。