我正在学习Python,主要用于文本挖掘,遵循(http://textminingonline.com/training-word2vec-model-on-english-wikipedia-by-gensim)的指导。我想从api返回的xml中提取维基百科英文文本。但是,会出现错误:
print(globals()['__doc__'] %locals())
TypeError: unsupported operand type(s) for %: 'NoneType' and 'dict'".
任何人都可以提供有关如何解决此问题的任何提示吗?我是否需要将outp
和inp
替换为文件地址?
提前致谢。我附上了代码:
import os
import logging
import sys
from gensim.corpora import WikiCorpus
if __name__=='__main__':
program = os.path.basename(sys.argv[0])
logger = logging.getLogger(program)
logging.basicConfig(format='%(asctime)s: %(levelname)s: %(message)s')
logging.root.setLevel(level=logging.INFO)
if len(sys.argv) < 3:
print(globals()['__doc__'] %locals())
sys.exit(1)
inp, outp = sys.argv[1:3]
space = ' '
i = 0
output = open(outp, 'w')
wiki = WikiCorpus(inp, lemmatize=False, dictionary={})
for text in wiki.get_texts():
output.write(space.join(text) + '\n')
i = i + 1
if i % 10000 == 0:
logger.info('Saved ' + str(i) + ' articles')
output.close()
logger.info('Finished ' + str(i) + ' articles')