以编号格式转换stanford依赖项

时间:2016-09-06 05:01:41

标签: python nlp nltk stanford-nlp dependency-parsing

我正在使用Stanford dependency parser,我得到了句子的以下输出

  

我在睡梦中拍了一头大象

>>>python dep_parsing.py 
[((u'shot', u'VBD'), u'nsubj', (u'I', u'PRP')), ((u'shot', u'VBD'), u'dobj', (u'elephant', u'NN')), ((u'elephant', u'NN'), u'det', (u'an', u'DT')), ((u'shot', u'VBD'), u'nmod', (u'sleep', u'NN')), ((u'sleep', u'NN'), u'case', (u'in', u'IN')), ((u'sleep', u'NN'), u'nmod:poss', (u'my', u'PRP$'))]

但是,我希望编号的标记作为输出it is here

nsubj(shot-2, I-1)
  root(ROOT-0, shot-2)
  det(elephant-4, an-3)
  dobj(shot-2, elephant-4)
  case(sleep-7, in-5)
  nmod:poss(sleep-7, my-6)
  nmod(shot-2, sleep-7)

这是我的代码,直到现在。

  from nltk.parse.stanford import StanfordDependencyParser
  stanford_parser_dir = 'stanford-parser/'
  eng_model_path = stanford_parser_dir + "stanford-parser-models/edu/stanford/nlp/models/lexparser/englishRNN.ser.gz"
  my_path_to_models_jar = stanford_parser_dir + "stanford-parser-3.5.2-models.jar"
  my_path_to_jar = stanford_parser_dir + "stanford-parser.jar"

  dependency_parser = StanfordDependencyParser(path_to_jar=my_path_to_jar, path_to_models_jar=my_path_to_models_jar)

  result = dependency_parser.raw_parse('I shot an elephant in my sleep')
  dep = result.next()
  a = list(dep.triples())
  print a

我怎么能有这样的输出?

1 个答案:

答案 0 :(得分:0)

编写遍历树的递归函数。作为第一遍,只需尝试将数字分配给单词。