解析xml DTD文件

时间:2014-03-31 10:02:35

标签: python xml-parsing yacc lexer pyparsing

我在实现解析器时很新,我正在尝试解析xml DTD文件,为它生成一个无上下文语法。我试过pyparsing和yacc,但我仍然可以得到任何结果。所以如果有人可以提供一些技巧或示例代码来编写这样的解析器,我将不胜感激。以下是DTD文件示例:

<!DOCTYPE PcSpecs [
<!ELEMENT PCS (PC*)>
<!ELEMENT PC (MODEL, PRICE, PROCESSOR, RAM, DISK+)>
<!ELEMENT MODEL (\#PCDATA)>
<!ELEMENT PRICE (\#PCDATA)>
<!ELEMENT PROCESSOR (MANF, MODEL, SPEED)>
<!ELEMENT MANF (\#PCDATA)>
<!ELEMENT MODEL (\#PCDATA)>
<!ELEMENT SPEED (\#PCDATA)>
<!ELEMENT RAM (\#PCDATA)>
<!ELEMENT DISK (HARDDISK | CD | DVD)>
<!ELEMENT HARDDISK (MANF, MODEL, SIZE)>
<!ELEMENT SIZE (\#PCDATA)>
<!ELEMENT CD (SPEED)>
<!ELEMENT DVD (SPEED)>
]>

提前致谢。

1 个答案:

答案 0 :(得分:1)

这里有一个开始,它会将数据解析为ParseResults数据结构,然后您可以遍历并为定义的doctype创建解析器:

from pyparsing import *

LT,GT,EXCLAM,LBRACK,RBRACK,LPAR,RPAR = map(Suppress,"<>![]()")
DOCTYPE = Keyword("DOCTYPE").suppress()
ELEMENT = Keyword("ELEMENT").suppress()
ident = Word(alphas, alphanums+"_")
elementRef = Group(ident("name") + Optional(oneOf("* +")("rep")))
elementExpr = infixNotation(elementRef,
    [
    (',', 2, opAssoc.LEFT),
    ('|', 2, opAssoc.LEFT),
    ])
PCDATA = Literal(r"\#PCDATA")
elementDefn = Group(LT+EXCLAM + ELEMENT + ident("name") + 
                  LPAR + (elementExpr | PCDATA("PCDATA"))("contents") + RPAR + GT)
doctypeDefn = LT+EXCLAM + DOCTYPE + ident("name") + 
                    LBRACK + ZeroOrMore(elementDefn)("elements") + RBRACK + GT

我开始只使用delimitedList作为每个ELEMENT定义中的元素列表,但后来我注意到了&#39;,&#39;和&#39; |&#39;实际上是运算符,而不仅仅是分隔符,甚至可以混合,如在&#34; A,B,C | D,E&#34;中。所以我使用了pyparsing的infixNotation助手来允许这些定义。

使用您的输入样本,我可以解析并显示结果:

doctype = doctypeDefn.parseString(sample)
print doctype.dump()
for elem in doctype.elements:
    print elem.dump()

,并提供:

['PcSpecs', ['PCS', ['PC', '*']], ['PC', [['MODEL'], ...
- elements: [['PCS', ['PC', '*']], ['PC', [['MODEL'], ...
- name: PcSpecs
['PCS', ['PC', '*']]
- contents: ['PC', '*']
  - name: PC
  - rep: *
- name: PCS
['PC', [['MODEL'], ',', ['PRICE'], ',', ['PROCESSOR'], ',', ['RAM'], ',', ['DISK', '+']]]
- contents: [['MODEL'], ',', ['PRICE'], ',', ['PROCESSOR'], ',', ['RAM'], ',', ['DISK', '+']]
- name: PC
['MODEL', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: MODEL
['PRICE', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: PRICE
['PROCESSOR', [['MANF'], ',', ['MODEL'], ',', ['SPEED']]]
- contents: [['MANF'], ',', ['MODEL'], ',', ['SPEED']]
- name: PROCESSOR
['MANF', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: MANF
['MODEL', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: MODEL
['SPEED', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: SPEED
['RAM', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: RAM
['DISK', [['HARDDISK'], '|', ['CD'], '|', ['DVD']]]
- contents: [['HARDDISK'], '|', ['CD'], '|', ['DVD']]
- name: DISK
['HARDDISK', [['MANF'], ',', ['MODEL'], ',', ['SIZE']]]
- contents: [['MANF'], ',', ['MODEL'], ',', ['SIZE']]
- name: HARDDISK
['SIZE', '\\#PCDATA']
- PCDATA: \#PCDATA
- contents: \#PCDATA
- name: SIZE
['CD', ['SPEED']]
- contents: ['SPEED']
  - name: SPEED
- name: CD
['DVD', ['SPEED']]
- contents: ['SPEED']
  - name: SPEED
- name: DVD