AttributeError:'str'对象没有属性'xpath'

时间:2019-08-08 17:19:50

标签: python xpath scrapy

使用Python 3,Scrapy 1.7.3来 正在使用以下链接Scrapy - Extract items from table

但是它给了我AttributeError错误:'str'对象没有属性'xpath'

    <table border="1" cellspacing="0" class="GridViewStyle" id="ctl00_BodyContents_subheading_gridview" rules="all" style="border-collapse:collapse;">
<tbody><tr class="GridViewHeaderStyle" style="background-color:#66B6F4;">
<th scope="col">
<span id="ctl00_BodyContents_subheading_gridview_ctl01_SUBHEADING_CODES_HEADING" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">HS-Code</span>
</th><th scope="col">
<span id="ctl00_BodyContents_subheading_gridview_ctl01_SUBHEADING_DESCRIPTION_HEADING" style="padding:20px 20px 20px 5px;font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;margin:2px">Item Description</span>
</th>
</tr><tr class="GridViewRowStyle">
<td style="width:15%;">
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl02_SUBHEADING_CODES" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td><td style="width:85%;">
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl02_SUBHEADING_DESCRIPTION" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td>
</tr><tr class="GridViewAlternatingRowStyle">
<td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl03_SUBHEADING_CODES" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td><td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl03_SUBHEADING_DESCRIPTION" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td>
</tr><tr class="GridViewRowStyle">
<td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl04_SUBHEADING_CODES" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td><td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl04_SUBHEADING_DESCRIPTION" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td>
</tr><tr class="GridViewAlternatingRowStyle">
<td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl05_SUBHEADING_CODES" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td><td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl05_SUBHEADING_DESCRIPTION" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td>
</tr><tr class="GridViewRowStyle">
<td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl06_SUBHEADING_CODES" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td><td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl06_SUBHEADING_DESCRIPTION" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td>
</tr><tr class="GridViewAlternatingRowStyle">
<td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl07_SUBHEADING_CODES" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td><td>
<a href="http://link.domain" id="ctl00_BodyContents_subheading_gridview_ctl07_SUBHEADING_DESCRIPTION" style="font-family: Helvetica Neue,Helvetica,Arial,sans-serif !important;font-size: 14px;">value1</a>
</td>
</tr>
</tbody></table>

草率代码

# -*- coding: utf-8 -*-
import scrapy
class CybexbotSpider(scrapy.Spider): 
   name = 'cybexbot'
   allowed_domains = ['http://links.com']
   start_urls = ['http://links.com']
   def parse(self, response):
       data=response.xpath('//tr[contains(@class,"GridView")]').extract()
       for d in data[1:]:
         print(type(d))
         temp=dict()
         temp['Code']=d.xpath('tr//td[1]/a/text()').extract()
         temp['Desc']=d.xpath('tr//td[2]/a/text()').extract()
         yield temp

创建临时字典并产生其值

我遇到的错误是

  temp['Code']=d.xpath('tr//td[1]/a/text()').extract()
AttributeError: 'str' object has no attribute 'xpath'

2 个答案:

答案 0 :(得分:1)

尝试一下:

import scrapy
class CybexbotSpider(scrapy.Spider): 
   name = 'cybexbot'
   allowed_domains = ['http://links.com']
   start_urls = ['http://links.com']
   def parse(self, response):
       data=response.xpath('//tr[contains(@class,"GridView")]')
       for d in data[1:]:
         print(type(d))
         temp=dict()
         temp['Code']=d.xpath('tr//td[1]/a/text()').extract()
         temp['Desc']=d.xpath('tr//td[2]/a/text()').extract()
         yield temp

一旦提取它,它就会变成一个字符串,因此库将无法再对其进行处理

答案 1 :(得分:1)

我相信您需要这样的东西(注意我如何使用 relative XPath来获取值):

New-AzureRmResourceGroupDeployment -ResourceGroupName $resourceGroupName -Name $deploymentName -TemplateFile $templateFilePath -TemplateParameterObject $params;