Scrapy X路径:仅在循环中获得第一项

时间:2019-07-03 15:25:34

标签: python-3.x xpath web-scraping scrapy

我正在尝试获取此页面每个元素的详细信息:https://www.mrlodge.de/wohnungen/

我经常通过for循环来做到这一点。但是这一次它只返回第一个元素。循环中必须有一个问题,因为当我使用getall()而不是get()时,我得到了我需要的所有详细信息,但没有排序。

请帮助

Varien/Autoload.php

2 个答案:

答案 0 :(得分:1)

尝试使用

//div[contains(@class,'mrlobject-row')] 

代替

//div[@class='mrl-ft-results mrlobject-list']

获得所需的结果。

答案 1 :(得分:1)

您需要更改第一个xpath查询

class MrlodgeSpiderSpider(scrapy.Spider):
    name = 'mrlodge_spider'

    payload = '''
    {mrl_ft%5Bfd%5D%5Bdate_from%5D=&mrl_ft%5Bfd%5D%5Brent_from%5D=1000&mrl
    _ft%5Bfd%5D%5Brent_to%5D=8500&mrl_ft%5Bfd%5D%5Bpersons%5D=1&mrl_ft%5Bfd
    %5D%5Bkids%5D=0&mrl_ft%5Bfd%5D%5Brooms_from%5D=1&mrl_ft%5Bfd%5D%5Brooms
    _to%5D=9&mrl_ft%5Bfd%5D%5Barea_from%5D=20&mrl_ft%5Bfd%5D%5Barea_to%5D=4
    80&mrl_ft%5Bfd%5D%5Bsterm%5D=&mrl_ft%5Bfd%5D%5Bradius%5D=50&mrl_ft%5Bfd
    %5D%5Bmvv%5D=&mrl_ft%5Bfd%5D%5Bobjecttype_cb%5D%5B%5D=w&mrl_ft%5Bfd%5D%
    5Bobjecttype_cb%5D%5B%5D=h&mrl_ft%5Bpage%5D=1}
'''

    def start_requests(self):
        yield scrapy.Request(
            url='https://www.mrlodge.de/wohnungen/',
            method='POST',
            body=self.payload,
            headers={"content-type": "application/json"},
        )

    def parse(self, response):
        for apartment in response.xpath('//div[@class="mrlobject-list__item mrlobject-row"]'):
            yield {
                'info': apartment.xpath(".//div[@class='obj-smallinfo']/text()").get()
            }