Ruby mechanize没有获得完整的内容

时间:2012-02-20 02:24:29

标签: ruby mechanize

我正在使用Mechanize和Nokogiri从这两个站点解析一些乐透结果(它们非常相似): http://www1.caixa.gov.br/loterias/loterias/lotofacil/lotofacil_resultado.asp http://lotofacil.resultadoloteria.org/

这是我的代码:

require 'nokogiri'
require 'mechanize'

agent = Mechanize.new
agent.user_agent_alias = 'Mac Safari'
page = agent.get('http://lotofacil.resultadoloteria.org/')
doc = Nokogiri::HTML(page.body)
doc.xpath('//table[@class="tabela_jogo"]//span').each { |value| puts value }

第二个网站运行正常。结果:

<span id="lfacil1">01</span>
<span id="lfacil2">03</span>
<span id="lfacil3">05</span>
<span id="lfacil4">08</span>
<span id="lfacil5">10</span>
<span id="lfacil6">11</span>
<span id="lfacil7">13</span>
<span id="lfacil8">14</span>
<span id="lfacil9">15</span>
<span id="lfacil10">18</span>
<span id="lfacil11">20</span>
<span id="lfacil12">22</span>
<span id="lfacil13">23</span>
<span id="lfacil14">24</span>
<span id="lfacil15">25</span>

但我无法从第一个获得乐透号码。结果如下:

<span id="lfacil1"></span>
<span id="lfacil2"></span>
<span id="lfacil3"></span>
<span id="lfacil4"></span>
<span id="lfacil5"></span>
<span id="lfacil6"></span>
<span id="lfacil7"></span>
<span id="lfacil8"></span>
<span id="lfacil9"></span>
<span id="lfacil10"></span>
<span id="lfacil11"></span>
<span id="lfacil12"></span>
<span id="lfacil13"></span>
<span id="lfacil14"></span>
<span id="lfacil15"></span>
<span id="lfacil1_2"></span>
<span id="lfacil2_2"></span>
<span id="lfacil3_2"></span>
<span id="lfacil4_2"></span>
<span id="lfacil5_2"></span>
<span id="lfacil6_2"></span>
<span id="lfacil7_2"></span>
<span id="lfacil8_2"></span>
<span id="lfacil9_2"></span>
<span id="lfacil10_2"></span>
<span id="lfacil11_2"></span>
<span id="lfacil12_2"></span>
<span id="lfacil13_2"></span>
<span id="lfacil14_2"></span>
<span id="lfacil15_2"></span>

我认为是Mechanize的一部分,因为p page.body也返回没有乐透号码的内容。有什么想法吗?

感谢。 :)

1 个答案:

答案 0 :(得分:2)

那是因为他们不在那里。我发现它们适合你:

page = agent.get('http://www1.caixa.gov.br/loterias/loterias/lotofacil/lotofacil_pesquisa_new.asp')
numbers = page.body.split('|')[3..17]

也代替了这个:

doc = Nokogiri::HTML(page.body)

机械化已经为你解决了这个问题:

doc = page.parser