使用python的Web抓取:urlopen返回HTTP错误403:禁止

时间:2020-03-27 00:13:13

标签: python python-requests urlopen

我正在尝试使用urlopen从Fragantica.com下载数据,但是即使更改了用户代理并添加了标头,也会出现错误(“ HTTP错误403:禁止访问”)。 我也从这里尝试过代码,但没有成功(http://wolfprojects.altervista.org/changeua.php#problem)。

这是我的代码:

import urllib.request

user_agent = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_2) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/13.0.4 Safari/605.1.15'



url = "https://www.fragrantica.com/perfume/Tom-Ford/Tobacco-Vanille-1825.html"
headers={'User-Agent':user_agent,} 

request=urllib.request.Request(url,None,headers) #The assembled request
response = urllib.request.urlopen(request)
data = response.read() # The data u need

这是我遇到的错误:HTTPError:HTTP错误403:禁止

1 个答案:

答案 0 :(得分:0)

您可能需要指定更多标题,请尝试以下操作:

import urllib.request    

url = "https://www.fragrantica.com/perfume/Tom-Ford/Tobacco-Vanille-1825.html"
headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_2) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/13.0.4 Safari/605.1.15',
       'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
       'Accept-Charset': 'ISO-8859-1,utf-8;q=0.7,*;q=0.3',
       'Accept-Encoding': 'none',
       'Accept-Language': 'en-US,en;q=0.8',
       'Connection': 'keep-alive'} 

request=urllib.request.Request(url=url, headers=headers) #The assembled request
response = urllib.request.urlopen(request)
data = response.read() # The data u need