为什么我收到此 Python 脚本的连接拒绝异常？

Question

我正在编写一个 Python 脚本来使用请求模块从 azlyrics 中获取歌曲的歌词。这是我写的脚本：

import requests, re
from bs4 import BeautifulSoup as bs
url = "http://search.azlyrics.com/search.php"
payload = {'q' : 'shape of you'}
r = requests.get(url, params = payload)
soup = bs(r.text,"html.parser")
try:
    link = soup.find('a', {'href':re.compile('http://www.azlyrics.com/lyrics/edsheeran/shapeofyou.html')})['href']
    link = link.replace('http', 'https')
    print(link)
    raw_data = requests.get(link)
except Exception as e: 
    print(e)

但我收到一个异常说明：

Max retries exceeded with url: /lyrics/edsheeran/shapeofyou.html (Caused by NewConnectionError('<requests.packages.urllib3.connection.VerifiedHTTPSConnection object at 0x7fbda00b37f0>: Failed to establish a new connection: [Errno 111] Connection refused',))

我在 Internet 上看到我可能尝试发送太多请求。所以我让脚本休眠了一段时间：

import requests, re
from bs4 import BeautifulSoup as bs
from time import sleep
url = "http://search.azlyrics.com/search.php"
payload = {'q' : 'shape of you'}
r = requests.get(url, params = payload)
soup = bs(r.text,"html.parser")
try:
    link = soup.find('a', {'href':re.compile('http://www.azlyrics.com/lyrics/edsheeran/shapeofyou.html')})['href']
    link = link.replace('http', 'https')
    sleep(60)
    print(link)
    raw_data = requests.get(link)
except Exception as e: 
    print(e)

但运气不好！

所以我尝试了 urllib.request

import requests, re
from bs4 import BeautifulSoup as bs
from time import sleep
from urllib.request import urlopen
url = "http://search.azlyrics.com/search.php"
payload = {'q' : 'shape of you'}
r = requests.get(url, params = payload)
soup = bs(r.text,"html.parser")
try:
    link = soup.find('a', {'href':re.compile('http://www.azlyrics.com/lyrics/edsheeran/shapeofyou.html')})['href']
    link = link.replace('http', 'https')
    sleep(60)
    print(link)
    raw_data = urlopen(link).read()
except Exception as e: 
    print(e)

但随后得到不同的异常说明：

<urlopen error [Errno 111] Connection refused>

任何人都可以告诉我它有什么问题以及如何解决它吗？

Answer 1

在您的网络浏览器中尝试一下；当您尝试访问 http://www.azlyrics.com/lyrics/edsheeran/shapeofyou.html it'll work fine, but when you try to visit https://www.azlyrics.com/lyrics/edsheeran/shapeofyou.html 时，它不起作用。

所以请删除您的 link = link.replace('http', 'https') 行并重试。

为什么我收到此 Python 脚本的连接拒绝异常？

Why am I getting connection refused exception for this Python script?

python

urllib

beautifulsoup

web-scraping

python-requests