如何从 IMDB 网站上抓取制作公司的名称
How do I web scrape the names of the production companies from IMDB website
我需要抓取一些电影的制片公司的名称。我一直尝试使用锚标记 a
和包含名称的 class 但它不会 return 生产公司。
URL : https://www.imdb.com/title/tt0473553/?ref_=fn_al_tt_1
这是我要抓取的网站的 HTML 部分:
<section class="ipc-page-section ipc-page-section--base">
<div data-testid="title-details-section" class="styles__MetaDataContainer-sc-12uhu9s-0 cgqHBf">
<ul>
<li role="presentation" class="ipc-metadata-list__item ipc-metadata-list-item--link" data-testid="title-details-companies"><a class="ipc-metadata-list-item__label ipc-metadata-list-item__label--link" rel="" href="/title/tt0473553/companycredits?ref_=tt_dt_co" target="">Production companies</a>
<div class="ipc-metadata-list-item__content-container">
<ul class="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base" role="presentation">
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">IDT Entertainment</a>
</li>
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0142161?ref_=tt_dt_co_2">New Arc Entertainment</a>
</li>
</ul>
</div>
</li>
</ul>
</div>
</section>
这是我尝试过的:
import requests
from bs4 import BeautifulSoup
movie_url="https://www.imdb.com/title/tt0473553/?ref_=fn_al_tt_1"
movie_page = requests.get(movie_url)
soup = BeautifulSoup(page.text, 'html.parser')
#movies_comp = soup.find_all("li", class_="ipc-inline-list__item")
movies_comp = soup.find_all("a", class_="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link")
print(movies_comp)
我没有得到理想的输出。我期望它 return 的输出是这样的:
['IDT Entertainment', 'New Arc Entertainment']
以下是您可以尝试的方法:
import requests
from bs4 import BeautifulSoup
page=requests.get("https://www.imdb.com/title/tt0473553/?ref_=fn_al_tt_1")
page="""
<section class="ipc-page-section ipc-page-section--base">
<div data-testid="title-details-section" class="styles__MetaDataContainer-sc-12uhu9s-0 cgqHBf">
<ul>
<li role="presentation" class="ipc-metadata-list__item ipc-metadata-list-item--link" data-testid="title-details-companies"><a class="ipc-metadata-list-item__label ipc-metadata-list-item__label--link" rel="" href="/title/tt0473553/companycredits?ref_=tt_dt_co" target="">Production companies</a>
<div class="ipc-metadata-list-item__content-container">
<ul class="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base" role="presentation">
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">IDT Entertainment</a>
</li>
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0142161?ref_=tt_dt_co_2">New Arc Entertainment</a>
</li>
</ul>
</div>
</li>
</ul>
</div>
</section>
"""
soup=BeautifulSoup(page,"lxml")
# To understand this is then structur of the data you want to extract :
# <li role="presentation" class="ipc-metadata-list__item ipc-metadata-list-item--link" data-testid="title-details-companies">
# <ul class="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base" role="presentation"><li role="presentation" class="ipc-inline-list__item"><a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">
# <a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">IDT Entertainment</a>
# <a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0142161?ref_=tt_dt_co_2">New Arc Entertainment</a>
print([a.text for a in soup.find("li",attrs={'class':r'ipc-metadata-list__item ipc-metadata-list-item--link','data-testid':r'title-details-companies'})
.find("ul",class_="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base")
.find_all("a")])
输出:
['IDT Entertainment', 'New Arc Entertainment']
<a>
class
所以,你得到了多个。
我需要抓取一些电影的制片公司的名称。我一直尝试使用锚标记 a
和包含名称的 class 但它不会 return 生产公司。
URL : https://www.imdb.com/title/tt0473553/?ref_=fn_al_tt_1
这是我要抓取的网站的 HTML 部分:
<section class="ipc-page-section ipc-page-section--base">
<div data-testid="title-details-section" class="styles__MetaDataContainer-sc-12uhu9s-0 cgqHBf">
<ul>
<li role="presentation" class="ipc-metadata-list__item ipc-metadata-list-item--link" data-testid="title-details-companies"><a class="ipc-metadata-list-item__label ipc-metadata-list-item__label--link" rel="" href="/title/tt0473553/companycredits?ref_=tt_dt_co" target="">Production companies</a>
<div class="ipc-metadata-list-item__content-container">
<ul class="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base" role="presentation">
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">IDT Entertainment</a>
</li>
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0142161?ref_=tt_dt_co_2">New Arc Entertainment</a>
</li>
</ul>
</div>
</li>
</ul>
</div>
</section>
这是我尝试过的:
import requests
from bs4 import BeautifulSoup
movie_url="https://www.imdb.com/title/tt0473553/?ref_=fn_al_tt_1"
movie_page = requests.get(movie_url)
soup = BeautifulSoup(page.text, 'html.parser')
#movies_comp = soup.find_all("li", class_="ipc-inline-list__item")
movies_comp = soup.find_all("a", class_="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link")
print(movies_comp)
我没有得到理想的输出。我期望它 return 的输出是这样的:
['IDT Entertainment', 'New Arc Entertainment']
以下是您可以尝试的方法:
import requests
from bs4 import BeautifulSoup
page=requests.get("https://www.imdb.com/title/tt0473553/?ref_=fn_al_tt_1")
page="""
<section class="ipc-page-section ipc-page-section--base">
<div data-testid="title-details-section" class="styles__MetaDataContainer-sc-12uhu9s-0 cgqHBf">
<ul>
<li role="presentation" class="ipc-metadata-list__item ipc-metadata-list-item--link" data-testid="title-details-companies"><a class="ipc-metadata-list-item__label ipc-metadata-list-item__label--link" rel="" href="/title/tt0473553/companycredits?ref_=tt_dt_co" target="">Production companies</a>
<div class="ipc-metadata-list-item__content-container">
<ul class="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base" role="presentation">
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">IDT Entertainment</a>
</li>
<li role="presentation" class="ipc-inline-list__item">
<a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0142161?ref_=tt_dt_co_2">New Arc Entertainment</a>
</li>
</ul>
</div>
</li>
</ul>
</div>
</section>
"""
soup=BeautifulSoup(page,"lxml")
# To understand this is then structur of the data you want to extract :
# <li role="presentation" class="ipc-metadata-list__item ipc-metadata-list-item--link" data-testid="title-details-companies">
# <ul class="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base" role="presentation"><li role="presentation" class="ipc-inline-list__item"><a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">
# <a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0136980?ref_=tt_dt_co_1">IDT Entertainment</a>
# <a class="ipc-metadata-list-item__list-content-item ipc-metadata-list-item__list-content-item--link" rel="" href="/company/co0142161?ref_=tt_dt_co_2">New Arc Entertainment</a>
print([a.text for a in soup.find("li",attrs={'class':r'ipc-metadata-list__item ipc-metadata-list-item--link','data-testid':r'title-details-companies'})
.find("ul",class_="ipc-inline-list ipc-inline-list--show-dividers ipc-inline-list--inline ipc-metadata-list-item__list-content base")
.find_all("a")])
输出:
['IDT Entertainment', 'New Arc Entertainment']
<a>
class
所以,你得到了多个。