selenium 在一个特殊 class、href link 下搜索项目

selenium search the item under one special class, href link

我正在尝试为网站 https://web.archive.org/web/*/https://cd.lianjia.com/ 获取一些 link,我想获取每个日期的 link,

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.support.ui import WebDriverWait

path = 'C:\Windows\chromedriver.exe'
driver = webdriver.Chrome(path)
driver.implicitly_wait(10)
driver.get('https://web.archive.org/web/*/https://cd.lianjia.com/ ')


element =driver.find_element_by_xpath('/html/body/div[4]/div[3]/div/div[2]/span[21]')
actions = ActionChains(driver)
actions.move_to_element(element).click().perform()

因为它有几年的数据,所以我尝试了这个特定的年份,2016,现在我试图获取每个日期的 href link,

  day = driver.find_elements_by_css_selector("div.calendar-day")
for a in day:
    print(a.text)

但它只有 return 天数

      3
5
6
7
27
6
14
25
26
28
7
8
21
27
1
9
25
15
19
21
4
17
21
27
13
29
15
19
20
26
2
4

然后我尝试使用这段代码找到 link

date = driver.find_elements_by_css_selector("a[href* = 'web']")
date = driver.find_elements_by_css_selector("a[href* = 'lianjia']")

但是找不到link,谁能帮我解决这个问题,我在这里停了五天

你的做法是正确的。但需要一些调整。您将需要直接定位元素并使用 get attribute 方法提取 hrefs。

days = driver.find_elements_by_css_selector("div.calendar-day > a")
for day in days:
    print(day.get_attribute('href'))

输出:

https://web.archive.org/web/20160103/https://cd.lianjia.com/
https://web.archive.org/web/20160105/https://cd.lianjia.com/
https://web.archive.org/web/20160106/https://cd.lianjia.com/
https://web.archive.org/web/20160107/https://cd.lianjia.com/
https://web.archive.org/web/20160127/https://cd.lianjia.com/
https://web.archive.org/web/20160306/https://cd.lianjia.com/
https://web.archive.org/web/20160314/https://cd.lianjia.com/
https://web.archive.org/web/20160325/https://cd.lianjia.com/
https://web.archive.org/web/20160326/https://cd.lianjia.com/
https://web.archive.org/web/20160328/https://cd.lianjia.com/
https://web.archive.org/web/20160407/https://cd.lianjia.com/
https://web.archive.org/web/20160408/https://cd.lianjia.com/
https://web.archive.org/web/20160421/https://cd.lianjia.com/
https://web.archive.org/web/20160427/https://cd.lianjia.com/
https://web.archive.org/web/20160501/https://cd.lianjia.com/
https://web.archive.org/web/20160509/https://cd.lianjia.com/
https://web.archive.org/web/20160525/https://cd.lianjia.com/
https://web.archive.org/web/20160615/https://cd.lianjia.com/
https://web.archive.org/web/20160619/https://cd.lianjia.com/
https://web.archive.org/web/20160621/https://cd.lianjia.com/
https://web.archive.org/web/20160704/https://cd.lianjia.com/
https://web.archive.org/web/20160717/https://cd.lianjia.com/
https://web.archive.org/web/20160721/https://cd.lianjia.com/
https://web.archive.org/web/20160927/https://cd.lianjia.com/
https://web.archive.org/web/20161013/https://cd.lianjia.com/
https://web.archive.org/web/20161029/https://cd.lianjia.com/
https://web.archive.org/web/20161115/https://cd.lianjia.com/
https://web.archive.org/web/20161119/https://cd.lianjia.com/
https://web.archive.org/web/20161120/https://cd.lianjia.com/
https://web.archive.org/web/20161126/https://cd.lianjia.com/
https://web.archive.org/web/20161202/https://cd.lianjia.com/
https://web.archive.org/web/20161204/https://cd.lianjia.com/