
本文详解如何利用 BeautifulSoup 的 select 和 select_one 方法,从 tarifaluzhora.es 网站中结构化提取每小时电价(含时间区间、价格数值及颜色等级),并整理为 Pandas DataFrame,便于后续集成至 WhatsApp 自动化推送系统。
本文详解如何利用 beautifulsoup 的 `select` 和 `select_one` 方法,从 tarifaluzhora.es 网站中结构化提取每小时电价(含时间区间、价格数值及颜色等级),并整理为 pandas dataframe,便于后续集成至 whatsapp 自动化推送系统。
在网页爬虫开发中,精准定位并提取目标字段(如时间与价格)是关键一步。针对 https://www.php.cn/link/d0b416a4ccac7ca256e9e0b8d137ec0b 这类结构清晰但嵌套较深的电价页面,直接使用 find_all('div', class_=...) 容易匹配到冗余容器,反而增加解析难度。更高效的方式是基于语义化 HTML 属性(如 itemprop)进行 CSS 选择器定位。
该网站使用 Microdata 标准标记关键信息:
- 时间区间由 <span itemprop="description"> 包裹(如 20:00 - 21:00);
- 价格数值由 <span itemprop="price"> 包裹(如 0.06643 €/kWh);
- 颜色等级(high/low/default)则隐含在上级 <div> 的 class 属性中(如 template-tlh__background-color-high)。
以下代码实现了稳健、可维护的数据提取逻辑:
import pandas as pd
import requests
from bs4 import BeautifulSoup
url = "https://www.php.cn/link/d0b416a4ccac7ca256e9e0b8d137ec0b"
response = requests.get(url)
response.raise_for_status() # 确保请求成功
soup = BeautifulSoup(response.content, "html.parser")
all_data = []
# 使用 CSS 选择器精准定位:排除嵌套的 .row,只选含 price 微数据的顶层行
for row in soup.select(".row:not(:has(.row)):has([itemprop=price])"):
time_elem = row.select_one('[itemprop="description"]')
price_elem = row.select_one('[itemprop="price"]')
if not (time_elem and price_elem):
continue # 跳过缺失任一字段的异常项
# 通过向上查找获取颜色标识(基于 class 名中的关键词)
circle_div = time_elem.find_previous("div", class_=lambda c: c and "circle" in c)
color_class = circle_div["class"][-1] if circle_div else "default"
color = "high" if "high" in color_class else \
"low" if "low" in color_class else "default"
# 清理价格文本中的换行符,确保格式统一
price_text = price_elem.get_text(strip=True).replace("\n", " ")
all_data.append({
"Time": time_elem.get_text(strip=True),
"Price": price_text,
"Color": color
})
df = pd.DataFrame(all_data)
print(df.head(10)) # 查看前10条记录✅ 关键技巧说明:
- ".row:not(:has(.row)):has([itemprop=price])" 是核心选择器:它排除了嵌套的 .row(避免误匹配子容器),同时确保只选取包含 itemprop="price" 的有效行块,大幅提升鲁棒性;
- find_previous("div", class_=...) 比硬编码层级遍历更可靠,能适应 HTML 结构微调;
- get_text(strip=True) 替代 .text 可自动去除首尾空白与多余换行,提升数据洁净度;
- 添加 response.raise_for_status() 和空值校验,使脚本在生产环境中更健壮。
? 注意事项:
- 该网站无反爬机制,但建议添加 headers={'User-Agent': 'Mozilla/5.0...'} 模拟浏览器访问,避免被服务端拦截;
- 若需定时抓取(如每小时更新),应加入异常重试与日志记录;
- 后续对接 whatsapp-web.js 时,可将 df[df['Color']=='high'] 筛选出高价时段,生成自然语言消息(如 "⚠️ 高价时段提醒:20:00-21:00 为 0.06643 €/kWh"),再通过 WebSocket 发送。
此方案兼顾简洁性与可扩展性,是典型「语义优先」式网页解析的最佳实践。

















