
本文详解如何绕过前端隐藏逻辑,通过捕获浏览器网络请求中的 api 响应,精准提取 maroof.sa(https://maroof.sa/businesses)页面中所有商家的动态生成详情链接。
本文详解如何绕过前端隐藏逻辑,通过捕获浏览器网络请求中的 api 响应,精准提取 maroof.sa(https://maroof.sa/businesses)页面中所有商家的动态生成详情链接。
Maroof.sa 的商家卡片(如“Marwa Store”)看似仅展示名称与基础信息,但其跳转链接并未直接存在于 HTML 中——它由前端 JavaScript 动态拼接生成,且真实数据来源于异步 API 请求 https://maroof.sa/business/search。因此,传统 DOM 解析(如 find_elements(By.CSS_SELECTOR, 'div.storeCard'))无法获取链接,必须转向网络层抓取。
核心原理:监听并解析 XHR 请求响应
Selenium 4+ 支持 Chrome DevTools Protocol(CDP),可通过 performance 日志捕获所有网络请求。关键步骤如下:
- 启用性能日志记录;
- 等待页面加载并触发搜索请求(通常在访问 /businesses 后自动发起);
- 筛选包含 business/search 的 Network.responseReceived 日志项;
- 使用 Network.getResponseBody 获取原始 JSON 响应;
- 从中提取 items[].id,拼接为标准详情链接:https://maroof.sa/details/{id}。
完整可运行代码(含反检测优化)
import json
import time
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
def setup_stealth_driver():
options = Options()
options.set_capability('goog:loggingPrefs', {'performance': 'ALL'})
# 关键反检测配置
options.add_argument("--no-sandbox")
options.add_argument("--disable-gpu")
options.add_argument('--disable-blink-features=AutomationControlled')
options.add_argument('--disable-dev-shm-usage')
options.add_argument("--enable-javascript")
options.add_argument("--enable-cookies")
options.add_experimental_option("useAutomationExtension", False)
options.add_experimental_option("excludeSwitches", ["enable-automation"])
# 隐藏 WebDriver 特征(可选增强)
options.add_argument("--disable-extensions")
options.add_argument("--disable-plugins-discovery")
return webdriver.Chrome(options=options)
def extract_business_links():
driver = setup_stealth_driver()
url = "https://maroof.sa/businesses"
try:
driver.get(url)
time.sleep(4) # 确保 search 请求已发出
# 获取性能日志
logs = driver.get_log("performance")
target_api = "business/search"
links = []
for log in logs:
try:
message = json.loads(log["message"])
if "Network.responseReceived" not in message["message"]["method"]:
continue
response = message["message"]["params"]["response"]
if target_api not in response["url"]:
continue
# 获取响应体
request_id = message["message"]["params"]["requestId"]
body = driver.execute_cdp_cmd(
"Network.getResponseBody",
{"requestId": request_id}
)
data = json.loads(body["body"])
# 提取所有商家 ID 并构造链接
for item in data.get("items", []):
link = f"{url}/details/{item['id']}"
links.append(link)
print(link)
except (KeyError, json.JSONDecodeError, Exception) as e:
continue # 跳过无效日志项
return links
finally:
driver.quit()
# 执行提取
if __name__ == "__main__":
business_links = extract_business_links()
print(f"\n✅ 成功提取 {len(business_links)} 个商家链接")注意事项与建议
-
稳定性问题:Maroof.sa 启用了强反爬机制(如 Cloudflare 或自定义 Bot 检测),即使使用 stealth 配置,仍可能被拦截。建议:
- 添加随机延迟(time.sleep(3–6));
- 使用可信住宅代理(如 BrightData、Oxylabs);
- 考虑 undetected-chromedriver 替代方案(需单独安装);
- API 响应结构变化风险:若 items 字段名或 URL 模式变更(如 /details/ → /store/),需及时调整代码;
- 法律与合规性:请遵守 robots.txt(https://www.php.cn/link/93eeb56ae2930647f2a428006b81d7ec)及网站服务条款,仅用于个人学习或授权用途;
- 备用方案:若 CDP 方案失效,可尝试直接调用 https://maroof.sa/business/search 接口(需分析请求头、Cookie 及 CSRF Token,复杂度更高)。
通过该方法,你不再依赖不可靠的 DOM 结构,而是直击数据源头,实现稳定、可扩展的链接提取。


















