必须显式设置 timeout 并拆分为连接与读取超时,区分 HTTP 状态码业务含义,用 logging+SMTP 实现邮件告警,schedule 适合简单轮询但需进程守护,故障定义须与业务对齐。

用 requests 发起请求并捕获超时与连接异常
状态码检测的前提是请求能发出去、有响应;但生产环境里最常卡在 DNS 解析失败、目标服务宕机或网络策略拦截,这时 requests.get() 会卡住或抛出异常,而不是返回 response.status_code。必须显式设置 timeout,且推荐拆成连接超时和读取超时:
try:
resp = requests.get(url, timeout=(3, 5)) # (connect_timeout, read_timeout)
except requests.exceptions.Timeout:
# 处理超时:可能是服务无响应或网络拥塞
except requests.exceptions.ConnectionError:
# 处理无法建立 TCP 连接(如端口关闭、域名不可达)
except requests.exceptions.RequestException as e:
# 兜底其他异常(如 SSL 验证失败)
不设 timeout 是脚本在告警系统中“静默失效”的头号原因。
区分 HTTP 状态码的业务含义而非仅看是否为 200
很多接口返回 200 但 body 里是 {"code": 500, "msg": "DB unavailable"};也有健康检查接口约定返回 204 或 302 表示正常。只判断 resp.status_code == 200 容易漏报。
- 先检查
resp.status_code是否在预期范围内(如[200, 204, 302]) - 再根据接口文档解析响应体,校验关键字段(如
resp.json().get("status") == "ok") - 对 HEAD 请求,
status_code就是全部依据,无需解析 body
requests 会自动跟随,若需检测原始响应码,加参数 allow_redirects=False。用 logging + smtplib 实现轻量级邮件告警
告警不能只靠 print,得持久化记录并触达人。Python 标准库足够支撑基础场景:
import logging
from logging.handlers import SMTPHandler
<p>mail_handler = SMTPHandler(
mailhost=("smtp.example.com", 587),
fromaddr="alert@myapp.com",
toaddrs=["ops@myapp.com"],
subject="API Health Check Alert",
credentials=("user", "app_password"), # 避免用邮箱明文密码
secure=() # 启用 TLS
)
mail_handler.setLevel(logging.ERROR)
logging.getLogger().addHandler(mail_handler)</p><h1>检测到异常时:</h1><p>logging.error(f"URL {url} returned {resp.status_code}, elapsed {resp.elapsed.total_seconds():.2f}s")
注意:SMTP 认证必须用应用专用密码(如 Gmail 的 App Password),且部分企业邮件网关会拦截未认证或低信誉 IP 发出的告警邮件。
用 schedule 库替代 crontab 做简单轮询调度
如果脚本只跑在单台服务器上,用 schedule 比配 crontab 更直观可控:
import schedule
import time
<p>def check_api():
try:
resp = requests.get("<a href="https://www.php.cn/link/1c9261a8e404615b4bb8bdaee81fa0d7">https://www.php.cn/link/1c9261a8e404615b4bb8bdaee81fa0d7</a>", timeout=(3, 5))
if resp.status_code not in [200, 204]:
logging.error(f"Unhealthy status: {resp.status_code}")
except Exception as e:
logging.exception("API check failed")</p><p>schedule.every(30).seconds.do(check_api) # 每30秒执行一次</p><div class="aritcle_card flexRow">
<div class="artcardd flexRow">
<a class="aritcle_card_img" href="/xiazai/skill7554" title="Galileo python sdk"><img
src="https://img.php.cn/upload/skill/000/000/081/179154376921467.jpg" alt="Galileo python sdk" onerror="this.onerror='';this.src='/static/lhimages/moren/morentu.png'" ></a>
<div class="aritcle_card_info flexColumn">
<a href="/xiazai/skill7554" title="Galileo python sdk">Galileo python sdk</a>
<p>Galileo AI 平台 Python SDK 完整参考,用于评估、监控和保护 GenAI 应用,适用于构建 Python 应用。</p>
</div>
<a href="/xiazai/skill7554" title="Galileo python sdk" class="aritcle_card_btn flexRow flexcenter"><b></b><span>下载</span> </a>
</div>
</div><p>while True:
schedule.run_pending()
time.sleep(1)
但要注意:schedule 不是守护进程,退出终端就会停;真正上线建议用 systemd 或 supervisord 管理进程。另外,高频轮询(如 <10 秒)可能被目标接口限流,需确认对方 SLA。
真正难的不是写通这个脚本,而是把「什么算故障」定义清楚:是连续 3 次失败才告警?还是某次耗时超过 2 秒就触发?这些阈值必须和业务方对齐,否则告警要么淹没人,要么形同虚设。

















