Prometheus需借助exporter采集网络丢包与延迟数据,推荐使用ping_exporter实现端到端监控,配置目标、采集间隔后通过PromQL和Grafana分析probe_success、probe_loss_ratio及probe_rtt_*等指标。

用 ping_exporter 监控端到端延迟与丢包率
ping_exporter 是最常用、轻量且开箱即用的方案,适合监控目标主机或服务的可达性、RTT 和丢包情况。
- 部署方式简单:下载二进制文件,启动时指定目标(如 ping_exporter --targets=192.168.1.10,google.com --interval=15s)
- 暴露指标包括:probe_success(是否成功)、probe_duration_seconds(耗时)、probe_loss_ratio(丢包率)、probe_rtt_*_seconds(best/avg/worst/ stddev)
- 在 Prometheus 中添加 job 配置,指向 ping_exporter 的 /metrics 接口(默认端口 9427)
- 注意:避免高频探测(如
用 node_exporter + 自定义脚本补全内核级丢包指标
node_exporter 默认不暴露丢包指标,但 Linux 内核通过 /proc/net/snmp 和 /proc/net/dev 提供了各层丢包统计,可配合 shell 脚本导出。
- 重点关注字段:IpExtInNoRoutes(无路由丢包)、TcpRetransSegs(TCP重传段)、netstat -s | grep -i "retransmitted\|drop"
- 编写简单脚本读取并格式化为 Prometheus 文本格式(如 node_network_receive_drop_total{device="eth0"}),通过 textfile collector 或独立 HTTP endpoint 暴露
- 配合 node_exporter 的 --collector.textfile.directory 使用,无需额外服务
用 blackbox_exporter 实现多协议黑盒探测
blackbox_exporter 更适合模拟真实用户行为,支持 ICMP、TCP、HTTP、DNS 等多种探测类型,能发现路由、防火墙、TLS、DNS 等环节问题。
- 配置示例中启用 icmp 模块后,可获取 probe_icmp_duration_seconds 和 probe_success
- 丢包判断逻辑由 exporter 内部实现:连续多次 ping 失败才标记 probe_success=0,同时记录失败次数(probe_icmp_failure_reason)
- 适合跨网络边界监控(如云上服务访问公网地址),也便于做路径对比(正向+反向 MTR 数据可导入 Prometheus 分析)
PromQL 查询与告警关键写法
有了指标,还需用对 PromQL 才能真正定位问题。
- 丢包率突增: rate(probe_loss_ratio[5m]) > 0.05(5% 以上持续丢包)
- 延迟异常: avg_over_time(probe_rtt_average_seconds[10m]) > 200(平均 RTT 超 200ms)
- 区分方向:给 target 加 label,如 {target="api.example.com:443",job="blackbox-https"},避免混淆
- 结合 node_exporter 的 node_network_receive_errs_total 和 node_network_transmit_drop_total,判断是接收还是发送侧丢包


















