要准确采集Nginx全链路耗时分布,需用nginx-prometheus-exporter暴露直方图指标,确保access_log中$request_time保留毫秒精度、按路径分流日志,并在Grafana中用histogram_quantile()计算P50/P95/P99。

要让 Prometheus 准确采集 Nginx 请求处理的全链路耗时分布(比如 P50/P95/P99),关键不是只记一个数,而是把原始耗时数据以直方图形式暴露出来,并确保时间粒度够细、来源分得清。
用 nginx-prometheus-exporter 暴露直方图指标
原生 Nginx 不输出 Prometheus 格式指标,必须借助 nginx-prometheus-exporter。它能解析 access log 或读取 Nginx 内置状态页,生成带分桶(bucket)的直方图指标:
- 启用
--enable-upstream-stats:为代理类请求暴露nginx_upstream_response_msecs_bucket,按后端节点和状态码分组 - 启用
--enable-static-stats:为静态资源暴露nginx_static_request_seconds_bucket,区分命中缓存/未命中 - 若使用
nginx-module-vts,可直接通过/status/format/prometheus端点获取更细粒度的直方图(含 location、upstream、cache 状态维度)
日志格式必须包含毫秒级原始值
Exporter 要靠日志反推分布,所以 access_log 中的 $request_time 必须保留三位小数(即毫秒精度),且不能被截断或四舍五入:
- 定义 log_format 时写成:
log_format main '$remote_addr - $remote_user [$time_local] "$request" $status $body_bytes_sent "$http_referer" "$http_user_agent" $request_time $upstream_response_time'; - 确保
$request_time是独立字段(不与其它变量拼接),方便 awk 或 exporter 按空格/制表符准确提取 - 避免用
combined等内置格式覆盖,否则可能缺失或错位
按路径和用途分流日志,避免分布失真
把 API、静态文件、健康检查混在一条日志里统计 P95,结果毫无意义。应在 location 块中分离日志目标:
-
location /api/ { access_log /var/log/nginx/api.log api_log; }—— 启用$upstream_response_time和$upstream_addr -
location ~* \.(js|css|png)$ { access_log /var/log/nginx/static.log static_log; }—— 加$sent_http_cache_control判断是否缓存命中 -
location = /healthz { access_log /var/log/nginx/health.log health_log; }—— 只记$request_time和$msec,排除业务干扰
在 Grafana 中计算真实分位数
Prometheus 本身不存原始耗时,而是聚合后的直方图计数。要用 histogram_quantile() 函数还原分布:
- 查 API 接口 P95 耗时:
histogram_quantile(0.95, sum(rate(nginx_upstream_response_msecs_bucket[1h])) by (le, upstream)) - 对比不同后端节点抖动:
stddev by (upstream)(rate(nginx_upstream_response_msecs_sum[1h]) / rate(nginx_upstream_response_msecs_count[1h])) - 识别缓存收益:对比
nginx_static_request_seconds_bucket{cache="hit"}和{cache="miss"}的 P90 差值
不需要写 Lua 或改编译参数,靠标准模块 + 合理日志结构 + 直方图暴露,就能拿到可落地的全链路耗时分布。


















