直接用 $request_time 记录请求总耗时并结合日志分析可快速定位慢接口,需结合路径、状态码及分布特征判断问题;$request_time 包含 Nginx 处理与 upstream 等待时间,对比 $upstream_response_time 可区分瓶颈所在。

直接用 $request_time 记录每个请求的总耗时,再结合日志分析,就能快速定位慢接口。关键不是单纯看数值大小,而是结合路径、状态码和分布特征来判断是否真有问题。
在 access_log 中记录 request_time
确保 Nginx 日志格式包含 $request_time,推荐使用带毫秒精度的格式:
log_format main '$remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'"$http_referer" "$http_user_agent" '
'$request_time $upstream_response_time';
access_log /var/log/nginx/access.log main;
注意:$request_time 是从接收第一个字节到发送完响应的总时间(单位:秒,精度为毫秒),它包含 Nginx 自身处理时间,也包括 upstream 等待时间;而 $upstream_response_time 只反映后端响应耗时,两者对比能帮你区分是 Nginx 层瓶颈还是后端问题。
按接口路径统计平均/最大耗时
用日志分析工具(如 awk、goaccess 或 ELK)提取高频路径的耗时统计。例如用 awk 快速查看耗时 top 5 的接口:
-
awk '{print $7, $9}' /var/log/nginx/access.log | sort -k2 -nr | head -n 5(假设 $7 是 $request,$9 是 $request_time) - 更准确的做法是按
$request的前两段(如GET /api/user/profile)分组,计算平均值和 P95/P99 - 重点关注那些 QPS 较高但平均
$request_time > 500ms的路径,它们对用户体验影响更大
结合状态码和 upstream_response_time 排查根因
单看 $request_time 高,并不等于后端慢。需交叉验证:
- 如果
$request_time高但$upstream_response_time很小(比如 < 10ms),说明慢在 Nginx 本层:可能是大文件传输、gzip 压缩、SSL 握手、或正则 rewrite 开销大 - 如果两者都高,且
$status == 200,大概率是后端逻辑慢,建议同步查后端应用日志和 profile - 若
$status == 502/504且$upstream_response_time接近或超过proxy_read_timeout,说明后端超时,需检查服务健康、连接池或数据库慢查询
设置动态告警与采样分析
线上不宜全量记录所有字段,可通过条件日志减少开销:
- 只对耗时超过阈值(如 1s)的请求记录详细信息:
log_if = $request_time > 1.0; - 用 log_format 定义 debug 格式,配合 map 指令做条件输出,避免影响常规日志性能
- 搭配 Prometheus + nginx-vts-exporter 或自定义 Lua 脚本,实现接口维度的实时 P95 耗时监控


















