精准记录Nginx缓存失效与重试需采集$upstream_cache_status、$upstream_http_cache_control、$upstream_addr、$upstream_response_time、$status和$upstream_status等关键字段,配合定制cache_debug日志格式及X-Cache-Status等响应头,并监控error.log中的超时、连接关闭、非法header警告。

明确记录哪些关键状态
真正反映缓存失效与重试的字段只有几个,必须显式采集:
-
$upstream_cache_status:核心指标,值为
MISS(首次请求未缓存)、EXPIRED(缓存存在但过期)、STALE(返回陈旧内容)、UPDATING(正在后台刷新)、BYPASS(被绕过)——这些直接对应失效场景 -
$upstream_http_cache_control 和 $upstream_http_expires:上游返回的原始缓存控制头,用于判断是否因后端返回
no-cache或非法值导致失效 -
$upstream_addr 与 $upstream_response_time:确认是否发生回源,以及回源耗时是否异常(如超时触发
stale回退) - $status 与 $upstream_status:区分是 Nginx 返回 502/504(回源失败),还是后端返回 200 但缓存未写入(如因响应头非法)
配置专用缓存调试日志格式
在 http 块中定义清晰、紧凑的日志格式,避免冗余字段干扰分析:
log_format cache_debug '$remote_addr [$time_local] "$request" '
'$status $upstream_status $upstream_cache_status '
'$upstream_response_time $body_bytes_sent '
'$upstream_http_cache_control "$upstream_http_content_type"';
然后在对应 location 中启用:
access_log /var/log/nginx/cache_debug.log cache_debug;
这样每行日志就能一眼看出:谁、什么时候、请求什么、缓存状态如何、回源是否成功、后端给了什么缓存指令。
用响应头辅助人工验证与监控
仅靠 access_log 不够直观,建议在响应中注入可读性更强的状态标识:
add_header X-Cache-Status $upstream_cache_status; add_header X-Cache-Age $upstream_http_age; add_header X-Cache-Valid $upstream_http_cache_control;
配合 curl -I 即可快速判断单次请求的缓存行为。若看到 X-Cache-Status: EXPIRED 同时 X-Cache-Valid 为空或含 max-age=0,基本锁定是后端响应头问题;若为 UPDATING 且 X-Cache-Age 显示旧值,则静默刷新已生效。
捕获后台重试失败与异常
后台刷新(proxy_cache_background_update on)失败时,Nginx 不会改变主响应,但会在 error.log 记录警告。需重点关注:
-
upstream timed out(后台请求超时) -
upstream prematurely closed connection(后端中断) -
upstream sent invalid header(畸形 Header 导致无法解析缓存指令,最常见于失效风暴源头)
可在 error.log 中定期 grep 这些关键词,结合时间戳与请求路径,定位重试薄弱点。


















