error_log是排查Nginx负载均衡配置错误最直接线索,典型错误如“no live upstreams”“Connection refused”“timed out”等精准指向upstream不可用、网络不通或超时,配合warn级别日志与access_log中upstream_addr/upstream_status可快速定位根因。

error_log 是排查 Nginx 后端负载均衡配置错误最直接、最可靠的线索来源之一。它不只记录崩溃或严重故障,更常在 upstream 配置不当、服务器不可达、健康检查失败等场景下输出关键提示,往往比 502/503 错误页面更早暴露问题本质。
关注 error_log 中与 upstream 相关的典型错误信息
当负载均衡配置有误时,Nginx 通常会在 error_log 中写入明确指向 upstream 模块的日志,例如:
- "no live upstreams while connecting to upstream":说明所有 upstream server 被标记为不可用(可能因 health_check 失败、max_fails 超限、或未配置 backup);
- "connect() failed (111: Connection refused) while connecting to upstream":后端服务未监听对应端口,或防火墙/网络策略阻断连接;
- "upstream timed out (110: Connection timed out) while connecting to upstream":后端响应过慢或未响应,需检查 proxy_connect_timeout 和后端服务状态;
- "invalid port in upstream" 或 "invalid number of arguments in 'server' directive":upstream 块语法错误,如漏写 port、多写分号、weight 值非法等。
确保 error_log 级别足够捕获 warning 及以上事件
默认 error_log 级别常设为 error,但很多配置类警告(如 server 指令解析异常、resolver 不可用、health check 初始化失败)仅在 warn 或 info 级别输出。建议在调试阶段临时提升日志级别:
error_log /var/log/nginx/error.log warn;
注意:生产环境避免长期使用 info 级别,以防日志量过大影响性能。
结合 access_log 与 error_log 定位具体请求路径
单看 error_log 只能知道“哪里出错”,要确认“哪类请求触发错误”,需关联 access_log 中的 upstream_addr 和 upstream_status 字段:
- upstream_addr 显示实际转发到的后端地址(含 IP+port),若为空或显示“none”,说明根本未进入 upstream 流程;
- upstream_status 记录后端返回状态码(如 200、502、504),配合 error_log 中的 connect/timeouts 可判断是网络层失败还是后端处理超时;
- 对特定 location 或 upstream name 添加 unique log_format,能快速过滤相关请求上下文。
验证 upstream 配置是否被正确加载
修改 nginx.conf 后,必须执行 nginx -t 检查语法,但即使语法通过,仍可能因以下原因导致 upstream 未生效:
- upstream 块定义在 server 块内(应置于 http 块顶层);
- proxy_pass 指向的 upstream 名称拼写错误(区分大小写);
- 使用了 resolver 但 DNS 解析失败,且未配置 valid 或 retry 参数,Nginx 会静默跳过该 server;
- 启用了 keepalive 连接池但后端不支持,可能导致部分请求失败,error_log 中可能出现 "upstream prematurely closed connection" 类提示。


















