必须显式注册指标(含NewGoCollector)、挂载promhttp.Handler()到/metrics、监听0.0.0.0,三者缺一不可;否则Prometheus抓取返回404、空页或target为down。

Go 服务暴露 /metrics 端点给 Prometheus 抓取,不是“写个 handler 就完事”,而是注册、挂载、网络三者必须同时合规——缺一不可。否则你会看到 target is down、no metrics found 或返回空页。
为什么 /metrics 返回空或 404
最常见原因是只挂了 promhttp.Handler(),但没注册任何指标。这个 handler 默认用 prometheus.DefaultRegisterer,而它初始是空的——连 go_goroutines 这类运行时指标都不会自动出现。
- 必须显式注册 Go 运行时指标:
prometheus.MustRegister(prometheus.NewGoCollector()),放在init()或main()开头 - 如果用了自定义 Registry(比如测试隔离),不能直接用
promhttp.Handler(),得改用promhttp.HandlerFor(reg, promhttp.HandlerOpts{}) - 别在 HTTP handler 里动态注册指标——
MustRegister重复调用会 panic,且未注册的指标不会出现在响应中
如何正确挂载 /metrics 路由
别用 http.HandleFunc("/metrics", promhttp.Handler().ServeHTTP),这是 v1.0 之前写法,已废弃,运行时会报 Handler() is deprecated。
- 标准写法就是:
http.Handle("/metrics", promhttp.Handler()) - 如果用
gorilla/mux或chi,要明确方法:router.Handle("/metrics", promhttp.Handler()).Methods("GET") - 禁止在
/metrics上加中间件(如鉴权、日志、CORS)——Prometheus 抓取不带 token,也不发Origin头,加了就 401 或 405 - 路径必须是
/metrics,除非你同步改了 Prometheus 的metrics_path配置
监听地址必须绑定 0.0.0.0,不能只绑 localhost
Prometheus 抓取是集群内其他节点发起的请求,如果 Go 服务只监听 127.0.0.1:8080,外部根本连不上,target 状态永远是 down。
Go 配置库,使用 spf13/viper — 分层优先级(flag > env >file > KV > default),提供 BindPFlag/BindPFlags、SetEnvPrefix + SetEnvKeyReplace 等功能。
立即学习“go语言免费学习笔记(深入)”;
- 启动时用:
http.ListenAndServe("0.0.0.0:8080", nil),而不是"localhost:8080" - 确认防火墙放行该端口(尤其 Kubernetes 中的 Pod 网络策略或云主机安全组)
- 若用反向代理(如 Nginx),确保透传
Accept: text/plain;version=0.0.4;q=1头,否则promhttp.Handler()可能返回 406
自定义指标注册的坑:Vec 类型和 label 基数
业务指标常用 CounterVec 或 HistogramVec,但它们极易因 label 控制不当导致采集失败或 Prometheus OOM。
-
WithLabelValues("GET", "200")传空字符串、label 数量不对、含非法字符(如/、{),会 panic:inconsistent label cardinality,且无法 recover - 高频打点场景下,避免每次调用
WithLabelValues();应提前缓存子指标:get200 := httpRequestsTotal.WithLabelValues("GET", "/", "200"),后续直接get200.Inc() -
HistogramVec的Buckets别照搬默认值(DefBuckets),Web 接口建议用[]float64{.01, .025, .05, .1, .25, .5, 1, 2, 5},否则_bucket时间序列爆炸 - ConstLabels 只用于静态值(如
version="v1.2.3"),绝不能塞请求级动态字段(如user_id)
真正卡住上线的往往不是代码逻辑,而是注册时机错、监听地址锁死、label 值失控这三类问题——它们不会报编译错误,但会让 /metrics 一直不可用,或者让 Prometheus 在几小时内内存飙到 30GB。

















