Prometheus 与监控体系
#监控 · #Prometheus · #Grafana · #PromQL · #告警 · #Metrics
Prometheus 是云原生时代的事实监控标准。它以拉取模式采集指标,用 PromQL 灵活查询,通过 Alertmanager 管理告警。配合 Grafana 可视化,构成完整的可观测性体系。
可观测性三支柱
mermaid
flowchart TB
subgraph Pillars["可观测性三支柱"]
subgraph M["Metrics (指标)"]
MT["Prometheus + Grafana"]
MD["'知道有什么问题'<br/>聚合、趋势、告警"]
end
subgraph T["Tracing (链路追踪)"]
TT["Jaeger / Zipkin"]
TD["'知道哪里有问题'<br/>链路、依赖、瓶颈"]
end
subgraph L["Logging (日志)"]
LT["Elasticsearch / Loki"]
LD["'知道为什么有问题'<br/>上下文、详情、排查"]
end
endPrometheus 架构
mermaid
flowchart TB
subgraph Targets["采集目标"]
T1["Node Exporter"]
T2["应用 /metrics"]
T3["MySQL Exporter"]
T4["K8s cAdvisor"]
end
subgraph Core["Prometheus Server"]
P1["Service Discovery<br/>(K8s/Consul/文件)"]
P2["Scrape<br/>(拉取指标)"]
P3["TSDB<br/>(时序存储)"]
P4["PromQL<br/>(查询引擎)"]
end
subgraph Alert["告警"]
AM["Alertmanager"]
R1["邮件"]
R2["企业微信"]
R3["PagerDuty"]
end
subgraph Viz["可视化"]
G["Grafana"]
end
Targets -->|"pull /metrics"| P2
P1 --> P2
P2 --> P3
P3 --> P4
P4 --> G
P4 --> AM
AM --> R1
AM --> R2
AM --> R3指标类型
| 类型 | 用途 | 示例 |
|---|---|---|
| Counter | 只增不减的计数 | 请求总数、错误数 |
| Gauge | 可增可减的值 | 内存使用、队列长度 |
| Histogram | 分桶统计分布 | 请求延迟分布 |
| Summary | 分位数统计 | P50/P90/P99 延迟 |
Go 应用暴露指标
go
import (
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
var (
httpRequestsTotal = prometheus.NewCounterVec(
prometheus.CounterOpts{
Name: "http_requests_total",
Help: "HTTP 请求总数",
},
[]string{"method", "path", "status"},
)
httpRequestDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Help: "HTTP 请求耗时分布",
Buckets: prometheus.DefBuckets, // .005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5, 10
},
[]string{"method", "path"},
)
)
func init() {
prometheus.MustRegister(httpRequestsTotal, httpRequestDuration)
}
func metricsMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
timer := prometheus.NewTimer(httpRequestDuration.WithLabelValues(r.Method, r.URL.Path))
defer timer.ObserveDuration()
wrapped := &responseWriter{ResponseWriter: w}
next.ServeHTTP(wrapped, r)
httpRequestsTotal.WithLabelValues(
r.Method, r.URL.Path, strconv.Itoa(wrapped.statusCode),
).Inc()
})
}
// 暴露 /metrics 端点
http.Handle("/metrics", promhttp.Handler())PromQL 查询
基础查询
promql
# 瞬时向量
http_requests_total{path="/api/users"}
# 速率(适合 Counter)
rate(http_requests_total[5m]) # 每秒请求数(5m 窗口)
irate(http_requests_total[5m]) # 更灵敏的速率
# 增长量
increase(http_requests_total[1h]) # 过去 1h 的请求增量
# 分位数(针对 Histogram)
histogram_quantile(0.99,
rate(http_request_duration_seconds_bucket[5m]))聚合运算
promql
# 所有接口的 QPS 总和
sum(rate(http_requests_total[5m]))
# 按 path 分组
sum by (path) (rate(http_requests_total[5m]))
# 错误率
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
# P99 延迟 > 1s 的接口
histogram_quantile(0.99,
rate(http_request_duration_seconds_bucket[5m])) > 1常用函数
| 函数 | 用途 |
|---|---|
rate(x[d]) | 增长率 |
increase(x[d]) | 增量 |
avg_over_time(x[d]) | 时间窗口平均值 |
quantile_over_time(0.99, x[d]) | 时间窗口分位数 |
predict_linear(x[d], t) | 线性预测 |
absent(x) | 判断指标是否存在 |
topk(5, x) | Top 5 |
bottomk(5, x) | Bottom 5 |
告警规则
yaml
# rules/alerts.yml
groups:
- name: http_alerts
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.05
for: 5m # 持续 5 分钟才触发
labels:
severity: critical
annotations:
summary: "HTTP 错误率过高"
description: "错误率 {{ $value | humanizePercentage }}"
- alert: HighLatency
expr: |
histogram_quantile(0.99,
rate(http_request_duration_seconds_bucket[5m]))
> 2
for: 10m
labels:
severity: warning
annotations:
summary: "P99 延迟 > 2s"
description: "P99 延迟 {{ $value }}s"Alertmanager 配置
yaml
# alertmanager.yml
route:
group_by: ['alertname']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: 'default'
receivers:
- name: 'default'
webhook_configs:
- url: 'http://webhook.example.com/alert'
wechat_configs:
- corp_id: 'xxx'
to_party: '1'
agent_id: '1000002'
api_secret: 'xxx'存储与性能
TSDB 原理
时序数据按时间分 block:
┌─────────┬─────────┬─────────┐
│ Block 1 │ Block 2 │ Block 3 │ → 每 2h 一个 block
│ 00:00 │ 02:00 │ 04:00 │
└─────────┘─────────┴─────────┘
每个 block 内:
- chunks/ → 原始数据压缩(XOR 压缩,10-20x)
- index → 倒排索引(查 label 快)
- meta.json → 元数据性能调优
yaml
# prometheus.yml
global:
scrape_interval: 15s # 采集间隔(不要太小)
evaluation_interval: 15s # 告警评估间隔
# 指标优化
- job_name: 'app'
scrape_interval: 30s # 低优先级 job 用更长间隔
metric_relabel_configs:
- source_labels: [__name__]
regex: 'go_gc_.*' # 去掉不需要的指标
action: dropCardinality 爆炸:label 组合过多导致内存暴增
# ❌ 高危:每个请求一个维度
http_requests_total{user_id="123456"} # 用户太多!
# ✅ 安全
http_requests_total{path="/api/users"} # path 有限Grafana 仪表盘
json
// 一个典型的 RED (Rate/Error/Duration) 面板
{
"panels": [
{
"title": "QPS",
"targets": [{
"expr": "sum(rate(http_requests_total[1m]))"
}]
},
{
"title": "Error Rate",
"targets": [{
"expr": "sum(rate(http_requests_total{status=~'5..'}[1m])) / sum(rate(http_requests_total[1m]))"
}]
},
{
"title": "P99 Latency",
"targets": [{
"expr": "histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[1m])) by (le))"
}]
}
]
}四个黄金信号
| 信号 | 指标 | PromQL |
|---|---|---|
| 延迟 | P99 请求延迟 | histogram_quantile(0.99, ...) |
| 流量 | QPS | rate(http_requests_total[1m]) |
| 错误 | 错误率 | 5xx / total |
| 饱和度 | 队列深度/CPU | go_goroutines / cpu_utilization |
参考
Prometheus vs 其他监控系统
| 维度 | Prometheus | Zabbix | Datadog | VictoriaMetrics |
|---|---|---|---|---|
| 数据模型 | 多维时序(label) | 固定 Host→Item | 多维时序 | 兼容 Prometheus |
| 采集方式 | Pull(拉取) | Agent Push | Agent Push | Pull + Push |
| 查询语言 | PromQL(强大) | 简单表达式 | 自定义 | MetricsQL(兼容 PromQL) |
| 存储 | 本地 TSDB(单机) | MySQL/PostgreSQL | SaaS 云端 | 分布式 TSDB |
| 扩展性 | 单机(需联邦/Thanos) | 中等(Proxy 分层) | 无限(SaaS) | 水平扩展 |
| 告警 | Alertmanager(灵活) | 内置(模板化) | 内置 + AI | 兼容 Alertmanager |
| 生态 | 云原生标准(K8s/Istio) | 传统运维 | 全栈 APM | Prometheus 兼容 |
| 成本 | 免费开源 | 免费开源 | 按量付费(贵) | 免费开源 |
| 适用 | 云原生/K8s/微服务 | 传统 IDC/网络设备 | 全栈可观测(预算充足) | 大规模 Prometheus 替代 |
选型建议:云原生环境首选 Prometheus + Grafana;数据量超过单机承载(百万级时序)时,用 VictoriaMetrics 或 Thanos 做长期存储;传统 IDC 监控(交换机/服务器硬件)用 Zabbix。
从告警到定位:完整排查流程
mermaid
flowchart TB
Alert["🔔 告警触发<br/>HighErrorRate > 5%"] --> Confirm{"确认告警<br/>是否误报?"}
Confirm -->|"误报"| Silence["静默/调整阈值"]
Confirm -->|"真实"| Scope["确定影响范围"]
Scope --> Dashboard["查看 Grafana 仪表盘<br/>哪些接口/实例受影响?"]
Dashboard --> Correlate["关联分析"]
Correlate --> CPU["CPU/内存/磁盘<br/>是否资源瓶颈?"]
Correlate --> Deps["依赖服务<br/>MySQL/Redis/下游 是否异常?"]
Correlate --> Deploy["最近是否有发布?<br/>(关联 CI/CD 事件)"]
CPU -->|"是"| ScaleUp["扩容/限流"]
Deps -->|"是"| DepFix["修复依赖<br/>(慢查询/连接池满)"]
Deploy -->|"是"| Rollback["回滚发布"]
CPU -->|"否"| Trace["查看链路追踪<br/>Jaeger/Zipkin"]
Deps -->|"否"| Trace
Deploy -->|"否"| Trace
Trace --> Logs["查看错误日志<br/>ELK/Loki"]
Logs --> RootCause["定位根因"]
RootCause --> Fix["修复 + 复盘"]排查实战 PromQL 模板
promql
# 第 1 步:哪些接口错误率高?
topk(5,
sum by (path) (rate(http_requests_total{status=~"5.."}[5m]))
/ sum by (path) (rate(http_requests_total[5m]))
)
# 第 2 步:哪些实例有问题?
sum by (instance) (rate(http_requests_total{status=~"5.."}[5m]))
# 第 3 步:是否资源瓶颈?
# CPU 使用率
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# 内存使用率
(node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes * 100
# 第 4 步:依赖是否异常?
# MySQL 慢查询
rate(mysql_global_status_slow_queries[5m])
# Redis 延迟
redis_commands_duration_seconds_total / redis_commands_processed_total
# 连接池使用率
go_sql_open_connections / go_sql_max_open_connections长期存储方案
Prometheus 本地 TSDB 默认只保留 15 天数据,且单机存储有上限。生产环境需要长期存储方案:
Thanos 架构
mermaid
flowchart TB
subgraph Prometheus["Prometheus 实例 (多个)"]
P1["Prometheus 1<br/>+ Thanos Sidecar"]
P2["Prometheus 2<br/>+ Thanos Sidecar"]
end
subgraph Thanos["Thanos 组件"]
Query["Thanos Query<br/>(统一查询入口)"]
Store["Thanos Store Gateway<br/>(读取对象存储)"]
Compact["Thanos Compactor<br/>(降采样 + 压缩)"]
end
subgraph Storage["对象存储"]
S3["S3 / COS / MinIO<br/>(廉价长期存储)"]
end
P1 -->|"上传 Block"| S3
P2 -->|"上传 Block"| S3
Store -->|"读取历史数据"| S3
Compact -->|"压缩/降采样"| S3
Query -->|"近期数据"| P1
Query -->|"近期数据"| P2
Query -->|"历史数据"| Store
Grafana["Grafana"] --> QueryThanos vs VictoriaMetrics
| 维度 | Thanos | VictoriaMetrics |
|---|---|---|
| 架构 | Sidecar 模式,组件多 | 单二进制/集群版 |
| 部署复杂度 | 高(5+ 组件) | 低(1 个二进制) |
| 查询性能 | 中(跨组件 RPC) | 高(本地存储优化) |
| 存储成本 | 低(对象存储) | 低(自研压缩,10x) |
| 兼容性 | 完全兼容 PromQL | MetricsQL(超集) |
| 降采样 | ✅ 5m/1h 自动降采样 | ✅ |
| 适用 | 已有多 Prometheus 实例 | 新建或替换 Prometheus |
RED / USE / 四个黄金信号实战
RED 方法(面向服务)
promql
# Rate: 每秒请求数
sum(rate(http_requests_total[5m]))
# Errors: 错误率
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# Duration: P99 延迟
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le))USE 方法(面向资源)
promql
# Utilization: CPU 使用率
1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m]))
# Saturation: CPU 运行队列长度(饱和度)
node_load1 / count(node_cpu_seconds_total{mode="idle"}) by (instance)
# Errors: 磁盘错误
rate(node_disk_io_time_weighted_seconds_total[5m])四个黄金信号告警模板
yaml
groups:
- name: golden_signals
rules:
# 延迟:P99 > 2s
- alert: HighLatency
expr: histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, job)) > 2
for: 5m
labels: { severity: warning }
# 流量:QPS 突降 50%(可能服务异常)
- alert: TrafficDrop
expr: sum(rate(http_requests_total[5m])) < sum(rate(http_requests_total[5m] offset 1h)) * 0.5
for: 5m
labels: { severity: critical }
# 错误:错误率 > 1%
- alert: HighErrorRate
expr: sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.01
for: 3m
labels: { severity: critical }
# 饱和度:内存使用 > 85%
- alert: HighMemoryUsage
expr: (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / node_memory_MemTotal_bytes > 0.85
for: 10m
labels: { severity: warning }
登录后即可发表评论 👇