TCS-IPTV 监控部署说明
组件
| 文件 | 说明 |
|---|---|
prometheus.yml |
Prometheus 采集配置,抓取 api-svc / chain-svc / hash-api / catalog-svc 的 /metrics |
grafana-dashboard.json |
Grafana 面板 JSON,覆盖 4 金指标 |
快速部署
1. Prometheus
# 启动 Prometheus(Docker)
docker run -d --name prometheus \
-p 9090:9090 \
-v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml \
prom/prometheus
2. Grafana
# 启动 Grafana(Docker)
docker run -d --name grafana \
-p 3000:3000 \
-v $(pwd)/grafana-dashboard.json:/etc/grafana/provisioning/dashboards/tcs-iptv.json \
grafana/grafana
3. 导入面板
- 打开 Grafana → http://localhost:3000
- Dashboards → Import → 上传
grafana-dashboard.json - 数据源选择 Prometheus(http://prometheus:9090)
面板指标说明
| 面板 | 指标 | Prometheus Metric |
|---|---|---|
| 请求延迟 | P50/P95/P99 | tcs_http_request_duration_seconds (Histogram) |
| 请求总量 | QPS by path | tcs_http_requests_total (Counter) |
| 错误率 | 4xx/5xx QPS | tcs_http_errors_total (Counter) |
| 在途请求 | 并发数 | tcs_http_in_flight_requests (Gauge) |
| 错误率百分比 | 错误/总请求 | 上述两个 Counter 计算 |
| 延迟热力图 | by path | Histogram bucket 热力图 |
告警规则建议
在 Prometheus 中配置 alerting_rules.yml:
groups:
- name: tcs-iptv-alerts
rules:
- alert: HighErrorRate
expr: sum(rate(tcs_http_errors_total[5m])) / sum(rate(tcs_http_requests_total[5m])) > 0.1
for: 5m
labels:
severity: critical
annotations:
summary: "错误率超过 10%"
- alert: HighLatency
expr: histogram_quantile(0.95, sum(rate(tcs_http_request_duration_seconds_bucket[5m])) by (le)) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "P95 延迟超过 2 秒"
- alert: HighInFlight
expr: tcs_http_in_flight_requests > 100
for: 2m
labels:
severity: warning
annotations:
summary: "在途请求超过 100"