diff --git a/README.md b/README.md index ac0af5b..aa4f4e8 100644 --- a/README.md +++ b/README.md @@ -241,6 +241,11 @@ flowchart LR - Query Performance: queries/sec, active queries, failed queries - MergeTree Storage: parts count, merge rate +Что алертится: +- Failed queries rate (`rate(ClickHouseProfileEvents_FailedQuery[5m]) > 0`) +- Memory Resident > 85% от `OSMemoryTotal` +- Active parts > 500 + Конфигурация provisioning находится в [`configs/grafana/provisioning/`](configs/grafana/provisioning/). Подробнее в [`docs/OPERATIONS.md`](docs/OPERATIONS.md#мониторинг). --- @@ -263,12 +268,11 @@ flowchart LR - **Трансформации**: DAG `etl_pipeline` (pre-check, batch STG→ODS→DDS→DM, валидация) - **Витрины**: VIEW в DM для бизнес-дашбордов (трафик, UTM, качество данных) - **Надёжность**: ошибки парсинга сохраняются в ODS, пайплайн не падает на "грязных" данных -- **Мониторинг**: Prometheus скрейпит ClickHouse метрики, Grafana дашборд для визуализации +- **Мониторинг**: Prometheus скрейпит ClickHouse метрики, Grafana дашборд и alert rules В планах (не требуется для MVP задания): - Инкрементальный batch (watermark вместо `full_refresh`) - DQ мониторинг по расписанию -- Алерты в Grafana --- diff --git a/configs/grafana/provisioning/alerting/clickhouse-alert-rules.yml b/configs/grafana/provisioning/alerting/clickhouse-alert-rules.yml new file mode 100644 index 0000000..bc3bc0c --- /dev/null +++ b/configs/grafana/provisioning/alerting/clickhouse-alert-rules.yml @@ -0,0 +1,268 @@ +apiVersion: 1 + +groups: + - orgId: 1 + name: clickhouse_health_group + folder: ClickHouse Alerts + interval: 1m + rules: + - uid: ch_failed_queries_rate + title: ClickHouse Failed Queries Rate + condition: C + data: + - refId: A + relativeTimeRange: + from: 300 + to: 0 + datasourceUid: prometheus_uid + model: + datasource: + type: prometheus + uid: prometheus_uid + editorMode: code + expr: rate(ClickHouseProfileEvents_FailedQuery[5m]) + instant: false + intervalMs: 1000 + legendFormat: __auto + maxDataPoints: 43200 + range: true + refId: A + - refId: B + datasourceUid: __expr__ + model: + conditions: + - evaluator: + params: [] + type: gt + operator: + type: and + query: + params: + - B + reducer: + params: [] + type: last + type: query + datasource: + type: __expr__ + uid: __expr__ + expression: A + intervalMs: 1000 + maxDataPoints: 43200 + reducer: last + refId: B + type: reduce + - refId: C + datasourceUid: __expr__ + model: + conditions: + - evaluator: + params: + - 0 + type: gt + operator: + type: and + query: + params: + - C + reducer: + params: [] + type: last + type: query + datasource: + type: __expr__ + uid: __expr__ + expression: B + intervalMs: 1000 + maxDataPoints: 43200 + refId: C + type: threshold + dashboardUid: clickhouse-overview + panelId: 8 + noDataState: NoData + execErrState: Error + for: 2m + annotations: + __dashboardUid__: clickhouse-overview + __panelId__: "8" + summary: "Есть ошибки запросов в ClickHouse" + runbook_url: "http://localhost:3000/d/clickhouse-overview/clickhouse-overview" + labels: + service: clickhouse + severity: warning + metric: failed_queries + isPaused: false + + - uid: ch_memory_resident_high + title: ClickHouse Memory Resident High + condition: C + data: + - refId: A + relativeTimeRange: + from: 300 + to: 0 + datasourceUid: prometheus_uid + model: + datasource: + type: prometheus + uid: prometheus_uid + editorMode: code + expr: 100 * ClickHouseAsyncMetrics_MemoryResident / ClickHouseAsyncMetrics_OSMemoryTotal + instant: false + intervalMs: 1000 + legendFormat: __auto + maxDataPoints: 43200 + range: true + refId: A + - refId: B + datasourceUid: __expr__ + model: + conditions: + - evaluator: + params: [] + type: gt + operator: + type: and + query: + params: + - B + reducer: + params: [] + type: last + type: query + datasource: + type: __expr__ + uid: __expr__ + expression: A + intervalMs: 1000 + maxDataPoints: 43200 + reducer: last + refId: B + type: reduce + - refId: C + datasourceUid: __expr__ + model: + conditions: + - evaluator: + params: + - 85 + type: gt + operator: + type: and + query: + params: + - C + reducer: + params: [] + type: last + type: query + datasource: + type: __expr__ + uid: __expr__ + expression: B + intervalMs: 1000 + maxDataPoints: 43200 + refId: C + type: threshold + dashboardUid: clickhouse-overview + panelId: 3 + noDataState: NoData + execErrState: Error + for: 5m + annotations: + __dashboardUid__: clickhouse-overview + __panelId__: "3" + summary: "Высокое потребление памяти ClickHouse (Memory Resident)" + runbook_url: "http://localhost:3000/d/clickhouse-overview/clickhouse-overview" + labels: + service: clickhouse + severity: critical + metric: memory_resident + isPaused: false + + - uid: ch_parts_active_high + title: ClickHouse Parts Active High + condition: C + data: + - refId: A + relativeTimeRange: + from: 300 + to: 0 + datasourceUid: prometheus_uid + model: + datasource: + type: prometheus + uid: prometheus_uid + editorMode: code + expr: ClickHouseMetrics_PartsActive + instant: false + intervalMs: 1000 + legendFormat: __auto + maxDataPoints: 43200 + range: true + refId: A + - refId: B + datasourceUid: __expr__ + model: + conditions: + - evaluator: + params: [] + type: gt + operator: + type: and + query: + params: + - B + reducer: + params: [] + type: last + type: query + datasource: + type: __expr__ + uid: __expr__ + expression: A + intervalMs: 1000 + maxDataPoints: 43200 + reducer: last + refId: B + type: reduce + - refId: C + datasourceUid: __expr__ + model: + conditions: + - evaluator: + params: + - 500 + type: gt + operator: + type: and + query: + params: + - C + reducer: + params: [] + type: last + type: query + datasource: + type: __expr__ + uid: __expr__ + expression: B + intervalMs: 1000 + maxDataPoints: 43200 + refId: C + type: threshold + dashboardUid: clickhouse-overview + panelId: 12 + noDataState: NoData + execErrState: Error + for: 10m + annotations: + __dashboardUid__: clickhouse-overview + __panelId__: "12" + summary: "Слишком много активных MergeTree parts" + runbook_url: "http://localhost:3000/d/clickhouse-overview/clickhouse-overview" + labels: + service: clickhouse + severity: warning + metric: parts_active + isPaused: false diff --git a/configs/grafana/provisioning/dashboards/clickhouse-overview.json b/configs/grafana/provisioning/dashboards/clickhouse-overview.json index 246e9e7..454acd4 100644 --- a/configs/grafana/provisioning/dashboards/clickhouse-overview.json +++ b/configs/grafana/provisioning/dashboards/clickhouse-overview.json @@ -38,7 +38,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -116,7 +116,7 @@ }, "targets": [ { - "expr": "rate(ClickHouseAsyncMetrics_CPUStealMicroseconds[1m]) / 10000", + "expr": "rate(ClickHouseProfileEvents_OSCPUVirtualTimeMicroseconds[1m]) / 1000000 / scalar(count({__name__=~\"ClickHouseAsyncMetrics_CPUFrequencyMHz_.*\"})) * 100", "refId": "A" } ], @@ -126,7 +126,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -210,7 +210,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -307,7 +307,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -391,7 +391,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -475,7 +475,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -536,7 +536,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -593,7 +593,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -667,7 +667,7 @@ }, "targets": [ { - "expr": "rate(ClickHouseProfileEvents_InsertedRowsCount[1m])", + "expr": "rate(ClickHouseProfileEvents_InsertedRows[1m])", "refId": "A" } ], @@ -690,7 +690,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -737,7 +737,7 @@ "pluginVersion": "11.5.2", "targets": [ { - "expr": "ClickHouseMetrics_Parts", + "expr": "ClickHouseAsyncMetrics_TotalPartsOfMergeTreeTables", "refId": "A" } ], @@ -747,7 +747,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -803,6 +803,7 @@ "gridPos": { "h": 8, "w": 16, + "x": 8, "y": 23 }, "id": 13, @@ -820,18 +821,18 @@ }, "targets": [ { - "expr": "ClickHouseMetrics_Parts", - "legendFormat": "{{ database }}.{{ table }}", + "expr": "{__name__=~\"ClickHouseMetrics_Parts(Active|Committed|Outdated|Deleting|PreActive|PreCommitted|Temporary|Wide|Compact|DeleteOnDestroy)\"}", + "legendFormat": "{{ __name__ }}", "refId": "A" } ], - "title": "Parts by Table", + "title": "Parts by State", "type": "timeseries" }, { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -888,7 +889,7 @@ { "datasource": { "type": "prometheus", - "uid": "${datasource}" + "uid": "prometheus_uid" }, "fieldConfig": { "defaults": { @@ -972,27 +973,13 @@ ], "refresh": "30s", "schemaVersion": 39, - "tags": ["clickhouse", "dwh", "monitoring"], + "tags": [ + "clickhouse", + "dwh", + "monitoring" + ], "templating": { - "list": [ - { - "current": { - "selected": false, - "text": "Prometheus", - "value": "Prometheus" - }, - "hide": 0, - "includeAll": false, - "multi": false, - "name": "datasource", - "options": [], - "query": "prometheus", - "refresh": 1, - "regex": "", - "skipUrlSync": false, - "type": "datasource" - } - ] + "list": [] }, "time": { "from": "now-1h", @@ -1002,6 +989,6 @@ "timezone": "Europe/Moscow", "title": "ClickHouse Overview", "uid": "clickhouse-overview", - "version": 2, + "version": 3, "weekStart": "" -} \ No newline at end of file +} diff --git a/configs/grafana/provisioning/datasources/prometheus.yml b/configs/grafana/provisioning/datasources/prometheus.yml index 9a5b209..6dc899b 100644 --- a/configs/grafana/provisioning/datasources/prometheus.yml +++ b/configs/grafana/provisioning/datasources/prometheus.yml @@ -7,6 +7,7 @@ apiVersion: 1 datasources: - name: Prometheus + uid: prometheus_uid type: prometheus access: proxy url: http://prometheus:9090 diff --git a/docs/OPERATIONS.md b/docs/OPERATIONS.md index 5de9f73..e98d88c 100644 --- a/docs/OPERATIONS.md +++ b/docs/OPERATIONS.md @@ -109,6 +109,7 @@ curl -s http://localhost:9090/api/v1/targets | grep -o '"health":"[^"]*"' - **Grafana provisioning** (`configs/grafana/provisioning/`): - Datasource Prometheus автоматически настроен - Dashboard "ClickHouse Overview" загружается при старте + - Alert rules для ClickHouse загружаются при старте ### Дашборд ClickHouse Overview @@ -118,7 +119,9 @@ URL: `http://localhost:3000/d/clickhouse-overview/clickhouse-overview` |--------|---------| | System Health | CPU Usage, Memory Resident, Memory Code | | Query Performance | Queries/sec, Active Queries, Failed Queries, Total Queries, Inserted Rows/sec | -| MergeTree Storage | Total Parts, Parts by Table, Total Merges, Merges/sec | +| MergeTree Storage | Total Parts, Parts by State, Total Merges, Merges/sec | + +Принятое решение по метрикам: сверили naming через Context7 (`/clickhouse/clickhouse-docs`, раздел Prometheus interface) и заменили недоступные в `25.1` серии на фактически экспортируемые (`ClickHouseProfileEvents_InsertedRows`, `ClickHouseAsyncMetrics_TotalPartsOfMergeTreeTables`, `ClickHouseMetrics_Parts*`). ### Проверка метрик @@ -130,6 +133,25 @@ curl -s "http://localhost:9090/api/v1/query?query=ClickHouseAsyncMetrics_MemoryR curl -s "http://localhost:9090/api/v1/query?query=ClickHouseProfileEvents_Query" ``` +### Алерты Grafana + +Provisioning-файл: `configs/grafana/provisioning/alerting/clickhouse-alert-rules.yml` + +Настроены правила: +- `ClickHouse Failed Queries Rate` — `rate(ClickHouseProfileEvents_FailedQuery[5m]) > 0` в течение `2m` +- `ClickHouse Memory Resident High` — `MemoryResident / OSMemoryTotal * 100 > 85` в течение `5m` +- `ClickHouse Parts Active High` — `ClickHouseMetrics_PartsActive > 500` в течение `10m` + +Проверка и reload без рестарта контейнера: + +```bash +# Список правил unified alerting +curl -s -u admin:admin http://localhost:3000/api/v1/provisioning/alert-rules + +# Принудительно перечитать provisioning alerting +curl -s -X POST -u admin:admin http://localhost:3000/api/admin/provisioning/alerting/reload +``` + ### Troubleshooting мониторинга - **"No data" в Grafana**: проверить, что Prometheus видит target (`Status -> Targets` в UI) diff --git a/docs/REPO_MAP.md b/docs/REPO_MAP.md index 2a4869b..e5acb07 100644 --- a/docs/REPO_MAP.md +++ b/docs/REPO_MAP.md @@ -32,6 +32,7 @@ - `data/*.jsonl` — исходные данные (могут быть грязными) - `configs/` — конфиги ClickHouse, Prometheus, Grafana +- `configs/grafana/provisioning/alerting/clickhouse-alert-rules.yml` — правила алертинга Grafana для ClickHouse ## Документация