- Why:
- Airflow task metrics were mapped to non-emitted StatsD keys
- reload-monitoring did not restart statsd-exporter after mapping changes
- What:
- update StatsD mapping for Airflow 2.10.5 metric names
- remove problematic catch-all mapping that produced inconsistent series
- restart statsd-exporter in reload-monitoring flow
- sync operations runbook and airflow monitoring plan with actual metrics
- Check:
- make reload-monitoring
- Prometheus targets: airflow/clickhouse/kafka are UP
- trigger ddl_init and verify airflow_task_duration_seconds_count
- verify airflow_task_success_total and airflow_task_failures_total in Prometheus
- Add statsd-exporter service to docker-compose.yml (prom/statsd-exporter:v0.27.1)
- Add StatsD env vars to airflow-default-env for metrics export
- Add airflow job to prometheus.yml scrape configs
- Add Airflow Overview dashboard (Grafana provisioning)
- Add Airflow alert rules: scheduler down, queue backlog, failures, parse time
- Add configs/statsd_mapping.yml for StatsD → Prometheus conversion
- Use Prometheus naming convention (_total for counters, _seconds for timers)
- Add monitoring plan at plans/monitoring_airflow_plan.md
- Update OPERATIONS.md and Makefile for airflow monitoring
Tested: all 3 jobs (airflow, clickhouse, kafka) showing UP in Prometheus,
metrics flowing (dagbag_size=3, executor slots, heartbeats with _total suffix),
all 4 alert rules loaded in Grafana