- Why:
- students hit permission denied after pull and grafana restart-loop with readonly db
- What:
- run grafana as default non-root user
- mount provisioning directory as read-only
- add troubleshooting for git permission issues and grafana volume reset
- normalize file modes for data jsonl and docs/DE-task.md to 100644
- Check:
- docker compose config
- docker compose up -d grafana
- curl -u admin:admin http://localhost:3000/api/health
- Add kafka-exporter service to docker-compose.yml
- Add kafka job to prometheus.yml scrape configs
- Add Kafka Overview dashboard (Grafana provisioning)
- Add Kafka alert rules (broker down, consumer lag, etc.)
- Add make reload-monitoring command for easy updates
- Update OPERATIONS.md with TL;DR and troubleshooting
API verified via Context7:
- /danielqsj/kafka_exporter for exporter config
- /prometheus/docs for scrape_configs format
- Why:
- student needs a simple way to apply Grafana/monitoring config updates after git pull
- What:
- add TL;DR block with minimal commands in monitoring section
- add detailed post-pull runbook for datasource/dashboard/alerting reload
- include clickhouse restart note for prometheus_ch.xml changes
- Check:
- reviewed commands and paths in docs/OPERATIONS.md
- Why:
- dashboard panels could resolve to stale datasource uid and show No data
- monitoring required proactive alerts for ClickHouse health signals
- What:
- pin dashboard panels to prometheus_uid and remove datasource templating variable
- fix PromQL metrics for CPU, inserted rows, and parts panels
- add provisioning alert rules for failed queries, memory resident, and active parts
- pin Prometheus datasource uid and update monitoring documentation
- Check:
- POST /api/admin/provisioning/datasources/reload
- POST /api/admin/provisioning/dashboards/reload
- POST /api/admin/provisioning/alerting/reload
- GET /api/v1/provisioning/alert-rules
- Why:
- commit messages with literal \n are hard to read in UI
- What:
- add explicit rule for multiline body formatting in CLI
- add correct examples with git commit -m and -F heredoc
- Check:
- reviewed new section in docs/COMMIT_RULES.md
- Why:\n - AGENTS.md became too large and mixed policy with operational details\n - context7 requirement was easy to miss in long text\n- What:\n - reduce AGENTS.md to a compact contributor contract\n - add explicit mandatory MCP Context7 workflow block\n - move runbook details to docs/OPERATIONS.md\n - move artifact map to docs/REPO_MAP.md\n- Check:\n - reviewed links and content after split\n - ensured only documentation files are included in commit
- Why:
- keep Airflow artifacts under a single airflow/ directory
- align repository layout with intended project structure
- What:
- move dags/ to airflow/dags/ and update compose mounts
- make SQL root resolution work in container and local runs
- update DAG path references in README, AGENTS, ARCHITECTURE, and plans
- remove tracked Python cache artifacts from old DAG location
- Check:
- airflow dags list
- airflow dags list-import-errors
- e2e success: ddl_init, kafka_load(limit=50), etl_pipeline
- Why:\n - User-facing docs mixed Airflow and legacy CLI ingest paths and caused confusion\n- What:\n - Rework README quick start and status to use DAG chain ddl_init -> kafka_load -> etl_pipeline\n - Rewrite runbook as canonical Airflow-first execution flow\n - Sync architecture diagrams/sequence and DQ wording with current SQL and DAG behavior\n- Check:\n - Verified updated sections and removed stale markers with rg in README.md, docs/ARCHITECTURE.md, plans/runbook.md
- Why:
- For DE task we only need full ingest or limit-based sample.
- load_* and full_load params were redundant and unclear in current flow.
- What:
- Remove full_load and load_* params from kafka_load DAG contract.
- Simplify kafka helpers (validate/check files) to fixed 4-stream ingest.
- Sync AGENTS, README, runbook, architecture and airflow plan docs.
- Check:
- python3 -m py_compile dags/kafka_load_dag.py dags/utils/kafka_helpers.py
- Airflow smoke/full runs: ddl_init -> kafka_load -> etl_pipeline (all success).
- Legacy path: make data && make transform (success).
- Why:
- Align with Conventional Commits specification for consistency
- English is standard for open-source and team collaboration
- What:
- Change primary language to English (Russian still allowed)
- Add type and scope reference tables
- Add both English and Russian body templates
- Add good/bad examples section
- Add quick reference for common commit types
- Check:
- File renders correctly in markdown viewer
- Examples follow the new format rules
- Зачем:
- унифицировать стиль коммитов для всех участников проекта
- Что сделано:
- добавлен документ docs/COMMIT_RULES.md с форматом и примерами
- добавлена ссылка на правила в AGENTS.md
- Проверка:
- проверен staged diff перед коммитом
Move DDL files from flat ddl/ directory to sql/ddl/ with layer-based
subdirectories (stg, ods, dds, dm). Move batch transformation SQL from
jobs/ to sql/ layer directories. Update scripts and documentation to
reflect new paths for improved organization and Airflow integration.
Update Airflow configuration to integrate with ClickHouse DWH instead of
PostgreSQL training database. Changes include:
- Switch Airflow dependencies from PostgreSQL to ClickHouse connector
- Update docker-compose to use ClickHouse connection and correct Dockerfile
- Refactor airflow/requirements.txt to include only essential packages
- Add DAGs directory for ETL pipeline orchestration
- Update documentation to reflect Airflow integration and access credentials
- Adjust service dependencies to wait for ClickHouse startup
Refine data flow diagrams and documentation to clarify error handling
in the ODS layer and partial data processing in the DDS layer. Add
detailed explanations for materialized views, batch SQL transformations,
and data quality metrics. Split DDS entity assembly diagrams for better
readability of event and click processing pipelines.
Add batch ETL pipeline with ODS→DDS→DM transformation jobs and scripts.
Create DDL infrastructure with automated database schema application.
Update Makefile with transform target for executing batch processes.
Rewrite README with complete Russian documentation including architecture
diagrams, quick start guide, and data flow visualization.