Commit Graph
146 Commits
Author SHA1 Message Date
ddadmin 869c189fe8 refactor(airflow): move DAGs to airflow/dags and update paths
- Why:
  - keep Airflow artifacts under a single airflow/ directory
  - align repository layout with intended project structure
- What:
  - move dags/ to airflow/dags/ and update compose mounts
  - make SQL root resolution work in container and local runs
  - update DAG path references in README, AGENTS, ARCHITECTURE, and plans
  - remove tracked Python cache artifacts from old DAG location
- Check:
  - airflow dags list
  - airflow dags list-import-errors
  - e2e success: ddl_init, kafka_load(limit=50), etl_pipeline
2026-02-08 19:34:10 +03:00
ddadmin 44a75691e8 feat(airflow): полная реализация фазы 2 — DAG для загрузки данных в Kafka
Слияние ветки с реализацией автоматизированной загрузки JSONL-файлов в Kafka
через Airflow DAG с валидацией, мониторингом и документацией.

- Что добавлено:
  - dags/kafka_load_dag.py: TaskGroup-пайплайн загрузки 4 потоков данных
  - dags/utils/kafka_helpers.py: хелперы для работы с Kafka (проверка,
    создание топиков, загрузка с лимитом)
  - airflow/requirements.txt: зависимость kafka-python==2.0.6
  - .gitignore: полноценный шаблон для ETL-проекта

- Параметры DAG:
  - limit: ограничение строк (0 = все)
  - reset_topics: пересоздание топиков перед загрузкой
  - load_browser/device/geo/location_events: выбор потоков

- Обновлена документация:
  - README.md, AGENTS.md, docs/ARCHITECTURE.md
  - plans/runbook.md, plans/airflow_dags_plan.md
2026-02-08 19:06:47 +03:00
ddadmin 60482af66b docs: обновлены Mermaid-диаграммы в README и ARCHITECTURE
- Упрощены диаграммы архитектуры, убраны эмодзи
- В README: компактная схема Airflow → Kafka → ClickHouse
- В ARCHITECTURE: обновлены общая схема, слои и сборка DDS
- Все диаграммы отражают 3 DAG: ddl_init, kafka_load, etl_pipeline
2026-02-08 19:03:11 +03:00
ddadmin 0b75da9c08 docs(docs): align docs with airflow-first ingest workflow
- Why:\n  - User-facing docs mixed Airflow and legacy CLI ingest paths and caused confusion\n- What:\n  - Rework README quick start and status to use DAG chain ddl_init -> kafka_load -> etl_pipeline\n  - Rewrite runbook as canonical Airflow-first execution flow\n  - Sync architecture diagrams/sequence and DQ wording with current SQL and DAG behavior\n- Check:\n  - Verified updated sections and removed stale markers with rg in README.md, docs/ARCHITECTURE.md, plans/runbook.md
2026-02-08 18:52:44 +03:00
ddadmin d7588a8caa chore: добавлен полноценный .gitignore для ETL-проекта
- Python: __pycache__, *.pyc, venv
- Docker: .env.local, volumes (clickhouse-data/, kafka-data/)
- Airflow: logs/, *.pid, airflow.db
- ClickHouse: логи сервера
- Kafka/Zookeeper: logs/, data dirs
- Superset: локальные БД
- IDE: VS Code (partial), PyCharm
- Секреты: *.pem, *.key, secrets/
- OS: .DS_Store, Thumbs.db
- Данные: *.csv.gz, *.parquet, архивы
2026-02-08 18:48:20 +03:00
ddadmin 1b9991f595 fix(airflow): simplify kafka_load params and align docs
- Why:

  - For DE task we only need full ingest or limit-based sample.

  - load_* and full_load params were redundant and unclear in current flow.

- What:

  - Remove full_load and load_* params from kafka_load DAG contract.

  - Simplify kafka helpers (validate/check files) to fixed 4-stream ingest.

  - Sync AGENTS, README, runbook, architecture and airflow plan docs.

- Check:

  - python3 -m py_compile dags/kafka_load_dag.py dags/utils/kafka_helpers.py

  - Airflow smoke/full runs: ddl_init -> kafka_load -> etl_pipeline (all success).

  - Legacy path: make data && make transform (success).
2026-02-08 18:33:54 +03:00
ddadmin 10f5bc3510 feat(airflow): реализован DAG kafka_load для загрузки в Kafka (фаза 2)
- Добавлен kafka-python==2.0.6 в airflow/requirements.txt
- Создан dags/utils/kafka_helpers.py с функциями:
  - check_kafka_ready() — проверка доступности брокера
  - prepare_topics() — создание/сброс топиков через KafkaAdminClient
  - load_jsonl() — загрузка данных через KafkaProducer (limit=0 = все)
  - validate_load_params(), check_input_files() — валидация
- Создан dags/kafka_load_dag.py с TaskGroup:
  - precheck: check_kafka, check_input_files, validate_load_params
  - ingest: prepare_topics, параллельная загрузка 4 потоков, verify_publish_counts
- Параметры DAG: limit (0 = все), reset_topics, load_* (выбор потоков)
- Обновлена документация: AGENTS.md, README.md, plans/runbook.md,
  plans/airflow_dags_plan.md, docs/ARCHITECTURE.md

Тестирование:
- Подключение к Kafka:  (kafka:29092 доступен, брокер 2.6.0)
- Загрузка данных:  (1000 сообщений — полный файл browser_events)
- Python синтаксис:  (py_compile проходит)
- Структура DAG:  (все 9 задач корректно определены)
2026-02-08 18:13:22 +03:00
ddadmin 12f35679f0 docs: update commit rules to support English language
- Why:
  - Align with Conventional Commits specification for consistency
  - English is standard for open-source and team collaboration
- What:
  - Change primary language to English (Russian still allowed)
  - Add type and scope reference tables
  - Add both English and Russian body templates
  - Add good/bad examples section
  - Add quick reference for common commit types
- Check:
  - File renders correctly in markdown viewer
  - Examples follow the new format rules
2026-02-08 17:25:04 +03:00
ddadmin 4819e10cf8 feat: изменён режим загрузки данных по умолчанию — теперь все записи
- По умолчанию bash ./scripts/load_kafka_data.sh
Resetting topics: browser_events location_events device_events geo_events
Loading mode: full (LIMIT=unset)
Bootstrap (inside container): kafka:29092
Publishing: data/browser_events.jsonl -> browser_events (full)
Publishing: data/device_events.jsonl -> device_events (full)
Publishing: data/geo_events.jsonl -> geo_events (full)
Publishing: data/location_events.jsonl -> location_events (full)
Done. загружает все записи из файлов (вместо 50 строк)
- Для ограничения используется bash ./scripts/load_kafka_data.sh
- Удалён устаревший параметр

Изменённые файлы:
- scripts/load_kafka_data.sh — обновлена логика и документация
- plans/runbook.md — обновлены примеры использования
- plans/kafka_ingest_plan.md — обновлён план реализации

Теперь:
- bash ./scripts/load_kafka_data.sh
Resetting topics: browser_events location_events device_events geo_events
Loading mode: full (LIMIT=unset)
Bootstrap (inside container): kafka:29092
Publishing: data/browser_events.jsonl -> browser_events (full)
Publishing: data/device_events.jsonl -> device_events (full)
Publishing: data/geo_events.jsonl -> geo_events (full)
Publishing: data/location_events.jsonl -> location_events (full)
Done. — все записи (4000 сообщений)
- bash ./scripts/load_kafka_data.sh
Resetting topics: browser_events location_events device_events geo_events
Loading mode: slice (LIMIT=50)
Bootstrap (inside container): kafka:29092
Publishing: data/browser_events.jsonl -> browser_events (first 50 lines)
Publishing: data/device_events.jsonl -> device_events (first 50 lines)
Publishing: data/geo_events.jsonl -> geo_events (first 50 lines)
Publishing: data/location_events.jsonl -> location_events (first 50 lines)
Done. — 50 строк каждого типа (200 сообщений)
2026-02-08 17:21:22 +03:00
ddadmin 33f8172892 Слой ODS теперь собирается в Airflow 2026-02-08 17:00:01 +03:00
ddadmin 60cb20406f docs(docs): добавить правила оформления коммитов
- Зачем:
  - унифицировать стиль коммитов для всех участников проекта
- Что сделано:
  - добавлен документ docs/COMMIT_RULES.md с форматом и примерами
  - добавлена ссылка на правила в AGENTS.md
- Проверка:
  - проверен staged diff перед коммитом
2026-02-08 16:42:43 +03:00
ddadmin 284dc3dc1f Move STG->ODS to Airflow batch and align monitoring 2026-02-08 16:36:01 +03:00
ddadmin a7963c17b3 fix: исправлен путь Kafka volume
- Изменен путь volume с /tmp/kraft-combined-logs на /var/lib/kafka/data
- Решена проблема с правами доступа при старте Kafka в KRaft mode
- Kafka теперь корректно инициализирует метаданные при первом запуске
2026-02-08 15:24:00 +03:00
ddadmin 413a711ddf Удаление ненужного файла 2026-02-08 15:16:52 +03:00
ddadmin 767198f6e2 Merge branch 'feature/airflow-orchestration' 2026-02-07 22:03:41 +03:00
ddadmin 225ae8bedb docs: update README and architecture docs for Airflow orchestration workflow 2026-02-07 22:03:13 +03:00
ddadmin d9b8a3f909 docs(airflow): update plugin and DAG docs 2026-02-07 21:57:15 +03:00
ddadmin 4d8f9d42f4 feat(infra): add clickhouse data persistence volume
Add persistent volume for ClickHouse to preserve data across container
restarts. The volume `clickhouse-data` is mounted to `/var/lib/clickhouse`,
ensuring data remains when containers are recreated.
2026-02-07 21:53:52 +03:00
ddadmin 14d16c5cae feat(airflow): implement dag orchestration for ddl and etl
Add comprehensive DAG implementation for ClickHouse schema initialization
and ETL pipeline orchestration. The ddl_init_dag manages database schema
creation across stg/ods/dds/dm layers with verification capabilities. The
etl_pipeline_dag implements full ODS to DDS to DM transformation flow with
data quality checks, branching logic for full/incremental loads, and
timeout handling for data availability.

Additional changes:
- Upgrade Airflow from 2.9.3 to 2.10.5
- Fix ClickHouse connection to use native protocol port 9000
- Mount SQL directory in docker-compose for DAG execution
- Update project requirements and documentation comments
- Remove unused pandas dependency
2026-02-07 21:52:31 +03:00
ddadmin de9d0429a2 fix(data): update dataset permissions for execution
Adjusting file modes on jsonl and markdown files to allow
execution within the ETL pipeline.
2026-02-07 21:36:55 +03:00
ddadmin 226807ecae chore(airflow): replace clickhouse-connect with airflow-clickhouse-plugin 2026-02-07 20:55:14 +03:00
ddadmin 9e340bb729 refactor(sql): reorganize sql files into structured directory hierarchy
Move DDL files from flat ddl/ directory to sql/ddl/ with layer-based
subdirectories (stg, ods, dds, dm). Move batch transformation SQL from
jobs/ to sql/ layer directories. Update scripts and documentation to
reflect new paths for improved organization and Airflow integration.
2026-02-07 20:51:51 +03:00
ddadmin 6466921bda docs(airflow): refactor dag implementation plan to separate kafka ingestion
Split Kafka ingestion into a dedicated `kafka_load` DAG to enable independent
experimentation with data loading without triggering the full ETL pipeline.
Restructure implementation phases: Stage 1 uses `make data` for MVP, Stage 2
adds the standalone Kafka DAG, Stage 3 adds monitoring. Update DAG numbering,
parameters, task groups, and acceptance criteria to reflect the new
architecture.
2026-02-07 19:44:37 +03:00
ddadmin 4da6a34e4c docs(airflow): update dag implementation plan for mvp
Consolidate the orchestration strategy by merging `kafka_load` and
`etl_batch_transform` into a unified `etl_pipeline`. Replace BashOperator
dependencies on Kafka CLI with PythonOperators utilizing `kafka-python`.
Add detailed technical specifications for helper functions, MVP stages,
and validation checks to align with current infrastructure constraints.
2026-02-07 00:03:32 +03:00
ddadmin cca315e4f6 docs(airflow): add migration plan for ETL orchestration
Detail the architecture for migrating ETL orchestration from make to
Airflow. Define DAG structures for database initialization, Kafka data
ingestion, batch transformation, and quality monitoring. Include
technical specifications, operator details, and implementation phases.
2026-02-06 23:42:50 +03:00
ddadmin fe9c15c0fe feat(airflow): configure ClickHouse connection and update infrastructure
Update Airflow configuration to integrate with ClickHouse DWH instead of
PostgreSQL training database. Changes include:

- Switch Airflow dependencies from PostgreSQL to ClickHouse connector
- Update docker-compose to use ClickHouse connection and correct Dockerfile
- Refactor airflow/requirements.txt to include only essential packages
- Add DAGs directory for ETL pipeline orchestration
- Update documentation to reflect Airflow integration and access credentials
- Adjust service dependencies to wait for ClickHouse startup
2026-02-06 23:34:45 +03:00
ddadmin d33cdb3fb0 feat(infra): add airflow orchestration services
Add Apache Airflow infrastructure with webserver, scheduler, and metadata
database to enable DAG-based pipeline orchestration. Includes optimized
requirements file and Docker configuration for Airflow 2.9.3.
2026-02-06 23:20:58 +03:00
ddadmin d7cff5ad1e docs: add russian comments to pipeline files 2026-02-06 22:45:39 +03:00
ddadmin 36139e0c78 docs: restructure AGENTS.md and add commenting conventions
- Update project structure (ddl/, jobs/, scripts/)
- Add Russian commenting conventions for SQL and Bash
- Reference example files for consistent style
2026-02-06 22:44:12 +03:00
ddadmin cbf5f22064 docs(architecture): update ODS error handling and DDS partial data support
Refine data flow diagrams and documentation to clarify error handling
in the ODS layer and partial data processing in the DDS layer. Add
detailed explanations for materialized views, batch SQL transformations,
and data quality metrics. Split DDS entity assembly diagrams for better
readability of event and click processing pipelines.
2026-02-06 22:21:16 +03:00
ddadmin e44b988d76 feat(data): add error handling for ODS layer and improve partial data support
Add materialized views to capture parsing errors from browser, location,
device, and geo raw staging tables and route them to dedicated error
tables in the ODS layer. Refactor DDS refresh logic to handle partial
data arrivals where device and geo events may arrive independently by
using a unified click_id source with LEFT JOINs. Add TRUNCATE command
to prevent duplicate data accumulation in DQ summary table.
2026-02-06 22:10:51 +03:00
ddadmin 6bbb26b9b3 feat(infra): implement batch transformation layer and comprehensive documentation
Add batch ETL pipeline with ODS→DDS→DM transformation jobs and scripts.
Create DDL infrastructure with automated database schema application.
Update Makefile with transform target for executing batch processes.
Rewrite README with complete Russian documentation including architecture
diagrams, quick start guide, and data flow visualization.
2026-02-06 21:58:17 +03:00
ddadmin 78b8b29fc7 docs(plans): update ddl architecture to use batch transforms for ods to dds
Replace materialized view joins with batch SQL transformations to avoid
consistency issues with out-of-order data. Document the reasoning for
using batch processing for ODS to DDS layer, including handling of
eventual consistency and versioning in ReplacingMergeTree. Update
data flow diagrams and remove MV creation DDL for DDS tables. Add
documentation for error handling tables and batch transformation jobs.
2026-02-05 23:07:57 +03:00
ddadmin c8564bfe03 docs(plans): add DDL modernization plan and restructure documentation
Add comprehensive plan for migrating executable DDL statements from markdown
to separate SQL files organized by layer. The plan outlines artifact structure,
execution requirements via make/Airflow, and environment parameters.

Existing inline DDL content is now marked as legacy in an appendix section,
providing clear separation between planned implementation and current state.
2026-02-05 22:37:51 +03:00
ddadmin a39dba6dae docs: add runbook reference and makefile commands
- Add reference to `plans/runbook.md` in key artifacts section
- Add reference to `plans/kafka_ingest_plan.md` in key artifacts section
- Document `make up`, `make ddl`, and `make data` commands in basic commands section
2026-02-05 22:12:51 +03:00
ddadmin a81a2c1b68 docs: add project readme 2026-02-05 22:11:10 +03:00
ddadmin eb80bc1870 feat(infra): add makefile and kafka data loading script
Add build automation via Makefile with targets for docker compose
management, DDL application, and data ingestion. Implement a robust bash
script for loading JSONL demo data into Kafka topics with configurable
options for limits, full dataset loading, and topic reset behavior.
2026-02-05 22:10:48 +03:00
ddadmin 380e8fcff3 docs(plans): fix kafka broker address and add data loading docs
Update Kafka broker address to use internal Docker network port (29092)
instead of external port (9092). Add kafka_ingest_plan.md with detailed
implementation strategy and runbook.md with user instructions.
2026-02-05 21:49:09 +03:00
ddadmin a2a1fdb785 feat(docker): add superset service
Add Apache Superset for data visualization to the docker-compose setup.
The custom Dockerfile installs additional tools and the clickhouse-connect
driver. The service is configured with health check, persistent volumes,
and environment variables.
2026-02-04 23:17:11 +03:00
ddadmin 60d02fdb28 fix(docker): correct network name and remove hardcoded container names
Fix network definition name from 'ch_replicated' to 'cs_dwh' to match
service references. Comment out hardcoded container names to allow
Docker to generate unique names automatically and avoid conflicts.
2026-02-04 23:06:03 +03:00
ddadmin f860bc47db Описание для роботов 2026-02-04 22:49:31 +03:00
ddadmin 06ffb0cf16 Наброски инфраструктуры 2026-02-04 22:47:19 +03:00
ddadmin df50e507e5 План трансформаций: JSON → витрины BI
- Описаны слои STG/ODS/DDS/DM и связи потоков (`event_id`/`click_id`)
- Добавлены DDL и MV-пайплайн для ingestion из Kafka (ClickHouse) + типизация/дедуп/DQ
- Добавлены витрины/VIEW для BI (Superset), mermaid-диаграмма и операционные заметки
2026-02-04 22:17:28 +03:00
ddadmin 70e12fbfb3 Текст задания 2026-02-04 22:11:07 +03:00
ddadmin aab9264121 Расширение исходных данных 2026-02-04 21:25:53 +03:00
ddadmin 06730fdf19 Исходные данные 2026-02-04 20:14:02 +03:00