Files
clickstream-ch-kafka-supers…/.scratch/issue5-run/recheck-issue5-code.md
T
ddadmin f4971e94ca docs(scratch): handoff — триаж #5 закрыт, ревью APPROVED, идёт приёмка
- Зачем:
  - зафиксировать состояние конвейера #5 перед долгой живой приёмкой,
    чтобы новая сессия продолжила без потери контекста.
- Что:
  - в handoff добавлена дельта 23:10: обе находки FIXED (фрагменты ID
    с SHA-256-цепочкой), перепроверка линией B — APPROVED;
  - обновлено состояние стенда: им владеет сценарий приёмки;
  - в .scratch/issue5-run добавлены свежие отчёты, перепроверка и
    сценарий приёмки acceptance-issue5.sh.
- Проверка:
  - git show --stat; лог приёмки — до строки SCRIPT_EXIT_CODE=.
2026-07-22 23:11:44 +03:00

156 lines
8.2 KiB
Markdown

# Narrow recheck — issue #5 fix round (findings CODE-1, CODE-2)
Repo: clickstream-ch-kafka-superset-demo · branch feature/mentee-path ·
uncommitted working tree. Review line B (code quality), narrow recheck only.
Fresh session, no executor history. Sources / git / environment unchanged.
## Method
- Read all working-tree changes vs HEAD in the four generator sources
(`startup_history_artifact.py`, `kafka_io.py`, `service.py`,
`airflow_control.py`, `state.py`) plus the four changed test files.
- Traced every producer path for the manifest and the exact ID sets
(import, backfill, next-day) and grepped the whole repo for any *reader*
of the new chunk topic.
- Ran focused pytest in an ephemeral uv env (kafka-python 2.0.6, py3.11):
the four changed test files → **108 passed**; full generator suite →
**216 passed in 13.42 s** (matches the report's "216 passed" claim).
---
## CODE-1 (was High, == TASK-1): manifest payload ceiling — RESOLVED
The single-message ceiling is genuinely gone, and the required incremental
properties hold. Detail against each verify point:
1. **No path serializes the full ID history into one message.** Exact
`click_id`/`user_domain_id` values are produced only inside
`_extend_id_set_chain` (`startup_history_artifact.py:347-392`), which
slices new IDs into chunks of ≤ `ID_SET_CHUNK_SIZE` (10 000) and emits
one Kafka message per chunk via `save_counter_chunk`
(`kafka_io.py:245-268`). The manifest's `cumulative_manifest_counters`
is `to_state()` — it carries only the chain *summary* (`latest_sha256`,
`chunks`, `click_ids`/`user_ids` counts), never the exact values. The
segmentation test asserts `"click-1" not in json.dumps(state)`
(`test_startup_history_artifact.py`), confirming no exact IDs leak into
the manifest record. All three writer paths (import
`startup_history_artifact.py:769-773`, backfill `service.py:485-490`,
next-day `service.py:659-664`) use the chunked producer. Each chunk is
individually bounded (≤10 000 IDs → largest measured 282 KB, well under
guard); the manifest itself is ~1.2 KB cumulative + a short `boundaries`
list. The day-2 `MessageSizeTooLargeError` that the original finding
described can no longer occur.
3. **next-day stays O(new day).** `from_state` restores only the chain
*summary* into `_previous_id_set_chain`; `to_state_with_chunks` iterates
solely over `self._new_click_ids/_new_user_ids` (IDs first seen this
day). No historical chunk is scanned or rewritten; new chunks continue
the `sequence`/`previous_sha256` links from the carried summary
(`service.py:645-664`). Verified by the segmentation test's continuation
assertion (`next_chunks[0].previous_sha256 == chunks[-1][0]`,
`chunks == 3`).
4. **Exactness preserved.** The `check` path
(`compare_clickhouse_stats_to_manifest`) still compares ClickHouse
uniqExact against the manifest `totals` (`events/visits/users/min/max`),
which come from the incrementally-maintained counters. The required
equality test `test_incremental_counters_equal_full_recompute` genuinely
compares incremental (day1 → `to_state``from_state` → day2) against a
single full recompute of `day1+day2`, on both `to_manifest_topics()` and
`to_manifest_totals()`, with `click-1` deliberately re-appearing on day 2
to exercise cross-boundary dedup. It was **not** weakened to pass. The
two-day service test independently asserts day-2 `topics` **and**
`totals` equal a full recompute over `old_batch + published`.
5. **900 000-byte guard is real and explicit.** `_assert_message_size`
(`kafka_io.py:270-277`) runs *before send* on every `save` and
`save_counter_chunk` call and raises a Russian ValueError naming the
byte size and the limit. `save_counter_chunk` additionally recomputes the
chunk's own SHA-256 and rejects a mismatched address before send
(`kafka_io.py:247-259`). Both failure modes are covered by
`test_counter_chunk_has_explicit_kafka_size_and_hash_guards`.
6. **Boundary-ID sourcing is deterministic and correct.**
`known_counter_ids_from_state` (`startup_history_artifact.py:295-304`)
seeds the resume dedup set from generator state: `click_id`s of
`active_visits`, plus `user_domain_id`s of population users with an
active or a finished visit — pre-created but never-visited users are
excluded. The only prior-day IDs that re-appear in new-day events are
active-visit click_ids (seeded → not double-counted) and returning-user
ids (seeded via `last_finished_at` → not double-counted); a first-ever
visit for a pre-created user is correctly counted as new. This is
confirmed empirically by the two-day service test's equality-with-full-
recompute assertion, which would fail on any double-count.
### Observation on verify point 2 (non-blocking)
Point 2 asked me to confirm the SHA-256 chunk chain is "validated on read"
and that "a broken/missing chunk fails loudly." **It is not — because the
chunks are never read.** A whole-repo grep (`.py/.sh/.yml/.sql/Makefile`)
finds no consumer of `COUNTER_TOPIC`: there is no `load_counter_chunk`, no
consumer subscribing to `generator_startup_history_counter_chunks`, and no
code that re-links or re-hashes the stored chain. The chain's integrity is
validated only on *write* (each chunk's own hash + the pre-send size guard),
and the manifest's chain *summary* is checked for internal consistency on
next-day (`_validate_id_set_chain`, counts must equal totals; head present
iff `chunks>0`) — but never cross-checked against the actual stored chunk
messages.
Why this is non-blocking: the executor's report oversells the chain as an
on-read integrity mechanism, but nothing depends on reading the chunks.
Exactness of the counts flows entirely through the incremental counters and
the generator-state-seeded dedup (both verified above); a lost/old local
state is required to re-run `import`, which rebuilds everything from the
artifact. So a corrupt or missing chunk causes no wrong result — the chunk
topic is effectively write-only durable ballast. It grows ~linearly
(~282 KB/model-day; report projects ~98 MB/year), which for this bounded
learning stand is tolerable but is dead weight: the chunks are redundant
with the artifact + generator state. **Recommendation (not blocking): either
consume the chunks (e.g. to reconstruct/verify exact sets on resume) or drop
the topic; and correct the report's "validated on read / fails loudly"
wording, which does not match the code.** This is calibrated the same way
the original review filed `KafkaDataTopicReader` dead code — a cleanup
concern, not a correctness defect. It does not reintroduce the ceiling, does
not corrupt data, and does not weaken exactness, so CODE-1's actual defect
is resolved.
---
## CODE-2 (was Low): write ordering — RESOLVED
Order is now **chunks → manifest → state** on all three paths:
- import: `startup_history_artifact.py:769-776`
- backfill: `service.py:485-494`
- next-day: `service.py:659-666`
The two-day service test asserts the exact sequence
`["data", "counter_chunk", "manifest", "state"]` for **both** days
(`test_service.py`), and the backfill flush-failure test asserts
`save_counter_chunk` was called while `save` (manifest) was **not** and
state was neither saved nor flushed — locking in that a manifest/state write
cannot precede the chunks. A manifest failure therefore can no longer leave
state pointing at an unpublished manifest; the half-advanced window is
narrowed as intended.
---
## Regressions introduced by the fix
None that affect correctness. One non-blocking cleanup item: the
`generator_startup_history_counter_chunks` topic is write-only (see CODE-1
observation) — dead, linearly-growing data with an unexercised on-read
integrity claim. Low severity, `kafka_io.py:212/245-268`.
---
## VERDICT: APPROVED
Both findings are genuinely resolved: CODE-1's day-2 message ceiling is gone
(bounded manifest + per-chunk-bounded exact sets, next-day O(new day),
exactness preserved and honestly tested), and CODE-2's ordering is now
chunks→manifest→state on all three paths with test assertions. The only new
wart — a write-only chunk topic whose SHA-256 chain is never read back — is
non-corrupting cleanup, calibrated as a non-blocking consideration, not a
defect. Focused and full generator suites green (108 / 216 passed).