docs(scratch): handoff — триаж #5 закрыт, ревью APPROVED, идёт приёмка
- Зачем: - зафиксировать состояние конвейера #5 перед долгой живой приёмкой, чтобы новая сессия продолжила без потери контекста. - Что: - в handoff добавлена дельта 23:10: обе находки FIXED (фрагменты ID с SHA-256-цепочкой), перепроверка линией B — APPROVED; - обновлено состояние стенда: им владеет сценарий приёмки; - в .scratch/issue5-run добавлены свежие отчёты, перепроверка и сценарий приёмки acceptance-issue5.sh. - Проверка: - git show --stat; лог приёмки — до строки SCRIPT_EXIT_CODE=.
This commit is contained in:
@@ -0,0 +1,155 @@
|
||||
# Narrow recheck — issue #5 fix round (findings CODE-1, CODE-2)
|
||||
|
||||
Repo: clickstream-ch-kafka-superset-demo · branch feature/mentee-path ·
|
||||
uncommitted working tree. Review line B (code quality), narrow recheck only.
|
||||
Fresh session, no executor history. Sources / git / environment unchanged.
|
||||
|
||||
## Method
|
||||
|
||||
- Read all working-tree changes vs HEAD in the four generator sources
|
||||
(`startup_history_artifact.py`, `kafka_io.py`, `service.py`,
|
||||
`airflow_control.py`, `state.py`) plus the four changed test files.
|
||||
- Traced every producer path for the manifest and the exact ID sets
|
||||
(import, backfill, next-day) and grepped the whole repo for any *reader*
|
||||
of the new chunk topic.
|
||||
- Ran focused pytest in an ephemeral uv env (kafka-python 2.0.6, py3.11):
|
||||
the four changed test files → **108 passed**; full generator suite →
|
||||
**216 passed in 13.42 s** (matches the report's "216 passed" claim).
|
||||
|
||||
---
|
||||
|
||||
## CODE-1 (was High, == TASK-1): manifest payload ceiling — RESOLVED
|
||||
|
||||
The single-message ceiling is genuinely gone, and the required incremental
|
||||
properties hold. Detail against each verify point:
|
||||
|
||||
1. **No path serializes the full ID history into one message.** Exact
|
||||
`click_id`/`user_domain_id` values are produced only inside
|
||||
`_extend_id_set_chain` (`startup_history_artifact.py:347-392`), which
|
||||
slices new IDs into chunks of ≤ `ID_SET_CHUNK_SIZE` (10 000) and emits
|
||||
one Kafka message per chunk via `save_counter_chunk`
|
||||
(`kafka_io.py:245-268`). The manifest's `cumulative_manifest_counters`
|
||||
is `to_state()` — it carries only the chain *summary* (`latest_sha256`,
|
||||
`chunks`, `click_ids`/`user_ids` counts), never the exact values. The
|
||||
segmentation test asserts `"click-1" not in json.dumps(state)`
|
||||
(`test_startup_history_artifact.py`), confirming no exact IDs leak into
|
||||
the manifest record. All three writer paths (import
|
||||
`startup_history_artifact.py:769-773`, backfill `service.py:485-490`,
|
||||
next-day `service.py:659-664`) use the chunked producer. Each chunk is
|
||||
individually bounded (≤10 000 IDs → largest measured 282 KB, well under
|
||||
guard); the manifest itself is ~1.2 KB cumulative + a short `boundaries`
|
||||
list. The day-2 `MessageSizeTooLargeError` that the original finding
|
||||
described can no longer occur.
|
||||
|
||||
3. **next-day stays O(new day).** `from_state` restores only the chain
|
||||
*summary* into `_previous_id_set_chain`; `to_state_with_chunks` iterates
|
||||
solely over `self._new_click_ids/_new_user_ids` (IDs first seen this
|
||||
day). No historical chunk is scanned or rewritten; new chunks continue
|
||||
the `sequence`/`previous_sha256` links from the carried summary
|
||||
(`service.py:645-664`). Verified by the segmentation test's continuation
|
||||
assertion (`next_chunks[0].previous_sha256 == chunks[-1][0]`,
|
||||
`chunks == 3`).
|
||||
|
||||
4. **Exactness preserved.** The `check` path
|
||||
(`compare_clickhouse_stats_to_manifest`) still compares ClickHouse
|
||||
uniqExact against the manifest `totals` (`events/visits/users/min/max`),
|
||||
which come from the incrementally-maintained counters. The required
|
||||
equality test `test_incremental_counters_equal_full_recompute` genuinely
|
||||
compares incremental (day1 → `to_state` → `from_state` → day2) against a
|
||||
single full recompute of `day1+day2`, on both `to_manifest_topics()` and
|
||||
`to_manifest_totals()`, with `click-1` deliberately re-appearing on day 2
|
||||
to exercise cross-boundary dedup. It was **not** weakened to pass. The
|
||||
two-day service test independently asserts day-2 `topics` **and**
|
||||
`totals` equal a full recompute over `old_batch + published`.
|
||||
|
||||
5. **900 000-byte guard is real and explicit.** `_assert_message_size`
|
||||
(`kafka_io.py:270-277`) runs *before send* on every `save` and
|
||||
`save_counter_chunk` call and raises a Russian ValueError naming the
|
||||
byte size and the limit. `save_counter_chunk` additionally recomputes the
|
||||
chunk's own SHA-256 and rejects a mismatched address before send
|
||||
(`kafka_io.py:247-259`). Both failure modes are covered by
|
||||
`test_counter_chunk_has_explicit_kafka_size_and_hash_guards`.
|
||||
|
||||
6. **Boundary-ID sourcing is deterministic and correct.**
|
||||
`known_counter_ids_from_state` (`startup_history_artifact.py:295-304`)
|
||||
seeds the resume dedup set from generator state: `click_id`s of
|
||||
`active_visits`, plus `user_domain_id`s of population users with an
|
||||
active or a finished visit — pre-created but never-visited users are
|
||||
excluded. The only prior-day IDs that re-appear in new-day events are
|
||||
active-visit click_ids (seeded → not double-counted) and returning-user
|
||||
ids (seeded via `last_finished_at` → not double-counted); a first-ever
|
||||
visit for a pre-created user is correctly counted as new. This is
|
||||
confirmed empirically by the two-day service test's equality-with-full-
|
||||
recompute assertion, which would fail on any double-count.
|
||||
|
||||
### Observation on verify point 2 (non-blocking)
|
||||
|
||||
Point 2 asked me to confirm the SHA-256 chunk chain is "validated on read"
|
||||
and that "a broken/missing chunk fails loudly." **It is not — because the
|
||||
chunks are never read.** A whole-repo grep (`.py/.sh/.yml/.sql/Makefile`)
|
||||
finds no consumer of `COUNTER_TOPIC`: there is no `load_counter_chunk`, no
|
||||
consumer subscribing to `generator_startup_history_counter_chunks`, and no
|
||||
code that re-links or re-hashes the stored chain. The chain's integrity is
|
||||
validated only on *write* (each chunk's own hash + the pre-send size guard),
|
||||
and the manifest's chain *summary* is checked for internal consistency on
|
||||
next-day (`_validate_id_set_chain`, counts must equal totals; head present
|
||||
iff `chunks>0`) — but never cross-checked against the actual stored chunk
|
||||
messages.
|
||||
|
||||
Why this is non-blocking: the executor's report oversells the chain as an
|
||||
on-read integrity mechanism, but nothing depends on reading the chunks.
|
||||
Exactness of the counts flows entirely through the incremental counters and
|
||||
the generator-state-seeded dedup (both verified above); a lost/old local
|
||||
state is required to re-run `import`, which rebuilds everything from the
|
||||
artifact. So a corrupt or missing chunk causes no wrong result — the chunk
|
||||
topic is effectively write-only durable ballast. It grows ~linearly
|
||||
(~282 KB/model-day; report projects ~98 MB/year), which for this bounded
|
||||
learning stand is tolerable but is dead weight: the chunks are redundant
|
||||
with the artifact + generator state. **Recommendation (not blocking): either
|
||||
consume the chunks (e.g. to reconstruct/verify exact sets on resume) or drop
|
||||
the topic; and correct the report's "validated on read / fails loudly"
|
||||
wording, which does not match the code.** This is calibrated the same way
|
||||
the original review filed `KafkaDataTopicReader` dead code — a cleanup
|
||||
concern, not a correctness defect. It does not reintroduce the ceiling, does
|
||||
not corrupt data, and does not weaken exactness, so CODE-1's actual defect
|
||||
is resolved.
|
||||
|
||||
---
|
||||
|
||||
## CODE-2 (was Low): write ordering — RESOLVED
|
||||
|
||||
Order is now **chunks → manifest → state** on all three paths:
|
||||
|
||||
- import: `startup_history_artifact.py:769-776`
|
||||
- backfill: `service.py:485-494`
|
||||
- next-day: `service.py:659-666`
|
||||
|
||||
The two-day service test asserts the exact sequence
|
||||
`["data", "counter_chunk", "manifest", "state"]` for **both** days
|
||||
(`test_service.py`), and the backfill flush-failure test asserts
|
||||
`save_counter_chunk` was called while `save` (manifest) was **not** and
|
||||
state was neither saved nor flushed — locking in that a manifest/state write
|
||||
cannot precede the chunks. A manifest failure therefore can no longer leave
|
||||
state pointing at an unpublished manifest; the half-advanced window is
|
||||
narrowed as intended.
|
||||
|
||||
---
|
||||
|
||||
## Regressions introduced by the fix
|
||||
|
||||
None that affect correctness. One non-blocking cleanup item: the
|
||||
`generator_startup_history_counter_chunks` topic is write-only (see CODE-1
|
||||
observation) — dead, linearly-growing data with an unexercised on-read
|
||||
integrity claim. Low severity, `kafka_io.py:212/245-268`.
|
||||
|
||||
---
|
||||
|
||||
## VERDICT: APPROVED
|
||||
|
||||
Both findings are genuinely resolved: CODE-1's day-2 message ceiling is gone
|
||||
(bounded manifest + per-chunk-bounded exact sets, next-day O(new day),
|
||||
exactness preserved and honestly tested), and CODE-2's ordering is now
|
||||
chunks→manifest→state on all three paths with test assertions. The only new
|
||||
wart — a write-only chunk topic whose SHA-256 chain is never read back — is
|
||||
non-corrupting cleanup, calibrated as a non-blocking consideration, not a
|
||||
defect. Focused and full generator suites green (108 / 216 passed).
|
||||
Reference in New Issue
Block a user