- Зачем: - зафиксировать состояние конвейера #5 перед долгой живой приёмкой, чтобы новая сессия продолжила без потери контекста. - Что: - в handoff добавлена дельта 23:10: обе находки FIXED (фрагменты ID с SHA-256-цепочкой), перепроверка линией B — APPROVED; - обновлено состояние стенда: им владеет сценарий приёмки; - в .scratch/issue5-run добавлены свежие отчёты, перепроверка и сценарий приёмки acceptance-issue5.sh. - Проверка: - git show --stat; лог приёмки — до строки SCRIPT_EXIT_CODE=.
8.2 KiB
Narrow recheck — issue #5 fix round (findings CODE-1, CODE-2)
Repo: clickstream-ch-kafka-superset-demo · branch feature/mentee-path · uncommitted working tree. Review line B (code quality), narrow recheck only. Fresh session, no executor history. Sources / git / environment unchanged.
Method
- Read all working-tree changes vs HEAD in the four generator sources
(
startup_history_artifact.py,kafka_io.py,service.py,airflow_control.py,state.py) plus the four changed test files. - Traced every producer path for the manifest and the exact ID sets (import, backfill, next-day) and grepped the whole repo for any reader of the new chunk topic.
- Ran focused pytest in an ephemeral uv env (kafka-python 2.0.6, py3.11): the four changed test files → 108 passed; full generator suite → 216 passed in 13.42 s (matches the report's "216 passed" claim).
CODE-1 (was High, == TASK-1): manifest payload ceiling — RESOLVED
The single-message ceiling is genuinely gone, and the required incremental properties hold. Detail against each verify point:
-
No path serializes the full ID history into one message. Exact
click_id/user_domain_idvalues are produced only inside_extend_id_set_chain(startup_history_artifact.py:347-392), which slices new IDs into chunks of ≤ID_SET_CHUNK_SIZE(10 000) and emits one Kafka message per chunk viasave_counter_chunk(kafka_io.py:245-268). The manifest'scumulative_manifest_countersisto_state()— it carries only the chain summary (latest_sha256,chunks,click_ids/user_idscounts), never the exact values. The segmentation test asserts"click-1" not in json.dumps(state)(test_startup_history_artifact.py), confirming no exact IDs leak into the manifest record. All three writer paths (importstartup_history_artifact.py:769-773, backfillservice.py:485-490, next-dayservice.py:659-664) use the chunked producer. Each chunk is individually bounded (≤10 000 IDs → largest measured 282 KB, well under guard); the manifest itself is ~1.2 KB cumulative + a shortboundarieslist. The day-2MessageSizeTooLargeErrorthat the original finding described can no longer occur. -
next-day stays O(new day).
from_staterestores only the chain summary into_previous_id_set_chain;to_state_with_chunksiterates solely overself._new_click_ids/_new_user_ids(IDs first seen this day). No historical chunk is scanned or rewritten; new chunks continue thesequence/previous_sha256links from the carried summary (service.py:645-664). Verified by the segmentation test's continuation assertion (next_chunks[0].previous_sha256 == chunks[-1][0],chunks == 3). -
Exactness preserved. The
checkpath (compare_clickhouse_stats_to_manifest) still compares ClickHouse uniqExact against the manifesttotals(events/visits/users/min/max), which come from the incrementally-maintained counters. The required equality testtest_incremental_counters_equal_full_recomputegenuinely compares incremental (day1 →to_state→from_state→ day2) against a single full recompute ofday1+day2, on bothto_manifest_topics()andto_manifest_totals(), withclick-1deliberately re-appearing on day 2 to exercise cross-boundary dedup. It was not weakened to pass. The two-day service test independently asserts day-2topicsandtotalsequal a full recompute overold_batch + published. -
900 000-byte guard is real and explicit.
_assert_message_size(kafka_io.py:270-277) runs before send on everysaveandsave_counter_chunkcall and raises a Russian ValueError naming the byte size and the limit.save_counter_chunkadditionally recomputes the chunk's own SHA-256 and rejects a mismatched address before send (kafka_io.py:247-259). Both failure modes are covered bytest_counter_chunk_has_explicit_kafka_size_and_hash_guards. -
Boundary-ID sourcing is deterministic and correct.
known_counter_ids_from_state(startup_history_artifact.py:295-304) seeds the resume dedup set from generator state:click_ids ofactive_visits, plususer_domain_ids of population users with an active or a finished visit — pre-created but never-visited users are excluded. The only prior-day IDs that re-appear in new-day events are active-visit click_ids (seeded → not double-counted) and returning-user ids (seeded vialast_finished_at→ not double-counted); a first-ever visit for a pre-created user is correctly counted as new. This is confirmed empirically by the two-day service test's equality-with-full- recompute assertion, which would fail on any double-count.
Observation on verify point 2 (non-blocking)
Point 2 asked me to confirm the SHA-256 chunk chain is "validated on read"
and that "a broken/missing chunk fails loudly." It is not — because the
chunks are never read. A whole-repo grep (.py/.sh/.yml/.sql/Makefile)
finds no consumer of COUNTER_TOPIC: there is no load_counter_chunk, no
consumer subscribing to generator_startup_history_counter_chunks, and no
code that re-links or re-hashes the stored chain. The chain's integrity is
validated only on write (each chunk's own hash + the pre-send size guard),
and the manifest's chain summary is checked for internal consistency on
next-day (_validate_id_set_chain, counts must equal totals; head present
iff chunks>0) — but never cross-checked against the actual stored chunk
messages.
Why this is non-blocking: the executor's report oversells the chain as an
on-read integrity mechanism, but nothing depends on reading the chunks.
Exactness of the counts flows entirely through the incremental counters and
the generator-state-seeded dedup (both verified above); a lost/old local
state is required to re-run import, which rebuilds everything from the
artifact. So a corrupt or missing chunk causes no wrong result — the chunk
topic is effectively write-only durable ballast. It grows ~linearly
(~282 KB/model-day; report projects ~98 MB/year), which for this bounded
learning stand is tolerable but is dead weight: the chunks are redundant
with the artifact + generator state. Recommendation (not blocking): either
consume the chunks (e.g. to reconstruct/verify exact sets on resume) or drop
the topic; and correct the report's "validated on read / fails loudly"
wording, which does not match the code. This is calibrated the same way
the original review filed KafkaDataTopicReader dead code — a cleanup
concern, not a correctness defect. It does not reintroduce the ceiling, does
not corrupt data, and does not weaken exactness, so CODE-1's actual defect
is resolved.
CODE-2 (was Low): write ordering — RESOLVED
Order is now chunks → manifest → state on all three paths:
- import:
startup_history_artifact.py:769-776 - backfill:
service.py:485-494 - next-day:
service.py:659-666
The two-day service test asserts the exact sequence
["data", "counter_chunk", "manifest", "state"] for both days
(test_service.py), and the backfill flush-failure test asserts
save_counter_chunk was called while save (manifest) was not and
state was neither saved nor flushed — locking in that a manifest/state write
cannot precede the chunks. A manifest failure therefore can no longer leave
state pointing at an unpublished manifest; the half-advanced window is
narrowed as intended.
Regressions introduced by the fix
None that affect correctness. One non-blocking cleanup item: the
generator_startup_history_counter_chunks topic is write-only (see CODE-1
observation) — dead, linearly-growing data with an unexercised on-read
integrity claim. Low severity, kafka_io.py:212/245-268.
VERDICT: APPROVED
Both findings are genuinely resolved: CODE-1's day-2 message ceiling is gone (bounded manifest + per-chunk-bounded exact sets, next-day O(new day), exactness preserved and honestly tested), and CODE-2's ordering is now chunks→manifest→state on all three paths with test assertions. The only new wart — a write-only chunk topic whose SHA-256 chain is never read back — is non-corrupting cleanup, calibrated as a non-blocking consideration, not a defect. Focused and full generator suites green (108 / 216 passed).