- Зачем: - зафиксировать состояние конвейера #5 перед долгой живой приёмкой, чтобы новая сессия продолжила без потери контекста. - Что: - в handoff добавлена дельта 23:10: обе находки FIXED (фрагменты ID с SHA-256-цепочкой), перепроверка линией B — APPROVED; - обновлено состояние стенда: им владеет сценарий приёмки; - в .scratch/issue5-run добавлены свежие отчёты, перепроверка и сценарий приёмки acceptance-issue5.sh. - Проверка: - git show --stat; лог приёмки — до строки SCRIPT_EXIT_CODE=.
10 KiB
Code-correctness review — issue #5 (incremental cumulative manifest counters)
Repo: clickstream-ch-kafka-superset-demo · branch feature/mentee-path · uncommitted working tree. Reviewer: independent, fresh session. Focus: code correctness and cross-module invariants.
Reviewed scope
Working-tree changes vs HEAD:
generator/src/clickstream_generator/startup_history_artifact.py(rolling checksum,to_state/from_state, id-set encode/decode, legacy-checksum fallback, import seeding, artifact stripping)generator/src/clickstream_generator/state.py(cumulative_manifest_countersfield, to_dict/from_dict)generator/src/clickstream_generator/service.py(next-day incremental path, removal ofKafkaDataTopicReaderfrom next-day, guards)generator/src/clickstream_generator/airflow_control.py(assert_next_day_snapshotcumulative guards)- tests: test_startup_history_artifact.py, test_state.py, test_service.py, test_airflow_control.py
- docs: ARCHITECTURE.md, OPERATIONS.md, runbooks/startup-history.md
Checks performed
- Read all changed sources + related generation/runtime/kafka_io to confirm batch composition and Kafka limits.
- Ran targeted new tests and the full generator suite in an ephemeral uv env
(
PYTHONPATH=src, kafka-python 2.0.6): 214 passed (matches report claim). Docker-basedmake test/make lintnot runnable here (no docker socket), consistent with the executor's note. - Measured the encoded manifest id-set size for realistic uuid4 populations (script in scratchpad).
- Loaded and validated the REAL shipped artifact
data/startup_history/reference-world.json.xzto test old-format compatibility empirically (12.4 s validate, OK).
Findings
CODE-1 — High — manifest message grows with world age and breaks the Kafka 1 MB limit after ~1 next-day
File: startup_history_artifact.py:196-207 (to_state embeds _encode_id_set(self.click_ids/user_ids)), :480 (build_manifest puts cumulative_manifest_counters into the manifest), :602 (import), consumed by kafka_io.py:219-237 (manifest producer, default max_request_size) and :365-374 (ensure_topics, no max.message.bytes).
Violated criterion: acceptance "Generation+manifest task on day N takes ~the same time as day 1" and the issue goal "next-day stops getting more expensive as the world ages" — the age-dependent cost is not removed, only moved from read-time to a hard write-time ceiling.
Failure scenario + evidence:
- The exact
click_id/user_domain_idsets are stored, per event, inside the single manifest Kafka message; they grow monotonically with world age. - The shipped reference world already holds 26,083 visits + 4,056 users spanning 3 model days (verified: totals
events=280437, visits=26083, users=4056, max_event_timestamp=2026-01-03). - Measured encoded size (zlib level 9 + base64 over uuid4): reference world id-sets ≈ 0.81 MB already; after ~one next-day (~34.8k visits) ≈ 1.1 MB; a couple of days later multiple MB. Measurement script:
scratchpad/measure.py. - No size override exists anywhere: producer
max_request_size(kafka-python default 1 048 576), brokermessage.max.bytes, manifest topicmax.message.bytes, and consumermax_partition_fetch_bytesare all at defaults (grep acrossgenerator/,infra/, compose returned nothing). The manifest producer (KafkaStartupHistoryManifest.save) does client-side size validation, so a >1 MB manifest raisesMessageSizeTooLargeErroronsend. - Net: the very first next-day after importing the shipped reference world lands at/over the 1 MB boundary; the second next-day exceeds it and
manifest_manager.save()fails. The feature therefore breaks on roughly the day it is meant to make cheap, andmake generated-history-chain-checkon day seams would fail once the manifest cannot be persisted.
Minimal way to verify: uv run python scratchpad/measure.py (shows 0.81 MB at 26 083 visits, 1.59 MB at ~52 k); grep -rn 'max_request_size\|message.max.bytes\|max.message.bytes\|max_partition_fetch_bytes' generator infra docker-compose*.yml → no hits; note the reference totals above.
Fix direction (not prescriptive): raise the four Kafka size limits with real headroom (producer + broker + topic + consumer) and/or stop embedding the full exact sets in a single compact-topic message (chunked/segmented store, or a side artifact). This must be resolved or explicitly accepted as a bounded-world limitation before live acceptance — it is currently a silent day-2 break.
CODE-2 — Low — next-day persists state before manifest, widening the partial-write window
File: service.py:641-644 (state.save → state.flush → manifest.save → manifest.flush), vs import order startup_history_artifact.py:637-639 (manifest first, then state).
Violated criterion: state/manifest coupling robustness under crash / failed write.
Failure scenario + evidence: the new day's events are published and flushed at service.py:614 before either compact write. State (which carries only the small manifest_sha256 reference) is then saved and flushed first; if the manifest write then fails — which CODE-1 makes likely — or the process dies between, state points at the new cumulative reference while the manifest is stale. This is caught fail-safe on the next run: _run_next_day/assert_next_day_snapshot raise IncompatibleStateError/RuntimeError because cumulative_counter_reference(counter_state) != counter_reference (guards at service.py:530-556, airflow_control.py:176-202). So there is no silent corruption, but the stand is left half-advanced (new events already in Kafka) and requires a clean re-import; manifest-first ordering would shrink the window. The mismatch path IS exercised by tests (test_airflow_control.py:152-155, test_service.py next-day guards). Non-corrupting, hence Low.
Minimal way to verify: read service.py:614-644; observe events flushed then state-before-manifest; the guard/error path is covered by the added tests.
Non-blocking considerations
-
Rolling checksum math is sound: per-event SHA-256 digest summed mod 2^256 is associative and commutative, so restore-then-append equals a single full recompute regardless of batch boundaries or order. Serialization (
json.dumps(sort_keys=True, ensure_ascii=True)) is byte-identical to the pre-existing/legacy checksum input (confirmed by diff). Min/max are order-independent aggregates; visit counts are set-based. The invariant "incremental == full recompute" genuinely holds. I traced the device/geo timestamp derivation (ManifestCounters.add_batch,runtime.py:487-488releases a page-view's 4 topic events together) and confirmed per-batch vs single-batch derivation converge to the same topic min/max. -
Old-format compatibility works on the REAL artifact, not just synthetically: the shipped
reference-world.json.xzstores the legacy concatenated SHA-256 (verifiedstored == legacy: True,stored == new-sum: False), andvalidate_startup_history_artifactpasses via the_LegacyManifestCountersfallback in_validate_manifest_topics. Import then recomputes new-sum counters and overwritesmanifest["topics"]/totals+ cumulative section before writing to Kafka, so the post-import stand is fully new-format and self-consistent. No false-fail. False-pass is negligible: acceptance requires totals (recomputed from events) to match AND (new-sum OR legacy checksum) to match, so a tamper that preserves all totals and collides a checksum is required. -
Equality test is honest:
test_incremental_counters_equal_full_recomputecomparesto_manifest_topics()(rows, per-topic checksum, min/max) ANDto_manifest_totals()(events/visits/users/min/max), round-trips throughto_state/from_state, and its two-day fixture makesclick-1span both days — exercising the cross-batch device/geo derivation, not just self-contained batches. -
Compatibility test is synthetic:
test_legacy_artifact_checksum_remains_validreconstructs the legacy checksum in-test rather than pinning the real fixture. The real file passes (I verified manually), so the risk is covered empirically today, but the suite would not catch a future regression against the shipped artifact. A fast smoke test loadingreference-world.json.xzand asserting validity would guard it (caveat: ~33 MB xz, ~12 s validate — heavier than repo "small slice" test norms). -
to_state()is recomputed twice per next-day and per backfill (once for the state reference hash atservice.py:629/469, once insidebuild_manifestat:480), each re-running zlib level 9 over the full ~1 MB id-set. Deterministic in-process so correct, but wasteful; import correctly computes it once and reuses (startup_history_artifact.py:595). -
manifest["topics"]is duplicated insidemanifest["cumulative_manifest_counters"]["topics"]— the same per-topic block is stored twice in the message, adding to CODE-1's size pressure. -
KafkaDataTopicReader(kafka_io.py:284) is now dead code — the class remains but has no callers after the next-day rewrite. Cleanup opportunity. -
Russian error texts for the old-state failure are clear and actionable (e.g. "накопительные счётчики state и manifest не совпадают; повторно запустите import эталонного мира") — they name the fix (re-run import) as required by the mandate.
-
_decode_id_setis robust against corrupt payloads: catches ValueError/TypeError/binascii.Error/zlib.error, enforces list-of-str, and rejects duplicates, all raising clean Russian ValueErrors surfaced asIncompatibleStateError.
VERDICT: CHANGES_REQUESTED
CODE-1 (High) requires a change: the incremental manifest embeds unbounded, world-age-growing exact ID sets in one Kafka message and breaks the default 1 MB producer/broker limit after roughly one next-day on the shipped reference world, silently reintroducing the age-dependent failure the issue set out to remove. CODE-2 (Low) is a fail-safe ordering weakness. The checksum math, incremental-vs-full equality, and real old-artifact compatibility are all correct and verified.