Average messages
1.157M/s100B messages/day ÷ 86,400 seconds/dayPreparing chapter
The curriculum shell is ready while the requested chapter is being prepared. You can wait here or return to the design library.
WhatsApp · Big Data Engineering · Durable handoff, bounded replay
Design WhatsApp's asynchronous event handoff, domain topics, partition keys, capacity envelope, producer guarantees, and replay windows without turning Kafka into the messaging database.
01 · Event handoff
The ingestion boundary converts committed operational changes into durable analytical events. Inspect each stage to see exactly what Kafka’s acknowledgement does—and does not—mean.
Continue the event-source contract
Ingestion preserves the stable event_id, source event_time, contract version, tokenized entity keys, producer and privacy class defined in Event Sources. It adds trustworthy ingestion time and Kafka coordinates.
02 · Topics + partition keys
There is no universal partition key. Separate domains by criticality and volume, preserve only the order a consumer can defend, and make hot-group behavior explicit.
03 · Capacity
Reuse the Requirements baseline, calculate two independent partition floors, then distribute the result across regional domain clusters instead of inventing one giant global Kafka cluster.
Fleet-wide peak planning envelope
Decimal units · illustrative interview assumptions · validate with production-format benchmarks
Average messages
1.157M/s100B messages/day ÷ 86,400 seconds/dayPeak messages
5.787M/s1.157M messages/s × 5 peak-to-average ratioPeak analytical events
21.412M/s5.787M peak messages/s × 3.7 events/messagePeak logical ingress
13.918 GB/s21.412M events/s × 650 bytes/event ÷ 10⁹Floor A · byte throughput
The 10 MB/s allowance is a conservative benchmark input for the chosen record size, compression, acknowledgement and replication settings—not a Kafka constant.
Floor B · event processing
Benchmark the slowest important consumer with real dedupe state, schema decoding, checkpointing and sink latency.
Defendable partition answer
max(1,810 byte lanes, 557 processing lanes) = 1,810 minimum fleet-wide lanes
This is a throughput floor, not one topic’s partition count. Allocate regional peak shares, then size each message, delivery, call and telemetry topic independently. Add enough partitions for broker/AZ loss, hot keys, consumer parallelism and forecast growth.
Broker write check
54.28 GB/s13.918 GB/s logical × RF3 × 1.30 headroomAlso model consumer egress, replica catch-up, disk, leader imbalance and rebalance traffic before choosing broker counts.
04 · Reliability
Kafka provides durable at-least-once transport. End-to-end correctness comes from the source commit, stable identities, checkpointed consumers, idempotent effects, reconciliation, and explicit workload isolation.
Producer
acks=all · idempotence · compression · bounded retry
Broker
RF3 · min ISR · quotas · rack/AZ awareness
Consumer
checkpoint · dedupe · idempotent sink · lag SLO
05 · Retention + replay
Retention follows event criticality, volume and recovery time. Kafka handles bounded hot replay; immutable Bronze storage supports economical history and backfills after the log window closes.
Topic family
Retention
3–7 days
Policy
deleteEnough hot replay for delivery metrics and incident recovery; Bronze holds longer history.
Topic family
Retention
7–14 days
Policy
deleteA longer diagnostic window for retry, expiry and regional failure analysis.
Topic family
Retention
6–48 hours
Policy
deleteExtreme volume; land raw evidence quickly and keep Kafka focused on transit.
Topic family
Retention
Current + tombstones
Policy
compact,deleteDistribute the latest keyed state while preserving controlled deletion history.
Topic family
Retention
7–30 days
Policy
deleteKeep failures long enough to correct and replay, with restricted access and reason codes.
Controlled replay
Identify the topic, partitions, offsets, event versions, regions and time range affected.
Replay with a separately throttled consumer group so recovery cannot starve live delivery metrics.
Run corrected code into a new table or serving version; never overwrite the current product in place.
Compare source, Kafka, Bronze and output counts, then atomically promote or roll back the new version.
If the required offsets have expired, read the immutable Bronze records with their original topic, partition, offset and schema ID. Kafka is the hot transport log; Bronze is the long-lived replay evidence.