Reliability telemetry
Track accepted, routed, delivered, read, retried, and failed outcomes across regions and devices.
Preparing chapter
The curriculum shell is ready while the requested chapter is being prepared. You can wait here or return to the design library.
WhatsApp · Big Data Engineering · Workload before tools
Translate 100 billion daily messages into peak event throughput, storage, freshness, privacy, replay, and recovery requirements.
01 · Functional scope
Messaging features provide context. The interview focus is the data work required to observe, analyze, govern, and safely improve the system.
Track accepted, routed, delivered, read, retried, and failed outcomes across regions and devices.
Compute end-to-end delivery latency and service-stage latency at p50, p95, and p99.
Measure upload success, transfer duration, call setup, jitter, loss, bitrate, and disconnects.
Detect regional degradation, queue growth, hot groups, broker lag, and abnormal traffic in seconds.
Publish privacy-safe DAU, adoption, experiment, reliability, and capacity datasets.
Reprocess history after a bug, late event, contract change, or metric-definition correction.
Boundary: operational metadata and aggregates only; encrypted message bodies are not an analytics source.
02 · Freshness
Incident detection
Live reliability
Experiment guardrails
Product analytics
Certified history
03 · Capacity + sizing
The source anchor is 100B messages/day. Every result below follows from the visible assumptions, uses decimal units, and separates logical volume from physical replication.
Derived capacity envelope
Decimal units · illustrative interview assumptions · before protocol overhead
Average message ingress
1.16M/s100B messages/day ÷ 86,400 seconds/day = 1.157M messages/sPeak message ingress
5.79M/s1.157M messages/s × 5 = 5.787M messages/sAverage analytical events
4.28M/s370B events/day ÷ 86,400 seconds/day = 4.282M events/sPeak analytical events
21.41M/s5.787M messages/s × 3.7 = 21.412M events/sPeak logical ingress
13.92 GB/s21.412M events/s × 650 bytes/event ÷ 10⁹ = 13.918 GB/sPlanned broker writes
54.28 GB/s13.918 GB/s × 3 replicas × 1.30 = 54.280 GB/sCapacity formula ledger
messages/day × events/messageLifecycle amplification: accepted, routed, delivered, read, retry, device, group, media, and system facts.
events/day × average event bytesBefore compression, replication, indexes, table metadata, and snapshots.
logical TB/day × retention × RFCapacity floor; add filesystem, protocol, rebalance, and operational reserve separately.
logical TB/day ÷ compression60.1 TB/day × 365 days ÷ 1,000 TB/PB = 21.95 PB/year before snapshots and serving copies.
peak MB/s × headroom ÷ tested partition MB/sA throughput floor, not a final topic count. Validate the 10 MB/s budget on the real broker, record, key, and replication mix.
peak events/s × headroom ÷ tested task rateBenchmark stateful jobs with real skew, checkpointing, serialization, and sink latency before provisioning.
messages/day × envelope bytes/messageAt RF3: 300 TB/day physical. Keep only delivery-required hot data; this is separate from analytics and media.
peak logical GB/s × RF × headroomExcludes protocol overhead, cross-AZ topology effects, replication catch-up, and backfill traffic.
04 · From math to architecture
Capacity math earns its place when it determines partitioning, cluster isolation, state limits, storage layout, or recovery strategy.
Separate critical lifecycle, high-volume telemetry, and bulk backfill traffic so one workload cannot consume every broker or quota.
Key only where ordering is required, salt aggregation keys, monitor partition imbalance, and give hot-key jobs explicit state limits.
Use incremental checkpoints, durable remote state, bounded windows, and a tested restart time rather than promising magical exactly-once behavior.
Write staged files, commit Iceberg snapshots atomically, compact small files, and evolve partitions without rewriting all history.
Reprocess from Kafka or immutable raw files into staging, compare results, then publish a new version without starving live consumers.
Tokenize early, minimize columns, enforce purpose-bound access, propagate deletion, audit reads, and expire data by product policy.