Commerce lifecycle truth
Preserve customer, catalog, cart, order-item, payment, inventory, fulfillment, return, and seller facts at their correct grains.
Preparing chapter
The curriculum shell is ready while the requested chapter is being prepared. You can wait here or return to the design library.
Amazon · Big Data Engineering · Workload, guarantees, and scale before tools
Translate orders, behavior, payments, inventory, sellers, and campaign peaks into explicit functions, freshness targets, correctness rules, capacity estimates, and architecture decisions.
Functional scope
Product features provide context. The requirements chapter focuses on the data products, guarantees, and recovery behavior the platform must own.
Preserve customer, catalog, cart, order-item, payment, inventory, fulfillment, return, and seller facts at their correct grains.
Expose current order failures, payment health, inventory pressure, fulfillment delay, and pipeline freshness within seconds to minutes.
Publish auditable GMV, captures, settlements, fees, taxes, refunds, chargebacks, and seller payouts without double counting.
Provide tenant-isolated order, conversion, inventory, fulfillment, return, and settlement products with clear completeness status.
Create point-in-time-correct behavioral, transaction, and risk features for online decisions and reproducible training.
Support funnels, cohorts, experiments, search quality, supply planning, executive BI, and long-range trend analysis.
Rebuild affected history after late data, a producer defect, a rule change, a failed job, or an incorrect metric release.
Classify PII and payment fields, enforce seller isolation, retain lineage, audit access, and propagate retention and deletion.
Boundary: checkout, payment authorization, inventory reservation, and fulfillment execution remain owned by transactional services; the data platform observes and reconciles them.
Freshness + correctness
One pipeline should not be forced to satisfy fraud, operations, sellers, product analytics, executive BI, and finance with the same latency and correctness contract.
Fraud features
Order operations
Inventory + recommendations
Seller order feed
Funnel + search performance
Executive KPIs
Reconciled finance
Capacity + sizing
These are rounded interview assumptions from the Amazon reference brief, not claims about Amazon internals. The calculator separates average load, campaign peaks, logical volume, replication, and headroom.
Inputs used by every equation
Derived capacity envelope
Decimal units · core behavior + order-domain events · before logs, CDC, catalog, and partner feeds
Average orders
116/sCalculation
10,000,000 orders/day ÷ 86,400 sec/day = 115.74 orders/secCampaign order peak
2.3K/sCalculation
115.74 orders/sec average × 20 peak = 2,314.81 orders/secPeak order events
69.4K/sCalculation
10,000,000 orders/day × 30 events/order = 300,000,000 events/day → ÷ 86,400 = 3,472.22 events/sec → × 20 = 69,444.44 events/secPeak behavior events
347.2K/sCalculation
3,000,000,000 events/day ÷ 86,400 = 34,722.22 events/sec → × 10 = 347,222.22 events/secPeak logical ingress
0.417 GB/sCalculation
(69,444.44 order + 347,222.22 behavior) events/sec × 1000 bytes/event ÷ 1,000,000,000 = 0.417 GB/secPlanned broker writes
1.625 GB/sCalculation
0.417 GB/sec logical × 3 replicas × 1.30 headroom = 1.625 GB/secModeled core data
3.30 TB/day(300,000,000 order + 3,000,000,000 behavior) events/day × 1000 B ÷ 10¹² = 3.30 TB/day
Kafka 7-day physical
69.30 TB3.30 TB/day × 7 days × 3 replicas = 69.30 TB
Core annual history
1.20 PB/year3.30 TB/day × 365 days ÷ 1,000 TB/PB = 1.20 PB/year
From requirements to architecture
Numbers and SLOs matter only when they determine isolation, keys, state, storage, publication, governance, or recovery.
Separate order and payment lifecycles, high-volume behavior, CDC, partner feeds, and replay traffic so a campaign cannot starve critical facts.
Partition only where lifecycle order is required, retain source sequence, and monitor hot products, sellers, and campaign partitions.
Serve provisional fraud and operations signals quickly, then correct them through complete lakehouse replay and financial reconciliation.
Capacity follows lifecycle events, impressions, telemetry, and peak skew—not the average order count alone.
Replay immutable Bronze into staging, compare counts and money totals, pass quality gates, and atomically publish a new Iceberg snapshot.
Tokenize early, isolate sellers, restrict columns and rows, audit reads, honor residency, and propagate retention and deletion.