Profile
The course begins with the fundamentals of scalable data systems. You'll learn what qualifies as big data and why conventional approaches fail as volumes climb.
Distributed storage and processing are then examined. The focus is on how data is partitioned, replicated, and coordinated across machines.
You'll also consider how different processing models change the system's design. The choice between one mode and another affects everything downstream.
The middle sessions move to pipeline design. You will trace the path from data source to storage to analysis, and see how each stage affects reliability.
Particular attention goes to scalability trade-offs. You'll compare vertical scaling with horizontal scaling, and consider when each approach is appropriate.
The course closes with an applied exercise. In it, you assess a hypothetical system, identify its constraints, and outline a scalable redesign.
By the end, you should be able to reason clearly about big data engineering choices. You'll know what questions to ask before choosing an architecture, and what pitfalls to avoid.