← All challenges
Level · 30 daysSoon

30 Days of Big Data

When data won’t fit in memory — distributed processing with Spark and the patterns of big-data pipelines.

Coming soon. Here’s the planned curriculum. Meanwhile, 7 Days of ML is live and runnable right now.
  1. Module 1Why distributed computing
  2. Module 2MapReduce thinking
  3. Module 3Spark fundamentals
  4. Module 4DataFrames at scale
  5. Module 5Partitioning & shuffles
  6. Module 6A big-data pipeline