Data at scale¶
What changes when data no longer fits on one machine: the MapReduce model, distributed storage and compute, link analysis, pattern mining and recommendation. Every algorithm runs first in plain Python with a trace of each step, then in a scale-out form.
This part builds on Machine learning.
2 of 7 topics ready, listed in reading order
MapReduceMap, shuffle and reduce, combiners and partitioners, with a small engine that prints every intermediate step.Ready
Distributed storage and computeGFS and HDFS, YARN, Spark's RDDs and DAGs, and today's lakehouse formats.Planned
PageRankThe random surfer, damping and teleportation, power iteration and a MapReduce formulation.ReadySearch enginesInverted indexes, tf-idf and BM25, combined with link-based ranking.Planned
Frequent pattern miningSupport, confidence and lift, Apriori, FP-growth and distributed mining.Planned
Recommender systemsCollaborative filtering, MinHash and LSH, matrix factorization and embedding retrieval.Planned
The analytics workflowAn end-to-end analysis of an open dataset, from cleaning to verified conclusions.Planned