Machine learning¶
Classical supervised and unsupervised learning, built from scratch and then checked against scikit-learn. The part ends with the evaluation mistakes that make good-looking numbers meaningless.
This part builds on Foundations.
13 of 13 topics ready, listed in reading order
The machine learning landscapeSupervised, unsupervised and reinforcement learning, the modelling workflow and how the topics connect.ReadyLinear regressionLeast squares, the normal equation, gradient descent, feature scaling and ridge regularization.ReadyLogistic regressionThe sigmoid, cross-entropy and its gradient, softmax regression and regularization.ReadyEvaluation metricsConfusion matrices, precision, recall, F1 averaging, ROC and PR curves, imbalance and cross-validation.ReadyNaive BayesBayes' rule as a classifier, multinomial and binary variants, smoothing and log-space arithmetic.Readyk-nearest neighboursDistance metrics, choosing k, scaling and the curse of dimensionality.ReadyDecision treesEntropy, information gain, gain ratio and Gini, with a traced ID3 build.ReadyEnsemblesBagging, random forests, AdaBoost round by round, gradient boosting and stacking.ReadySupport vector machinesThe margin, hinge loss, soft margins and kernels.ReadyClusteringK-means and k-means++, silhouette and elbow, hierarchical, mean shift and spectral clustering.ReadyDimensionality reductionPCA through the covariance matrix and the SVD, explained variance, and a look at t-SNE and UMAP.ReadyAnomaly detectionStatistical baselines, Isolation Forest from scratch and reconstruction-error detectors.ReadyData leakage and pitfallsLeaked labels, preprocessing fitted before the split, augmentation at test time and target leakage, each with a runnable demo.Ready