Transformers and large language models¶
Attention computed by hand, a small GPT trained from scratch, and the techniques that turn a pretrained model into a useful system: decoding, fine-tuning, preference tuning, retrieval and efficient inference.
This part builds on Neural networks and Natural language processing.
3 of 12 topics ready, listed in reading order
AttentionScaled dot-product attention by hand, masks, multiple heads and why the scaling matters.Planned
The transformer architectureEncoder and decoder blocks, positional encodings, normalization placement and model families.Planned
A tiny GPTA decoder-only language model built and trained from scratch on a CPU.ReadyDecoding strategiesGreedy, beam, temperature, top-k, top-p, constrained decoding and self-consistency.ReadyTransfer learningPretraining and fine-tuning with BERT-style encoders.Planned
Fine-tuning and LoRAFull fine-tuning of a small model, LoRA and QLoRA.Planned
Preference tuningReward models, RLHF with PPO and direct preference optimization.Planned
Retrieval-augmented generationChunking, embeddings, vector search, prompt assembly and evaluation.ReadyInference efficiencyKV caches, grouped-query attention, paged and flash attention, quantization.Planned
Mixture of expertsSparse expert layers, routing and load balancing.Planned
Prompting and structured outputPrompt patterns, conversation state, JSON output and multi-agent setups.Planned
Multimodal modelsContrastive image-text models, captioning and vision-language models.Planned