Natural language processing¶
Language before transformers: text processing, statistical language models, classic classifiers, word representations and recurrent sequence models.
This part builds on Machine learning and Neural networks.
0 of 10 topics ready, listed in reading order
Text processingRegular expressions, tokenization, byte-pair encoding, normalization, stemming and lemmatization.Planned
Edit distanceMinimum edit distance by dynamic programming with alignments.Planned
Grammars and parsingContext-free grammars, ambiguity, CKY and dependency parsing.Planned
N-gram language modelsCounting, smoothing, interpolation, perplexity and sampling.Planned
Text classificationSentiment analysis with naive Bayes, logistic regression and lexicons.Planned
Text representationsBag of words, tf-idf, word2vec and GloVe, and contextual embeddings.Planned
Recurrent networksRNNs with backpropagation through time, LSTMs and GRUs.Planned
Sequence to sequenceEncoder-decoder models, attention and beam search.Planned
Dialogue systemsFrom ELIZA to frame-based and dialogue-state systems.Planned
Low-resource languagesWhat changes for languages with little data, using Uzbek as the example.Planned