ELMo and ULMFiT: Transfer Learning for NLP
Covers ELMo and ULMFiT, the advance methods that established transfer learning for NLP in 2018.
110 items · Page 2 of 3
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.

Covers ELMo and ULMFiT, the advance methods that established transfer learning for NLP in 2018.
Covers OpenAI's GPT-1 and GPT-2 models. Explains how autoregressive pretraining with transformers enabled transfer learning across NLP tasks.
Covers BERT (Bidirectional Encoder Representations from Transformers), including masked language modeling, bidirectional context understanding.
Explains how XLNet, RoBERTa, and ALBERT refined BERT through permutation language modeling, optimized training procedures, and architectural efficiency.
Covers preference-based learning, the framework developed by Christiano et al. in 2017.
Covers the Transformer architecture, including self-attention mechanisms, multi-head attention, positional encodings.
Covers Wikidata, the collaborative multilingual knowledge base launched in 2012. Explains how Wikidata changed structured knowledge representation.
Covers FastText and subword tokenization, including character n-gram embeddings, handling out-of-vocabulary words, morphological processing.
Covers residual connections, the architectural innovation that solved the vanishing gradient problem in deep networks.
Covers Google's transition to neural machine translation in 2016. Explains how GNMT replaced statistical phrase-based methods with end-to-end neural networks.
Covers sequence-to-sequence neural machine translation, the 2014 advance that changed translation from statistical pipelines to end-to-end neural models.
Covers GloVe (Global Vectors) and the Adam optimizer, two early 2014 developments that changed neural language processing.
The application of deep neural networks to speech recognition in 2012, led by Geoffrey Hinton and his colleagues, marked a major advance.
Covers Memory Networks, the 2014 advance that introduced external memory to neural networks.
Covers neural information retrieval, the advance approach that learned semantic representations for queries and documents.
Covers layer normalization, the normalization technique that computes statistics across features for each example.
Covers word2vec, the advance method for learning dense vector representations of words.
Covers SQuAD (Stanford Question Answering Dataset), the benchmark that established reading comprehension as a flagship NLP task.
DeepMind's WaveNet changed text-to-speech synthesis in 2016 by generating raw audio waveforms directly using neural networks.
Examines IBM Watson's historic victory on Jeopardy! in February 2011, examining the system's architecture, multi-hypothesis answer generation.
Freebase introduced a collaborative knowledge graph for representing entities and relationships. Covers its schema, community editing, applications, and legacy.
Covers Latent Dirichlet Allocation (LDA), the advance Bayesian probabilistic model that changed topic modeling.
Examines Yoshua Bengio's early 2003 Neural Probabilistic Language Model that changed NLP by learning dense, continuous word embeddings.
In 2005, the PropBank project at the University of Pennsylvania added semantic role labels to the Penn Treebank.
Traces statistical parsing's major shift from rule-based to data-driven approaches. Explains how Michael Collins's 1997 parser.
How phrase-based translation (2003) extended IBM statistical MT to phrase-level learning, capturing idioms and collocations.
How Maximum Entropy models and Support Vector Machines changed NLP in 1996 by enabling flexible feature integration for sequence labeling, text classification.
In 1998, Charles Fillmore's FrameNet project at ICSI Berkeley released the first large-scale computational resource based on frame semantics.
Examines John Searle's influential 1980 thought experiment challenging strong AI.
Examines William Woods's influential 1970 parsing formalism that extended finite-state machines with registers, recursion, and actions.
Latent Semantic Analysis uses matrix factorization to uncover hidden relationships between terms and documents for retrieval, clustering, and topic discovery.
Conceptual Dependency maps sentences to primitive actions and roles. Covers Roger Schank's theory, semantic equivalence, inference, and question answering.
Andrew Viterbi's 1967 dynamic programming algorithm finds the most likely hidden-state sequence, with applications in HMMs, tagging, and speech recognition.
The 1954 Georgetown-IBM demonstration marked an important moment in computational linguistics.
Covers BM25, the major probabilistic ranking algorithm that changed information retrieval. Explains how BM25 solved TF-IDF's limitations.
Traces Richard Montague's major framework for formal natural language semantics. Explains how Montague Grammar introduced compositionality, intensional logic.
Michael Lesk's 1983 algorithm resolves word senses by comparing dictionary definitions with surrounding context. Covers the method, code, and limitations.
Explains how Gerard Salton's Vector Space Model and TF-IDF weighting changed information retrieval in 1968.
Examines Noam Chomsky's early 1957 work "Syntactic Structures" that changed linguistics, challenged behaviorism.
In 2002, IBM researchers introduced BLEU (Bilingual Evaluation Understudy), revolutionizing machine translation evaluation.
Conditional Random Fields model structured predictions without assuming independent labels. Covers the 2001 paper, feature design, inference, and later uses.
Natural language processing underwent a fundamental shift from symbolic rules to statistical learning.
Claude Shannon's 1948 work on information theory introduced n-gram models, one of the most core concepts in natural language processing.
In 1950, Alan Turing proposed a deceptively simple test for machine intelligence, originally called the Imitation Game.
Joseph Weizenbaum's ELIZA, created in 1966, became the first computer program to hold something resembling a conversation.
Hidden Markov Models changed speech recognition in the 1970s by introducing a clever probabilistic approach.
In 1958, Frank Rosenblatt created the perceptron at Cornell Aeronautical Laboratory, the first artificial neural network.
In 1968, Terry Winograd's SHRDLU system demonstrated a major approach to natural language understanding by grounding language in a simulated blocks world.
Create a free account to keep your reading organized, join thoughtful discussions, and get more from every chapter.
Free to join · Takes less than a minute
Why
I created this space so readers can learn together, ask questions, and make sense of difficult ideas.
