AI Timeline
Three waves, fifty milestones: what each wave bet on, where it hit a wall, what it left behind. Every row names its source and the day it was checked.
First wave: symbols and expert systems (1943–1997)
The bet of this wave: intelligence = search and reasoning over symbols, and people can write knowledge as rules and hand it to the machine. It hit the same wall twice. Combinatorial explosion kept search stuck in toy worlds. Common sense is too large and too implicit to write by hand (the knowledge acquisition bottleneck). The 1973 Lighthill report and the 1974 DARPA cuts were the first winter. The 1987 collapse of the Lisp machine market and the exposed maintenance cost of expert systems were the second. What it left behind did not disappear: search, planning and knowledge representation are still in the agent toolbox today. They are just no longer the lead.
1943
McCulloch & Pitts neuron model — writes the neuron as a threshold logic unit and proves that networks of them can compute any logical function. "Brain" and "computation" enter the same mathematics for the first time, and both the neural network path and the symbolic logic path branch from here.
A Logical Calculus of the Ideas Immanent in Nervous ActivitySource · checked 2026-10-08
1950
Turing, "Computing Machinery and Intelligence" — replaces the unanswerable "can machines think?" with an operational imitation game, and predicts machines that learn.
Computing Machinery and IntelligenceSource · checked 2026-10-08
1956
Dartmouth summer workshop — names the field artificial intelligence. The proposal's bet: every aspect of intelligence can in principle be described so precisely that a machine can simulate it.
1958
Rosenblatt's perceptron — weights are learned from examples, with a convergence proof, and it was built in hardware. The first learning classifier and the start of connectionism.
The Perceptron: A Probabilistic Model for Information Storage and Organization in the BrainSource · checked 2026-10-08
1959
Samuel's checkers program — tunes its evaluation function through self-play and ends up better than its author. The term "machine learning" comes from here.
1965–1976
DENDRAL → MYCIN — rules from chemists and doctors replace general search, and diagnosis rivals experts. Domain knowledge matters more than the reasoning algorithm, and knowledge engineering methodology starts here. MYCIN was never used clinically (liability and integration problems).
1966
ELIZA — pattern matching creates the illusion of being understood, and even the author was startled by how invested users became. The gap between "looks like it understands" and "really understands" is seen for the first time.
1969
Minsky & Papert, "Perceptrons" — proves a single-layer perceptron cannot even represent XOR, and nobody knew how to train multiple layers. Funding moves to the symbolic camp, and neural networks go quiet for over ten years.
Perceptrons: An Introduction to Computational GeometrySource · checked 2026-10-08
1970–1972
SHRDLU — natural language understanding in a blocks world. The peak demo of the symbolic camp, and it also shows this path cannot leave the toy world.
1973–1974
Lighthill report and DARPA cuts — the British government judges that AI failed to deliver (combinatorial explosion is the core argument), and the US cuts funding in step: the first winter.
1980
XCON / R1 (DEC) — an expert system that configures VAX orders saves money at scale for the first time and sets off the 1980s expert system industry.
1984
Cyc starts — Lenat bets on hand-coding all common sense, and keeps at it for nearly forty years. The most thorough experiment of the symbolic camp, and its outcome is the best textbook on why knowledge cannot be written by hand.
1987
Lisp machine market collapses, second winter — special-purpose hardware is replaced by general workstations, and the maintenance cost and brittleness of expert systems are exposed. Japan's Fifth Generation Computer project (1982–1992) also missed its goals. First wave: symbols and expert systems
1997
Deep Blue beats Kasparov — search + special-purpose hardware + a hand-built evaluation function wins at chess. It proves brute force works in a narrow domain, and it is the closing high point of the symbolic search path.
Second wave: statistical learning and deep learning (1986–2016)
The bet changed: knowledge is learned from data, not written by hand. Probability replaces logic. Features are not designed by hand either; the network learns the representation. Backpropagation existed in 1986, so why did deep learning only take off in 2012? What was missing was not the algorithm but three things arriving together: million-scale labeled data (ImageNet, 2009), cheap parallel compute (GPU + CUDA, from 2007), and small tricks that make deep networks trainable (ReLU, dropout, good initialization). AlexNet in 2012 put all three together, and over the next four years vision, speech, translation and games fell one by one.
1986
Backpropagation (Rumelhart, Hinton, Williams) — multi-layer networks can be trained on error, and hidden layers learn representations on their own. It answers the question left by "Perceptrons", and connectionism revives.
Learning representations by back-propagating errorsSource · checked 2026-10-08
1988
Pearl, "Probabilistic Reasoning in Intelligent Systems" — Bayesian networks make reasoning under uncertainty computable, and probability returns to the AI mainstream.
1989 / 1998
LeCun's convolutional networks: zip code recognition → LeNet-5 — puts priors into the network structure (local receptive fields, weight sharing) instead of into rules, and end-to-end gradient training yields a system that can read checks. MNIST becomes the standard dataset.
Gradient-Based Learning Applied to Document RecognitionSource · checked 2026-10-08
1995
SVM (Cortes & Vapnik) — maximum margin + kernel trick, a classifier backed by generalization theory, dominant in the 2000s.
1997
LSTM — gated units solve vanishing gradients, so recurrent networks can remember over long distances. The workhorse of speech and translation in the 2010s.
2003
Neural probabilistic language model (Bengio) — learns a continuous vector for each word, and a neural network predicts the next word. It sidesteps the sparsity problem of n-grams and is a common ancestor of word vectors and LLM.
A Neural Probabilistic Language ModelSource · checked 2026-10-08
2006
Deep belief networks (Hinton) — layer-by-layer pretraining makes deep networks trainable. "Deep learning" becomes a term and draws funding again.
A fast learning algorithm for deep belief netsSource · checked 2026-10-08
2009
ImageNet — ten million-scale human-labeled images. The bet is on data scale, not algorithms, and the 2012 breakthrough happens on it.
2012
AlexNet — a deep convolutional network trained on two GPUs cuts the ImageNet error rate by ten points in one step, and the old path of hand-made features + SVM is dropped. The breakout point of deep learning.
ImageNet Classification with Deep Convolutional Neural NetworksSource · checked 2026-10-08
2013
word2vec — a shallow model learns word vectors on a billion words, and vector arithmetic reveals semantic structure. "Pretrain representations without supervision first" becomes routine in NLP.
Efficient Estimation of Word Representations in Vector SpaceSource · checked 2026-10-08
2013
DQN — a convolutional network learns Atari straight from pixels, one set of hyperparameters across dozens of games. Deep learning and reinforcement learning merge.
Playing Atari with Deep Reinforcement LearningSource · checked 2026-10-08
2014
GAN — a generator and a discriminator train against each other, and it generates without writing a likelihood. The start of the deep generative model branch.
2014
seq2seq — an encoder-decoder LSTM does translation end to end, replacing IBM's 1990s statistical translation. It exposes the fixed-length vector bottleneck.
Sequence to Sequence Learning with Neural NetworksSource · checked 2026-10-08
2014
Attention mechanism (Bahdanau, Cho, Bengio) — at decode time, look back and weight each position of the source sentence, which fixes the drop in quality on long sentences. The direct predecessor of the Transformer.
Neural Machine Translation by Jointly Learning to Align and TranslateSource · checked 2026-10-08
2015
ResNet — residual connections make hundred-layer networks trainable, and it beats humans on ImageNet. Residuals become standard in every deep network from then on.
Deep Residual Learning for Image RecognitionSource · checked 2026-10-08
2016
AlphaGo beats Lee Sedol — deep networks + tree search + self-play (this line started with TD-Gammon's backgammon in 1992). Go was thought to be another ten years away. The public image of deep learning shifts from "recognizes images" to "makes decisions".
Mastering the game of Go with deep neural networks and tree searchSource · checked 2026-10-08
Third wave: scaling and LLM (2017–2026)
The bet: the same skeleton (the Transformer) + more data and compute = predictable growth in capability, and one general pretrained model replaces all task-specific models. Scaling became the theme for three reasons. The Transformer removed recurrence, so training is fully parallel and any amount of compute can be used. GPU clouds turned compute into something you can buy. Scaling laws turned "throw ten times more at it" from faith into arithmetic. After 2022 two more axes were added: alignment (RLHF/DPO decide whether the model "listens") and inference-time compute (o1/R1 show that letting the model "think longer" also buys capability). This wave has not hit a wall yet, but benchmark saturation, running out of data and compute cost are the three walls it is approaching.
2017
Transformer — drops recurrence and keeps only attention, so training is fully parallel. The skeleton of every large model since.
2017
Deep reinforcement learning from human preferences (Christiano et al.) — trains a reward model from pairwise comparisons, and humans see less than one percent of the interactions. The prototype of RLHF. In the same year AlphaZero shows that three board games can be mastered with zero human game records.
Deep Reinforcement Learning from Human PreferencesSource · checked 2026-10-08
2018
ELMo / GPT-1 / BERT — pretraining + fine-tuning becomes the default paradigm in NLP, and within a year BERT pushes GLUE past the human baseline.
BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingSource · checked 2026-10-08
2019
GPT-2 — ten times larger, and it does many tasks without fine-tuning: "tasks are learned along the way". Its staged release of weights makes release safety an issue for the first time.
Language Models are Unsupervised Multitask LearnersSource · checked 2026-10-08
2019
The Bitter Lesson — Sutton says, from seventy years of experience, that human knowledge written in works in the short term, but in the long term loses to general methods driven by compute. The creed of the scaling wave.
2020
Scaling Laws (Kaplan et al.) — loss follows a power law in parameters, data and compute each, holding across seven orders of magnitude. Investment becomes predictable.
Scaling Laws for Neural Language ModelsSource · checked 2026-10-08
2020
GPT-3 — 175 billion parameters, does new tasks from a few examples in the prompt with no weight updates. Prompting becomes a way to program, and the LLM startup rush forms around its API.
Language Models are Few-Shot LearnersSource · checked 2026-10-08
2020–2022
ViT · DDPM · CLIP · Stable Diffusion — the Transformer enters vision, diffusion models become usable, images and text share one vector space, and latent-space diffusion runs text-to-image on consumer graphics cards with open weights. GANs leave the stage.
Learning Transferable Visual Models From Natural Language SupervisionHigh-Resolution Image Synthesis with Latent Diffusion ModelsSource · checked 2026-10-08
2021
AlphaFold 2 — protein structure prediction reaches experimental accuracy, and a fifty-year problem is largely solved. Deep learning enters scientific discovery.
Highly accurate protein structure prediction with AlphaFoldSource · checked 2026-10-08
2022
Chain-of-thought prompting — have the model write out reasoning steps before answering, and it only works on large models. It sparks the "emergent abilities" debate, and two years later reasoning models move it from prompt into training.
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsSource · checked 2026-10-08
2022
InstructGPT and Constitutional AI — SFT + reward model + PPO make a 1.3 billion model preferred over the 175 billion one, the recipe of ChatGPT. In the same year Anthropic replaces part of the human labeling with principles written in text + AI feedback.
Training language models to follow instructions with human feedbackConstitutional AI: Harmlessness from AI FeedbackSource · checked 2026-10-08
2022
Chinchilla — parameters and tokens should scale in the same proportion, and earlier large models were mostly undertrained. It rewrites the scaling recipe, and "smaller model, more data" becomes the mainstream.
Training Compute-Optimal Large Language ModelsSource · checked 2026-10-08
2022-11
ChatGPT — a hundred million users in two months. AI goes from papers to a consumer product and an industry topic. The public start of the third wave.
2023
LLaMA · Llama 2 · Mistral 7B — frontier-class open-weight models. The whole ecosystem of llama.cpp, quantization and local inference grows from here, and Europe's Mistral shows a small model + good data can beat larger models.
LLaMA: Open and Efficient Foundation Language ModelsMistral AISource · checked 2026-10-08
2023
GPT-4 — multimodal, top 10% on professional exams, and the report says outright that architecture and data are not disclosed. The dividing line where frontier labs turn from "publishing papers" to "shipping products".
2023
DPO — folds the reward model and PPO into one classification loss, and preference alignment takes a few dozen lines of code. Alignment goes from big-lab engineering to everyday open-source work.
Direct Preference Optimization: Your Language Model is Secretly a Reward ModelSource · checked 2026-10-08
2024
o1 and MCP — long chains of thought trained with reinforcement learning make "how long it thinks" a new scaling axis (inference-time compute). In November of the same year the Model Context Protocol starts to standardize how models connect to tools. Reasoning and tools are the two cornerstones of the agent era. Third wave: scaling and LLM MCP: one standard socket for every system
2024
Nobel Prize in Physics (Hopfield, Hinton) and in Chemistry (Hassabis, Jumper, Baker) — formal recognition of neural networks and AlphaFold by mainstream science.
2025-01
DeepSeek-R1 — reasoning elicited by pure reinforcement learning, with open weights and published cost. Reasoning models are not one company's secret, and the compute narrative gets discounted for the first time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningSource · checked 2026-10-08
2026
Benchmark saturation becomes the bottleneck — Stanford AI Index 2026: SWE-bench Verified rises from about 60% to nearly 100% in one year, and industry produces over 90% of notable models. Evals cannot keep up with the models, and ARC-AGI moves to interactive environments.
Labs and companies
The organisations the timeline names; the ones in this site's hiring data link to their company page.
Amazon
1994LayersL6
One of the main employers for Applied Scientist roles. It publishes its hiring process and Leadership Principles as pages, which makes them primary source material for behavioral interview prep.
Anthropic
2021LayersL6L5L4Hiring data
A frontier lab, and the maker of Claude. For the career layer, what matters is that it wrote its policy on candidates' use of AI as a public page. The hiring page itself is the best prep material.
In the history layer it stands for the branch of "alignment as a core research direction". It split off from OpenAI in 2021. Constitutional AI: Harmlessness from AI Feedback replaces most human preference labeling with a set of written principles, and has the model critique and rewrite its own answers. Beyond the Claude series, it released the protocol in MCP:给每个系统一个统一的插口, which made "models connecting to tools" an industry standard. In Opinions and predictions, Dario Amodei's long essay is the most often cited judgment for the frontier layer.
- Related
- Constitutional AI: Harmlessness from AI Feedback (released) · MCP: One Standard Socket for Every System (released) · Behavioral Rounds and the Values Round (set)
- Linked from
- Dario Amodei (works at) · Jack Clark (works at) · Jared Kaplan (works at) · Constitutional AI: Harmlessness from AI Feedback (released) · Anthropic 博客(news · engineering · research) (released)
- Sources
- Official hiring page: no degree required. It suggests putting independent research / blog / open source at the very top of your résumé. (checked 2026-10-08) · Candidate AI use policy. The page is marked last updated 2025-07-10. (checked 2026-10-08)
Bell Labs
1925LayersL5
In the 1990s, LeCun's convolutional networks and Vapnik's SVM competed in the same building. The two routes of the second wave (neural networks vs. statistical learning) met here.
- Related
- Yann LeCun (works at) · Vladimir Vapnik (works at) · Gradient-Based Learning Applied to Document Recognition (released) · Support-Vector Networks (released)
- Linked from
- Vladimir Vapnik (works at) · Yann LeCun (works at) · Gradient-Based Learning Applied to Document Recognition (works at) · Support-Vector Networks (works at)
- Sources
- Bell Labs official website (checked 2026-10-08)
Carnegie Mellon University
1956LayersL5
Newell and Simon's Logic Theorist (1956) was the first AI program. XCON, speech recognition (Sphinx) and self-driving (NavLab) all came out of this line of work.
- Related
- The first wave: symbols and expert systems (proposed) · The Second Wave: Statistical Learning and Deep Learning (discusses)
- Sources
- CMU School of Computer Science (checked 2026-10-08)
Cursor
2022LayersL6L4Hiring data
An AI editor company (legal entity: Anysphere). It is one of the fastest-hiring AI-native startups. Its hiring page is a good sample of "work-sample style" role requirements.
- Sources
- Official hiring page. It lists roles only and says nothing about the process. (checked 2026-10-08)
DARPA
1958LayersL5
The main funder of AI in the 1960s and 70s. Its 1974 cuts were the US side of the first AI winter. In the 1980s its Strategic Computing Initiative gave expert systems another push. When you read AI history, its name stands for the funding curve.
- Related
- The first wave: symbols and expert systems (funded) · The first wave: symbols and expert systems (discusses)
- Sources
- DARPA website (checked 2026-10-08)
DeepSeek
2023LayersL5
A lab in Hangzhou funded by the quant fund High-Flyer. Its R1, released in January 2025 with open weights and a publicly stated low training cost, showed that reasoning models are not one company's secret.
- Related
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released) · Liang Wenfeng (works at)
- Linked from
- Liang Wenfeng (works at) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released)
- Sources
- DeepSeek official website (checked 2026-10-08)
Epoch AI
2022LayersL5L3
A nonprofit research group that tracks AI trends. It publishes open datasets on training compute, model size, chip cost, and when we run out of data. When I talk about scaling and compute economics, I check the numbers here first.
- Related
- Epoch AI 数据中心 (released) · Pretraining and scaling (measures) · The economics of compute: where training money goes, how inference is priced (discusses)
- Linked from
- Epoch AI 数据中心 (released)
- Sources
- Official site (checked 2026-10-08)
1998LayersL5
Google Brain (2011) turned deep learning into infrastructure. The Transformer, BERT and TPU all came from there. In 2023 it merged with DeepMind to form Google DeepMind, which entered the race with Gemini.
- Related
- Efficient Estimation of Word Representations in Vector Space (released) · Sequence to Sequence Learning with Neural Networks (released) · Attention Is All You Need (released) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Jeff Dean (works at)
- Linked from
- Ashish Vaswani (works at) · Geoffrey Hinton (works at) · Ian Goodfellow (works at) · Ilya Sutskever (works at) · Jeff Dean (works at) · Noam Shazeer (works at) · Tomáš Mikolov (works at) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Sequence to Sequence Learning with Neural Networks (released) · Efficient Estimation of Word Representations in Vector Space (released)
- Sources
- Google Research official website (checked 2026-10-08)
Google DeepMind
2010LayersL6L5L2
Google's AI lab (DeepMind was founded in 2010 and merged with Google Brain in 2023). It has public official hiring pages for research roles and research engineering roles.
In terms of history, it is the standard-bearer of the second half of the second wave: Playing Atari with Deep Reinforcement Learning taught a network to play Atari from pixels. Mastering the game of Go with deep neural networks and tree search used self-play plus search to beat top human players. Highly accurate protein structure prediction with AlphaFold applied the same approach to protein structures and won a Nobel Prize. Training Compute-Optimal Large Language Models corrected the recipe for scaling: for the same compute, data has to grow together with parameters. After the 2023 merger with Google Brain, the Gemini series put it back on the front line of LLM competition.
- Related
- Playing Atari with Deep Reinforcement Learning (released) · Mastering the game of Go with deep neural networks and tree search (released) · Highly accurate protein structure prediction with AlphaFold (released) · Training Compute-Optimal Large Language Models (released)
- Linked from
- Arthur Mensch (works at) · David Silver (works at) · Demis Hassabis (works at) · Highly accurate protein structure prediction with AlphaFold (released) · Mastering the game of Go with deep neural networks and tree search (released) · Training Compute-Optimal Large Language Models (released) · Deep Reinforcement Learning from Human Preferences (released) · Playing Atari with Deep Reinforcement Learning (released) · Google DeepMind 博客 (released)
- Sources
- Official hiring page. It lays out a four-stage process and says it is tuned per role. (checked 2026-10-08) · Google's How we hire page. The body text is rendered by script. (checked 2026-10-08)
Hugging Face
2016LayersL5
The transformers library and the model hub standardized how pretrained models are distributed. It is the "GitHub" of the open-weights ecosystem.
- Related
- Clément Delangue (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Clément Delangue (works at) · Hugging Face Daily Papers (released)
- Sources
- Hugging Face official website (checked 2026-10-08)
IBM Research
1945LayersL5
Samuel's checkers program (1959), Deep Blue (1997), Watson (2011), and statistical machine translation (1990s) are all here. In each of the three waves, it stood at the peak of the mainstream method of its day.
- Related
- Arthur Samuel (works at) · The first wave: symbols and expert systems (discusses)
- Linked from
- Arthur Samuel (works at)
- Sources
- IBM history page: Deep Blue (checked 2026-10-08)
IDSIA
1988LayersL5
A small lab in Lugano, Switzerland, and the birthplace of the LSTM. Around 2011 it won several vision competitions with GPU convolutional networks, before AlexNet.
- Related
- Jürgen Schmidhuber (works at) · Long Short-Term Memory (released)
- Linked from
- Jürgen Schmidhuber (works at)
- Sources
- IDSIA website (checked 2026-10-08)
Meta
2004LayersL6L5
One of the big-tech employers with the most AI roles. Besides FAIR, it has a superintelligence lab. Its careers page is a good sample of how a big company titles its AI roles.
Meta AI (FAIR)
2013LayersL5
FAIR, which LeCun founded in 2013. PyTorch and LLaMA both came out of it. It is the biggest driver of the open-weights route.
- Related
- LLaMA: Open and Efficient Foundation Language Models (released) · Yann LeCun (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Kaiming He (works at) · Tomáš Mikolov (works at) · Yann LeCun (works at) · LLaMA: Open and Efficient Foundation Language Models (released) · Meta AI 博客 (released)
- Sources
- Meta AI official website (checked 2026-10-08)
Microsoft
1975LayersL5
ResNet came out of Microsoft Research Asia. Since 2019 Microsoft has invested in OpenAI and been its exclusive compute provider, and Copilot made code generation a mass-market product for the first time.
- Related
- Deep Residual Learning for Image Recognition (released) · OpenAI (funded) · Kaiming He (works at)
- Linked from
- Kaiming He (works at) · Deep Residual Learning for Image Recognition (released)
- Sources
- Microsoft Research website (checked 2026-10-08)
Mila – Quebec AI Institute
1993LayersL5
Bengio's lab. In 2014 it produced both the attention mechanism and GANs in the same year. One of the three places (Toronto, Montreal, Edmonton) where Canada kept deep learning alive in academia.
- Related
- Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (released) · Generative Adversarial Nets (released)
- Linked from
- Ian Goodfellow (works at) · Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (works at)
- Sources
- Mila website (checked 2026-10-08)
Mistral AI
2023LayersL5Hiring data
Founded in Paris in 2023. Mistral 7B and Mixtral showed that a small model with good data can beat larger models. The main player in Europe's open-weights approach.
- Related
- Arthur Mensch (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Arthur Mensch (works at)
- Sources
- Mistral AI official website (checked 2026-10-08)
MIT AI Lab / CSAIL
1959LayersL5
The lab Minsky and McCarthy founded in 1959. It was the stronghold of the symbolic school: ELIZA, SHRDLU and the Lisp machines all came out of it. It merged into CSAIL in 2003.
- Related
- The first wave: symbols and expert systems (proposed) · Perceptrons: An Introduction to Computational Geometry (funded) · Marvin Minsky (works at)
- Linked from
- John McCarthy (works at) · Joseph Weizenbaum (works at) · Kaiming He (works at) · Marvin Minsky (works at) · Terry Winograd (works at)
- Sources
- CSAIL website (checked 2026-10-08)
NVIDIA
1993LayersL5
CUDA (2007) turned the GPU into a general-purpose compute device. After AlexNet, it became the physical foundation of deep learning. The compute economics of the third wave revolve around its supply cycle.
- Related
- Jensen Huang (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Jensen Huang (works at)
- Sources
- NVIDIA research page (checked 2026-10-08)
OpenAI
2015LayersL6L5Hiring data
A frontier lab, the author of the GPT series. Research engineer roles are its main hiring line, and its official interview guide is public.
In the history layer, it is the lead of the third wave. Improving Language Understanding by Generative Pre-Training through Language Models are Few-Shot Learners showed that the road of "the same objective, scaled up ten times" could keep going. Scaling Laws for Neural Language Models wrote that down as a formula. Training language models to follow instructions with human feedback taught the continuation machine to follow instructions. ChatGPT then put it in everyone's hands. The later GPT-4 Technical Report no longer discloses the architecture, which marks research moving from papers to products. For the nodes, see The third wave: scaling and LLM and Pretraining and scaling.
- Related
- Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Learning Transferable Visual Models From Natural Language Supervision (released)
- Linked from
- Alec Radford (works at) · Andrej Karpathy (works at) · Daniel Kokotajlo (works at) · Dario Amodei (works at) · Ilya Sutskever (works at) · Jared Kaplan (works at) · Paul Christiano (works at) · Sam Altman (works at) · Learning Transferable Visual Models From Natural Language Supervision (released) · Deep Reinforcement Learning from Human Preferences (released) · Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Microsoft (funded) · OpenAI 博客 (released)
- Sources
- Official interview guide. The fetch tool got a 403, but the search index confirms the page exists (including multilingual versions). (checked 2026-10-08)
Stanford AI Lab (SAIL)
1963LayersL5
Founded by McCarthy in 1963. Expert systems (DENDRAL, MYCIN) came out of it. Forty years later, ImageNet and DPO came from here too. It shows up in all the three waves.
- Related
- The first wave: symbols and expert systems (proposed) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released) · ImageNet (proposed)
- Linked from
- Andrew Ng (works at) · Edward Feigenbaum (works at) · Fei-Fei Li (works at) · John McCarthy (works at) · Terry Winograd (works at) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released)
- Sources
- SAIL website (checked 2026-10-08)
University of Toronto
1827LayersL5
The home of Hinton's lab, where deep learning was kept alive through the winter. Deep belief networks (2006) and AlexNet (2012) were both built here.
- Related
- Geoffrey Hinton (works at) · ImageNet Classification with Deep Convolutional Neural Networks (released) · A fast learning algorithm for deep belief nets (released)
- Linked from
- Alex Krizhevsky (works at) · Geoffrey Hinton (works at) · A fast learning algorithm for deep belief nets (works at)
- Sources
- University of Toronto, Department of Computer Science (checked 2026-10-08)
Benchmarks
What the field measured itself by, and when each stopped being hard.
MNIST
1998LayersL5L2
What it measures: 10-class classification of 70,000 handwritten digits, the "Hello World" of the second wave. When it saturated: by the early 2010s the error rate was below 0.3%, close to the label noise. Now it is only used for teaching and sanity checks.
- Related
- Classical machine learning: learning from examples (measures) · Yann LeCun (proposed) · Gradient-Based Learning Applied to Document Recognition (discusses)
- Linked from
- Gradient-Based Learning Applied to Document Recognition (measures)
- Sources
- MNIST entry (LeCun's original site is no longer reliable) (checked 2026-10-08)
ImageNet
2009LayersL5L2
What it measures: image classification over 1,000 classes and about a million images (ILSVRC). When it saturated: in 2012, AlexNet cut the top-5 error rate to 15%. In 2015, ResNet got below the human level of about 5%. In 2017, the challenge was discontinued. It is the prototype of "dataset-driven breakthroughs."
- Related
- Classical machine learning: learning from examples (measures) · Fei-Fei Li (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (discusses) · Deep Residual Learning for Image Recognition (discusses)
- Linked from
- Fei-Fei Li (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (measures) · Deep Residual Learning for Image Recognition (measures) · Stanford AI Lab (SAIL) (proposed)
- Sources
- ILSVRC challenge page (2010–2017) (checked 2026-10-08)
SQuAD
2016LayersL5L2
What it measures: reading comprehension, extracting answers from Wikipedia paragraphs. When it saturated: BERT beat human EM/F1 in 2018. After version 2.0 added "unanswerable" questions, models beat humans on that too in 2019. Since then, reading comprehension has not been a standalone benchmark.
- Related
- LLM: how the model works (measures) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (discusses)
- Linked from
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (measures)
- Sources
- Official leaderboard (checked 2026-10-08)
GLUE / SuperGLUE
2018LayersL5L2
What it measures: the average score over nine sentence-level NLP tasks. When it saturated: it was released in 2018, BERT pushed it past the human baseline within a year, it was replaced by the harder SuperGLUE in 2019, and SuperGLUE was also surpassed in 2021. The first "saturated in one year" case of the pretrained-model era.
- Related
- LLM: how the model works (measures) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (discusses)
- Linked from
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (measures)
- Sources
- The SuperGLUE paper (2019) explains that GLUE had been passed by models beating the human baseline. (checked 2026-10-08)
ARC-AGI
2019LayersL5L2
What it measures: grid puzzles where you infer a rule from a few examples, designed so you can't memorize them. When it saturated: the 2019 version was cracked by reasoning models with heavy compute in late 2024. In 2025 it was replaced by ARC-AGI-2. In 2026 frontier models passed 90%, and ARC-AGI-3 moves to interactive environments. It is the proving ground for the "generalization vs memorization" debate.
- Related
- LLM: how the model works (measures) · Evals: knowing whether it got better (measures) · The third wave: scaling and LLM (discusses)
- Sources
- ARC Prize official site: ARC-AGI-2 (2025) page (checked 2026-10-08)
MMLU
2020LayersL5L2
What it measures: multiple-choice questions across 57 subjects, from elementary math to professional law. When it saturated: GPT-4 (2023) scored 86%. In 2024, frontier models crowded into 88–90%, and some questions turned out to be wrong. MMLU-Pro appeared because of this. It is the yardstick for the "knowledge" axis in the LLM era, and it has been saturated for four years.
- Related
- LLM: how the model works (measures) · GPT-4 Technical Report (discusses)
- Linked from
- GPT-4 Technical Report (measures)
- Sources
- The MMLU-Pro paper (2024) says MMLU performance has plateaued. (checked 2026-10-08)
GSM8K
2021LayersL5L2
What it measures: 8.5k grade-school word problems that test multi-step arithmetic reasoning. When it saturated: chain-of-thought prompting first showed its power on it (2022). In 2024, frontier models passed 95%. After that, people moved to competition problems such as MATH and AIME.
- Related
- LLM: how the model works (measures) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (discusses)
- Linked from
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (measures)
- Sources
- Paper page (OpenAI, 2021) (checked 2026-10-08)
HumanEval
2021LayersL5L2
What it measures: 164 hand-written Python function problems, scored by unit test pass rate. When it saturated: 28.8% at the Codex release, 67% for GPT-4 (2023), and over 90% for frontier models in 2024. It was the first yardstick for code ability. SWE-bench took over after it.
- Related
- LLM: how the model works (measures) · GPT-4 Technical Report (discusses)
- Linked from
- GPT-4 Technical Report (measures)
- Sources
- Codex paper page (2021) (checked 2026-10-08)
SWE-bench
2023LayersL5L2
What it measures: real GitHub issues, where the model has to change code in a repo and pass the tests. It was the first agent-level benchmark. When it saturated: the Verified subset was about 50% in 2024 and 70–80% in 2025. The AI Index 2026 report says it is now close to 100%, and training data contamination means the scores can no longer be trusted.
- Related
- Agent: a model with tools in a loop (measures) · LLM: how the model works (measures) · The third wave: scaling and LLM (discusses)
- Sources
- Official leaderboard; AI Index 2026 says Verified went from about 60% to nearly 100% within a year (checked 2026-10-08)
Translated from the author's Chinese notes by a model; the Chinese page is the original.