Skip to main content

Papers and venues

The papers the atlas keeps coming back to, in date order: what each solved, what it overturned or extended, and where it appeared.

Papers

In date order.

A Logical Calculus of the Ideas Immanent in Nervous Activity

1943LayersL1L5

It abstracts the neuron into a logic unit that only "fires once it reaches a threshold," and proves that networks built from such units can express any propositional logic. It was the first time "what the brain does" and "what a computer can compute" were put into one mathematical framework. Both neural networks and symbolic logic branch off from here.

Related
Warren McCulloch (proposed) · The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (extends) · The first wave: symbols and expert systems (discusses)
Linked from
Warren McCulloch (proposed) · The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (overturned)
Sources
Bulletin of Mathematical Biophysics 5, 1943 (checked 2026-10-08)

Computing Machinery and Intelligence

1950LayersL5

Swap the unanswerable question "can machines think?" for a test you can actually run: if you can't tell the machine from a human in a text conversation, stop arguing about the definition. The end of the paper also predicts machines that learn, at a time when not a single learning program existed yet.

Related
Alan Turing (proposed) · The first wave: symbols and expert systems (discusses) · Nature (published in)
Linked from
Alan Turing (proposed)
Sources
Mind LIX(236), 1950 (checked 2026-10-08)

The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain

1958LayersL1L5

In McCulloch–Pitts units, a person sets the weights. The perceptron learns the weights from examples, comes with a convergence proof, and was also built as hardware. It was the first real learning classifier, and it was the one that Perceptrons knocked down ten years later.

Related
Frank Rosenblatt (proposed) · Classical machine learning: learning from examples (solves) · A Logical Calculus of the Ideas Immanent in Nervous Activity (overturned)
Linked from
Frank Rosenblatt (proposed) · A Logical Calculus of the Ideas Immanent in Nervous Activity (extends) · Perceptrons: An Introduction to Computational Geometry (overturned)
Sources
Psychological Review 65(6), 1958, course mirror at UPenn (checked 2026-10-08)

Perceptrons: An Introduction to Computational Geometry

1969LayersL1L5

The book proved mathematically that a single-layer perceptron cannot represent even a simple function like XOR, and nobody knew how to train multiple layers. It pushed funding and talent toward the symbolic camp, and neural network research went quiet for more than a decade. Its argument was only worked around once backpropagation became widespread.

Related
Marvin Minsky (proposed) · The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (overturned) · The first wave: symbols and expert systems (discusses)
Linked from
Marvin Minsky (proposed) · Learning representations by back-propagating errors (overturned) · MIT AI Lab / CSAIL (funded)
Sources
MIT Press book page (checked 2026-10-08)

Learning representations by back-propagating errors

1986LayersL1L5

It applies the chain rule to multi-layer networks, so the weights of the hidden layers can also be adjusted by the error. It also shows that the hidden layers learn useful internal representations on their own. It answers the question Perceptrons left open: how do you train more than one layer? It is the technical starting point of the second wave.

Related
Geoffrey Hinton (proposed) · Math: learn it when you need it (solves) · Perceptrons: An Introduction to Computational Geometry (overturned) · Nature (published in)
Linked from
Geoffrey Hinton (proposed) · Nature (released)
Sources
Nature 323, 1986 (checked 2026-10-08)

Support-Vector Networks

1995LayersL1L5

It turns classification into finding the maximum-margin hyperplane, then uses the kernel trick to bring linear methods into a high-dimensional feature space, with generalization theory to back it up. It was the most reliable classifier of the 2000s, and deep learning only clearly beat it on large-scale vision tasks in 2012.

Related
Vladimir Vapnik (proposed) · Classical machine learning: learning from examples (solves) · Bell Labs (works at)
Linked from
Vladimir Vapnik (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (overturned) · Bell Labs (released)
Sources
Machine Learning 20, 1995 (checked 2026-10-08)

Long Short-Term Memory

1997LayersL1L5

When a recurrent network backpropagates through time, the gradient vanishes or explodes, so information from far back is not learned. LSTM uses gated units to open a channel that lets information flow stably over many steps. Nearly all speech recognition and neural translation in the 2010s relied on it. Before the Transformer appeared, it was the default answer for sequence modeling.

Related
Sepp Hochreiter (proposed) · Jürgen Schmidhuber (proposed) · Transformer (solves)
Linked from
Jürgen Schmidhuber (proposed) · Sepp Hochreiter (proposed) · IDSIA (released)
Sources
Neural Computation 9(8), 1997 (checked 2026-10-08)

Gradient-Based Learning Applied to Document Recognition

1998LayersL1L5

It combined convolution, pooling and gradient training into one end-to-end handwriting recognition system, and it was actually deployed to read checks. It showed that putting priors into the network structure (local receptive fields, weight sharing) instead of into rules was workable. AlexNet is just a scaled-up version of it.

Related
Yann LeCun (proposed) · Classical machine learning: learning from examples (solves) · MNIST (measures) · Bell Labs (works at)
Linked from
Yann LeCun (proposed) · Bell Labs (released) · MNIST (discusses)
Sources
Proceedings of the IEEE 86(11), 1998 (checked 2026-10-08)

A Neural Probabilistic Language Model

2003LayersL2L5

The problem with n-gram language models is that words have no similarity to each other, so any combination never seen in training gets probability zero. This paper has each word learn a continuous vector, then uses a neural network to predict the next word, so similar words share statistics automatically. Both word vectors and today's LLM pretraining objective come from here.

Related
Yoshua Bengio (proposed) · Embedding: turning meaning into vectors (solves) · Efficient Estimation of Word Representations in Vector Space (extends) · JMLR (published in)
Linked from
Yoshua Bengio (proposed) · JMLR (released)
Sources
JMLR 3, 2003 (checked 2026-10-08)

A fast learning algorithm for deep belief nets

2006LayersL1L5

Multi-layer networks were thought to be untrainable in the 2000s. This paper used layer-by-layer unsupervised pretraining to give a deep network a good starting point, then fine-tuned the whole thing. It showed that "deep" could be trained. It made "deep learning" a term and brought funding back to neural networks. Later people found that with ReLU, GPU and big data, the pretraining step could be dropped.

Related
Geoffrey Hinton (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (extends) · University of Toronto (works at)
Linked from
Geoffrey Hinton (proposed) · University of Toronto (released)
Sources
Neural Computation 18(7), 2006 (checked 2026-10-08)

ImageNet Classification with Deep Convolutional Neural Networks

2012LayersL1L5

A deep convolutional network trained on two consumer-grade GPUs cut the error rate on the ImageNet challenge by ten percentage points in one step. The runner-up still took the old route of hand-built features plus SVM classifiers. This is the moment deep learning took off: the methods had existed for a long time, but the dataset and the GPU were the two things that only came together in 2012.

Related
Alex Krizhevsky (proposed) · Ilya Sutskever (proposed) · Geoffrey Hinton (proposed) · ImageNet (measures) · NeurIPS (published in) · Support-Vector Networks (overturned)
Linked from
Alex Krizhevsky (proposed) · Geoffrey Hinton (proposed) · Ilya Sutskever (proposed) · A fast learning algorithm for deep belief nets (extends) · NeurIPS (released) · University of Toronto (released) · ImageNet (discusses)
Sources
NeurIPS 2012 paper page (checked 2026-10-08)

Playing Atari with Deep Reinforcement Learning

2013LayersL2L5

A convolutional network reads screen pixels directly and outputs action values. The same network and hyperparameters play dozens of Atari games. It joined deep learning to reinforcement learning, and it was the most convincing demo DeepMind had before Google acquired it.

Related
David Silver (proposed) · Google DeepMind (released) · Post-training: shaping with rewards (solves) · Mastering the game of Go with deep neural networks and tree search (extends)
Linked from
David Silver (proposed) · Mastering the game of Go with deep neural networks and tree search (extends) · Google DeepMind (released)
Sources
arXiv 1312.5602,2013-12-19 (checked 2026-10-08)

Efficient Estimation of Word Representations in Vector Space

2013LayersL2L5

Cut the language model down to one shallow prediction task. In return, you can train word vectors on a billion words, and vector subtraction gives analogies like "king − man + woman ≈ queen". It made "pretrain representations on unlabeled text first" standard practice in NLP. It was the most important round of pretraining before BERT.

Related
Tomáš Mikolov (proposed) · Embedding: turning meaning into vectors (solves) · Google (released) · ICLR (published in)
Linked from
Tomáš Mikolov (proposed) · A Neural Probabilistic Language Model (extends) · ICLR (released) · Google (released)
Sources
arXiv 1301.3781, first version 2013-01-16 (checked 2026-10-08)

Generative Adversarial Nets

2014LayersL1L5

Train a generator and a discriminator against each other, and the model learns to produce realistic samples without ever writing down an explicit likelihood function. It started a whole line of deep generative models (image synthesis, style transfer, deepfakes), until diffusion models replaced it after 2021.

Related
Ian Goodfellow (proposed) · Yoshua Bengio (proposed) · NeurIPS (published in) · High-Resolution Image Synthesis with Latent Diffusion Models (extends)
Linked from
Ian Goodfellow (proposed) · Yoshua Bengio (proposed) · High-Resolution Image Synthesis with Latent Diffusion Models (overturned) · NeurIPS (released) · Mila – Quebec AI Institute (released)
Sources
arXiv 1406.2661,2014-06-10 (checked 2026-10-08)

Sequence to Sequence Learning with Neural Networks

2014LayersL2L5

One LSTM compresses the whole sentence into a vector. Another LSTM decodes the translation from that vector. The model trains end to end, with no word alignment and no grammar rules. It turned neural machine translation from a paper into a system you could ship. It also exposed a bottleneck: a fixed-length vector cannot hold a long sentence. The attention mechanism was invented to fix that bottleneck.

Related
Ilya Sutskever (proposed) · Google (released) · NeurIPS (published in) · Neural Machine Translation by Jointly Learning to Align and Translate (extends)
Linked from
Ilya Sutskever (proposed) · Neural Machine Translation by Jointly Learning to Align and Translate (solves) · NeurIPS (released) · Google (released)
Sources
arXiv 1409.3215,2014-09-10 (checked 2026-10-08)

Neural Machine Translation by Jointly Learning to Align and Translate

2014LayersL2L5

When decoding each word, the model no longer looks at a single compressed vector. It looks back and gives every position in the source sentence a weight, then takes the weighted sum. This is attention. It fixed the drop in quality that seq2seq models had on long sentences. Three years later, the Transformer pushed the idea of "keep only attention" to the extreme.

Related
Yoshua Bengio (proposed) · Sequence to Sequence Learning with Neural Networks (solves) · Attention Is All You Need (extends) · ICLR (published in) · Mila – Quebec AI Institute (works at)
Linked from
Yoshua Bengio (proposed) · Attention Is All You Need (overturned) · Sequence to Sequence Learning with Neural Networks (extends) · ICLR (released) · Mila – Quebec AI Institute (released)
Sources
arXiv 1409.0473,2014-09-01 (checked 2026-10-08)

Deep learning

2015LayersL1L5

A survey written three years after AlexNet by three future Turing Award winners. It covers why learning representations layer by layer beat hand-built features, what convolutional and recurrent networks each solve, and why unsupervised learning would be the next step. It is the people who did the work summing up the second wave themselves.

Related
Yann LeCun (proposed) · Yoshua Bengio (proposed) · Geoffrey Hinton (proposed) · Nature (published in) · The Second Wave: Statistical Learning and Deep Learning (discusses)
Linked from
Yann LeCun (proposed) · Nature (released)
Sources
Nature 521, 2015-05-27; OpenAlex 2026-10-08 cited 85,418 times (checked 2026-10-08)

Deep Residual Learning for Image Recognition

2015LayersL1L5

When a network gets deep, it stops training well. That isn't overfitting. The optimizer just can't make progress. Residual connections let each layer learn a correction relative to the previous layer, so a network with over a hundred layers still converges. It was the first to beat human annotators on ImageNet. Residual connections have been standard in every deep network since, and every Transformer block has one.

Related
Kaiming He (proposed) · Microsoft (released) · ImageNet (measures) · CVPR (published in) · Transformer (solves)
Linked from
Kaiming He (proposed) · CVPR (released) · Microsoft (released) · ImageNet (discusses)
Sources
arXiv 1512.03385,2015-12-10 (checked 2026-10-08)

Mastering the game of Go with deep neural networks and tree search

2016LayersL2L5

A policy network and a value network prune the Monte Carlo tree search. Self-play reinforcement learning then trains the networks stronger and stronger. In March 2016 it beat Lee Sedol 4:1. Go was thought to be another decade away. It changed the public image of deep learning from "recognizing images" to "making decisions".

Related
David Silver (proposed) · Demis Hassabis (proposed) · Google DeepMind (released) · Nature (published in) · Post-training: shaping with rewards (solves) · Playing Atari with Deep Reinforcement Learning (extends)
Linked from
David Silver (proposed) · Demis Hassabis (proposed) · Playing Atari with Deep Reinforcement Learning (extends) · The Bitter Lesson (extends) · Nature (released) · Google DeepMind (released)
Sources
Nature 529, 2016-01-27 (checked 2026-10-08)

Attention Is All You Need

2017LayersL2L5

It removed all the recurrence and convolution from the sequence-to-sequence models of the time and kept only attention. In exchange, training could be fully parallel. On machine translation it set new results with less training time. It is where the three waves begin: GPT, BERT and ViT are all variants of this skeleton. See the node at Transformer.

Related
Transformer (solves) · Ashish Vaswani (proposed) · NeurIPS (published in) · Neural Machine Translation by Jointly Learning to Align and Translate (overturned)
Linked from
Ashish Vaswani (proposed) · Noam Shazeer (proposed) · Highly accurate protein structure prediction with AlphaFold (extends) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (extends) · Neural Machine Translation by Jointly Learning to Align and Translate (extends) · NeurIPS (released) · Google (released)
Sources
arXiv 1706.03762, first version 2017-06-12 (checked 2026-10-08)

Deep Reinforcement Learning from Human Preferences

2017LayersL2L5

Many tasks have no reward function you can write down, but a person can compare two pieces of behavior and say which is better. This paper trains a reward model on those pairwise comparisons, then uses it for reinforcement learning. People only need to look at less than one percent of the interactions. It is the prototype of RLHF. Five years later it was carried over almost unchanged to language models and became InstructGPT.

Related
Paul Christiano (proposed) · OpenAI (released) · Google DeepMind (released) · NeurIPS (published in) · Training language models to follow instructions with human feedback (extends) · Post-training: shaping with rewards (solves)
Linked from
Paul Christiano (proposed) · Training language models to follow instructions with human feedback (extends)
Sources
arXiv 1706.03741,2017-06-12 (checked 2026-10-08)

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

2018LayersL2L5

It trains a bidirectional Transformer encoder with the objective of masking some words and guessing them. After fine-tuning, it swept GLUE and SQuAD. It made the pretrained model the default starting point for every NLP task. It also pushed benchmarks like GLUE past human level within a year.

Related
Google (released) · GLUE / SuperGLUE (measures) · SQuAD (measures) · ACL (published in) · LLM: how the model works (solves) · Attention Is All You Need (extends)
Linked from
ACL (released) · Google (released) · GLUE / SuperGLUE (discusses) · SQuAD (discusses)
Sources
arXiv 1810.04805,2018-10-11 (checked 2026-10-08)

Improving Language Understanding by Generative Pre-Training

2018LayersL2L5

First train a decoder-only Transformer as a language model on a large amount of unlabeled text, then fine-tune it for each downstream task. One model family got the best results at the time on several NLP benchmarks. It set the "generative pre-training + fine-tuning" route, and the GPT series name starts here.

Related
Alec Radford (proposed) · Ilya Sutskever (proposed) · OpenAI (released) · Language Models are Unsupervised Multitask Learners (extends) · LLM: how the model works (solves)
Linked from
Alec Radford (proposed) · Ilya Sutskever (proposed) · OpenAI (released)
Sources
OpenAI release page, 2018-06-11 (checked 2026-10-08)

Language Models are Unsupervised Multitask Learners

2019LayersL2L5

It scales GPT-1 up by ten times and trains only on web text. With no fine-tuning at all, it can do summarization, translation, and question answering. It was the first sign that once a language model is big enough, the tasks are learned along the way. The staged release of the weights also made the safety of model releases a public issue for the first time.

Related
Alec Radford (proposed) · OpenAI (released) · Language Models are Few-Shot Learners (extends) · LLM: how the model works (solves)
Linked from
Alec Radford (proposed) · Improving Language Understanding by Generative Pre-Training (extends) · OpenAI (released)
Sources
OpenAI release page, 2019-02-14 (checked 2026-10-08)

The Bitter Lesson

2019LayersL5

The lesson of seventy years: every time people write human knowledge into a system, it works in the short term. In the long term, it loses to general methods built on raw compute. Chess, speech and vision all followed this script. It is the shortest note on the three waves, and the creed of this scaling wave.

Related
Richard Sutton (proposed) · AI History: How to Read the Three Waves (discusses) · Scaling Laws for Neural Language Models (extends) · Mastering the game of Go with deep neural networks and tree search (extends)
Linked from
Richard Sutton (proposed)
Sources
Original PDF from the UT Austin course archive (the author's site, incompleteideas.net, only has http). Original dated 2019-03-13. (checked 2026-10-08)

Language Models are Few-Shot Learners

2020LayersL2L5

175 billion parameters, and not a single weight updated. Just put a few examples in the prompt and it does a new task. It made the prompt a way of programming, and it was the first time scaling laws paid off in public view. The two years of LLM startups that followed all grew up around its API.

Related
OpenAI (released) · NeurIPS (published in) · Prompting: instructions you can measure (solves) · Training language models to follow instructions with human feedback (extends)
Linked from
Language Models are Unsupervised Multitask Learners (extends) · GPT-4 Technical Report (extends) · NeurIPS (released) · OpenAI (released)
Sources
arXiv 2005.14165,2020-05-28 (checked 2026-10-08)

Scaling Laws for Neural Language Models

2020LayersL2L5

Language model loss falls as a power law in parameter count, data size and compute, each on its own. This holds across seven orders of magnitude, and model shape barely matters. It turns "should we spend ten times more compute" from a matter of faith into something you can calculate. The investment logic of the three waves' third wave rests on this curve.

Related
Jared Kaplan (proposed) · Dario Amodei (proposed) · OpenAI (released) · Training Compute-Optimal Large Language Models (extends) · LLM: how the model works (solves) · The third wave: scaling and LLM (discusses)
Linked from
Dario Amodei (proposed) · Jared Kaplan (proposed) · Training Compute-Optimal Large Language Models (overturned) · The Bitter Lesson (extends) · arXiv (released) · OpenAI (released)
Sources
arXiv 2001.08361,2020-01-23 (checked 2026-10-08)

Highly accurate protein structure prediction with AlphaFold

2021LayersL5

It predicts a protein's 3D structure directly from its amino acid sequence, at an accuracy on par with experimental methods. This largely solved a biology problem that had stood for fifty years. It marked deep learning moving beyond language and images into scientific discovery, and it won the 2024 Nobel Prize in Chemistry.

Related
Demis Hassabis (proposed) · Google DeepMind (released) · Nature (published in) · Attention Is All You Need (extends)
Linked from
Demis Hassabis (proposed) · Nature (released) · Google DeepMind (released)
Sources
Nature 596, 2021-07-15 (checked 2026-10-08)

Learning Transferable Visual Models From Natural Language Supervision

2021LayersL2L5

Contrastive learning on 400 million image-text pairs puts images and text into the same vector space, so it can do zero-shot classification without labels. It is the eye that text-to-image models use to understand text, and it is the common starting point for multimodal models.

Related
Alec Radford (proposed) · OpenAI (released) · ICML (published in) · High-Resolution Image Synthesis with Latent Diffusion Models (extends) · Embedding: turning meaning into vectors (solves)
Linked from
Alec Radford (proposed) · High-Resolution Image Synthesis with Latent Diffusion Models (extends) · ICML (released) · OpenAI (released)
Sources
arXiv 2103.00020,2021-02-26 (checked 2026-10-08)

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

2022LayersL4L5

Put a few examples with reasoning steps in the prompt, and a large model's scores on arithmetic and commonsense reasoning jump by a lot. Small models get nothing from it. It turned "make the model think before it answers" into a technique, and it set off the argument over whether "emergent abilities" are real. Two years later, reasoning models moved this from the prompt into training.

Related
Google (released) · NeurIPS (published in) · Prompting: instructions you can measure (solves) · GSM8K (measures) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (extends)
Linked from
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (extends) · Google (released) · GSM8K (discusses)
Sources
arXiv 2201.11903,2022-01-28 (checked 2026-10-08)

Training Compute-Optimal Large Language Models

2022LayersL2L5

With the same compute, parameters and training tokens should grow in equal proportion. Earlier large models were mostly fed too little data. The 70B-parameter Chinchilla beat the 280B Gopher by using four times the data. It rewrote the scaling recipe. After LLaMA, "smaller models fed more data" became the mainstream approach.

Related
Google DeepMind (released) · Scaling Laws for Neural Language Models (overturned) · NeurIPS (published in) · LLaMA: Open and Efficient Foundation Language Models (extends) · LLM: how the model works (solves)
Linked from
LLaMA: Open and Efficient Foundation Language Models (extends) · Scaling Laws for Neural Language Models (extends) · Google DeepMind (released)
Sources
arXiv 2203.15556,2022-03-29 (checked 2026-10-08)

Constitutional AI: Harmlessness from AI Feedback

2022LayersL2L5

It uses a set of written principles to have the model critique and revise its own answers. Then it uses preferences given by AI to replace most human harmlessness labels. This cuts the reliance on human labeling. It also makes "which values the model is aligned to" a document you can read, for the first time.

Related
Anthropic (released) · Training language models to follow instructions with human feedback (extends) · Post-training: shaping with rewards (solves)
Linked from
Anthropic (released)
Sources
arXiv 2212.08073,2022-12-15 (checked 2026-10-08)

Training language models to follow instructions with human feedback

2022LayersL2L5

A pretrained model only continues text. It does not follow instructions. First do supervised fine-tuning on demonstrations. Then train a reward model on human preferences and optimize with PPO. The 1.3B-parameter result was preferred by labelers over the original 175B model. This three-step recipe is the base of ChatGPT, and it is the reference point for all the "alignment" work that followed.

Related
Paul Christiano (proposed) · OpenAI (released) · Post-training: shaping with rewards (solves) · Deep Reinforcement Learning from Human Preferences (extends) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (extends) · NeurIPS (published in)
Linked from
Paul Christiano (proposed) · Constitutional AI: Harmlessness from AI Feedback (extends) · Deep Reinforcement Learning from Human Preferences (extends) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (overturned) · Language Models are Few-Shot Learners (extends) · OpenAI (released)
Sources
arXiv 2203.02155,2022-03-04 (checked 2026-10-08)

High-Resolution Image Synthesis with Latent Diffusion Models

2022LayersL2L5

It moves the diffusion process from pixel space into a compressed latent space. That cuts the cost of training and generation by an order of magnitude, so it runs on a consumer GPU. Stable Diffusion, with its open weights, is this paper in practice. Image generation no longer belonged only to companies with the compute.

Related
CVPR (published in) · Learning Transferable Visual Models From Natural Language Supervision (extends) · Generative Adversarial Nets (overturned) · The third wave: scaling and LLM (discusses)
Linked from
Learning Transferable Visual Models From Natural Language Supervision (extends) · Generative Adversarial Nets (extends) · CVPR (released)
Sources
arXiv 2112.10752, 2021-12-20; Stable Diffusion open weights released 2022-08 (checked 2026-10-08)

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

2023LayersL2L5

RLHF's reward model and PPO steps can be collapsed into a single classification loss, trained directly on preference pairs with supervised learning. The results hold up and the implementation is only a few dozen lines. It moved preference alignment from big-lab engineering to everyday work in the open-source community.

Related
Stanford AI Lab (SAIL) (released) · Training language models to follow instructions with human feedback (overturned) · Post-training: shaping with rewards (solves) · NeurIPS (published in)
Linked from
Training language models to follow instructions with human feedback (extends) · Stanford AI Lab (SAIL) (released)
Sources
arXiv 2305.18290,2023-05-29 (checked 2026-10-08)

GPT-4 Technical Report

2023LayersL2L5

Multimodal input, top 10% on professional exams such as the bar exam, and the report says outright that it does not disclose parameter count, architecture or data. It is both a step up in capability and the point where frontier labs moved from "publishing papers" to "shipping products"; the "technical report" replaced the paper from then on.

Related
OpenAI (released) · MMLU (measures) · HumanEval (measures) · Language Models are Few-Shot Learners (extends) · The third wave: scaling and LLM (discusses)
Linked from
OpenAI (released) · HumanEval (discusses) · MMLU (discusses)
Sources
arXiv 2303.08774,2023-03-15 (checked 2026-10-08)

LLaMA: Open and Efficient Foundation Language Models

2023LayersL2L5

Trained only on public data, and following the Chinchilla approach of training a smaller model for longer, the 13B-parameter model beat GPT-3 on most benchmark tests. The weights leaked within a week. Half a year later Llama 2 was officially opened for commercial use. The whole open-source ecosystem of llama.cpp, quantization and local inference grew out of it.

Related
Meta AI (FAIR) (released) · Training Compute-Optimal Large Language Models (extends) · LLM: how the model works (solves) · The third wave: scaling and LLM (discusses)
Linked from
Training Compute-Optimal Large Language Models (extends) · arXiv (released) · Meta AI (FAIR) (released)
Sources
arXiv 2302.13971, 2023-02-27; for Llama 2 see arXiv 2307.09288 (checked 2026-10-08)

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

2025LayersL2L5

Reinforcement learning with only verifiable rewards, and no demonstrations: the model grows long chains of thought and self-checking on its own. The weights are open and the training cost is public. It shows that an o1-style reasoning model is not one lab's secret, and it puts the new scaling axis of inference-time compute in the hands of the open-source community.

Related
DeepSeek (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (extends) · Post-training: shaping with rewards (solves) · The third wave: scaling and LLM (discusses) · Nature (published in)
Linked from
Liang Wenfeng (proposed) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (extends) · DeepSeek (released)
Sources
arXiv 2501.12948,2025-01-22 (checked 2026-10-08)

Venues

Where the papers above appeared.

ACL

1962LayersL5

The annual meeting of the Association for Computational Linguistics, the home conference of NLP. BERT was published at its North American chapter, NAACL. After LLM took off, a lot of its topics moved to NeurIPS and ICLR.

Related
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · The Second Wave: Statistical Learning and Deep Learning (discusses)
Linked from
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (published in)
Sources
Official site (checked 2026-10-08)

arXiv

1991LayersL5

Where the third wave is really published: papers go on arXiv first and then to a conference. Since GPT-2, many frontier-lab "technical reports" are never submitted to a conference at all. The "first-version date" on the timeline refers to this.

Related
Scaling Laws for Neural Language Models (released) · LLaMA: Open and Efficient Foundation Language Models (released) · The third wave: scaling and LLM (discusses)
Sources
Official website (checked 2026-10-08)

CVPR

1983LayersL5

The top conference for computer vision. ResNet and latent diffusion models were published here. From 2012 to 2016, vision was the main battlefield of deep learning, so reading its papers is reading the history of those years.

Related
Deep Residual Learning for Image Recognition (released) · High-Resolution Image Synthesis with Latent Diffusion Models (released)
Linked from
High-Resolution Image Synthesis with Latent Diffusion Models (published in) · Deep Residual Learning for Image Recognition (published in)
Sources
Official website (checked 2026-10-08)

ICLR

2013LayersL5

Bengio and LeCun founded it in 2013 specifically for "representation learning", with open review. word2vec and the attention mechanism were papers at its first and second editions. A conference of the deep learning era.

Related
Efficient Estimation of Word Representations in Vector Space (released) · Neural Machine Translation by Jointly Learning to Align and Translate (released)
Linked from
Neural Machine Translation by Jointly Learning to Align and Translate (published in) · Efficient Estimation of Word Representations in Vector Space (published in)
Sources
Official website (checked 2026-10-08)

ICML

1980LayersL5

The longest-running machine learning conference (1980). It leans toward methods and theory. CLIP was published here.

Related
Learning Transferable Visual Models From Natural Language Supervision (released) · The Second Wave: Statistical Learning and Deep Learning (discusses)
Linked from
Learning Transferable Visual Models From Natural Language Supervision (published in)
Sources
Official website (checked 2026-10-08)

JMLR

2000LayersL5

An open-access journal founded in 2000 after the editorial board of Machine Learning left together. Bengio's neural language model (2003) was published here.

Related
A Neural Probabilistic Language Model (released)
Linked from
A Neural Probabilistic Language Model (published in)
Sources
Official website (checked 2026-10-08)

Nature

1869LayersL5

The landmark outlet for AI entering mainstream science: backpropagation in 1986, the deep learning review in 2015, AlphaGo in 2016 and AlphaFold in 2021 were all published here.

Related
Learning representations by back-propagating errors (released) · Mastering the game of Go with deep neural networks and tree search (released) · Highly accurate protein structure prediction with AlphaFold (released) · Deep learning (released)
Linked from
Highly accurate protein structure prediction with AlphaFold (published in) · Mastering the game of Go with deep neural networks and tree search (published in) · Learning representations by back-propagating errors (published in) · Deep learning (published in) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (published in) · Computing Machinery and Intelligence (published in)
Sources
Official website (checked 2026-10-08)

NeurIPS

1987LayersL5

The top machine learning conference, held every December since 1987. AlexNet, GAN, seq2seq, Transformer and GPT-3 were all published here. Read a year of NeurIPS papers and you can guess the products two years out.

Related
ImageNet Classification with Deep Convolutional Neural Networks (released) · Generative Adversarial Nets (released) · Sequence to Sequence Learning with Neural Networks (released) · Attention Is All You Need (released) · Language Models are Few-Shot Learners (released)
Linked from
ImageNet Classification with Deep Convolutional Neural Networks (published in) · Attention Is All You Need (published in) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (published in) · Training Compute-Optimal Large Language Models (published in) · Deep Reinforcement Learning from Human Preferences (published in) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (published in) · Generative Adversarial Nets (published in) · Language Models are Few-Shot Learners (published in) · Training language models to follow instructions with human feedback (published in) · Sequence to Sequence Learning with Neural Networks (published in)
Sources
Official website (checked 2026-10-08)

Translated from the author's Chinese notes by a model; the Chinese page is the original.