Papers and venues
The papers the atlas keeps coming back to, in date order: what each solved, what it overturned or extended, and where it appeared.
Papers
In date order.
A Logical Calculus of the Ideas Immanent in Nervous Activity
1943LayersL1L5
It abstracts the neuron into a logic unit that only "fires once it reaches a threshold," and proves that networks built from such units can express any propositional logic. It was the first time "what the brain does" and "what a computer can compute" were put into one mathematical framework. Both neural networks and symbolic logic branch off from here.
- Related
- Warren McCulloch (proposed) · The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (extends) · The first wave: symbols and expert systems (discusses)
- Linked from
- Warren McCulloch (proposed) · The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (overturned)
- Sources
- Bulletin of Mathematical Biophysics 5, 1943 (checked 2026-10-08)
Computing Machinery and Intelligence
1950LayersL5
Swap the unanswerable question "can machines think?" for a test you can actually run: if you can't tell the machine from a human in a text conversation, stop arguing about the definition. The end of the paper also predicts machines that learn, at a time when not a single learning program existed yet.
- Related
- Alan Turing (proposed) · The first wave: symbols and expert systems (discusses) · Nature (published in)
- Linked from
- Alan Turing (proposed)
- Sources
- Mind LIX(236), 1950 (checked 2026-10-08)
The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain
1958LayersL1L5
In McCulloch–Pitts units, a person sets the weights. The perceptron learns the weights from examples, comes with a convergence proof, and was also built as hardware. It was the first real learning classifier, and it was the one that Perceptrons knocked down ten years later.
- Related
- Frank Rosenblatt (proposed) · Classical machine learning: learning from examples (solves) · A Logical Calculus of the Ideas Immanent in Nervous Activity (overturned)
- Linked from
- Frank Rosenblatt (proposed) · A Logical Calculus of the Ideas Immanent in Nervous Activity (extends) · Perceptrons: An Introduction to Computational Geometry (overturned)
- Sources
- Psychological Review 65(6), 1958, course mirror at UPenn (checked 2026-10-08)
Perceptrons: An Introduction to Computational Geometry
1969LayersL1L5
The book proved mathematically that a single-layer perceptron cannot represent even a simple function like XOR, and nobody knew how to train multiple layers. It pushed funding and talent toward the symbolic camp, and neural network research went quiet for more than a decade. Its argument was only worked around once backpropagation became widespread.
- Related
- Marvin Minsky (proposed) · The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (overturned) · The first wave: symbols and expert systems (discusses)
- Linked from
- Marvin Minsky (proposed) · Learning representations by back-propagating errors (overturned) · MIT AI Lab / CSAIL (funded)
- Sources
- MIT Press book page (checked 2026-10-08)
Learning representations by back-propagating errors
1986LayersL1L5
It applies the chain rule to multi-layer networks, so the weights of the hidden layers can also be adjusted by the error. It also shows that the hidden layers learn useful internal representations on their own. It answers the question Perceptrons left open: how do you train more than one layer? It is the technical starting point of the second wave.
- Related
- Geoffrey Hinton (proposed) · Math: learn it when you need it (solves) · Perceptrons: An Introduction to Computational Geometry (overturned) · Nature (published in)
- Linked from
- Geoffrey Hinton (proposed) · Nature (released)
- Sources
- Nature 323, 1986 (checked 2026-10-08)
Support-Vector Networks
1995LayersL1L5
It turns classification into finding the maximum-margin hyperplane, then uses the kernel trick to bring linear methods into a high-dimensional feature space, with generalization theory to back it up. It was the most reliable classifier of the 2000s, and deep learning only clearly beat it on large-scale vision tasks in 2012.
- Related
- Vladimir Vapnik (proposed) · Classical machine learning: learning from examples (solves) · Bell Labs (works at)
- Linked from
- Vladimir Vapnik (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (overturned) · Bell Labs (released)
- Sources
- Machine Learning 20, 1995 (checked 2026-10-08)
Long Short-Term Memory
1997LayersL1L5
When a recurrent network backpropagates through time, the gradient vanishes or explodes, so information from far back is not learned. LSTM uses gated units to open a channel that lets information flow stably over many steps. Nearly all speech recognition and neural translation in the 2010s relied on it. Before the Transformer appeared, it was the default answer for sequence modeling.
- Related
- Sepp Hochreiter (proposed) · Jürgen Schmidhuber (proposed) · Transformer (solves)
- Linked from
- Jürgen Schmidhuber (proposed) · Sepp Hochreiter (proposed) · IDSIA (released)
- Sources
- Neural Computation 9(8), 1997 (checked 2026-10-08)
Gradient-Based Learning Applied to Document Recognition
1998LayersL1L5
It combined convolution, pooling and gradient training into one end-to-end handwriting recognition system, and it was actually deployed to read checks. It showed that putting priors into the network structure (local receptive fields, weight sharing) instead of into rules was workable. AlexNet is just a scaled-up version of it.
- Related
- Yann LeCun (proposed) · Classical machine learning: learning from examples (solves) · MNIST (measures) · Bell Labs (works at)
- Linked from
- Yann LeCun (proposed) · Bell Labs (released) · MNIST (discusses)
- Sources
- Proceedings of the IEEE 86(11), 1998 (checked 2026-10-08)
A Neural Probabilistic Language Model
2003LayersL2L5
The problem with n-gram language models is that words have no similarity to each other, so any combination never seen in training gets probability zero. This paper has each word learn a continuous vector, then uses a neural network to predict the next word, so similar words share statistics automatically. Both word vectors and today's LLM pretraining objective come from here.
- Related
- Yoshua Bengio (proposed) · Embedding: turning meaning into vectors (solves) · Efficient Estimation of Word Representations in Vector Space (extends) · JMLR (published in)
- Linked from
- Yoshua Bengio (proposed) · JMLR (released)
- Sources
- JMLR 3, 2003 (checked 2026-10-08)
A fast learning algorithm for deep belief nets
2006LayersL1L5
Multi-layer networks were thought to be untrainable in the 2000s. This paper used layer-by-layer unsupervised pretraining to give a deep network a good starting point, then fine-tuned the whole thing. It showed that "deep" could be trained. It made "deep learning" a term and brought funding back to neural networks. Later people found that with ReLU, GPU and big data, the pretraining step could be dropped.
- Related
- Geoffrey Hinton (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (extends) · University of Toronto (works at)
- Linked from
- Geoffrey Hinton (proposed) · University of Toronto (released)
- Sources
- Neural Computation 18(7), 2006 (checked 2026-10-08)
ImageNet Classification with Deep Convolutional Neural Networks
2012LayersL1L5
A deep convolutional network trained on two consumer-grade GPUs cut the error rate on the ImageNet challenge by ten percentage points in one step. The runner-up still took the old route of hand-built features plus SVM classifiers. This is the moment deep learning took off: the methods had existed for a long time, but the dataset and the GPU were the two things that only came together in 2012.
- Related
- Alex Krizhevsky (proposed) · Ilya Sutskever (proposed) · Geoffrey Hinton (proposed) · ImageNet (measures) · NeurIPS (published in) · Support-Vector Networks (overturned)
- Linked from
- Alex Krizhevsky (proposed) · Geoffrey Hinton (proposed) · Ilya Sutskever (proposed) · A fast learning algorithm for deep belief nets (extends) · NeurIPS (released) · University of Toronto (released) · ImageNet (discusses)
- Sources
- NeurIPS 2012 paper page (checked 2026-10-08)
Playing Atari with Deep Reinforcement Learning
2013LayersL2L5
A convolutional network reads screen pixels directly and outputs action values. The same network and hyperparameters play dozens of Atari games. It joined deep learning to reinforcement learning, and it was the most convincing demo DeepMind had before Google acquired it.
- Related
- David Silver (proposed) · Google DeepMind (released) · Post-training: shaping with rewards (solves) · Mastering the game of Go with deep neural networks and tree search (extends)
- Linked from
- David Silver (proposed) · Mastering the game of Go with deep neural networks and tree search (extends) · Google DeepMind (released)
- Sources
- arXiv 1312.5602,2013-12-19 (checked 2026-10-08)
Efficient Estimation of Word Representations in Vector Space
2013LayersL2L5
Cut the language model down to one shallow prediction task. In return, you can train word vectors on a billion words, and vector subtraction gives analogies like "king − man + woman ≈ queen". It made "pretrain representations on unlabeled text first" standard practice in NLP. It was the most important round of pretraining before BERT.
- Related
- Tomáš Mikolov (proposed) · Embedding: turning meaning into vectors (solves) · Google (released) · ICLR (published in)
- Linked from
- Tomáš Mikolov (proposed) · A Neural Probabilistic Language Model (extends) · ICLR (released) · Google (released)
- Sources
- arXiv 1301.3781, first version 2013-01-16 (checked 2026-10-08)
Generative Adversarial Nets
2014LayersL1L5
Train a generator and a discriminator against each other, and the model learns to produce realistic samples without ever writing down an explicit likelihood function. It started a whole line of deep generative models (image synthesis, style transfer, deepfakes), until diffusion models replaced it after 2021.
- Related
- Ian Goodfellow (proposed) · Yoshua Bengio (proposed) · NeurIPS (published in) · High-Resolution Image Synthesis with Latent Diffusion Models (extends)
- Linked from
- Ian Goodfellow (proposed) · Yoshua Bengio (proposed) · High-Resolution Image Synthesis with Latent Diffusion Models (overturned) · NeurIPS (released) · Mila – Quebec AI Institute (released)
- Sources
- arXiv 1406.2661,2014-06-10 (checked 2026-10-08)
Sequence to Sequence Learning with Neural Networks
2014LayersL2L5
One LSTM compresses the whole sentence into a vector. Another LSTM decodes the translation from that vector. The model trains end to end, with no word alignment and no grammar rules. It turned neural machine translation from a paper into a system you could ship. It also exposed a bottleneck: a fixed-length vector cannot hold a long sentence. The attention mechanism was invented to fix that bottleneck.
- Related
- Ilya Sutskever (proposed) · Google (released) · NeurIPS (published in) · Neural Machine Translation by Jointly Learning to Align and Translate (extends)
- Linked from
- Ilya Sutskever (proposed) · Neural Machine Translation by Jointly Learning to Align and Translate (solves) · NeurIPS (released) · Google (released)
- Sources
- arXiv 1409.3215,2014-09-10 (checked 2026-10-08)
Neural Machine Translation by Jointly Learning to Align and Translate
2014LayersL2L5
When decoding each word, the model no longer looks at a single compressed vector. It looks back and gives every position in the source sentence a weight, then takes the weighted sum. This is attention. It fixed the drop in quality that seq2seq models had on long sentences. Three years later, the Transformer pushed the idea of "keep only attention" to the extreme.
- Related
- Yoshua Bengio (proposed) · Sequence to Sequence Learning with Neural Networks (solves) · Attention Is All You Need (extends) · ICLR (published in) · Mila – Quebec AI Institute (works at)
- Linked from
- Yoshua Bengio (proposed) · Attention Is All You Need (overturned) · Sequence to Sequence Learning with Neural Networks (extends) · ICLR (released) · Mila – Quebec AI Institute (released)
- Sources
- arXiv 1409.0473,2014-09-01 (checked 2026-10-08)
Deep learning
2015LayersL1L5
A survey written three years after AlexNet by three future Turing Award winners. It covers why learning representations layer by layer beat hand-built features, what convolutional and recurrent networks each solve, and why unsupervised learning would be the next step. It is the people who did the work summing up the second wave themselves.
- Related
- Yann LeCun (proposed) · Yoshua Bengio (proposed) · Geoffrey Hinton (proposed) · Nature (published in) · The Second Wave: Statistical Learning and Deep Learning (discusses)
- Linked from
- Yann LeCun (proposed) · Nature (released)
- Sources
- Nature 521, 2015-05-27; OpenAlex 2026-10-08 cited 85,418 times (checked 2026-10-08)
Deep Residual Learning for Image Recognition
2015LayersL1L5
When a network gets deep, it stops training well. That isn't overfitting. The optimizer just can't make progress. Residual connections let each layer learn a correction relative to the previous layer, so a network with over a hundred layers still converges. It was the first to beat human annotators on ImageNet. Residual connections have been standard in every deep network since, and every Transformer block has one.
- Related
- Kaiming He (proposed) · Microsoft (released) · ImageNet (measures) · CVPR (published in) · Transformer (solves)
- Linked from
- Kaiming He (proposed) · CVPR (released) · Microsoft (released) · ImageNet (discusses)
- Sources
- arXiv 1512.03385,2015-12-10 (checked 2026-10-08)
Mastering the game of Go with deep neural networks and tree search
2016LayersL2L5
A policy network and a value network prune the Monte Carlo tree search. Self-play reinforcement learning then trains the networks stronger and stronger. In March 2016 it beat Lee Sedol 4:1. Go was thought to be another decade away. It changed the public image of deep learning from "recognizing images" to "making decisions".
- Related
- David Silver (proposed) · Demis Hassabis (proposed) · Google DeepMind (released) · Nature (published in) · Post-training: shaping with rewards (solves) · Playing Atari with Deep Reinforcement Learning (extends)
- Linked from
- David Silver (proposed) · Demis Hassabis (proposed) · Playing Atari with Deep Reinforcement Learning (extends) · The Bitter Lesson (extends) · Nature (released) · Google DeepMind (released)
- Sources
- Nature 529, 2016-01-27 (checked 2026-10-08)
Attention Is All You Need
2017LayersL2L5
It removed all the recurrence and convolution from the sequence-to-sequence models of the time and kept only attention. In exchange, training could be fully parallel. On machine translation it set new results with less training time. It is where the three waves begin: GPT, BERT and ViT are all variants of this skeleton. See the node at Transformer.
- Related
- Transformer (solves) · Ashish Vaswani (proposed) · NeurIPS (published in) · Neural Machine Translation by Jointly Learning to Align and Translate (overturned)
- Linked from
- Ashish Vaswani (proposed) · Noam Shazeer (proposed) · Highly accurate protein structure prediction with AlphaFold (extends) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (extends) · Neural Machine Translation by Jointly Learning to Align and Translate (extends) · NeurIPS (released) · Google (released)
- Sources
- arXiv 1706.03762, first version 2017-06-12 (checked 2026-10-08)
Deep Reinforcement Learning from Human Preferences
2017LayersL2L5
Many tasks have no reward function you can write down, but a person can compare two pieces of behavior and say which is better. This paper trains a reward model on those pairwise comparisons, then uses it for reinforcement learning. People only need to look at less than one percent of the interactions. It is the prototype of RLHF. Five years later it was carried over almost unchanged to language models and became InstructGPT.
- Related
- Paul Christiano (proposed) · OpenAI (released) · Google DeepMind (released) · NeurIPS (published in) · Training language models to follow instructions with human feedback (extends) · Post-training: shaping with rewards (solves)
- Linked from
- Paul Christiano (proposed) · Training language models to follow instructions with human feedback (extends)
- Sources
- arXiv 1706.03741,2017-06-12 (checked 2026-10-08)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
2018LayersL2L5
It trains a bidirectional Transformer encoder with the objective of masking some words and guessing them. After fine-tuning, it swept GLUE and SQuAD. It made the pretrained model the default starting point for every NLP task. It also pushed benchmarks like GLUE past human level within a year.
- Related
- Google (released) · GLUE / SuperGLUE (measures) · SQuAD (measures) · ACL (published in) · LLM: how the model works (solves) · Attention Is All You Need (extends)
- Linked from
- ACL (released) · Google (released) · GLUE / SuperGLUE (discusses) · SQuAD (discusses)
- Sources
- arXiv 1810.04805,2018-10-11 (checked 2026-10-08)
Improving Language Understanding by Generative Pre-Training
2018LayersL2L5
First train a decoder-only Transformer as a language model on a large amount of unlabeled text, then fine-tune it for each downstream task. One model family got the best results at the time on several NLP benchmarks. It set the "generative pre-training + fine-tuning" route, and the GPT series name starts here.
- Related
- Alec Radford (proposed) · Ilya Sutskever (proposed) · OpenAI (released) · Language Models are Unsupervised Multitask Learners (extends) · LLM: how the model works (solves)
- Linked from
- Alec Radford (proposed) · Ilya Sutskever (proposed) · OpenAI (released)
- Sources
- OpenAI release page, 2018-06-11 (checked 2026-10-08)
Language Models are Unsupervised Multitask Learners
2019LayersL2L5
It scales GPT-1 up by ten times and trains only on web text. With no fine-tuning at all, it can do summarization, translation, and question answering. It was the first sign that once a language model is big enough, the tasks are learned along the way. The staged release of the weights also made the safety of model releases a public issue for the first time.
- Related
- Alec Radford (proposed) · OpenAI (released) · Language Models are Few-Shot Learners (extends) · LLM: how the model works (solves)
- Linked from
- Alec Radford (proposed) · Improving Language Understanding by Generative Pre-Training (extends) · OpenAI (released)
- Sources
- OpenAI release page, 2019-02-14 (checked 2026-10-08)
The Bitter Lesson
2019LayersL5
The lesson of seventy years: every time people write human knowledge into a system, it works in the short term. In the long term, it loses to general methods built on raw compute. Chess, speech and vision all followed this script. It is the shortest note on the three waves, and the creed of this scaling wave.
- Related
- Richard Sutton (proposed) · AI History: How to Read the Three Waves (discusses) · Scaling Laws for Neural Language Models (extends) · Mastering the game of Go with deep neural networks and tree search (extends)
- Linked from
- Richard Sutton (proposed)
- Sources
- Original PDF from the UT Austin course archive (the author's site, incompleteideas.net, only has http). Original dated 2019-03-13. (checked 2026-10-08)
Language Models are Few-Shot Learners
2020LayersL2L5
175 billion parameters, and not a single weight updated. Just put a few examples in the prompt and it does a new task. It made the prompt a way of programming, and it was the first time scaling laws paid off in public view. The two years of LLM startups that followed all grew up around its API.
- Related
- OpenAI (released) · NeurIPS (published in) · Prompting: instructions you can measure (solves) · Training language models to follow instructions with human feedback (extends)
- Linked from
- Language Models are Unsupervised Multitask Learners (extends) · GPT-4 Technical Report (extends) · NeurIPS (released) · OpenAI (released)
- Sources
- arXiv 2005.14165,2020-05-28 (checked 2026-10-08)
Scaling Laws for Neural Language Models
2020LayersL2L5
Language model loss falls as a power law in parameter count, data size and compute, each on its own. This holds across seven orders of magnitude, and model shape barely matters. It turns "should we spend ten times more compute" from a matter of faith into something you can calculate. The investment logic of the three waves' third wave rests on this curve.
- Related
- Jared Kaplan (proposed) · Dario Amodei (proposed) · OpenAI (released) · Training Compute-Optimal Large Language Models (extends) · LLM: how the model works (solves) · The third wave: scaling and LLM (discusses)
- Linked from
- Dario Amodei (proposed) · Jared Kaplan (proposed) · Training Compute-Optimal Large Language Models (overturned) · The Bitter Lesson (extends) · arXiv (released) · OpenAI (released)
- Sources
- arXiv 2001.08361,2020-01-23 (checked 2026-10-08)
Highly accurate protein structure prediction with AlphaFold
2021LayersL5
It predicts a protein's 3D structure directly from its amino acid sequence, at an accuracy on par with experimental methods. This largely solved a biology problem that had stood for fifty years. It marked deep learning moving beyond language and images into scientific discovery, and it won the 2024 Nobel Prize in Chemistry.
- Related
- Demis Hassabis (proposed) · Google DeepMind (released) · Nature (published in) · Attention Is All You Need (extends)
- Linked from
- Demis Hassabis (proposed) · Nature (released) · Google DeepMind (released)
- Sources
- Nature 596, 2021-07-15 (checked 2026-10-08)
Learning Transferable Visual Models From Natural Language Supervision
2021LayersL2L5
Contrastive learning on 400 million image-text pairs puts images and text into the same vector space, so it can do zero-shot classification without labels. It is the eye that text-to-image models use to understand text, and it is the common starting point for multimodal models.
- Related
- Alec Radford (proposed) · OpenAI (released) · ICML (published in) · High-Resolution Image Synthesis with Latent Diffusion Models (extends) · Embedding: turning meaning into vectors (solves)
- Linked from
- Alec Radford (proposed) · High-Resolution Image Synthesis with Latent Diffusion Models (extends) · ICML (released) · OpenAI (released)
- Sources
- arXiv 2103.00020,2021-02-26 (checked 2026-10-08)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
2022LayersL4L5
Put a few examples with reasoning steps in the prompt, and a large model's scores on arithmetic and commonsense reasoning jump by a lot. Small models get nothing from it. It turned "make the model think before it answers" into a technique, and it set off the argument over whether "emergent abilities" are real. Two years later, reasoning models moved this from the prompt into training.
- Related
- Google (released) · NeurIPS (published in) · Prompting: instructions you can measure (solves) · GSM8K (measures) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (extends)
- Linked from
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (extends) · Google (released) · GSM8K (discusses)
- Sources
- arXiv 2201.11903,2022-01-28 (checked 2026-10-08)
Training Compute-Optimal Large Language Models
2022LayersL2L5
With the same compute, parameters and training tokens should grow in equal proportion. Earlier large models were mostly fed too little data. The 70B-parameter Chinchilla beat the 280B Gopher by using four times the data. It rewrote the scaling recipe. After LLaMA, "smaller models fed more data" became the mainstream approach.
- Related
- Google DeepMind (released) · Scaling Laws for Neural Language Models (overturned) · NeurIPS (published in) · LLaMA: Open and Efficient Foundation Language Models (extends) · LLM: how the model works (solves)
- Linked from
- LLaMA: Open and Efficient Foundation Language Models (extends) · Scaling Laws for Neural Language Models (extends) · Google DeepMind (released)
- Sources
- arXiv 2203.15556,2022-03-29 (checked 2026-10-08)
Constitutional AI: Harmlessness from AI Feedback
2022LayersL2L5
It uses a set of written principles to have the model critique and revise its own answers. Then it uses preferences given by AI to replace most human harmlessness labels. This cuts the reliance on human labeling. It also makes "which values the model is aligned to" a document you can read, for the first time.
- Related
- Anthropic (released) · Training language models to follow instructions with human feedback (extends) · Post-training: shaping with rewards (solves)
- Linked from
- Anthropic (released)
- Sources
- arXiv 2212.08073,2022-12-15 (checked 2026-10-08)
Training language models to follow instructions with human feedback
2022LayersL2L5
A pretrained model only continues text. It does not follow instructions. First do supervised fine-tuning on demonstrations. Then train a reward model on human preferences and optimize with PPO. The 1.3B-parameter result was preferred by labelers over the original 175B model. This three-step recipe is the base of ChatGPT, and it is the reference point for all the "alignment" work that followed.
- Related
- Paul Christiano (proposed) · OpenAI (released) · Post-training: shaping with rewards (solves) · Deep Reinforcement Learning from Human Preferences (extends) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (extends) · NeurIPS (published in)
- Linked from
- Paul Christiano (proposed) · Constitutional AI: Harmlessness from AI Feedback (extends) · Deep Reinforcement Learning from Human Preferences (extends) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (overturned) · Language Models are Few-Shot Learners (extends) · OpenAI (released)
- Sources
- arXiv 2203.02155,2022-03-04 (checked 2026-10-08)
High-Resolution Image Synthesis with Latent Diffusion Models
2022LayersL2L5
It moves the diffusion process from pixel space into a compressed latent space. That cuts the cost of training and generation by an order of magnitude, so it runs on a consumer GPU. Stable Diffusion, with its open weights, is this paper in practice. Image generation no longer belonged only to companies with the compute.
- Related
- CVPR (published in) · Learning Transferable Visual Models From Natural Language Supervision (extends) · Generative Adversarial Nets (overturned) · The third wave: scaling and LLM (discusses)
- Linked from
- Learning Transferable Visual Models From Natural Language Supervision (extends) · Generative Adversarial Nets (extends) · CVPR (released)
- Sources
- arXiv 2112.10752, 2021-12-20; Stable Diffusion open weights released 2022-08 (checked 2026-10-08)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
2023LayersL2L5
RLHF's reward model and PPO steps can be collapsed into a single classification loss, trained directly on preference pairs with supervised learning. The results hold up and the implementation is only a few dozen lines. It moved preference alignment from big-lab engineering to everyday work in the open-source community.
- Related
- Stanford AI Lab (SAIL) (released) · Training language models to follow instructions with human feedback (overturned) · Post-training: shaping with rewards (solves) · NeurIPS (published in)
- Linked from
- Training language models to follow instructions with human feedback (extends) · Stanford AI Lab (SAIL) (released)
- Sources
- arXiv 2305.18290,2023-05-29 (checked 2026-10-08)
GPT-4 Technical Report
2023LayersL2L5
Multimodal input, top 10% on professional exams such as the bar exam, and the report says outright that it does not disclose parameter count, architecture or data. It is both a step up in capability and the point where frontier labs moved from "publishing papers" to "shipping products"; the "technical report" replaced the paper from then on.
- Related
- OpenAI (released) · MMLU (measures) · HumanEval (measures) · Language Models are Few-Shot Learners (extends) · The third wave: scaling and LLM (discusses)
- Linked from
- OpenAI (released) · HumanEval (discusses) · MMLU (discusses)
- Sources
- arXiv 2303.08774,2023-03-15 (checked 2026-10-08)
LLaMA: Open and Efficient Foundation Language Models
2023LayersL2L5
Trained only on public data, and following the Chinchilla approach of training a smaller model for longer, the 13B-parameter model beat GPT-3 on most benchmark tests. The weights leaked within a week. Half a year later Llama 2 was officially opened for commercial use. The whole open-source ecosystem of llama.cpp, quantization and local inference grew out of it.
- Related
- Meta AI (FAIR) (released) · Training Compute-Optimal Large Language Models (extends) · LLM: how the model works (solves) · The third wave: scaling and LLM (discusses)
- Linked from
- Training Compute-Optimal Large Language Models (extends) · arXiv (released) · Meta AI (FAIR) (released)
- Sources
- arXiv 2302.13971, 2023-02-27; for Llama 2 see arXiv 2307.09288 (checked 2026-10-08)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
2025LayersL2L5
Reinforcement learning with only verifiable rewards, and no demonstrations: the model grows long chains of thought and self-checking on its own. The weights are open and the training cost is public. It shows that an o1-style reasoning model is not one lab's secret, and it puts the new scaling axis of inference-time compute in the hands of the open-source community.
- Related
- DeepSeek (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (extends) · Post-training: shaping with rewards (solves) · The third wave: scaling and LLM (discusses) · Nature (published in)
- Linked from
- Liang Wenfeng (proposed) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (extends) · DeepSeek (released)
- Sources
- arXiv 2501.12948,2025-01-22 (checked 2026-10-08)
Venues
Where the papers above appeared.
ACL
1962LayersL5
The annual meeting of the Association for Computational Linguistics, the home conference of NLP. BERT was published at its North American chapter, NAACL. After LLM took off, a lot of its topics moved to NeurIPS and ICLR.
- Related
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · The Second Wave: Statistical Learning and Deep Learning (discusses)
- Linked from
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (published in)
- Sources
- Official site (checked 2026-10-08)
arXiv
1991LayersL5
Where the third wave is really published: papers go on arXiv first and then to a conference. Since GPT-2, many frontier-lab "technical reports" are never submitted to a conference at all. The "first-version date" on the timeline refers to this.
- Related
- Scaling Laws for Neural Language Models (released) · LLaMA: Open and Efficient Foundation Language Models (released) · The third wave: scaling and LLM (discusses)
- Sources
- Official website (checked 2026-10-08)
CVPR
1983LayersL5
The top conference for computer vision. ResNet and latent diffusion models were published here. From 2012 to 2016, vision was the main battlefield of deep learning, so reading its papers is reading the history of those years.
- Related
- Deep Residual Learning for Image Recognition (released) · High-Resolution Image Synthesis with Latent Diffusion Models (released)
- Linked from
- High-Resolution Image Synthesis with Latent Diffusion Models (published in) · Deep Residual Learning for Image Recognition (published in)
- Sources
- Official website (checked 2026-10-08)
ICLR
2013LayersL5
Bengio and LeCun founded it in 2013 specifically for "representation learning", with open review. word2vec and the attention mechanism were papers at its first and second editions. A conference of the deep learning era.
- Related
- Efficient Estimation of Word Representations in Vector Space (released) · Neural Machine Translation by Jointly Learning to Align and Translate (released)
- Linked from
- Neural Machine Translation by Jointly Learning to Align and Translate (published in) · Efficient Estimation of Word Representations in Vector Space (published in)
- Sources
- Official website (checked 2026-10-08)
ICML
1980LayersL5
The longest-running machine learning conference (1980). It leans toward methods and theory. CLIP was published here.
- Related
- Learning Transferable Visual Models From Natural Language Supervision (released) · The Second Wave: Statistical Learning and Deep Learning (discusses)
- Linked from
- Learning Transferable Visual Models From Natural Language Supervision (published in)
- Sources
- Official website (checked 2026-10-08)
JMLR
2000LayersL5
An open-access journal founded in 2000 after the editorial board of Machine Learning left together. Bengio's neural language model (2003) was published here.
- Related
- A Neural Probabilistic Language Model (released)
- Linked from
- A Neural Probabilistic Language Model (published in)
- Sources
- Official website (checked 2026-10-08)
Nature
1869LayersL5
The landmark outlet for AI entering mainstream science: backpropagation in 1986, the deep learning review in 2015, AlphaGo in 2016 and AlphaFold in 2021 were all published here.
- Related
- Learning representations by back-propagating errors (released) · Mastering the game of Go with deep neural networks and tree search (released) · Highly accurate protein structure prediction with AlphaFold (released) · Deep learning (released)
- Linked from
- Highly accurate protein structure prediction with AlphaFold (published in) · Mastering the game of Go with deep neural networks and tree search (published in) · Learning representations by back-propagating errors (published in) · Deep learning (published in) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (published in) · Computing Machinery and Intelligence (published in)
- Sources
- Official website (checked 2026-10-08)
NeurIPS
1987LayersL5
The top machine learning conference, held every December since 1987. AlexNet, GAN, seq2seq, Transformer and GPT-3 were all published here. Read a year of NeurIPS papers and you can guess the products two years out.
- Related
- ImageNet Classification with Deep Convolutional Neural Networks (released) · Generative Adversarial Nets (released) · Sequence to Sequence Learning with Neural Networks (released) · Attention Is All You Need (released) · Language Models are Few-Shot Learners (released)
- Linked from
- ImageNet Classification with Deep Convolutional Neural Networks (published in) · Attention Is All You Need (published in) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (published in) · Training Compute-Optimal Large Language Models (published in) · Deep Reinforcement Learning from Human Preferences (published in) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (published in) · Generative Adversarial Nets (published in) · Language Models are Few-Shot Learners (published in) · Training language models to follow instructions with human feedback (published in) · Sequence to Sequence Learning with Neural Networks (published in)
- Sources
- Official website (checked 2026-10-08)
Translated from the author's Chinese notes by a model; the Chinese page is the original.