Skip to main content

AI Timeline

Three waves, fifty milestones: what each wave bet on, where it hit a wall, what it left behind. Every row names its source and the day it was checked.

First wave: symbols and expert systems (1943–1997)

The bet of this wave: intelligence = search and reasoning over symbols, and people can write knowledge as rules and hand it to the machine. It hit the same wall twice. Combinatorial explosion kept search stuck in toy worlds. Common sense is too large and too implicit to write by hand (the knowledge acquisition bottleneck). The 1973 Lighthill report and the 1974 DARPA cuts were the first winter. The 1987 collapse of the Lisp machine market and the exposed maintenance cost of expert systems were the second. What it left behind did not disappear: search, planning and knowledge representation are still in the agent toolbox today. They are just no longer the lead.

  1. 1943

    McCulloch & Pitts neuron model — writes the neuron as a threshold logic unit and proves that networks of them can compute any logical function. "Brain" and "computation" enter the same mathematics for the first time, and both the neural network path and the symbolic logic path branch from here.

    A Logical Calculus of the Ideas Immanent in Nervous ActivitySource · checked 2026-10-08

  2. 1950

    Turing, "Computing Machinery and Intelligence" — replaces the unanswerable "can machines think?" with an operational imitation game, and predicts machines that learn.

    Computing Machinery and IntelligenceSource · checked 2026-10-08

  3. 1956

    Dartmouth summer workshop — names the field artificial intelligence. The proposal's bet: every aspect of intelligence can in principle be described so precisely that a machine can simulate it.

    John McCarthySource · checked 2026-10-08

  4. 1958

    Rosenblatt's perceptron — weights are learned from examples, with a convergence proof, and it was built in hardware. The first learning classifier and the start of connectionism.

    The Perceptron: A Probabilistic Model for Information Storage and Organization in the BrainSource · checked 2026-10-08

  5. 1959

    Samuel's checkers program — tunes its evaluation function through self-play and ends up better than its author. The term "machine learning" comes from here.

    Arthur SamuelSource · checked 2026-10-08

  6. 1965–1976

    DENDRAL → MYCIN — rules from chemists and doctors replace general search, and diagnosis rivals experts. Domain knowledge matters more than the reasoning algorithm, and knowledge engineering methodology starts here. MYCIN was never used clinically (liability and integration problems).

    Edward FeigenbaumSource · checked 2026-10-08

  7. 1966

    ELIZA — pattern matching creates the illusion of being understood, and even the author was startled by how invested users became. The gap between "looks like it understands" and "really understands" is seen for the first time.

    Joseph WeizenbaumSource · checked 2026-10-08

  8. 1969

    Minsky & Papert, "Perceptrons" — proves a single-layer perceptron cannot even represent XOR, and nobody knew how to train multiple layers. Funding moves to the symbolic camp, and neural networks go quiet for over ten years.

    Perceptrons: An Introduction to Computational GeometrySource · checked 2026-10-08

  9. 1970–1972

    SHRDLU — natural language understanding in a blocks world. The peak demo of the symbolic camp, and it also shows this path cannot leave the toy world.

    Terry WinogradSource · checked 2026-10-08

  10. 1973–1974

    Lighthill report and DARPA cuts — the British government judges that AI failed to deliver (combinatorial explosion is the core argument), and the US cuts funding in step: the first winter.

    DARPASource · checked 2026-10-08

  11. 1980

    XCON / R1 (DEC) — an expert system that configures VAX orders saves money at scale for the first time and sets off the 1980s expert system industry.

    Carnegie Mellon UniversitySource · checked 2026-10-08

  12. 1984

    Cyc starts — Lenat bets on hand-coding all common sense, and keeps at it for nearly forty years. The most thorough experiment of the symbolic camp, and its outcome is the best textbook on why knowledge cannot be written by hand.

    Douglas LenatSource · checked 2026-10-08

  13. 1987

    Lisp machine market collapses, second winter — special-purpose hardware is replaced by general workstations, and the maintenance cost and brittleness of expert systems are exposed. Japan's Fifth Generation Computer project (1982–1992) also missed its goals. First wave: symbols and expert systems

    Source · checked 2026-10-08

  14. 1997

    Deep Blue beats Kasparov — search + special-purpose hardware + a hand-built evaluation function wins at chess. It proves brute force works in a narrow domain, and it is the closing high point of the symbolic search path.

    IBM ResearchSource · checked 2026-10-08

Second wave: statistical learning and deep learning (1986–2016)

The bet changed: knowledge is learned from data, not written by hand. Probability replaces logic. Features are not designed by hand either; the network learns the representation. Backpropagation existed in 1986, so why did deep learning only take off in 2012? What was missing was not the algorithm but three things arriving together: million-scale labeled data (ImageNet, 2009), cheap parallel compute (GPU + CUDA, from 2007), and small tricks that make deep networks trainable (ReLU, dropout, good initialization). AlexNet in 2012 put all three together, and over the next four years vision, speech, translation and games fell one by one.

  1. 1986

    Backpropagation (Rumelhart, Hinton, Williams) — multi-layer networks can be trained on error, and hidden layers learn representations on their own. It answers the question left by "Perceptrons", and connectionism revives.

    Learning representations by back-propagating errorsSource · checked 2026-10-08

  2. 1988

    Pearl, "Probabilistic Reasoning in Intelligent Systems" — Bayesian networks make reasoning under uncertainty computable, and probability returns to the AI mainstream.

    Judea PearlSource · checked 2026-10-08

  3. 1989 / 1998

    LeCun's convolutional networks: zip code recognition → LeNet-5 — puts priors into the network structure (local receptive fields, weight sharing) instead of into rules, and end-to-end gradient training yields a system that can read checks. MNIST becomes the standard dataset.

    Gradient-Based Learning Applied to Document RecognitionSource · checked 2026-10-08

  4. 1995

    SVM (Cortes & Vapnik) — maximum margin + kernel trick, a classifier backed by generalization theory, dominant in the 2000s.

    Support-Vector NetworksSource · checked 2026-10-08

  5. 1997

    LSTM — gated units solve vanishing gradients, so recurrent networks can remember over long distances. The workhorse of speech and translation in the 2010s.

    Long Short-Term MemorySource · checked 2026-10-08

  6. 2003

    Neural probabilistic language model (Bengio) — learns a continuous vector for each word, and a neural network predicts the next word. It sidesteps the sparsity problem of n-grams and is a common ancestor of word vectors and LLM.

    A Neural Probabilistic Language ModelSource · checked 2026-10-08

  7. 2006

    Deep belief networks (Hinton) — layer-by-layer pretraining makes deep networks trainable. "Deep learning" becomes a term and draws funding again.

    A fast learning algorithm for deep belief netsSource · checked 2026-10-08

  8. 2009

    ImageNet — ten million-scale human-labeled images. The bet is on data scale, not algorithms, and the 2012 breakthrough happens on it.

    ImageNetSource · checked 2026-10-08

  9. 2012

    AlexNet — a deep convolutional network trained on two GPUs cuts the ImageNet error rate by ten points in one step, and the old path of hand-made features + SVM is dropped. The breakout point of deep learning.

    ImageNet Classification with Deep Convolutional Neural NetworksSource · checked 2026-10-08

  10. 2013

    word2vec — a shallow model learns word vectors on a billion words, and vector arithmetic reveals semantic structure. "Pretrain representations without supervision first" becomes routine in NLP.

    Efficient Estimation of Word Representations in Vector SpaceSource · checked 2026-10-08

  11. 2013

    DQN — a convolutional network learns Atari straight from pixels, one set of hyperparameters across dozens of games. Deep learning and reinforcement learning merge.

    Playing Atari with Deep Reinforcement LearningSource · checked 2026-10-08

  12. 2014

    GAN — a generator and a discriminator train against each other, and it generates without writing a likelihood. The start of the deep generative model branch.

    Generative Adversarial NetsSource · checked 2026-10-08

  13. 2014

    seq2seq — an encoder-decoder LSTM does translation end to end, replacing IBM's 1990s statistical translation. It exposes the fixed-length vector bottleneck.

    Sequence to Sequence Learning with Neural NetworksSource · checked 2026-10-08

  14. 2014

    Attention mechanism (Bahdanau, Cho, Bengio) — at decode time, look back and weight each position of the source sentence, which fixes the drop in quality on long sentences. The direct predecessor of the Transformer.

    Neural Machine Translation by Jointly Learning to Align and TranslateSource · checked 2026-10-08

  15. 2015

    ResNet — residual connections make hundred-layer networks trainable, and it beats humans on ImageNet. Residuals become standard in every deep network from then on.

    Deep Residual Learning for Image RecognitionSource · checked 2026-10-08

  16. 2016

    AlphaGo beats Lee Sedol — deep networks + tree search + self-play (this line started with TD-Gammon's backgammon in 1992). Go was thought to be another ten years away. The public image of deep learning shifts from "recognizes images" to "makes decisions".

    Mastering the game of Go with deep neural networks and tree searchSource · checked 2026-10-08

Third wave: scaling and LLM (2017–2026)

The bet: the same skeleton (the Transformer) + more data and compute = predictable growth in capability, and one general pretrained model replaces all task-specific models. Scaling became the theme for three reasons. The Transformer removed recurrence, so training is fully parallel and any amount of compute can be used. GPU clouds turned compute into something you can buy. Scaling laws turned "throw ten times more at it" from faith into arithmetic. After 2022 two more axes were added: alignment (RLHF/DPO decide whether the model "listens") and inference-time compute (o1/R1 show that letting the model "think longer" also buys capability). This wave has not hit a wall yet, but benchmark saturation, running out of data and compute cost are the three walls it is approaching.

  1. 2017

    Transformer — drops recurrence and keeps only attention, so training is fully parallel. The skeleton of every large model since.

    Attention Is All You NeedSource · checked 2026-10-08

  2. 2017

    Deep reinforcement learning from human preferences (Christiano et al.) — trains a reward model from pairwise comparisons, and humans see less than one percent of the interactions. The prototype of RLHF. In the same year AlphaZero shows that three board games can be mastered with zero human game records.

    Deep Reinforcement Learning from Human PreferencesSource · checked 2026-10-08

  3. 2018

    ELMo / GPT-1 / BERT — pretraining + fine-tuning becomes the default paradigm in NLP, and within a year BERT pushes GLUE past the human baseline.

    BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingSource · checked 2026-10-08

  4. 2019

    GPT-2 — ten times larger, and it does many tasks without fine-tuning: "tasks are learned along the way". Its staged release of weights makes release safety an issue for the first time.

    Language Models are Unsupervised Multitask LearnersSource · checked 2026-10-08

  5. 2019

    The Bitter Lesson — Sutton says, from seventy years of experience, that human knowledge written in works in the short term, but in the long term loses to general methods driven by compute. The creed of the scaling wave.

    The Bitter LessonSource · checked 2026-10-08

  6. 2020

    Scaling Laws (Kaplan et al.) — loss follows a power law in parameters, data and compute each, holding across seven orders of magnitude. Investment becomes predictable.

    Scaling Laws for Neural Language ModelsSource · checked 2026-10-08

  7. 2020

    GPT-3 — 175 billion parameters, does new tasks from a few examples in the prompt with no weight updates. Prompting becomes a way to program, and the LLM startup rush forms around its API.

    Language Models are Few-Shot LearnersSource · checked 2026-10-08

  8. 2020–2022

    ViT · DDPM · CLIP · Stable Diffusion — the Transformer enters vision, diffusion models become usable, images and text share one vector space, and latent-space diffusion runs text-to-image on consumer graphics cards with open weights. GANs leave the stage.

    Learning Transferable Visual Models From Natural Language SupervisionHigh-Resolution Image Synthesis with Latent Diffusion ModelsSource · checked 2026-10-08

  9. 2021

    AlphaFold 2 — protein structure prediction reaches experimental accuracy, and a fifty-year problem is largely solved. Deep learning enters scientific discovery.

    Highly accurate protein structure prediction with AlphaFoldSource · checked 2026-10-08

  10. 2022

    Chain-of-thought prompting — have the model write out reasoning steps before answering, and it only works on large models. It sparks the "emergent abilities" debate, and two years later reasoning models move it from prompt into training.

    Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsSource · checked 2026-10-08

  11. 2022

    InstructGPT and Constitutional AI — SFT + reward model + PPO make a 1.3 billion model preferred over the 175 billion one, the recipe of ChatGPT. In the same year Anthropic replaces part of the human labeling with principles written in text + AI feedback.

    Training language models to follow instructions with human feedbackConstitutional AI: Harmlessness from AI FeedbackSource · checked 2026-10-08

  12. 2022

    Chinchilla — parameters and tokens should scale in the same proportion, and earlier large models were mostly undertrained. It rewrites the scaling recipe, and "smaller model, more data" becomes the mainstream.

    Training Compute-Optimal Large Language ModelsSource · checked 2026-10-08

  13. 2022-11

    ChatGPT — a hundred million users in two months. AI goes from papers to a consumer product and an industry topic. The public start of the third wave.

    OpenAISource · checked 2026-10-08

  14. 2023

    LLaMA · Llama 2 · Mistral 7B — frontier-class open-weight models. The whole ecosystem of llama.cpp, quantization and local inference grows from here, and Europe's Mistral shows a small model + good data can beat larger models.

    LLaMA: Open and Efficient Foundation Language ModelsMistral AISource · checked 2026-10-08

  15. 2023

    GPT-4 — multimodal, top 10% on professional exams, and the report says outright that architecture and data are not disclosed. The dividing line where frontier labs turn from "publishing papers" to "shipping products".

    GPT-4 Technical ReportSource · checked 2026-10-08

  16. 2023

    DPO — folds the reward model and PPO into one classification loss, and preference alignment takes a few dozen lines of code. Alignment goes from big-lab engineering to everyday open-source work.

    Direct Preference Optimization: Your Language Model is Secretly a Reward ModelSource · checked 2026-10-08

  17. 2024

    o1 and MCP — long chains of thought trained with reinforcement learning make "how long it thinks" a new scaling axis (inference-time compute). In November of the same year the Model Context Protocol starts to standardize how models connect to tools. Reasoning and tools are the two cornerstones of the agent era. Third wave: scaling and LLM MCP: one standard socket for every system

    Source · checked 2026-10-08

  18. 2024

    Nobel Prize in Physics (Hopfield, Hinton) and in Chemistry (Hassabis, Jumper, Baker) — formal recognition of neural networks and AlphaFold by mainstream science.

    Geoffrey HintonDemis HassabisSource · checked 2026-10-08

  19. 2025-01

    DeepSeek-R1 — reasoning elicited by pure reinforcement learning, with open weights and published cost. Reasoning models are not one company's secret, and the compute narrative gets discounted for the first time.

    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningSource · checked 2026-10-08

  20. 2026

    Benchmark saturation becomes the bottleneck — Stanford AI Index 2026: SWE-bench Verified rises from about 60% to nearly 100% in one year, and industry produces over 90% of notable models. Evals cannot keep up with the models, and ARC-AGI moves to interactive environments.

    SWE-benchARC-AGISource · checked 2026-10-08

Labs and companies

The organisations the timeline names; the ones in this site's hiring data link to their company page.

Amazon

1994LayersL6

One of the main employers for Applied Scientist roles. It publishes its hiring process and Leadership Principles as pages, which makes them primary source material for behavioral interview prep.

Related
Behavioral Rounds and the Values Round (set)
Sources
Official How we hire page: apply → assessment → phone screen → interview loop, with the Bar Raiser and Leadership Principles covered in the same place (checked 2026-10-08)

Anthropic

2021LayersL6L5L4Hiring data

A frontier lab, and the maker of Claude. For the career layer, what matters is that it wrote its policy on candidates' use of AI as a public page. The hiring page itself is the best prep material.

In the history layer it stands for the branch of "alignment as a core research direction". It split off from OpenAI in 2021. Constitutional AI: Harmlessness from AI Feedback replaces most human preference labeling with a set of written principles, and has the model critique and rewrite its own answers. Beyond the Claude series, it released the protocol in MCP:给每个系统一个统一的插口, which made "models connecting to tools" an industry standard. In Opinions and predictions, Dario Amodei's long essay is the most often cited judgment for the frontier layer.

Related
Constitutional AI: Harmlessness from AI Feedback (released) · MCP: One Standard Socket for Every System (released) · Behavioral Rounds and the Values Round (set)
Linked from
Dario Amodei (works at) · Jack Clark (works at) · Jared Kaplan (works at) · Constitutional AI: Harmlessness from AI Feedback (released) · Anthropic 博客(news · engineering · research) (released)
Sources
Official hiring page: no degree required. It suggests putting independent research / blog / open source at the very top of your résumé. (checked 2026-10-08) · Candidate AI use policy. The page is marked last updated 2025-07-10. (checked 2026-10-08)

Bell Labs

1925LayersL5

In the 1990s, LeCun's convolutional networks and Vapnik's SVM competed in the same building. The two routes of the second wave (neural networks vs. statistical learning) met here.

Related
Yann LeCun (works at) · Vladimir Vapnik (works at) · Gradient-Based Learning Applied to Document Recognition (released) · Support-Vector Networks (released)
Linked from
Vladimir Vapnik (works at) · Yann LeCun (works at) · Gradient-Based Learning Applied to Document Recognition (works at) · Support-Vector Networks (works at)
Sources
Bell Labs official website (checked 2026-10-08)

Carnegie Mellon University

1956LayersL5

Newell and Simon's Logic Theorist (1956) was the first AI program. XCON, speech recognition (Sphinx) and self-driving (NavLab) all came out of this line of work.

Related
The first wave: symbols and expert systems (proposed) · The Second Wave: Statistical Learning and Deep Learning (discusses)
Sources
CMU School of Computer Science (checked 2026-10-08)

Cursor

2022LayersL6L4Hiring data

An AI editor company (legal entity: Anysphere). It is one of the fastest-hiring AI-native startups. Its hiring page is a good sample of "work-sample style" role requirements.

Sources
Official hiring page. It lists roles only and says nothing about the process. (checked 2026-10-08)

DARPA

1958LayersL5

The main funder of AI in the 1960s and 70s. Its 1974 cuts were the US side of the first AI winter. In the 1980s its Strategic Computing Initiative gave expert systems another push. When you read AI history, its name stands for the funding curve.

Related
The first wave: symbols and expert systems (funded) · The first wave: symbols and expert systems (discusses)
Sources
DARPA website (checked 2026-10-08)

DeepSeek

2023LayersL5

A lab in Hangzhou funded by the quant fund High-Flyer. Its R1, released in January 2025 with open weights and a publicly stated low training cost, showed that reasoning models are not one company's secret.

Related
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released) · Liang Wenfeng (works at)
Linked from
Liang Wenfeng (works at) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released)
Sources
DeepSeek official website (checked 2026-10-08)

Epoch AI

2022LayersL5L3

A nonprofit research group that tracks AI trends. It publishes open datasets on training compute, model size, chip cost, and when we run out of data. When I talk about scaling and compute economics, I check the numbers here first.

Related
Epoch AI 数据中心 (released) · Pretraining and scaling (measures) · The economics of compute: where training money goes, how inference is priced (discusses)
Linked from
Epoch AI 数据中心 (released)
Sources
Official site (checked 2026-10-08)

Google

1998LayersL5

Google Brain (2011) turned deep learning into infrastructure. The Transformer, BERT and TPU all came from there. In 2023 it merged with DeepMind to form Google DeepMind, which entered the race with Gemini.

Related
Efficient Estimation of Word Representations in Vector Space (released) · Sequence to Sequence Learning with Neural Networks (released) · Attention Is All You Need (released) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Jeff Dean (works at)
Linked from
Ashish Vaswani (works at) · Geoffrey Hinton (works at) · Ian Goodfellow (works at) · Ilya Sutskever (works at) · Jeff Dean (works at) · Noam Shazeer (works at) · Tomáš Mikolov (works at) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Sequence to Sequence Learning with Neural Networks (released) · Efficient Estimation of Word Representations in Vector Space (released)
Sources
Google Research official website (checked 2026-10-08)

Google DeepMind

2010LayersL6L5L2

Google's AI lab (DeepMind was founded in 2010 and merged with Google Brain in 2023). It has public official hiring pages for research roles and research engineering roles.

In terms of history, it is the standard-bearer of the second half of the second wave: Playing Atari with Deep Reinforcement Learning taught a network to play Atari from pixels. Mastering the game of Go with deep neural networks and tree search used self-play plus search to beat top human players. Highly accurate protein structure prediction with AlphaFold applied the same approach to protein structures and won a Nobel Prize. Training Compute-Optimal Large Language Models corrected the recipe for scaling: for the same compute, data has to grow together with parameters. After the 2023 merger with Google Brain, the Gemini series put it back on the front line of LLM competition.

Related
Playing Atari with Deep Reinforcement Learning (released) · Mastering the game of Go with deep neural networks and tree search (released) · Highly accurate protein structure prediction with AlphaFold (released) · Training Compute-Optimal Large Language Models (released)
Linked from
Arthur Mensch (works at) · David Silver (works at) · Demis Hassabis (works at) · Highly accurate protein structure prediction with AlphaFold (released) · Mastering the game of Go with deep neural networks and tree search (released) · Training Compute-Optimal Large Language Models (released) · Deep Reinforcement Learning from Human Preferences (released) · Playing Atari with Deep Reinforcement Learning (released) · Google DeepMind 博客 (released)
Sources
Official hiring page. It lays out a four-stage process and says it is tuned per role. (checked 2026-10-08) · Google's How we hire page. The body text is rendered by script. (checked 2026-10-08)

Hugging Face

2016LayersL5

The transformers library and the model hub standardized how pretrained models are distributed. It is the "GitHub" of the open-weights ecosystem.

Related
Clément Delangue (works at) · The third wave: scaling and LLM (discusses)
Linked from
Clément Delangue (works at) · Hugging Face Daily Papers (released)
Sources
Hugging Face official website (checked 2026-10-08)

IBM Research

1945LayersL5

Samuel's checkers program (1959), Deep Blue (1997), Watson (2011), and statistical machine translation (1990s) are all here. In each of the three waves, it stood at the peak of the mainstream method of its day.

Related
Arthur Samuel (works at) · The first wave: symbols and expert systems (discusses)
Linked from
Arthur Samuel (works at)
Sources
IBM history page: Deep Blue (checked 2026-10-08)

IDSIA

1988LayersL5

A small lab in Lugano, Switzerland, and the birthplace of the LSTM. Around 2011 it won several vision competitions with GPU convolutional networks, before AlexNet.

Related
Jürgen Schmidhuber (works at) · Long Short-Term Memory (released)
Linked from
Jürgen Schmidhuber (works at)
Sources
IDSIA website (checked 2026-10-08)

Meta

2004LayersL6L5

One of the big-tech employers with the most AI roles. Besides FAIR, it has a superintelligence lab. Its careers page is a good sample of how a big company titles its AI roles.

Sources
Official careers site. It returned 429 / 400 to the scraping tool, so I could not verify the candidate guide page. (checked 2026-10-08)

Meta AI (FAIR)

2013LayersL5

FAIR, which LeCun founded in 2013. PyTorch and LLaMA both came out of it. It is the biggest driver of the open-weights route.

Related
LLaMA: Open and Efficient Foundation Language Models (released) · Yann LeCun (works at) · The third wave: scaling and LLM (discusses)
Linked from
Kaiming He (works at) · Tomáš Mikolov (works at) · Yann LeCun (works at) · LLaMA: Open and Efficient Foundation Language Models (released) · Meta AI 博客 (released)
Sources
Meta AI official website (checked 2026-10-08)

Microsoft

1975LayersL5

ResNet came out of Microsoft Research Asia. Since 2019 Microsoft has invested in OpenAI and been its exclusive compute provider, and Copilot made code generation a mass-market product for the first time.

Related
Deep Residual Learning for Image Recognition (released) · OpenAI (funded) · Kaiming He (works at)
Linked from
Kaiming He (works at) · Deep Residual Learning for Image Recognition (released)
Sources
Microsoft Research website (checked 2026-10-08)

Mila – Quebec AI Institute

1993LayersL5

Bengio's lab. In 2014 it produced both the attention mechanism and GANs in the same year. One of the three places (Toronto, Montreal, Edmonton) where Canada kept deep learning alive in academia.

Related
Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (released) · Generative Adversarial Nets (released)
Linked from
Ian Goodfellow (works at) · Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (works at)
Sources
Mila website (checked 2026-10-08)

Mistral AI

2023LayersL5Hiring data

Founded in Paris in 2023. Mistral 7B and Mixtral showed that a small model with good data can beat larger models. The main player in Europe's open-weights approach.

Related
Arthur Mensch (works at) · The third wave: scaling and LLM (discusses)
Linked from
Arthur Mensch (works at)
Sources
Mistral AI official website (checked 2026-10-08)

MIT AI Lab / CSAIL

1959LayersL5

The lab Minsky and McCarthy founded in 1959. It was the stronghold of the symbolic school: ELIZA, SHRDLU and the Lisp machines all came out of it. It merged into CSAIL in 2003.

Related
The first wave: symbols and expert systems (proposed) · Perceptrons: An Introduction to Computational Geometry (funded) · Marvin Minsky (works at)
Linked from
John McCarthy (works at) · Joseph Weizenbaum (works at) · Kaiming He (works at) · Marvin Minsky (works at) · Terry Winograd (works at)
Sources
CSAIL website (checked 2026-10-08)

NVIDIA

1993LayersL5

CUDA (2007) turned the GPU into a general-purpose compute device. After AlexNet, it became the physical foundation of deep learning. The compute economics of the third wave revolve around its supply cycle.

Related
Jensen Huang (works at) · The third wave: scaling and LLM (discusses)
Linked from
Jensen Huang (works at)
Sources
NVIDIA research page (checked 2026-10-08)

OpenAI

2015LayersL6L5Hiring data

A frontier lab, the author of the GPT series. Research engineer roles are its main hiring line, and its official interview guide is public.

In the history layer, it is the lead of the third wave. Improving Language Understanding by Generative Pre-Training through Language Models are Few-Shot Learners showed that the road of "the same objective, scaled up ten times" could keep going. Scaling Laws for Neural Language Models wrote that down as a formula. Training language models to follow instructions with human feedback taught the continuation machine to follow instructions. ChatGPT then put it in everyone's hands. The later GPT-4 Technical Report no longer discloses the architecture, which marks research moving from papers to products. For the nodes, see The third wave: scaling and LLM and Pretraining and scaling.

Related
Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Learning Transferable Visual Models From Natural Language Supervision (released)
Linked from
Alec Radford (works at) · Andrej Karpathy (works at) · Daniel Kokotajlo (works at) · Dario Amodei (works at) · Ilya Sutskever (works at) · Jared Kaplan (works at) · Paul Christiano (works at) · Sam Altman (works at) · Learning Transferable Visual Models From Natural Language Supervision (released) · Deep Reinforcement Learning from Human Preferences (released) · Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Microsoft (funded) · OpenAI 博客 (released)
Sources
Official interview guide. The fetch tool got a 403, but the search index confirms the page exists (including multilingual versions). (checked 2026-10-08)

Stanford AI Lab (SAIL)

1963LayersL5

Founded by McCarthy in 1963. Expert systems (DENDRAL, MYCIN) came out of it. Forty years later, ImageNet and DPO came from here too. It shows up in all the three waves.

Related
The first wave: symbols and expert systems (proposed) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released) · ImageNet (proposed)
Linked from
Andrew Ng (works at) · Edward Feigenbaum (works at) · Fei-Fei Li (works at) · John McCarthy (works at) · Terry Winograd (works at) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released)
Sources
SAIL website (checked 2026-10-08)

University of Toronto

1827LayersL5

The home of Hinton's lab, where deep learning was kept alive through the winter. Deep belief networks (2006) and AlexNet (2012) were both built here.

Related
Geoffrey Hinton (works at) · ImageNet Classification with Deep Convolutional Neural Networks (released) · A fast learning algorithm for deep belief nets (released)
Linked from
Alex Krizhevsky (works at) · Geoffrey Hinton (works at) · A fast learning algorithm for deep belief nets (works at)
Sources
University of Toronto, Department of Computer Science (checked 2026-10-08)

Benchmarks

What the field measured itself by, and when each stopped being hard.

MNIST

1998LayersL5L2

What it measures: 10-class classification of 70,000 handwritten digits, the "Hello World" of the second wave. When it saturated: by the early 2010s the error rate was below 0.3%, close to the label noise. Now it is only used for teaching and sanity checks.

Related
Classical machine learning: learning from examples (measures) · Yann LeCun (proposed) · Gradient-Based Learning Applied to Document Recognition (discusses)
Linked from
Gradient-Based Learning Applied to Document Recognition (measures)
Sources
MNIST entry (LeCun's original site is no longer reliable) (checked 2026-10-08)

ImageNet

2009LayersL5L2

What it measures: image classification over 1,000 classes and about a million images (ILSVRC). When it saturated: in 2012, AlexNet cut the top-5 error rate to 15%. In 2015, ResNet got below the human level of about 5%. In 2017, the challenge was discontinued. It is the prototype of "dataset-driven breakthroughs."

Related
Classical machine learning: learning from examples (measures) · Fei-Fei Li (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (discusses) · Deep Residual Learning for Image Recognition (discusses)
Linked from
Fei-Fei Li (proposed) · ImageNet Classification with Deep Convolutional Neural Networks (measures) · Deep Residual Learning for Image Recognition (measures) · Stanford AI Lab (SAIL) (proposed)
Sources
ILSVRC challenge page (2010–2017) (checked 2026-10-08)

SQuAD

2016LayersL5L2

What it measures: reading comprehension, extracting answers from Wikipedia paragraphs. When it saturated: BERT beat human EM/F1 in 2018. After version 2.0 added "unanswerable" questions, models beat humans on that too in 2019. Since then, reading comprehension has not been a standalone benchmark.

Related
LLM: how the model works (measures) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (discusses)
Linked from
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (measures)
Sources
Official leaderboard (checked 2026-10-08)

GLUE / SuperGLUE

2018LayersL5L2

What it measures: the average score over nine sentence-level NLP tasks. When it saturated: it was released in 2018, BERT pushed it past the human baseline within a year, it was replaced by the harder SuperGLUE in 2019, and SuperGLUE was also surpassed in 2021. The first "saturated in one year" case of the pretrained-model era.

Related
LLM: how the model works (measures) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (discusses)
Linked from
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (measures)
Sources
The SuperGLUE paper (2019) explains that GLUE had been passed by models beating the human baseline. (checked 2026-10-08)

ARC-AGI

2019LayersL5L2

What it measures: grid puzzles where you infer a rule from a few examples, designed so you can't memorize them. When it saturated: the 2019 version was cracked by reasoning models with heavy compute in late 2024. In 2025 it was replaced by ARC-AGI-2. In 2026 frontier models passed 90%, and ARC-AGI-3 moves to interactive environments. It is the proving ground for the "generalization vs memorization" debate.

Related
LLM: how the model works (measures) · Evals: knowing whether it got better (measures) · The third wave: scaling and LLM (discusses)
Sources
ARC Prize official site: ARC-AGI-2 (2025) page (checked 2026-10-08)

MMLU

2020LayersL5L2

What it measures: multiple-choice questions across 57 subjects, from elementary math to professional law. When it saturated: GPT-4 (2023) scored 86%. In 2024, frontier models crowded into 88–90%, and some questions turned out to be wrong. MMLU-Pro appeared because of this. It is the yardstick for the "knowledge" axis in the LLM era, and it has been saturated for four years.

Related
LLM: how the model works (measures) · GPT-4 Technical Report (discusses)
Linked from
GPT-4 Technical Report (measures)
Sources
The MMLU-Pro paper (2024) says MMLU performance has plateaued. (checked 2026-10-08)

GSM8K

2021LayersL5L2

What it measures: 8.5k grade-school word problems that test multi-step arithmetic reasoning. When it saturated: chain-of-thought prompting first showed its power on it (2022). In 2024, frontier models passed 95%. After that, people moved to competition problems such as MATH and AIME.

Related
LLM: how the model works (measures) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (discusses)
Linked from
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (measures)
Sources
Paper page (OpenAI, 2021) (checked 2026-10-08)

HumanEval

2021LayersL5L2

What it measures: 164 hand-written Python function problems, scored by unit test pass rate. When it saturated: 28.8% at the Codex release, 67% for GPT-4 (2023), and over 90% for frontier models in 2024. It was the first yardstick for code ability. SWE-bench took over after it.

Related
LLM: how the model works (measures) · GPT-4 Technical Report (discusses)
Linked from
GPT-4 Technical Report (measures)
Sources
Codex paper page (2021) (checked 2026-10-08)

SWE-bench

2023LayersL5L2

What it measures: real GitHub issues, where the model has to change code in a repo and pass the tests. It was the first agent-level benchmark. When it saturated: the Verified subset was about 50% in 2024 and 70–80% in 2025. The AI Index 2026 report says it is now close to 100%, and training data contamination means the scores can no longer be trusted.

Related
Agent: a model with tools in a loop (measures) · LLM: how the model works (measures) · The third wave: scaling and LLM (discusses)
Sources
Official leaderboard; AI Index 2026 says Verified went from about 60% to nearly 100% within a year (checked 2026-10-08)

Translated from the author's Chinese notes by a model; the Chinese page is the original.