Skip to main content

Layer L5

Industry Context I: History

The layer

Why it's worth learning

History is the only stable thing in this layer. The frontier is a stream, opinions have expiry dates, and "what the three waves bet on and where they hit the wall" won't change in ten years. It is worth two things. One is the "why" in interviews: why Transformer and not LSTM, why RLHF and not direct fine-tuning. The answer is always the wall the previous wave hit. The other is a coordinate system for judging the frontier. When a new paper or model comes out, first ask which wave it belongs to, which wall it targets, and whether there is a precedent. With coordinates, the stream won't sweep you away.

The three waves: bets and turning points

Symbolic AI and expert systems (1943–1997) The bet: intelligence is search and reasoning over symbols, and humans write the knowledge as rules. It delivered (DENDRAL, XCON, Deep Blue) and hit two walls. Combinatorial explosion kept search inside toy worlds, and common sense could not be hand-written in full or maintained. The two winters (1974, 1987) were how funding reacted to these two walls. The turning point was 1986, when backpropagation made multi-layer networks trainable and revived the line that Perceptrons had held down for seventeen years.

Statistical learning and deep learning (1986–2016) The bet: both knowledge and representations can be learned from data. The key question is "why only in 2012": the algorithm was already reading zip codes in 1989. What was missing was ImageNet-scale data, GPU compute and a few training tricks all arriving at the same time. After AlexNet, vision, translation and games fell one by one over four years. AlphaGo was the peak, and the 2014 attention mechanism was the seed of the next wave.

Scaling and LLM (2017–now) The bet: the same backbone plus more compute buys predictable capability growth. Transformer made training fully parallel, scaling laws made the investment calculable, GPT-3 delivered prompting as programming, RLHF turned it into ChatGPT, LLaMA opened the weights, and o1/R1 added inference-time compute. It hasn't hit a wall yet, but three walls are already visible: benchmark saturation, running out of data, and compute cost.

How to use the timeline and the entity graph

timeline is the skeleton: fifty milestones, each with one sentence on "what it solved and what it overturned" plus a first-hand source. Read it wave by wave. After each wave, close the file and retell the bet and the reason for the wall. If you can't, go back. entities/ is the flesh: for papers, it says who they overturned (overturn / extend in edges). For people, it says only at which node they mattered. For benchmarks, it says "what it measures and when it saturated". The speed of saturation is itself a thermometer for the three waves (ImageNet eight years, GLUE one year, SWE-bench one year). In Obsidian, open the graph view from any entity and walk two steps along the edges. That is reading a piece of history.

Top picks

In reading order; each pick links to its node below.

  1. 1

    Artificial Intelligence: A Modern Approach (4th ed.), ch. 1 paid

    Thirty pages that tell each change of direction as "who bet on what, and why they lost" → AI History: How to Read the Three Waves

  2. 2

    Between the Booms: AI in Winter free

    Why the winters happened, and how the word "winter" got constructed → The first wave: symbols and expert systems

  3. 3

    Cyc: Obituary for the greatest monument to logical AGI free

    The most concrete dissection of the knowledge acquisition bottleneck → The first wave: symbols and expert systems

  4. 4

    The Quest for Artificial Intelligence: A History of Ideas and Achievements free

    A full history written by a first-generation researcher, with first-hand detail on the first wave of systems → AI History: How to Read the Three Waves

  5. 5
  6. 6

    Deep learning free

    A self-summary by three people who were there, the standard structure for "why deep learning works" → The Second Wave: Statistical Learning and Deep Learning

  7. 7

    AlphaGo - The Movie free

    The peak of the second wave, turned from a one-line description into a story you can retell → The Second Wave: Statistical Learning and Deep Learning

  8. 8

    [1hr Talk] Intro to Large Language Models free

    Pretraining, alignment and scaling told in one hour as a causal chain → The third wave: scaling and LLM

The nodes

AI History: How to Read the Three Waves

Half-life: stable

ai-historyai-history-deep-learningai-history-symbolicno prerequisite

Seventy-odd years of AI compress into three waves. Each one is a bet, a stretch of payoff, and a wall. Symbolic AI and expert systems bet that intelligence can be written by hand as rules. They hit combinatorial explosion and the common-sense bottleneck. Statistical learning and deep learning bet that both knowledge and representations can be learned from data. That paid off in 2012, when data, GPU and tricks all came together. Scaling and LLM bet that the same skeleton plus more compute buys predictable capability growth. It is still paying off. The walls are benchmark saturation, running out of data, and the cost of compute. I study this for the "why" questions in interviews (why Transformer, why RLHF) and for a coordinate system to judge the frontier: when something new appears, ask which wave it belongs to and which wall it targets. Milestones are in timeline. There is one node per wave: Wave 1: symbolic AI and expert systems · Wave 2: statistical learning and deep learning · Wave 3: scaling and LLM. The write-up is in the atlas.

Leads to
The Second Wave: Statistical Learning and Deep Learning · The first wave: symbols and expert systems
  • Artificial Intelligence: A Modern Approach (4th ed.), ch. 1 paid

    book · Ch. 1, about 2 hours · only ch. 1 — Thirty pages on each route change, 1943–2020: who bet on what, why they lost. Written by a symbolic insider, fairest to wave one. You can then name each wave's bet by era (first finish-line item).

  • The Quest for Artificial Intelligence: A History of Ideas and Achievements free

    book · Pick by wave, about 10 hours — Only full 1950–2008 history written by a first-generation researcher; free PDF on his Stanford page. Firsthand detail on DENDRAL, SHRDLU, expert systems: how they worked, why they stopped.

Hands-on project: Redraw the Three Waves in Your Own Words — Without looking at the timeline, write from memory each wave's bet, why it hit a wall, and what each of three key breakthroughs solved. Then check against timeline.md and mark what you got wrong or missed. Those errors are the focus of your next review.

Not picked (4)
  • Sutton《The Bitter Lesson》 — The author's site is http only, and this site links https only. The reading guide still sums up its argument in one sentence. Search for the original yourself.
  • Michael Wooldridge《A Brief History of AI》(2021) — Coverage overlaps AIMA chapter 1, and it is a paid book. One node keeps only one paid item.
  • Cade Metz《Genius Makers》(2021) — Good narrative, but it covers only second-wave figures. Only 8 HN mentions, so weak endorsement. Revisit under Not picked in the deep learning node.
  • Pamela McCorduck《Machines Who Think》(1979/2004) — Best first-wave oral history, but out of print and paid. 9 HN mentions, and I found no course syllabus citing it.

The first wave: symbols and expert systems

Half-life: stable

ai-historyai-history-symbolic

The 1956 Dartmouth bet was that every aspect of intelligence can be described precisely enough for a machine to simulate it. In practice that meant search and logical reasoning, later joined by hand-written domain knowledge (expert systems). It paid off: Samuel's checkers program, DENDRAL and XCON all really worked, and Deep Blue is the high point of this road. It hit two walls. Search could not leave toy worlds (combinatorial explosion, the core argument of the Lighthill report). And common sense is too large and too implicit to write by hand or to maintain (the knowledge acquisition bottleneck, with Cyc as the longest proof). Both winters (1974, 1987) were funding's reaction to these two walls. Today's agent tool calling, planning and the knowledge atlas are what it left behind, though no longer in the lead role. For milestones see timeline. Key people: John McCarthy, Marvin Minsky, Edward Feigenbaum, Douglas Lenat.

Builds on
AI History: How to Read the Three Waves
  • Between the Booms: AI in Winter free

    article · 30 minutes — A historian on the AI winters: why funding broke twice, how expert systems became a maintenance nightmare, how the word winter was coined afterward. Turns two dates into a causal story.

  • Cyc: Obituary for the greatest monument to logical AGI free

    article · 2–3 hours (long read) — Cyc's forty years of rise and fall: why hand-written common sense never ends and gets brittle, and why Lenat had no successor. The best dissection of the knowledge acquisition bottleneck.

Hands-on project: Take apart an expert system — Read the sample rules in any chapter of the MYCIN book. Write a small diagnosis program in under a hundred rules. Then give it three cases outside its boundary and record how it breaks. That is how the knowledge acquisition bottleneck feels.

Not picked (3)
  • 1973 年 Lighthill 辩论录像(BBC)与其整理稿仓库 — Primary source, but 64 points on HN and 37 repo stars. Not enough endorsement to clear the bar.
  • Buchanan & Shortliffe《Rule — Based Expert Systems: The MYCIN Experiments》(1984, free online from the author): first-hand technical detail on expert systems. HN 92, just short; project material, not a pick.
  • Pamela McCorduck《Machines Who Think》 — See AI history: how to read the three waves, under Not picked.

The Second Wave: Statistical Learning and Deep Learning

Half-life: stable

ai-historyai-history-deep-learningai-history-scaling

The bet changed from "hand-write the knowledge" to "learn from data". First came probability and statistical learning (Bayesian networks, SVM, statistical machine translation), which replaced logic. Then neural networks stopped hand-designing even the features. Backpropagation existed in 1986, but deep learning did not take off until 2012. What was missing was not the algorithm but three things arriving together: ImageNet-scale labeled data, cheap parallel compute from GPU, and a small set of tricks that made deep nets trainable. Once AlexNet put them together, four years later vision (ResNet), speech, translation (seq2seq + attention) and games (DQN, AlphaGo) fell one by one. The 2014 attention mechanism was already the seed of the next wave. Milestones are in the timeline. Key people: Geoffrey Hinton, Yann LeCun, Yoshua Bengio, Fei-Fei Li. Key papers: Learning representations by back-propagating errors, ImageNet Classification with Deep Convolutional Neural Networks, Deep Residual Learning for Image Recognition, Mastering the game of Go with deep neural networks and tree search.

Builds on
AI History: How to Read the Three Waves
Leads to
The third wave: scaling and LLM
  • Deep Neural Nets: 33 years ago and 33 years from now free

    article · 30 minutes — Rebuild LeCun's 1989 zip-code net in PyTorch, then add 30 years of improvements (Adam, augmentation, dropout). It answers why 2012: the algorithm was old; data and compute were missing.

  • AlphaGo - The Movie free

    video · 90 minutes — The second wave's climax as a documentary: why move 37 shocked players, how Lee Sedol won game 4, how unsure DeepMind was. Deep learning + search + self-play becomes a story you can retell.

  • Deep learning free

    paper · 1.5 hours — Three pioneers' review, 3 years after AlexNet: why learned representations beat handcrafted features, what conv and recurrent nets solve. Use its structure for 'why does deep learning work?'

Hands-on project: Reproduce 1989 — Follow Karpathy's post to get the 1989 network running, then change only one thing (10x the data, or switch to Adam, or add augmentation) and plot the test error for each change. That plot is your own evidence for why it took off in 2012.

Not picked (3)
  • Cade Metz《Genius Makers》(2021) — The most readable people's history of this wave (the Hinton auction, the DeepMind acquisition), but paywalled and only 8 HN comment mentions, below the endorsement bar.
  • Terrence Sejnowski《The Deep Learning Revolution》(2018) — Told from the participants' view, but zero HN discussion and paywalled.
  • Hinton 2024 年诺贝尔演讲 — A good oral history, but no community endorsement numbers to cite.

The third wave: scaling and LLM

Half-life: stable

ai-history-deep-learningtransformerai-history-scaling

赌注是「同一个骨架 + 更多数据和算力 = 可预测的能力增长」,一个通用预训练模型替代所有任务专用模型。scaling 能成为主题有三个前提:Transformer 去掉循环后训练完全并行,算力再多也用得上;GPU 云把算力变成可买的商品;scaling laws 让「再砸十倍」从信仰变成可算的账。GPT-3 兑现了提示即编程,InstructGPT 加上对齐让它变成 ChatGPT,Chinchilla 改写配方,LLaMA 把前沿模型开放出去,o1 和 DeepSeek-R1 又加了「推理期算力」这条新轴。它还没撞墙,但三堵墙已经看得见:基准饱和、高质量数据耗尽、算力成本。里程碑见 timeline;关键论文 Attention Is All You Need、Scaling Laws for Neural Language Models、Language Models are Few-Shot Learners、Training language models to follow instructions with human feedback、Training Compute-Optimal Large Language Models、LLaMA: Open and Efficient Foundation Language Models、DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning;Transformer 本身在 Transformer,后训练在 后训练:用奖励塑形。

Builds on
The Second Wave: Statistical Learning and Deep Learning · Transformer
  • [1hr Talk] Intro to Large Language Models free

    video · 60 minutes — Pretraining packs the internet into weights; fine-tuning and RLHF make an assistant; scaling laws explain the GPU rush. The 'LLM OS' close previews the agent wave. Scaling becomes a causal story.

Hands-on project: Verify a scaling law on small models — Use nanoGPT to train 4 model sizes on the same corpus (parameters doubling each step), with a fixed token count, and plot loss against parameter count on a log-log chart. Then double the token count and run it again to see how the slope changes. This is a hand-made version of the Kaplan and Chinchilla papers.

Not picked (4)
  • Ilya Sutskever 的 NeurIPS 2024 时间检验奖演讲(「预训练时代会结束」) — Only an unofficial re-upload on YouTube (seremot channel). The site embeds only videos the author published himself. Will add once there is an official upload.
  • Gwern《The Scaling Hypothesis》(2020) — The earliest and longest defense of the scaling belief, but its HN score is 97 (2022-05-30), three points short. The 146-point one is someone else's retelling.
  • Karpathy《Deep Dive into LLMs like ChatGPT》(2025,HN 582 分) — Better and longer, but it covers how an LLM works (L2), not history. Saved for LLM: How the Model Works.
  • Kaplan 等《Scaling Laws for Neural Language Models》 — 一手论文已是实体 Scaling Laws for Neural Language Models,不再占精选名额。

Self-check

  • Without looking at any material, state in three sentences the bet of each of the three waves, and the wall each one hit (or is approaching).
  • Given four milestones, backpropagation, AlexNet, Transformer and InstructGPT, say one sentence on each: what it solved and what it overturned.
  • Explain "why deep learning only took off in 2012" and "why scaling is the theme of the third wave", with at least two concrete reasons each.
  • Given a new paper or release from this week, place it in a wave, name the wall it targets, and give one precedent.

Translated from the author's Chinese notes by a model; the Chinese page is the original.