Skip to main content

Layer L5

Frontier: a stream, not a library

The layer

L5 has three parts: history (stable), frontier (a stream), opinions (dated). This page covers the last two. History lives in timeline and AI history: how to read the three waves.

Why the frontier is a stream

The half-life of the frontier is measured in weeks. A few months later, most of it is replaced by the next thing. Treat it as a library and you only pile up links you never read again, and you get a false feeling of "I've already learned this." Treat it as a stream: subscribe to a few filtered sources, read at a fixed time, leave no trace after reading, and note only what changed your judgment. Knowledge goes into nodes/. The hype flows away with the stream.

How to use the two tiers

  • Weekly reads: the frontier feed (≤ 5): weekly reads, at most five, covering five slots: papers, research and policy, post-training and open models, hands-on tests of the application layer, and people's opinions. To add a sixth, drop one first.
  • Look up when needed: frontier reference list (≤ 15): look up when needed, at most fifteen, each source tied to one kind of question (choosing a vendor, checking orders of magnitude, entering a new area, verifying numbers from a launch). No subscribing. The question comes before the source.
  • One media entity per source (entities/media/): who runs it, its stance, when it is worth reading, with edges to the nodes it often discusses.

The weekly time rule

The details are in Weekly reads: the frontier feed (≤ 5): 75 minutes every Saturday morning, in a fixed order, with the full podcast saved for the commute. When time is up, stop. Do not make up what you missed. The output is a single line: one opinion written into opinions, or "nothing worth noting this week." For a paper I want to read, I write down only the title. If I still remember it next week, I read it.

Why opinions carry dates

Someone says "AGI in five to ten years" every year. Without a date, it can never be wrong. Recording opinions is not collecting quotes. It is a ledger I can come back to and check against the answers: who, which day, what they said, when it can be verified. So I only take falsifiable claims that carry a time, a scale or a capability. I restate each claim in my own words, precisely enough to be judged right or wrong, and the source carries the date I checked it. A few years on, this page will tell me how much weight to give whose judgment.

How reviews write back

When an opinion comes due, write its status and one review line in opinions.md. If the result changed my judgment of a node, edit that node. If it changed my view of a person, add one sentence to the person entity. If the weekly stream shows that a pick has gone stale, note one line, then revise the nodes together once a quarter. Between the stream and the library, only this one narrow channel remains.

The nodes

Weekly reads: the frontier feed (≤ 5)

Half-life: a stream

llmfrontier-weekly

The frontier is a stream, not a library: subscribe, read for a fixed time, and don't save anything. This table is all of the weekly reads. Every other source is in Look up when needed: frontier reference list (≤ 15), so look up when needed. The five sources cover papers (HF), research and policy (Import AI), post-training and open models (Interconnects), hands-on testing at the application layer (Simon Willison), and people's views (Dwarkesh).

Source Form Rhythm Why this one How to read
Hugging Face Daily Papers Paper feed Daily, read once a week arXiv filtered by community votes. It took over from Papers with Code as the "what's hot lately" entry point Read only titles and vote counts, 10 minutes. Note one line for each paper I want to read, don't open it
Import AI newsletter One issue a week Each item says "why it matters". Two views, research and policy. The author is in a frontier lab Read the bold conclusion of each section, 20 minutes. Open at most one original
Interconnects newsletter One or two issues a week First-hand analysis of post-training and open models, independent judgment on release days Read the whole thing, 20 minutes. In a release week, read it before the press release
Simon Willison's Weblog Practitioner blog Daily Same-day tests of every new model, with costs and pitfalls. A first-hand L4 record Scan this week's entry titles, 15 minutes. Open only the hands-on tests
Dwarkesh Podcast podcast Two or three episodes a month Key people's dated judgments are spoken here Listen to one full episode a month (commute time, not counted in the weekly total). For the rest, read only the transcript subheadings, 10 minutes

Fixed weekly time

  • One slot: Saturday morning, 75 minutes in total (the sum of the items above; the full podcast episode goes in the commute and doesn't use it).
  • Order: HF titles → Import AI → Interconnects → Simon Willison → Dwarkesh subheadings. When time is up, stop. Don't make up what I didn't finish.
  • Don't save anything: no saved links, no clippings. Each week I write just one line to opinions: either a dated opinion, or "nothing worth noting this week". For papers I want to read, I note the title. If I still remember it next week, I go read it.
  • When to touch a node: only edit nodes/ when something I saw this week changes my judgment of a node (for example, a pick has gone stale). Do it in one batch each quarter, and leave it alone otherwise.
  • Missed an issue: skip it, don't go back. That's what a stream means: missing things is fine.
Builds on
LLM: how the model works
  • Import AI free

    article · 20 min / week — The week's papers and releases worth knowing. Each comes with why it matters and a policy view. The author works in a frontier lab and doesn't hype.

  • Interconnects free

    article · 20 min / week — First-hand analysis of post-training and open models. On big LLM release days, an independent take beats the press release.

  • Dwarkesh Podcast free

    video · Monthly full episode, ~2.5 h — Key people's dated judgments are mostly spoken here. It is the main source of opinions.

Not picked (6)
  • Latent Space — Gets the application-layer mood right, but it is long and the daily AINews is too dense. Moved to look up when needed.
  • Ahead of AI — Irregular long posts, not a weekly rhythm. Moved to look up when needed.
  • The Batch — Aimed at beginners. I don't rely on it to follow the frontier.
  • Lex Fridman Podcast — Wide guest list but shallow follow-up questions. For the few AI episodes, just listen to them directly.
  • No Priors — Investor view: more about products and funding than technology.
  • arXiv cs.CL / cs.LG 日更 — No filtering. Subscribing will drown you. Use it for lookups.

Look up when needed: frontier reference list (≤ 15)

Half-life: a stream

llmfrontier-on-demand

I don't subscribe to these sources or read them on a schedule. Each one answers one kind of question, and I go there when I have that question. The only things I read every week are the five in weekly reads: the frontier feed (≤ 5). Credibility has three levels: High = primary and independent; Medium = primary but with a stake, or independent but an estimate; Low = needs several sources that back each other up.

Source What to look up Credibility
arXiv cs.CL / cs.LG listing pages Find the original when you know the ID or title; browse the past week's papers in an area by keyword High (primary, not peer reviewed)
alphaXiv Which papers were discussed in the past week, and how the authors answer questions in the comments Medium (paper is primary, comments are secondary)
Anthropic blog (news · engineering · research) Engineering practice for agent / tool calling; model cards and eval tables for new models Medium (primary, has a stake)
OpenAI blog Original source for system cards, pricing, and work on reasoning and RL Medium (primary, marketing-facing)
Google DeepMind blog Gemini / Gemma technical reports, world models, RL and scientific applications Medium (primary, has a stake)
Meta AI blog Llama model cards and licenses, original papers on the JEPA line Medium (primary, has a stake)
LMArena leaderboard (arena.ai) A new model's position on the preference leaderboard and category boards, one or two weeks after release Medium (reflects preference, can be gamed)
Artificial Analysis Choosing a model / provider: capability index, throughput, latency and price on one page High (independent; the index uses custom weights)
Epoch AI data hub Historical data on training compute, data and hardware cost; check before writing an order-of-magnitude claim High
SemiAnalysis The economics of chips, interconnect and data centers; how much a lab spent on compute Medium (supply chain is primary, economics are estimates, paywalled)
Latent Space What people use now in an application-layer area (eval tools, agent frameworks); the start-of-year reading list Medium (good read of the mood, judgments sometimes hot)
Ahead of AI Layer-by-layer comparison of a batch of new model architectures; a month of papers walked through by theme High
Lil'Log (Lilian Weng's blog) Surveys before entering a new area (agent, diffusion, RL, hallucination) High
The Batch Wording for explaining this week's events to non-technical people High but shallow
r/LocalLLaMA Whether a model runs on your own hardware and how fast; community measurements when you doubt a release's numbers Low to medium (consensus is credible, a single post is not)

How to use it

  • Question before source: write down what you need to answer first (choose a provider? a survey of an area? the origin of a number?), then find the source in the "What to look up" column. Don't browse the other way round.
  • Write what you find back into the relevant node or opinions. Don't bookmark the source page itself.
  • Read all four lab blogs together: if you read only one for a given event, its stake will carry you off. For numbers, go by Artificial Analysis and the Epoch AI data hub.
Builds on
LLM: how the model works
  • Epoch AI free

    docs — Check it before writing any order-of-magnitude claim: time series of training compute, data and hardware cost. Methods are public and the data is downloadable.

  • Artificial Analysis free

    docs — The only page where you can see capability, speed and price together when choosing a model or provider, independent of what labs report about themselves.

  • Lil'Log free

    article — Before entering a new area, check whether she has written a survey. One post can replace years of tracking the main papers in that area.

Not picked (4)
  • Hacker News — Not a source but a tool for gathering evidence: this library's `evidence` scores come from it. Don't subscribe to it as a feed.
  • Semantic Scholar — Good for checking citation counts, but this library's picks use community endorsement as evidence. Citation counts are left to paper entities.
  • Lex Fridman Podcast · No Priors — See Not picked in weekly reads: the frontier feed (≤ 5).
  • 各实验室的 Discord — Too noisy. One visit when a model is released is enough, so it is not listed.

Media: who to follow

Twenty outlets, each with who runs it, its stance and when it is worth the time; the nodes above say which to read weekly and which on demand.

Ahead of AI

2022LayersL2L5

Sebastian Raschka (author of 《Build a Large Language Model (From Scratch)》)'s newsletter. Long posts, no fixed schedule. He either lines up the architectures of a batch of new models and compares them layer by layer, or walks through a month of papers by topic. His taste is the L2 teaching school: many diagrams, few formulas, and he has reproduced the work himself. Credibility: high. Look up when needed: come here when you want one readable comparison of an architecture change (MoE, attention variants, looped layers).

Related
Transformer (discusses) · LLM: how the model works (discusses) · Fine-tuning: teach it your task (discusses) · Sebastian Raschka (released)
Linked from
Sebastian Raschka (released)
Sources
About page: Over 220,000 subscribers (checked 2026-10-08) · Hacker News, 521 points (2026-09-09): post on GPT-6 Astra and the looped Transformer (checked 2026-10-08)

alphaXiv

2024LayersL2L5

A "comment section on top of arXiv" built by Stanford students: one discussion page per paper, a trending list, and an auto-generated overview. The authors often answer questions in the comments. Since Papers with Code stopped being maintained, if you want to find "which papers got discussed this past week", this and Hugging Face Daily Papers are the two replacements. Credibility: the papers are primary sources. The comments and auto-generated overviews are secondary. Use them to locate questions, not as conclusions.

Related
LLM: how the model works (discusses) · Agent: a model with tools in a loop (discusses)
Sources
Homepage, trending list + paper comments (checked 2026-10-08) · Hacker News, 549 points (2024-09-08) (checked 2026-10-08)

Anthropic 博客(news · engineering · research)

2021LayersL4L5

Three indexes: news covers models and products, engineering covers practical work on agent, tool calling, and Claude Code, and research covers interpretability, alignment, and model cards. The engineering blog is primary material for L4 (the pick in Agent:循环里带工具的模型 comes from here). Credibility: primary, but with a stance. For model cards and research, read the data. For product posts, read what they leave out. Look up when needed: on a new model's release day, check the eval table in the model card; when building an agent, check engineering.

Related
Agent: a model with tools in a loop (discusses) · MCP: One Standard Socket for Every System (discusses) · Calling a model API: token, streaming, tools, cost (discusses) · Anthropic (released)
Sources
Engineering blog index (checked 2026-10-08) · Research index (interpretability, alignment, model cards) (checked 2026-10-08) · Hacker News,673 分(2025-11-24):工程博客「Advanced Tool Use」 (checked 2026-10-08)

Artificial Analysis

2023LayersL3L5

An independent evaluation site. It runs models on the same set of benchmarks to produce an "intelligence index", then adds each API provider's throughput, first-token latency and price. It is the only place where I can see all three on one page when choosing a model and a provider. Credibility: higher than lab self-reports. The index is its own weighted definition, so scores are not comparable across version changes. Look up when needed: check the price and speed pages when choosing a model or provider; on release day, compare against the numbers the lab reports itself.

Related
LLM: how the model works (measures) · Serving: running models yourself (discusses) · Calling a model API: token, streaming, tools, cost (discusses)
Sources
Home page: the three charts for intelligence index, speed and price (checked 2026-10-08) · Hacker News, 916 points (2026-06-17): the post on GLM-5.2 topping open-weights models (checked 2026-10-08)

arXiv cs.CL / cs.LG 列表页

1991LayersL2L3L5

The original source for preprints. Several hundred papers a day, no filtering. Don't subscribe to the daily feed (it will drown you). Look up when needed: if you know the paper ID or title, go straight to the abs page. If you want to see what appeared in one area over the past week, use the listing page plus a keyword search. Credibility: the paper itself is a primary source, but it has not been peer reviewed. Before trusting a conclusion, check whether its eval setup and baselines are fair.

Related
LLM: how the model works (discusses) · Transformer (discusses) · Serving: running models yourself (discusses)
Sources
Recent listing page for cs.CL (checked 2026-10-08) · Recent listing page for cs.LG (checked 2026-10-08)

Google DeepMind 博客

2010LayersL2L5

The official outlet for Gemini / Gemma releases, world models (Genie), scientific applications (the AlphaFold line), and reinforcement learning. Compared with other labs, it has more science and RL content. Historical milestones (AlphaGo, AlphaFold) also have their original write-ups here. Credibility is the same as other lab blogs: first-hand, with a point of view. The technical reports are more trustworthy than the blog posts. Look up when needed: on a Gemini launch day, read the technical report; for RL and scientific applications, start from here to find the original papers.

Related
LLM: how the model works (discusses) · Post-training: shaping with rewards (discusses) · AI History: How to Read the Three Waves (discusses) · Google DeepMind (released)
Sources
Blog index (checked 2026-10-08) · Hacker News, 1812 points (2026-04-02): Gemma 4 release (checked 2026-10-08)

Dwarkesh Podcast

2020LayersL5

A long-form interview podcast hosted by Dwarkesh Patel. Guests are lab founders, chief scientists and economists. Each episode runs two to three hours. The host does thorough prep and keeps pushing until he gets a real answer. It is the main source of the views in opinions: most dated judgments from key people come from here. When it is worth your time: each month, pick one episode tied to a layer I care about and listen to the full version. For the rest, read only the subheadings of the transcript. When I hear a claim with a date attached, I note down one opinion.

Related
AI History: How to Read the Three Waves (discusses) · LLM: how the model works (discusses) · Agent: a model with tools in a loop (discusses) · Dwarkesh Patel (released)
Linked from
Dwarkesh Patel (released)
Sources
About page: Over 111,000 subscribers (Substack) (checked 2026-10-08) · Hacker News, 1212 points (2025-10-17): the Karpathy episode (checked 2026-10-08) · Hacker News, 450 points (2025-11-25): the Ilya Sutskever episode (checked 2026-10-08)

Epoch AI 数据中心

2022LayersL3L5

A nonprofit research institute, Epoch AI, maintains these datasets and charts: training compute, data volume, hardware price-performance and compute growth rates for models over the years, plus its own benchmarks such as FrontierMath. When you need numbers for "how fast has scaling actually gone," this is the most commonly cited primary compilation. Credibility: high. The methods are public and the data can be downloaded. One caution: it has benchmark partnerships with labs (the FrontierMath funding dispute with OpenAI). Look up when needed: check here before you write any judgment that carries an order of magnitude.

Related
Serving: running models yourself (discusses) · Evals: knowing whether it got better (discusses) · AI History: How to Read the Three Waves (discusses) · Epoch AI (released)
Linked from
Epoch AI (released)
Sources
Trends page: time series for training compute, data, hardware and cost (checked 2026-10-08) · Benchmark dashboard (checked 2026-10-08) · Hacker News, 480 points (2026-03-24): an open FrontierMath problem was solved (checked 2026-10-08) · Hacker News, 445 points (2026-05-24): a post on the cost breakdown of chips (checked 2026-10-08)

Hugging Face Daily Papers

2023LayersL2L4L5

A daily paper board run by Hugging Face. The community submits the arXiv papers worth reading that day and upvotes them into a ranking. Each paper comes with a short summary written by the authors or by AK. The taste leans toward open-source models, training and inference methods, and multimodal work, not theory. When it is worth your time: scan the titles once a week (titles and upvote counts only, 5–10 minutes), and log the one or two papers you really want to read in the "one-line record" of weekly reads: the frontier feed (≤ 5). It replaces the now-dormant Papers with Code as my entry point for "what's hot lately".

Related
LLM: how the model works (discusses) · Fine-tuning: teach it your task (discusses) · Agent: a model with tools in a loop (discusses) · Hugging Face (released)
Sources
Daily papers page, ranked by community upvotes, with an email digest you can subscribe to (checked 2026-10-08)

Import AI

2016LayersL2L5

Jack Clark (co-founder of Anthropic, formerly policy lead at OpenAI) has run this weekly newsletter since 2016. Each issue picks 5–8 papers or releases and writes a paragraph on each: why it matters. At the end there is a policy and safety angle and a piece of short fiction. His stance is "tech optimism, but serious about risk." He rarely hypes. When it is worth reading: read one issue every week. Read the bold conclusion of each paragraph first, then decide which originals to open. 20 minutes is enough.

Related
LLM: how the model works (discusses) · Evals: knowing whether it got better (discusses) · AI History: How to Read the Three Waves (discusses) · Jack Clark (released)
Linked from
Jack Clark (released)
Sources
About page: Over 143,000 subscribers (checked 2026-10-08) · Issue 455 (2026-05-04): his probability estimate for automated AI R&D, included in [[opinions]] (checked 2026-10-08)

Interconnects

2023LayersL2L5

A newsletter by Nathan Lambert (post-training lead at Ai2). It focuses on open models, post-training (RLHF / DPO / RL reasoning) and lab news. It is one of the few public analyses written by someone who does post-training at the front line. The taste leans technical and open-source, and he gives independent judgment on closed labs' launches. When to read it: one issue a week; when a major model ships (especially open weights), read his quick take, which is more useful than the press release.

Related
Fine-tuning: teach it your task (discusses) · Post-training: shaping with rewards (discusses) · LLM: how the model works (discusses) · Nathan Lambert (released)
Linked from
Nathan Lambert (released)
Sources
About page: Over 85,000 subscribers (checked 2026-10-08) · Hacker News, 367 points (2026-06-23): post on GLM-5.2 (checked 2026-10-08)

Latent Space

2023LayersL4L5

A newsletter + podcast on "AI Engineer" run by swyx and Alessio, plus a daily AINews roundup. The taste is application layer: agent frameworks, evals, RAG, inference providers. Many guests build products. Credibility: medium. It reads the mood of the industry well, but its technical judgment sometimes runs hot. Look up when needed: when I want to know what people use now in some application-layer area (eval tools, agent frameworks), I search its back issues. The reading list from the start of the year is a good starting point for L4.

Related
Agent: a model with tools in a loop (discusses) · RAG: Look Up First, Then Answer (discusses) · Evals: knowing whether it got better (discusses) · swyx (Shawn Wang) (released)
Linked from
swyx (Shawn Wang) (released)
Sources
About page: Over 203,000 subscribers (checked 2026-10-08) · Hacker News, 490 points (2025-01-13): AI Engineer Reading List (checked 2026-10-08)

Lil'Log(Lilian Weng 的博客)

2017LayersL2L4L5

Lilian Weng's personal blog. She posts a few times a year. Each post is a survey of one area: agent, the Transformer family, diffusion models, RL, hallucination. Reading one post is like reading the main line of papers in that area from the past several years. Credibility: high. The author does research on the front line. Look up when needed: before you enter a new area, check whether she has written a survey on it. Don't wait for updates. They are slow.

Related
Agent: a model with tools in a loop (discusses) · Transformer (discusses) · Post-training: shaping with rewards (discusses) · Lilian Weng (released)
Linked from
Lilian Weng (released)
Sources
Blog home page (checked 2026-10-08) · Hacker News, 334 points (2026-08-04): post on self-improving harness engineering (checked 2026-10-08) · Hacker News, 285 points (2023-06-27): survey of LLM-powered autonomous agents (checked 2026-10-08)

LMArena 榜(arena.ai)

2023LayersL2L5

An anonymous, double-blind head-to-head leaderboard where people vote and the votes become Elo scores. It measures "which answer people like better," not the ceiling of a model's ability. Credibility: medium. It reflects preference, and preference can be gamed with formatting and a flattering tone. Llama 4's leaderboard gaming in 2025 and the public criticism in early 2026 both show it should be only one reference among several. Look up when needed: after a new model ships, watch how its position changes over the next week or two, and check the category boards (code, long text). Don't look at the rank on the overall board.

Related
LLM: how the model works (measures) · Evals: knowing whether it got better (discusses)
Sources
Leaderboard page (lmarena.ai now redirects to arena.ai with a 301) (checked 2026-10-08) · Hacker News, 118 points (2023-05-25): the leaderboard's first launch (checked 2026-10-08) · Hacker News, 246 points (2026-01-07): a critique of it, "LMArena is a cancer on AI" (checked 2026-10-08)

Meta AI 博客

2013LayersL2L5

The official outlet for the Llama series and FAIR research (world models, JEPA, multimodal). It used to be the flag-bearer of open-weight models. Since Llama 4, Chinese models have taken a good share of that spot, so I read it more for the research line than for the leaderboards. Credibility: primary source, with a stake in the outcome. The model cards for open-weight models can be trusted. The comparison charts in launch posts need a third-party leaderboard next to them. Look up when needed: for a new Llama version, read the model card and the license terms. For progress on the JEPA line, find the original papers here.

Related
LLM: how the model works (discusses) · Fine-tuning: teach it your task (discusses) · Meta AI (FAIR) (released)
Sources
Blog index (returns 400 to scripted fetches; opens fine in a browser) (checked 2026-10-08) · Hacker News, 1235 points (2025-04-05): Llama 4 release (checked 2026-10-08)

OpenAI 博客

2015LayersL2L5

The official outlet for model releases, system cards and research posts. Credibility: primary source, but marketing-facing. Read the curves in a launch post alongside third-party leaderboards (LMArena leaderboard (arena.ai), Artificial Analysis). The system card is more trustworthy than the launch post. Look up when needed: on launch day, check the system card and the pricing page. Research posts (reasoning, RL) are still the original source for that area.

Related
LLM: how the model works (discusses) · Post-training: shaping with rewards (discusses) · Calling a model API: token, streaming, tools, cost (discusses) · OpenAI (released)
Sources
News index (script fetch returns 403; opens in a browser) (checked 2026-10-08) · Hacker News, 2279 points (2026-09-03): GPT-6 Astra launch post (checked 2026-10-08)

r/LocalLLaMA

2023LayersL3L4L5

A Reddit community for running local models. Within hours of a new open-weights release, there are quantized versions, VRAM numbers, and measurements of "how many token/s on my machine." It is also the first place where the exaggerations in a release post get called out. Credibility: low to medium. A single post is not trustworthy. The consensus across dozens of posts is. Look up when needed: when you want to run a model on your own hardware, or you doubt the numbers in a release, search the subreddit's posts from the past week.

Related
Serving: running models yourself (discusses) · Fine-tuning: teach it your task (discusses) · LLM: how the model works (discusses)
Sources
Subreddit front page (checked 2026-10-08) · Hacker News, 248 points (2025-08-11): the subreddit post "GPT-OSS-120B running on 8GB VRAM" was picked up on HN (checked 2026-10-08)

SemiAnalysis

2020LayersL3L5

Dylan Patel's semiconductor and AI infrastructure analysis: chips, interconnects, data center costs, and the economics of training and inference. Most of the in-depth content is paid. Credibility: his supply chain information is first-hand, but his model economics are estimates, and in late 2025 there were public questions about his conflicts of interest. Treat the numbers as references with error bars. Look up when needed: when you want to know what a particular chip or a particular lab's compute really costs, read the free summary first, then decide whether to pay for that piece.

Related
Serving: running models yourself (discusses) · Kubernetes: running it at the customer's site (discusses) · Dylan Patel (released)
Linked from
Dylan Patel (released)
Sources
About page: Over 320,000 subscribers (checked 2026-10-08) · Hacker News, 584 points (2026-08-25): an article on OpenAI's in-house chip (checked 2026-10-08)

Simon Willison's Weblog

2002LayersL4L5

Simon Willison (co-creator of Django, author of Datasette) writes a personal blog and updates it almost every day. On the day each new model ships, he runs it himself and posts the cost and the real results. His long-running coverage of prompt injection, tool calling, and agent security is a primary-source record for L4. His stance is a hands-on skeptic: try first, judge after. When to read: once a week, skim the titles of that week's posts. On a model release day, go straight to his hands-on tests. To check for pitfalls in an API, search his site first.

Related
Agent: a model with tools in a loop (discusses) · Calling a model API: token, streaming, tools, cost (discusses) · Prompting: instructions you can measure (discusses) · Simon Willison (released)
Linked from
Simon Willison (released)
Sources
Blog home page, updated daily, includes TIL and link blog (checked 2026-10-08) · Hacker News, 1094 points (2026-05-27): a post on product-market fit (checked 2026-10-08)

The Batch

2019LayersL5

A weekly newsletter from DeepLearning.AI. Andrew Ng opens with a letter. Then come four or five news items, each with a "why it matters" line. It is aimed at beginners to intermediate readers. The tone is calm and it does not chase hype. Credibility: high but shallow. Look up when needed: when I have to explain this week's event to someone non-technical, its wording works best. I don't rely on it to keep up with the frontier.

Related
LLM: how the model works (discusses) · AI History: How to Read the Three Waves (discusses) · Andrew Ng (released)
Linked from
Andrew Ng (released)
Sources
Newsletter index (checked 2026-10-08) · HN, 55 points (2025-01-30): the Andrew Ng issue on DeepSeek (checked 2026-10-08)

Opinions and predictions

L5's "future" is not knowledge. It is dated opinions. For each one I record the opinion, the person, the date, the source and the date it comes due for review. When it comes due, I come back and write status and review. The subscription lists are in 每周必看:前沿信息流(≤ 5) and 有事再查:前沿参考清单(≤ 15). The explanation of the layer is in atlas.

How to pick opinions

  • It must be falsifiable: I only take claims with a time, a scale or a capability ("60% chance by end of 2028", "obsolete in 3–5 years", "a decade"). I do not take slogans ("AI will change everything"). The more important the person, the more I pick the most specific thing they said, not the loudest.
  • Mostly one per person: if the same person keeps saying the same thing, I record only the earliest and most specific time. If they change their mind, I start a new entry, and the old one is still reviewed when it comes due.
  • The source is the page where the original words are: a podcast page, the original blog post, a paper. I use second-hand reports only when the original cannot be fetched, and I add the timestamp in the original recording at review time. source.date is the date I checked the evidence, not the publication date.
  • claim is in my own words: I paraphrase it precisely enough that I can judge it right or wrong a year later. I do not copy the original sentence.
  • About ten entries at most are open: past that, I review the oldest first, then add new ones.

How to review

  • The rule for due: if the opinion carries its own time, use that time (for a year, use 12-31 of that year). If it gives a range, take the lower bound for the first review, and if there is no conclusion, push due to the upper bound. If it gives no time, use the publication date + 12 months.
  • In the week it comes due, I do it in the weekly-feed slot. I find the hardest evidence of the time (data from Epoch AI 数据中心, the leaderboard at Artificial Analysis, released model cards), write one sentence of review, and change status to right / wrong / partial. If I cannot judge it, I write "cannot judge, because..." and push due back. I do not leave it blank.
  • Write the conclusion back. If it changes my judgment of some node, I edit that node's body or its Not picked list. If it changes my view of how far to trust a person, I add one sentence to the person entity's body.
  • At the end of each year I look over the reviewed entries and count who was right. That is the weight I give the "person" when picking opinions the next year.

Relation to the weekly feed

Opinions are a by-product of the weekly feed. The rule of 每周必看:前沿信息流(≤ 5) is to read and not bookmark, and to write only one line here each week. When I hit a dated claim while listening to a podcast or reading a newsletter, I record one on the spot. If there is none, I write "nothing worth recording this week". After a year, this page is my own ledger of who said what and when, and how it turned out, not a collection of other people's opinions.

  1. Sequence modeling can be faster and better using only attention, with no recurrence and no convolution.

    Ashish Vaswani2017-06-12Source · checked 2026-10-08Review due: 2020-06-12right

    Review: Within three years the Transformer became the default backbone for NLP, vision and speech. This entry is a sample of the notation.

  2. AI that is better than almost all humans at almost everything will most likely appear in 2026–2027. The cost is millions of chips and at least tens of billions of dollars.

    Dario Amodei2025-01-29Source · checked 2026-10-08Review due: 2027-12-31open

  3. The capabilities of AGI (AI that matches humans on any task) will arrive one by one over the next 5–10 years. We are not there yet.

    Demis Hassabis2025-03-17Source · checked 2026-10-08Review due: 2030-03-17open

  4. Global spending on data center construction will reach one trillion dollars a year in 2028, earlier than the 2030 he gave before.

    Jensen Huang2025-03-18Source · checked 2026-10-09Review due: 2028-12-31open

  5. The current LLM paradigm has only 3–5 years left. If JEPA-style world models succeed within that window, they will give a better paradigm for reasoning and planning and make the LLM route obsolete.

    Yann LeCun2025-04-02Source · checked 2026-10-08Review due: 2028-04-02open

  6. Around March 2027 a superhuman programmer appears that can finish multi-year software tasks at 80% reliability, and it goes all the way to superintelligence within 2027 (the author's modal year; the median is later).

    Daniel Kokotajlo2025-04-03Source · checked 2026-10-08Review due: 2027-12-31open

  7. AI that learns on the job like a person and can do any white-collar job has a 50% chance of arriving before 2032. An earlier milestone: a 50% chance by 2028 that it can do a small business's taxes end to end.

    Dwarkesh Patel2025-06-02Source · checked 2026-10-08Review due: 2028-12-31open

  8. In 2026 we will likely see systems that can come up with novel insights on their own. In 2027 we may see robots that can do work in the real world.

    Sam Altman2025-06-10Source · checked 2026-10-08Review due: 2027-12-31open

  9. This is "the decade of agents", not "the year of agents". Gaps like continual learning are solvable but hard, and closing them will take about ten years.

    Andrej Karpathy2025-10-17Source · checked 2026-10-08Review due: 2035-10-17open

  10. The scaling era (2020–2025) is over and we are back to the age of research. Current methods will stall after a while, because data is limited. Systems that learn like humans and then surpass them are still 5–20 years away.

    Ilya Sutskever2025-11-25Source · checked 2026-10-08Review due: 2030-11-25open

  11. By the end of 2028 there is a 60%+ chance of fully unattended AI R&D. About 30% in 2027, and none in 2026. If it does not happen, the current paradigm has a fundamental flaw.

    Jack Clark2026-05-04Source · checked 2026-10-08Review due: 2028-12-31open

Self-check

  • Without the notes, I can say which slot each of the five weekly sources covers, and why the sixth was rejected.
  • For each of the past four weeks I left one line in opinions (an opinion, or "none"), and my bookmarks gained no new AI links.
  • Pick any `open` opinion: I can say what evidence will settle it when due, and which row of Look up when needed: frontier reference list (≤ 15) holds it.
  • Given a new newsletter, in two minutes I can file it into one of the two tiers by "what it answers / credibility", or mark it Not picked, and give the reason.

Translated from the author's Chinese notes by a model; the Chinese page is the original.