Layer L5
Frontier: a stream, not a library
The layer
L5 has three parts: history (stable), frontier (a stream), opinions (dated). This page covers the last two. History lives in timeline and AI history: how to read the three waves.
Why the frontier is a stream
The half-life of the frontier is measured in weeks. A few months later, most of it is replaced by the next thing. Treat it as a library and you only pile up links you never read again, and you get a false feeling of "I've already learned this." Treat it as a stream: subscribe to a few filtered sources, read at a fixed time, leave no trace after reading, and note only what changed your judgment. Knowledge goes into nodes/. The hype flows away with the stream.
How to use the two tiers
- Weekly reads: the frontier feed (≤ 5): weekly reads, at most five, covering five slots: papers, research and policy, post-training and open models, hands-on tests of the application layer, and people's opinions. To add a sixth, drop one first.
- Look up when needed: frontier reference list (≤ 15): look up when needed, at most fifteen, each source tied to one kind of question (choosing a vendor, checking orders of magnitude, entering a new area, verifying numbers from a launch). No subscribing. The question comes before the source.
- One media entity per source (
entities/media/): who runs it, its stance, when it is worth reading, with edges to the nodes it often discusses.
The weekly time rule
The details are in Weekly reads: the frontier feed (≤ 5): 75 minutes every Saturday morning, in a fixed order, with the full podcast saved for the commute. When time is up, stop. Do not make up what you missed. The output is a single line: one opinion written into opinions, or "nothing worth noting this week." For a paper I want to read, I write down only the title. If I still remember it next week, I read it.
Why opinions carry dates
Someone says "AGI in five to ten years" every year. Without a date, it can never be wrong. Recording opinions is not collecting quotes. It is a ledger I can come back to and check against the answers: who, which day, what they said, when it can be verified. So I only take falsifiable claims that carry a time, a scale or a capability. I restate each claim in my own words, precisely enough to be judged right or wrong, and the source carries the date I checked it. A few years on, this page will tell me how much weight to give whose judgment.
How reviews write back
When an opinion comes due, write its status and one review line in opinions.md. If the result changed my judgment of a node, edit that node. If it changed my view of a person, add one sentence to the person entity. If the weekly stream shows that a pick has gone stale, note one line, then revise the nodes together once a quarter. Between the stream and the library, only this one narrow channel remains.
Top picks
In reading order; each pick links to its node below.
- 1
Import AI free
Weekly issue, each item with a "why it matters" line, with both a research view and a policy view → Weekly reads: the frontier feed (≤ 5)
- 2
Interconnects free
Front-line analysis of post-training and open models; read it first in release weeks → Weekly reads: the frontier feed (≤ 5)
- 3
Dwarkesh Podcast free
Where key people's dated judgments are stated. The home of opinions → Weekly reads: the frontier feed (≤ 5)
The nodes
Weekly reads: the frontier feed (≤ 5)
Half-life: a stream
The frontier is a stream, not a library: subscribe, read for a fixed time, and don't save anything. This table is all of the weekly reads. Every other source is in Look up when needed: frontier reference list (≤ 15), so look up when needed. The five sources cover papers (HF), research and policy (Import AI), post-training and open models (Interconnects), hands-on testing at the application layer (Simon Willison), and people's views (Dwarkesh).
| Source | Form | Rhythm | Why this one | How to read |
|---|---|---|---|---|
| Hugging Face Daily Papers | Paper feed | Daily, read once a week | arXiv filtered by community votes. It took over from Papers with Code as the "what's hot lately" entry point | Read only titles and vote counts, 10 minutes. Note one line for each paper I want to read, don't open it |
| Import AI | newsletter | One issue a week | Each item says "why it matters". Two views, research and policy. The author is in a frontier lab | Read the bold conclusion of each section, 20 minutes. Open at most one original |
| Interconnects | newsletter | One or two issues a week | First-hand analysis of post-training and open models, independent judgment on release days | Read the whole thing, 20 minutes. In a release week, read it before the press release |
| Simon Willison's Weblog | Practitioner blog | Daily | Same-day tests of every new model, with costs and pitfalls. A first-hand L4 record | Scan this week's entry titles, 15 minutes. Open only the hands-on tests |
| Dwarkesh Podcast | podcast | Two or three episodes a month | Key people's dated judgments are spoken here | Listen to one full episode a month (commute time, not counted in the weekly total). For the rest, read only the transcript subheadings, 10 minutes |
Fixed weekly time
- One slot: Saturday morning, 75 minutes in total (the sum of the items above; the full podcast episode goes in the commute and doesn't use it).
- Order: HF titles → Import AI → Interconnects → Simon Willison → Dwarkesh subheadings. When time is up, stop. Don't make up what I didn't finish.
- Don't save anything: no saved links, no clippings. Each week I write just one line to opinions: either a dated opinion, or "nothing worth noting this week". For papers I want to read, I note the title. If I still remember it next week, I go read it.
- When to touch a node: only edit
nodes/when something I saw this week changes my judgment of a node (for example, a pick has gone stale). Do it in one batch each quarter, and leave it alone otherwise. - Missed an issue: skip it, don't go back. That's what a stream means: missing things is fine.
- Builds on
- LLM: how the model works
- Import AI free
article · 20 min / week — The week's papers and releases worth knowing. Each comes with why it matters and a policy view. The author works in a frontier lab and doesn't hype.
- Interconnects free
article · 20 min / week — First-hand analysis of post-training and open models. On big LLM release days, an independent take beats the press release.
- Dwarkesh Podcast free
video · Monthly full episode, ~2.5 h — Key people's dated judgments are mostly spoken here. It is the main source of opinions.
Not picked (6)
- Latent Space — Gets the application-layer mood right, but it is long and the daily AINews is too dense. Moved to look up when needed.
- Ahead of AI — Irregular long posts, not a weekly rhythm. Moved to look up when needed.
- The Batch — Aimed at beginners. I don't rely on it to follow the frontier.
- Lex Fridman Podcast — Wide guest list but shallow follow-up questions. For the few AI episodes, just listen to them directly.
- No Priors — Investor view: more about products and funding than technology.
- arXiv cs.CL / cs.LG 日更 — No filtering. Subscribing will drown you. Use it for lookups.
Look up when needed: frontier reference list (≤ 15)
Half-life: a stream
I don't subscribe to these sources or read them on a schedule. Each one answers one kind of question, and I go there when I have that question. The only things I read every week are the five in weekly reads: the frontier feed (≤ 5). Credibility has three levels: High = primary and independent; Medium = primary but with a stake, or independent but an estimate; Low = needs several sources that back each other up.
| Source | What to look up | Credibility |
|---|---|---|
| arXiv cs.CL / cs.LG listing pages | Find the original when you know the ID or title; browse the past week's papers in an area by keyword | High (primary, not peer reviewed) |
| alphaXiv | Which papers were discussed in the past week, and how the authors answer questions in the comments | Medium (paper is primary, comments are secondary) |
| Anthropic blog (news · engineering · research) | Engineering practice for agent / tool calling; model cards and eval tables for new models | Medium (primary, has a stake) |
| OpenAI blog | Original source for system cards, pricing, and work on reasoning and RL | Medium (primary, marketing-facing) |
| Google DeepMind blog | Gemini / Gemma technical reports, world models, RL and scientific applications | Medium (primary, has a stake) |
| Meta AI blog | Llama model cards and licenses, original papers on the JEPA line | Medium (primary, has a stake) |
| LMArena leaderboard (arena.ai) | A new model's position on the preference leaderboard and category boards, one or two weeks after release | Medium (reflects preference, can be gamed) |
| Artificial Analysis | Choosing a model / provider: capability index, throughput, latency and price on one page | High (independent; the index uses custom weights) |
| Epoch AI data hub | Historical data on training compute, data and hardware cost; check before writing an order-of-magnitude claim | High |
| SemiAnalysis | The economics of chips, interconnect and data centers; how much a lab spent on compute | Medium (supply chain is primary, economics are estimates, paywalled) |
| Latent Space | What people use now in an application-layer area (eval tools, agent frameworks); the start-of-year reading list | Medium (good read of the mood, judgments sometimes hot) |
| Ahead of AI | Layer-by-layer comparison of a batch of new model architectures; a month of papers walked through by theme | High |
| Lil'Log (Lilian Weng's blog) | Surveys before entering a new area (agent, diffusion, RL, hallucination) | High |
| The Batch | Wording for explaining this week's events to non-technical people | High but shallow |
| r/LocalLLaMA | Whether a model runs on your own hardware and how fast; community measurements when you doubt a release's numbers | Low to medium (consensus is credible, a single post is not) |
How to use it
- Question before source: write down what you need to answer first (choose a provider? a survey of an area? the origin of a number?), then find the source in the "What to look up" column. Don't browse the other way round.
- Write what you find back into the relevant node or opinions. Don't bookmark the source page itself.
- Read all four lab blogs together: if you read only one for a given event, its stake will carry you off. For numbers, go by Artificial Analysis and the Epoch AI data hub.
- Builds on
- LLM: how the model works
- Epoch AI free
docs — Check it before writing any order-of-magnitude claim: time series of training compute, data and hardware cost. Methods are public and the data is downloadable.
- Artificial Analysis free
docs — The only page where you can see capability, speed and price together when choosing a model or provider, independent of what labs report about themselves.
- Lil'Log free
article — Before entering a new area, check whether she has written a survey. One post can replace years of tracking the main papers in that area.
Not picked (4)
- Hacker News — Not a source but a tool for gathering evidence: this library's `evidence` scores come from it. Don't subscribe to it as a feed.
- Semantic Scholar — Good for checking citation counts, but this library's picks use community endorsement as evidence. Citation counts are left to paper entities.
- Lex Fridman Podcast · No Priors — See Not picked in weekly reads: the frontier feed (≤ 5).
- 各实验室的 Discord — Too noisy. One visit when a model is released is enough, so it is not listed.
Media: who to follow
Twenty outlets, each with who runs it, its stance and when it is worth the time; the nodes above say which to read weekly and which on demand.
Ahead of AI
2022LayersL2L5
Sebastian Raschka (author of 《Build a Large Language Model (From Scratch)》)'s newsletter. Long posts, no fixed schedule. He either lines up the architectures of a batch of new models and compares them layer by layer, or walks through a month of papers by topic. His taste is the L2 teaching school: many diagrams, few formulas, and he has reproduced the work himself. Credibility: high. Look up when needed: come here when you want one readable comparison of an architecture change (MoE, attention variants, looped layers).
- Related
- Transformer (discusses) · LLM: how the model works (discusses) · Fine-tuning: teach it your task (discusses) · Sebastian Raschka (released)
- Linked from
- Sebastian Raschka (released)
- Sources
- About page: Over 220,000 subscribers (checked 2026-10-08) · Hacker News, 521 points (2026-09-09): post on GPT-6 Astra and the looped Transformer (checked 2026-10-08)
alphaXiv
2024LayersL2L5
A "comment section on top of arXiv" built by Stanford students: one discussion page per paper, a trending list, and an auto-generated overview. The authors often answer questions in the comments. Since Papers with Code stopped being maintained, if you want to find "which papers got discussed this past week", this and Hugging Face Daily Papers are the two replacements. Credibility: the papers are primary sources. The comments and auto-generated overviews are secondary. Use them to locate questions, not as conclusions.
- Related
- LLM: how the model works (discusses) · Agent: a model with tools in a loop (discusses)
- Sources
- Homepage, trending list + paper comments (checked 2026-10-08) · Hacker News, 549 points (2024-09-08) (checked 2026-10-08)
Anthropic 博客(news · engineering · research)
2021LayersL4L5
Three indexes: news covers models and products, engineering covers practical work on agent, tool calling, and Claude Code, and research covers interpretability, alignment, and model cards. The engineering blog is primary material for L4 (the pick in Agent:循环里带工具的模型 comes from here). Credibility: primary, but with a stance. For model cards and research, read the data. For product posts, read what they leave out. Look up when needed: on a new model's release day, check the eval table in the model card; when building an agent, check engineering.
- Related
- Agent: a model with tools in a loop (discusses) · MCP: One Standard Socket for Every System (discusses) · Calling a model API: token, streaming, tools, cost (discusses) · Anthropic (released)
- Sources
- Engineering blog index (checked 2026-10-08) · Research index (interpretability, alignment, model cards) (checked 2026-10-08) · Hacker News,673 分(2025-11-24):工程博客「Advanced Tool Use」 (checked 2026-10-08)
Artificial Analysis
2023LayersL3L5
An independent evaluation site. It runs models on the same set of benchmarks to produce an "intelligence index", then adds each API provider's throughput, first-token latency and price. It is the only place where I can see all three on one page when choosing a model and a provider. Credibility: higher than lab self-reports. The index is its own weighted definition, so scores are not comparable across version changes. Look up when needed: check the price and speed pages when choosing a model or provider; on release day, compare against the numbers the lab reports itself.
- Related
- LLM: how the model works (measures) · Serving: running models yourself (discusses) · Calling a model API: token, streaming, tools, cost (discusses)
- Sources
- Home page: the three charts for intelligence index, speed and price (checked 2026-10-08) · Hacker News, 916 points (2026-06-17): the post on GLM-5.2 topping open-weights models (checked 2026-10-08)
arXiv cs.CL / cs.LG 列表页
1991LayersL2L3L5
The original source for preprints. Several hundred papers a day, no filtering. Don't subscribe to the daily feed (it will drown you). Look up when needed: if you know the paper ID or title, go straight to the abs page. If you want to see what appeared in one area over the past week, use the listing page plus a keyword search. Credibility: the paper itself is a primary source, but it has not been peer reviewed. Before trusting a conclusion, check whether its eval setup and baselines are fair.
- Related
- LLM: how the model works (discusses) · Transformer (discusses) · Serving: running models yourself (discusses)
- Sources
- Recent listing page for cs.CL (checked 2026-10-08) · Recent listing page for cs.LG (checked 2026-10-08)
Google DeepMind 博客
2010LayersL2L5
The official outlet for Gemini / Gemma releases, world models (Genie), scientific applications (the AlphaFold line), and reinforcement learning. Compared with other labs, it has more science and RL content. Historical milestones (AlphaGo, AlphaFold) also have their original write-ups here. Credibility is the same as other lab blogs: first-hand, with a point of view. The technical reports are more trustworthy than the blog posts. Look up when needed: on a Gemini launch day, read the technical report; for RL and scientific applications, start from here to find the original papers.
- Related
- LLM: how the model works (discusses) · Post-training: shaping with rewards (discusses) · AI History: How to Read the Three Waves (discusses) · Google DeepMind (released)
- Sources
- Blog index (checked 2026-10-08) · Hacker News, 1812 points (2026-04-02): Gemma 4 release (checked 2026-10-08)
Dwarkesh Podcast
2020LayersL5
A long-form interview podcast hosted by Dwarkesh Patel. Guests are lab founders, chief scientists and economists. Each episode runs two to three hours. The host does thorough prep and keeps pushing until he gets a real answer. It is the main source of the views in opinions: most dated judgments from key people come from here. When it is worth your time: each month, pick one episode tied to a layer I care about and listen to the full version. For the rest, read only the subheadings of the transcript. When I hear a claim with a date attached, I note down one opinion.
- Related
- AI History: How to Read the Three Waves (discusses) · LLM: how the model works (discusses) · Agent: a model with tools in a loop (discusses) · Dwarkesh Patel (released)
- Linked from
- Dwarkesh Patel (released)
- Sources
- About page: Over 111,000 subscribers (Substack) (checked 2026-10-08) · Hacker News, 1212 points (2025-10-17): the Karpathy episode (checked 2026-10-08) · Hacker News, 450 points (2025-11-25): the Ilya Sutskever episode (checked 2026-10-08)
Epoch AI 数据中心
2022LayersL3L5
A nonprofit research institute, Epoch AI, maintains these datasets and charts: training compute, data volume, hardware price-performance and compute growth rates for models over the years, plus its own benchmarks such as FrontierMath. When you need numbers for "how fast has scaling actually gone," this is the most commonly cited primary compilation. Credibility: high. The methods are public and the data can be downloaded. One caution: it has benchmark partnerships with labs (the FrontierMath funding dispute with OpenAI). Look up when needed: check here before you write any judgment that carries an order of magnitude.
- Related
- Serving: running models yourself (discusses) · Evals: knowing whether it got better (discusses) · AI History: How to Read the Three Waves (discusses) · Epoch AI (released)
- Linked from
- Epoch AI (released)
- Sources
- Trends page: time series for training compute, data, hardware and cost (checked 2026-10-08) · Benchmark dashboard (checked 2026-10-08) · Hacker News, 480 points (2026-03-24): an open FrontierMath problem was solved (checked 2026-10-08) · Hacker News, 445 points (2026-05-24): a post on the cost breakdown of chips (checked 2026-10-08)
Hugging Face Daily Papers
2023LayersL2L4L5
A daily paper board run by Hugging Face. The community submits the arXiv papers worth reading that day and upvotes them into a ranking. Each paper comes with a short summary written by the authors or by AK. The taste leans toward open-source models, training and inference methods, and multimodal work, not theory. When it is worth your time: scan the titles once a week (titles and upvote counts only, 5–10 minutes), and log the one or two papers you really want to read in the "one-line record" of weekly reads: the frontier feed (≤ 5). It replaces the now-dormant Papers with Code as my entry point for "what's hot lately".
- Related
- LLM: how the model works (discusses) · Fine-tuning: teach it your task (discusses) · Agent: a model with tools in a loop (discusses) · Hugging Face (released)
- Sources
- Daily papers page, ranked by community upvotes, with an email digest you can subscribe to (checked 2026-10-08)
Import AI
2016LayersL2L5
Jack Clark (co-founder of Anthropic, formerly policy lead at OpenAI) has run this weekly newsletter since 2016. Each issue picks 5–8 papers or releases and writes a paragraph on each: why it matters. At the end there is a policy and safety angle and a piece of short fiction. His stance is "tech optimism, but serious about risk." He rarely hypes. When it is worth reading: read one issue every week. Read the bold conclusion of each paragraph first, then decide which originals to open. 20 minutes is enough.
- Related
- LLM: how the model works (discusses) · Evals: knowing whether it got better (discusses) · AI History: How to Read the Three Waves (discusses) · Jack Clark (released)
- Linked from
- Jack Clark (released)
- Sources
- About page: Over 143,000 subscribers (checked 2026-10-08) · Issue 455 (2026-05-04): his probability estimate for automated AI R&D, included in [[opinions]] (checked 2026-10-08)
Interconnects
2023LayersL2L5
A newsletter by Nathan Lambert (post-training lead at Ai2). It focuses on open models, post-training (RLHF / DPO / RL reasoning) and lab news. It is one of the few public analyses written by someone who does post-training at the front line. The taste leans technical and open-source, and he gives independent judgment on closed labs' launches. When to read it: one issue a week; when a major model ships (especially open weights), read his quick take, which is more useful than the press release.
- Related
- Fine-tuning: teach it your task (discusses) · Post-training: shaping with rewards (discusses) · LLM: how the model works (discusses) · Nathan Lambert (released)
- Linked from
- Nathan Lambert (released)
- Sources
- About page: Over 85,000 subscribers (checked 2026-10-08) · Hacker News, 367 points (2026-06-23): post on GLM-5.2 (checked 2026-10-08)
Latent Space
2023LayersL4L5
A newsletter + podcast on "AI Engineer" run by swyx and Alessio, plus a daily AINews roundup. The taste is application layer: agent frameworks, evals, RAG, inference providers. Many guests build products. Credibility: medium. It reads the mood of the industry well, but its technical judgment sometimes runs hot. Look up when needed: when I want to know what people use now in some application-layer area (eval tools, agent frameworks), I search its back issues. The reading list from the start of the year is a good starting point for L4.
- Related
- Agent: a model with tools in a loop (discusses) · RAG: Look Up First, Then Answer (discusses) · Evals: knowing whether it got better (discusses) · swyx (Shawn Wang) (released)
- Linked from
- swyx (Shawn Wang) (released)
- Sources
- About page: Over 203,000 subscribers (checked 2026-10-08) · Hacker News, 490 points (2025-01-13): AI Engineer Reading List (checked 2026-10-08)
Lil'Log(Lilian Weng 的博客)
2017LayersL2L4L5
Lilian Weng's personal blog. She posts a few times a year. Each post is a survey of one area: agent, the Transformer family, diffusion models, RL, hallucination. Reading one post is like reading the main line of papers in that area from the past several years. Credibility: high. The author does research on the front line. Look up when needed: before you enter a new area, check whether she has written a survey on it. Don't wait for updates. They are slow.
- Related
- Agent: a model with tools in a loop (discusses) · Transformer (discusses) · Post-training: shaping with rewards (discusses) · Lilian Weng (released)
- Linked from
- Lilian Weng (released)
- Sources
- Blog home page (checked 2026-10-08) · Hacker News, 334 points (2026-08-04): post on self-improving harness engineering (checked 2026-10-08) · Hacker News, 285 points (2023-06-27): survey of LLM-powered autonomous agents (checked 2026-10-08)
LMArena 榜(arena.ai)
2023LayersL2L5
An anonymous, double-blind head-to-head leaderboard where people vote and the votes become Elo scores. It measures "which answer people like better," not the ceiling of a model's ability. Credibility: medium. It reflects preference, and preference can be gamed with formatting and a flattering tone. Llama 4's leaderboard gaming in 2025 and the public criticism in early 2026 both show it should be only one reference among several. Look up when needed: after a new model ships, watch how its position changes over the next week or two, and check the category boards (code, long text). Don't look at the rank on the overall board.
- Related
- LLM: how the model works (measures) · Evals: knowing whether it got better (discusses)
- Sources
- Leaderboard page (lmarena.ai now redirects to arena.ai with a 301) (checked 2026-10-08) · Hacker News, 118 points (2023-05-25): the leaderboard's first launch (checked 2026-10-08) · Hacker News, 246 points (2026-01-07): a critique of it, "LMArena is a cancer on AI" (checked 2026-10-08)
Meta AI 博客
2013LayersL2L5
The official outlet for the Llama series and FAIR research (world models, JEPA, multimodal). It used to be the flag-bearer of open-weight models. Since Llama 4, Chinese models have taken a good share of that spot, so I read it more for the research line than for the leaderboards. Credibility: primary source, with a stake in the outcome. The model cards for open-weight models can be trusted. The comparison charts in launch posts need a third-party leaderboard next to them. Look up when needed: for a new Llama version, read the model card and the license terms. For progress on the JEPA line, find the original papers here.
- Related
- LLM: how the model works (discusses) · Fine-tuning: teach it your task (discusses) · Meta AI (FAIR) (released)
- Sources
- Blog index (returns 400 to scripted fetches; opens fine in a browser) (checked 2026-10-08) · Hacker News, 1235 points (2025-04-05): Llama 4 release (checked 2026-10-08)
OpenAI 博客
2015LayersL2L5
The official outlet for model releases, system cards and research posts. Credibility: primary source, but marketing-facing. Read the curves in a launch post alongside third-party leaderboards (LMArena leaderboard (arena.ai), Artificial Analysis). The system card is more trustworthy than the launch post. Look up when needed: on launch day, check the system card and the pricing page. Research posts (reasoning, RL) are still the original source for that area.
- Related
- LLM: how the model works (discusses) · Post-training: shaping with rewards (discusses) · Calling a model API: token, streaming, tools, cost (discusses) · OpenAI (released)
- Sources
- News index (script fetch returns 403; opens in a browser) (checked 2026-10-08) · Hacker News, 2279 points (2026-09-03): GPT-6 Astra launch post (checked 2026-10-08)
r/LocalLLaMA
2023LayersL3L4L5
A Reddit community for running local models. Within hours of a new open-weights release, there are quantized versions, VRAM numbers, and measurements of "how many token/s on my machine." It is also the first place where the exaggerations in a release post get called out. Credibility: low to medium. A single post is not trustworthy. The consensus across dozens of posts is. Look up when needed: when you want to run a model on your own hardware, or you doubt the numbers in a release, search the subreddit's posts from the past week.
- Related
- Serving: running models yourself (discusses) · Fine-tuning: teach it your task (discusses) · LLM: how the model works (discusses)
- Sources
- Subreddit front page (checked 2026-10-08) · Hacker News, 248 points (2025-08-11): the subreddit post "GPT-OSS-120B running on 8GB VRAM" was picked up on HN (checked 2026-10-08)
SemiAnalysis
2020LayersL3L5
Dylan Patel's semiconductor and AI infrastructure analysis: chips, interconnects, data center costs, and the economics of training and inference. Most of the in-depth content is paid. Credibility: his supply chain information is first-hand, but his model economics are estimates, and in late 2025 there were public questions about his conflicts of interest. Treat the numbers as references with error bars. Look up when needed: when you want to know what a particular chip or a particular lab's compute really costs, read the free summary first, then decide whether to pay for that piece.
- Related
- Serving: running models yourself (discusses) · Kubernetes: running it at the customer's site (discusses) · Dylan Patel (released)
- Linked from
- Dylan Patel (released)
- Sources
- About page: Over 320,000 subscribers (checked 2026-10-08) · Hacker News, 584 points (2026-08-25): an article on OpenAI's in-house chip (checked 2026-10-08)
Simon Willison's Weblog
2002LayersL4L5
Simon Willison (co-creator of Django, author of Datasette) writes a personal blog and updates it almost every day. On the day each new model ships, he runs it himself and posts the cost and the real results. His long-running coverage of prompt injection, tool calling, and agent security is a primary-source record for L4. His stance is a hands-on skeptic: try first, judge after. When to read: once a week, skim the titles of that week's posts. On a model release day, go straight to his hands-on tests. To check for pitfalls in an API, search his site first.
- Related
- Agent: a model with tools in a loop (discusses) · Calling a model API: token, streaming, tools, cost (discusses) · Prompting: instructions you can measure (discusses) · Simon Willison (released)
- Linked from
- Simon Willison (released)
- Sources
- Blog home page, updated daily, includes TIL and link blog (checked 2026-10-08) · Hacker News, 1094 points (2026-05-27): a post on product-market fit (checked 2026-10-08)
The Batch
2019LayersL5
A weekly newsletter from DeepLearning.AI. Andrew Ng opens with a letter. Then come four or five news items, each with a "why it matters" line. It is aimed at beginners to intermediate readers. The tone is calm and it does not chase hype. Credibility: high but shallow. Look up when needed: when I have to explain this week's event to someone non-technical, its wording works best. I don't rely on it to keep up with the frontier.
- Related
- LLM: how the model works (discusses) · AI History: How to Read the Three Waves (discusses) · Andrew Ng (released)
- Linked from
- Andrew Ng (released)
- Sources
- Newsletter index (checked 2026-10-08) · HN, 55 points (2025-01-30): the Andrew Ng issue on DeepSeek (checked 2026-10-08)
Opinions and predictions
L5's "future" is not knowledge. It is dated opinions. For each one I record the opinion, the person, the date, the source and the date it comes due for review. When it comes due, I come back and write status and review. The subscription lists are in 每周必看:前沿信息流(≤ 5) and 有事再查:前沿参考清单(≤ 15). The explanation of the layer is in atlas.
How to pick opinions
- It must be falsifiable: I only take claims with a time, a scale or a capability ("60% chance by end of 2028", "obsolete in 3–5 years", "a decade"). I do not take slogans ("AI will change everything"). The more important the person, the more I pick the most specific thing they said, not the loudest.
- Mostly one per person: if the same person keeps saying the same thing, I record only the earliest and most specific time. If they change their mind, I start a new entry, and the old one is still reviewed when it comes due.
- The source is the page where the original words are: a podcast page, the original blog post, a paper. I use second-hand reports only when the original cannot be fetched, and I add the timestamp in the original recording at review time.
source.dateis the date I checked the evidence, not the publication date. claimis in my own words: I paraphrase it precisely enough that I can judge it right or wrong a year later. I do not copy the original sentence.- About ten entries at most are
open: past that, I review the oldest first, then add new ones.
How to review
- The rule for
due: if the opinion carries its own time, use that time (for a year, use 12-31 of that year). If it gives a range, take the lower bound for the first review, and if there is no conclusion, pushdueto the upper bound. If it gives no time, use the publication date + 12 months. - In the week it comes due, I do it in the weekly-feed slot. I find the hardest evidence of the time (data from Epoch AI 数据中心, the leaderboard at Artificial Analysis, released model cards), write one sentence of
review, and changestatustoright/wrong/partial. If I cannot judge it, I write "cannot judge, because..." and pushdueback. I do not leave it blank. - Write the conclusion back. If it changes my judgment of some node, I edit that node's body or its Not picked list. If it changes my view of how far to trust a person, I add one sentence to the person entity's body.
- At the end of each year I look over the reviewed entries and count who was right. That is the weight I give the "person" when picking opinions the next year.
Relation to the weekly feed
Opinions are a by-product of the weekly feed. The rule of 每周必看:前沿信息流(≤ 5) is to read and not bookmark, and to write only one line here each week. When I hit a dated claim while listening to a podcast or reading a newsletter, I record one on the spot. If there is none, I write "nothing worth recording this week". After a year, this page is my own ledger of who said what and when, and how it turned out, not a collection of other people's opinions.
Sequence modeling can be faster and better using only attention, with no recurrence and no convolution.
Ashish Vaswani2017-06-12Source · checked 2026-10-08Review due: 2020-06-12right
Review: Within three years the Transformer became the default backbone for NLP, vision and speech. This entry is a sample of the notation.
AI that is better than almost all humans at almost everything will most likely appear in 2026–2027. The cost is millions of chips and at least tens of billions of dollars.
Dario Amodei2025-01-29Source · checked 2026-10-08Review due: 2027-12-31open
The capabilities of AGI (AI that matches humans on any task) will arrive one by one over the next 5–10 years. We are not there yet.
Demis Hassabis2025-03-17Source · checked 2026-10-08Review due: 2030-03-17open
Global spending on data center construction will reach one trillion dollars a year in 2028, earlier than the 2030 he gave before.
Jensen Huang2025-03-18Source · checked 2026-10-09Review due: 2028-12-31open
The current LLM paradigm has only 3–5 years left. If JEPA-style world models succeed within that window, they will give a better paradigm for reasoning and planning and make the LLM route obsolete.
Yann LeCun2025-04-02Source · checked 2026-10-08Review due: 2028-04-02open
Around March 2027 a superhuman programmer appears that can finish multi-year software tasks at 80% reliability, and it goes all the way to superintelligence within 2027 (the author's modal year; the median is later).
Daniel Kokotajlo2025-04-03Source · checked 2026-10-08Review due: 2027-12-31open
AI that learns on the job like a person and can do any white-collar job has a 50% chance of arriving before 2032. An earlier milestone: a 50% chance by 2028 that it can do a small business's taxes end to end.
Dwarkesh Patel2025-06-02Source · checked 2026-10-08Review due: 2028-12-31open
In 2026 we will likely see systems that can come up with novel insights on their own. In 2027 we may see robots that can do work in the real world.
Sam Altman2025-06-10Source · checked 2026-10-08Review due: 2027-12-31open
This is "the decade of agents", not "the year of agents". Gaps like continual learning are solvable but hard, and closing them will take about ten years.
Andrej Karpathy2025-10-17Source · checked 2026-10-08Review due: 2035-10-17open
The scaling era (2020–2025) is over and we are back to the age of research. Current methods will stall after a while, because data is limited. Systems that learn like humans and then surpass them are still 5–20 years away.
Ilya Sutskever2025-11-25Source · checked 2026-10-08Review due: 2030-11-25open
By the end of 2028 there is a 60%+ chance of fully unattended AI R&D. About 30% in 2027, and none in 2026. If it does not happen, the current paradigm has a fundamental flaw.
Jack Clark2026-05-04Source · checked 2026-10-08Review due: 2028-12-31open
Self-check
- Without the notes, I can say which slot each of the five weekly sources covers, and why the sixth was rejected.
- For each of the past four weeks I left one line in opinions (an opinion, or "none"), and my bookmarks gained no new AI links.
- Pick any `open` opinion: I can say what evidence will settle it when due, and which row of Look up when needed: frontier reference list (≤ 15) holds it.
- Given a new newsletter, in two minutes I can file it into one of the two tiers by "what it answers / credibility", or mark it Not picked, and give the reason.
Translated from the author's Chinese notes by a model; the Chinese page is the original.