Layer L4
Application engineering: from calling an API once to a production system with evals
The layer
This layer solves one problem: turning a model into a feature that can ship, be measured, and be changed. The topics are in one line, and each one is the wall the previous one hits in production.
How the line runs
- Calling model APIs: token, streaming, tools, cost: Everything starts with a message list going in and a token stream coming out. What you need to know: context window, billing, streaming, the request and response shapes of tool calls, the milliseconds and dollars of one call.
- Prompting: instructions you can measure / Structured output: making the model's answer consumable by programs: A prompt is an interface spec, not a spell. When you change one thing, you must be able to say how much the metric moved. The output must be consumable by programs, so schema constraints and a fallback chain follow right after.
- Embedding: turning meaning into vectors: The model doesn't know your data. Turn meaning into vectors and "search by meaning" becomes computing distance. Understand when to mix it with keyword search and how to measure recall.
- RAG: look up first, then answer: Look up first, then answer. The hard part is not hooking up a vector store. It is chunking, recall, rerank, the context budget, and a set of questions with gold answers.
- Agent: a model with tools in a loop: Model + tools + loop. The engineering is all at the boundaries: step limits, tool errors, context trimming, and when to use a fixed flow instead.
- MCP: one standard socket for every system: One standard socket for every system. The core of writing a server is describing operations so the model picks the right one. Whether you did it well is decided by the agent's tool-selection evals.
- Evals: knowing whether it got better: Without evals, changing a prompt or swapping a model is guessing. The eval set is built up one example at a time from production errors, and the judge must be checked against humans.
- Observability: knowing what happens in production, and getting back to evals: One trace per production call, plus quality signals, then sample back into error analysis. This loop keeps the eval set from stopping on launch day.
- Cost: every call has a price tag: Caching, batching, routing, fallback, limits. Measure the cost per successful result, not the average.
- Safety: the model acts on untrusted text: The model acts on untrusted text. Split apart the three elements: private data, untrusted content, and outbound sending. Cut capability, don't rely on detection alone.
TypeScript: the language where front ends and SDKs live is a side branch: the ecosystem's SDKs, frameworks, and MCP reference implementations are all in it. The goal is to read and modify it, not to learn the language.
Where they sit in production: model-apis, structured-output, cost, and safety stick to every call. embeddings and rag come before the call. agents and mcp string calls into a loop. evals come before release and observability after release, and the two join end to end.
How to learn a fast-changing layer
This layer has a half-life measured in months, so a tutorial is out of date once written. I collect only three kinds of things: official docs (what the vendor guarantees), open-source projects and eval tools (how others built it into a system), and engineering articles from practitioners (the traps they hit). The method is "one page of official docs + one project I built myself". The project field of each node is that project, and a node counts as done only after I finish it. Each quarter I re-check the picks, and tool picks past the freshness limit get replaced.
Top picks
In reading order; each pick links to its node below.
- 1
Claude Cookbooks — Anthropic free
Runnable notebooks: one each for tool use, JSON, caching, batching → Calling a model API: token, streaming, tools, cost
- 2
Prompting best practices — Anthropic free
The model vendor's prompt reference, rewritten with each release. Read the part that applies to all models first → Prompting: instructions you can measure
- 3
Structured outputs — Anthropic free
The official limits of schema constraints and strict tool parameters. They decide which layer you need to add your own fallback in → Structured output: making the model's answers consumable by programs
- 4
What are embeddings? — Vicki Boykis free
A short book from one-hot to Transformer embedding, with engineering → Embedding: turning meaning into vectors
- 5
Introducing Contextual Retrieval — Anthropic free
Hybrid search, rerank, and chunk context each come with measured numbers → RAG: Look Up First, Then Answer
- 6
Building effective agents — Anthropic free
The difference between workflows and agents, and when you need neither → Agent: a model with tools in a loop
- 7
Model Context Protocol — introduction and spec free
Read the protocol itself, not the framework wrappers around it → MCP: One Standard Socket for Every System
- 8
AI Evals: Everything You Need to Know — Hamel Husain and Shreya Shankar free
Start with error analysis on real traces, then build a judge that has been checked against humans → Evals: knowing whether it got better
- 9
Prompt caching — Anthropic free
The biggest built-in way to save money. The authoritative numbers are only on this page → Cost: every call has a price tag
- 10
My Lethal Trifecta talk at the Bay Area AI Security Meetup — Simon Willison free
One rule to remember how data gets stolen. When you design permission boundaries, remove one element first → Security: the model acts on text it should not trust
The nodes
Calling a model API: token, streaming, tools, cost
Half-life: fast-moving
The model never sees letters or words. The tokenizer cuts your text into tokens (a common word is one, a rare word is split into several pieces), and gives each one an ID. This is why models misspell words: they never saw the letters. Everything is counted in tokens. The context window is how much the model can read at once (your prompt and its answer added together). The price and the latency of a call both grow with it. Calling a model from code is mostly about managing this: reading the stream as each token arrives, letting the model call your tools, and asking for a JSON you can validate. These decide every engineering trade-off later in RAG, agent, and evals.
- Builds on
- LLM: how the model works
- Leads to
- Agent: a model with tools in a loop · Cost: every call has a price tag · Observability: know what is happening in production, and feed it back into evals · Prompting: instructions you can measure · Structured output: making the model's answers consumable by programs · TypeScript: the language of front ends and SDKs
- On the routes
- Foundations (Call a model from code)
repo — Runnable notebooks for the calls you will actually make — tool use, JSON output, prompt caching, batches — each a working example rather than a description.
video · 15 min · only 0:00–14:56 — What a token is, in a live tokenizer: why the same text costs different amounts on different models, and why models stumble on spelling and arithmetic.
Hands-on project: A client that knows what it costs — Call one model API with streaming, one tool and a JSON schema; log tokens and latency per call, and work out the cost of a thousand requests.
Not picked (2)
- OpenAI API reference — A similar doc from another provider. My main stack is Anthropic and OpenRouter, and the Cookbooks already cover the same call patterns.
- Anthropic SDK 的 README — Install and call examples. The Cookbooks are a superset of it.
Prompting: instructions you can measure
Half-life: fast-movingWho asks for it
For a single call, the prompt is the model's whole world: your instructions, a few examples, the documents it should use, and the question. Small changes in wording shift the result in ways you don't expect. So a good prompt is revised against a set of test cases you keep, never by the feel of one answer. The sign that you can write prompts is not fancy writing. It is being able to say how far the metric moved after changing one thing. It is the entry point to evals, and the cheapest lever in RAG and agent work.
- Builds on
- Calling a model API: token, streaming, tools, cost
- Leads to
- Agent: a model with tools in a loop · Evals: knowing whether it got better · RAG: Look Up First, Then Answer · Structured output: making the model's answers consumable by programs
- On the routes
- Foundations (Call a model from code)
docs — The model maker's reference, rewritten with each release so it fits the models you will call: instructions, examples, structure, output format, tools. Start with the techniques for all models.
Hands-on project: Improve one prompt with a test set, not by feel — Fifty cases with expected outputs, a baseline score, then each prompt change measured as a diff — keep the ones that move the number.
Not picked (3)
- Prompt engineering guide(OpenAI) — Same kind of official page. Keep one.
- DSPy — A framework that treats the prompt as a parameter to optimize. Build an eval set first, then come back. It goes after Evals: knowing whether it got better.
- 各种 prompt 技巧合集站 — Tutorial type. The fast-changing layer is not included.
Structured output: making the model's answers consumable by programs
Half-life: fast-moving
The model emits text. The program wants fields. There are four levels of constraint, from strong to weak: the vendor constrains decoding by schema (grammar), strict mode for tool parameters, a format given in the prompt plus local validation, and parsing plus retry only. All the engineering work is in the fallback chain: what the next level is when one fails, how failures are counted, which fields may be missing. Two more things are often ignored: the field names and descriptions in the schema are documentation the model reads, so how you write them directly affects accuracy; and "valid JSON" does not mean "correct content", which belongs to Evals: knowing whether it got better.
- Builds on
- Calling a model API: token, streaming, tools, cost · Prompting: instructions you can measure
docs — Official docs on JSON schema constraints and strict tool parameters: supported features, grammar caching, tool calls. Know what the vendor guarantees, then decide where your fallback goes.
- Instructor free
repo — Open-source thin layer that unifies schema validation, retry on failure and multiple vendors. Its validate-and-retry loop is the reference answer for writing your own fallback chain.
Hands-on project: An extractor with validation and a fallback chain — Extract fields from a batch of real inputs under a JSON schema constraint. When a schema feature is unsupported, fall back to prompt constraints plus local validation, then to retries. Record each level's hit rate, failure rate and cost per request.
Not picked (3)
- Structured Outputs guide(OpenAI) — Another official page on the same thing; my main stack is Anthropic and OpenRouter, so one vendor is enough.
- Pydantic AI — A whole agent framework. This node only needs the validation and fallback layer, and Instructor is thinner.
- JSON Schema 规范 — Theory, not fast-changing content for this layer; look up when needed.
Embedding: turning meaning into vectors
Half-life: fast-moving
Inside the model, every token becomes a vector: a long list of numbers, arranged so that things with similar meaning land near each other. Do the same for a whole piece of text and you can search by meaning. A question like "how do I get my money back" finds the refund policy, even though the two share no words. Keyword search is still stronger on exact names, codes and error messages, so good systems use both (hybrid). What you need to understand: what each is good at, when to mix them, how much data it takes before a vector index is worth it, and how to measure recall. It is the foundation of RAG.
- Builds on
- LLM: how the model works
- Leads to
- RAG: Look Up First, Then Answer
- On the routes
- Foundations (Search by meaning)
book — A short free book that goes from one-hot vectors to transformer embeddings with the engineering around them — the one explanation practitioners keep linking.
Hands-on project: Semantic search over your own documents — Embed a few thousand of your own documents into a vector index, then compare its top ten against keyword search on twenty queries you label.
Not picked (2)
- MTEB leaderboard — A stream you look up on the spot when choosing a model. Not bookmarked.
- Sentence — Transformers docs — tooling docs for self-hosting embedding models. Users go through a hosted API, so read it when you need it.
RAG: Look Up First, Then Answer
Half-life: fast-movingWho asks for it
Retrieval-augmented generation: before the model answers, your code searches your own documents and pastes the most relevant passages into the prompt. The model answers from them instead of from memory, and it can say where it found them. When a RAG system answers wrong, most of the time the right passage was never retrieved. So the first thing to measure is retrieval itself: what share of the time the passage that answers the question shows up in what you retrieved. The hard part isn't wiring up a vector store. It is chunking, recall, rerank, the context budget, and proving, with a set of questions that have reference answers, that it beats not retrieving.
- Builds on
- Embedding: turning meaning into vectors · Prompting: instructions you can measure
- On the routes
- Forward Deployed Engineer (Build on the customer's data and tools) · AI Engineer (Build LLM features)
article — Hybrid search (embeddings plus BM25), reranking and chunk context, each with a measured drop in retrieval failures — design with numbers.
Hands-on project: Hybrid retrieval with a labelled query set — BM25 plus vectors fused with reciprocal rank fusion, a reranker, and 50 hand-labelled queries; report recall@k and nDCG against keyword search alone.
Not picked (3)
- RAGAS — A metrics library for answer quality. I'll add it once I have the generation half.
- LlamaIndex / LangChain 的 RAG 教程 — A framework tutorial. It's in the fast-changing layer, so I don't pick it.
- 各家 reranker 文档(Cohere 等) — Tool docs. I'll look them up when I get to the rerank step.
Agent: a model with tools in a loop
Half-life: fast-movingWho asks for it
Give the model a list of tools (search, run code, call an API), then put it in a loop: it proposes which tool to use, your code runs it, the result goes back, and the model decides the next step, until the task is done. A plain loop you write yourself beats a fancy framework. The hard part is what surrounds the loop: a step limit, what to do when a tool errors, how to trim the context, when to use a fixed workflow instead of an agent, and finding which step it actually went wrong at. Evals should check whether it picked the right tool, whether the arguments were right, and how many steps it took.
- Builds on
- Calling a model API: token, streaming, tools, cost · Prompting: instructions you can measure
- Leads to
- MCP: One Standard Socket for Every System · Security: the model acts on text it should not trust
- On the routes
- Forward Deployed Engineer (Build on the customer's data and tools) · AI Engineer (Build LLM features)
article — Workflows versus agents, and when you need neither: the handful of patterns these systems are built from, named by a lab that ships them.
course — Hands-on: tool calling, a ReAct-style loop and multi-agent setups in code you run — the gap between reading about agents and writing one.
article — Context is the agent loop's most expensive variable: when to trim, compress, take notes or hand off to a sub-agent, sorted by problem — the checklist for your own loop's context budget.
Hands-on project: Write the agent loop yourself, then evaluate it — Messages API tool use with a step cap, tool-error retries and context trimming; score it on 30 real tasks for tool choice, argument correctness and steps taken.
Not picked (2)
- OpenAI Agents SDK / LangGraph — Frameworks. The project asks you to write the loop yourself. Look at frameworks after you have done that.
- Writing effective tools for agents(Anthropic,2025 — 09) — The content fits #557, but its top HN score was 3 and it missed the endorsement bar. Also noted under Not picked in MCP: one standard plug for every system.
MCP: One Standard Socket for Every System
Half-life: fast-movingWho asks for it
Model Context Protocol is the standard way to offer a system (a database, a ticketing system, an internal API) as a tool to any agent: you write the server once, and every assistant and agent that speaks MCP can connect to it. The craft is in the tool's name and description. An agent uses a tool well only if it can tell from the description when to pick it. The tool name, the parameter schema and the error messages are all documentation written for the model to read. You don't get to say whether it is good. The agent's tool-selection evals say so.
- Builds on
- Agent: a model with tools in a loop
- Leads to
- Security: the model acts on text it should not trust
- On the routes
- Forward Deployed Engineer (Build on the customer's data and tools) · AI Engineer (Build LLM features)
docs — The protocol itself — tools, resources, transports — straight from the spec rather than from a framework's wrapper around it.
Hands-on project: An MCP server for an API you know, measured by an agent — Expose five real operations as tools, then run an agent on 30 tasks and count how often it picks the right tool with the right arguments; rewrite descriptions until that number moves.
Not picked (2)
- MCP TypeScript SDK — The official implementation has 13,533 stars, but it is a tool for writing servers, not a concept to learn. One entry for the spec is enough.
- Writing effective tools for agents(Anthropic,2025 — 09) — It covers using an agent to iterate on tool descriptions, which is exactly the method in #557. Its top HN score is 3, below the endorsement line, so a mention in the body is enough.
Evals: knowing whether it got better
Half-life: fast-movingWho asks for it
The same question can get a different answer every time, so one test tells you nothing. An eval is a set of real cases plus a way to score each answer: exact checks where you can write them, a rubric or another model as judge where you can't. Run it on every change. Start by reading real outputs and naming the ways they fail. The eval set grows out of that list, and it is what tells you whether a change really helped or just moved the failures somewhere else. It is the scarcest skill in this layer and the one that best separates candidates.
- Builds on
- Prompting: instructions you can measure
- Leads to
- Fine-tuning: teach it your task · Observability: know what is happening in production, and feed it back into evals
- On the routes
- Foundations (Measure it) · Forward Deployed Engineer (Prove it works) · AI Engineer (Measure and improve)
article — Start from error analysis on real traces, not from a metric; then judges you check against people — the practice behind 'evaluation frameworks' in these postings.
- promptfoo free
repo — Runs an eval set as CI — cases in YAML, assertions by rule or by judge, red when the score drops — turning one-off eval studies into a regression suite. Open source; part of OpenAI since March 2026.
- Inspect free
repo — An evaluation institute's open-source framework: a clean task / dataset / solver / scorer split, with agent tasks and tool use first-class — the shape to grow a reusable harness in.
Hands-on project: An eval harness that gates a change — Label 100 real outputs by hand, build an LLM judge and measure its agreement with you, and make CI fail when a prompt or model change drops the score.
Not picked (4)
- Hamel Husain 与 Shreya Shankar 的 evals 课程 — Paid course. It is a tutorial, and the fast-moving layer takes only docs, projects, tools and articles. One FAQ entry already distils the course.
- Braintrust — Hosted platform. This node wants an open-source tool that can run in CI.
- Who Validates the Validators(Shankar 等,2024) — Original paper on aligning the judge with humans. It is theory, so it belongs in L5.
- OpenAI Evals — The repo leans toward benchmark work, not the shape of application evals.
Observability: know what is happening in production, and feed it back into evals
Half-life: fast-moving
Observability for an LLM feature is three things. First, one trace per call (prompt version, model, token count, latency, cost, tool-call sequence). Second, live quality signals (whether the user edited the result, retries, fallbacks triggered, complaints). Third, the loop: sample from the traces, label by hand, do error analysis, and turn the failure modes into new cases for Evals: knowing whether it got better. Logs that only record status codes are not enough. You need to be able to rebuild what the model saw and what it answered, without storing more than your privacy rules allow. Without this loop, the eval set stays frozen as it was on launch day.
- Langfuse free
repo — Open-source trace/session/scoring platform, self-hostable on k3s. Value is not the dashboard but production traces you can sample, score and feed back into evals; error analysis starts here.
Hands-on project: Add traces to a live LLM feature, then go from traces back to an eval set — Record one trace for every call of an LLM feature in production (input summary, prompt version, model, token count, latency, cost, whether the user edited the translation). A month later, sample 100 traces, do error analysis, and write the failure modes up as new eval cases.
Not picked (3)
- OpenTelemetry GenAI semantic conventions — Right direction (vendor-neutral span attributes), but the repo has 664 stars, short of the endorsement bar. When you wire up Langfuse, check its field names against it.
- LangSmith — Closed-source and hosted; self-hosting is enterprise-only. This layer prefers open source you can self-host.
- Phoenix(Arize) — Same kind of tool as Langfuse; pick one of the two.
Cost: every call has a price tag
Half-life: fast-moving
Cost = unit price × token × number of calls. The levers, ordered from cheapest to most expensive: shorten the prompt and the output; prompt caching (a shared prefix is paid at full price only once); the Batch API (async in exchange for half price); model routing (cheap model first, escalate when unsure); fallback when a vendor is down; and last, hard caps: per-user quotas, daily budgets, quote before executing. When you measure, look at p50 / p95 per request, not the average. Look even more at cost "per successful result": a cheap model that fails and retries three times is not cheap. Caching and routing change what the model sees, so after changing them, run Evals: Knowing Whether It Got Better again.
docs — Biggest vendor-native saving: which prefixes cache, per-model minimum length, write/hit price multipliers. The only authoritative numbers. Then compute savings for a shared long system prompt.
Hands-on project: Add prompt caching to a production call with a shared system prompt, and verify it against the bill — Put the shared system prompt and glossary into the cached prefix, run it for a week, and compare cached token share, cost per thousand calls and p50 latency. Move batch backfill jobs to the Batch API and compare again.
Not picked (3)
- Message Batches API docs(Anthropic) — Meant to include it, but its HN top score was 3, below the endorsement bar. Same docs site as prompt caching; read it right after that page. Already named in the body and project.
- OpenRouter docs — The main routing layer, but it is product docs, not a knowledge point. Routing and fallback are already my own implementation in `OpenRouterRouting.java`.
- 各家价格表 — A feed, not a bookmark. Look it up when doing the math.
Security: the model acts on text it should not trust
Half-life: fast-moving
Security here means application security for LLM apps, not alignment. Three questions: where does untrusted content come in (user input, retrieved documents, tool returns, web pages); what can the model do (least-privilege tools, human confirmation for irreversible actions, tiers by task); where can data go out (external links, markdown images, tool calls with parameters). Prompt injection has no complete fix, so you cut capability instead of relying only on detection: split the three elements apart, validate output before executing it, log every call and cap it. The author of an MCP server is also the author of an attack surface, because tool descriptions and return values are text the model will follow.
docs — Ten risk classes in one list: prompt injection, unsafe output handling, excessive agency, vector store weaknesses, unbounded consumption. Run it on your agent or MCP server; no class is missed.
article — One rule to remember: private data, untrusted content and external sending together let data be stolen. Ask which to remove when setting tool permissions; more reliable than stacking detectors.
Hands-on project: 给 ImageStep MCP 做一轮 prompt injection 红队 — Use promptfoo's red team to run injection cases against the 7 tools of an MCP server (instructions hidden in file names, URL pages, tool returns). Count how often the agent crosses a permission boundary or sends data out; fix, then run another round.
Not picked (3)
- 各家 prompt injection 防护指南(Anthropic / OpenAI 文档) — Mostly prompt-layer tactics; weaker than the two above, which argue from capability boundaries.
- Greshake 等《Not what you've signed up for》(2023) — The original paper on indirect injection. Theory; belongs to L5 frontier, not this layer.
- Lakera 等 guardrail 产品 — Detection products; they change fast. A stream, not something to keep.
TypeScript: the language of front ends and SDKs
Half-life: fast-movingWho asks for it
Customer integrations run into web front ends and SDKs, and most of those are written in TypeScript. Model vendors' SDKs, agent frameworks and MCP reference implementations are in it too. If you already write it, skip this. Its point at this layer is "read and change the ecosystem's code", not the language itself.
- Builds on
- Calling a model API: token, streaming, tools, cost
- On the routes
- Forward Deployed Engineer (Deploy into their environment)
docs — The language's own reference, short enough to read in an evening — enough to read and change the web and SDK code these roles ship to customers.
Hands-on project: Type an SDK for an API you use — Write a small typed client for a real API — request and response types, errors as values, one streaming endpoint — and publish it with its tests.
Not picked (2)
- Effective TypeScript — Paid book. Its depth goes beyond the "read it and change it" requirement.
- Total TypeScript — Course. We don't include tutorials in the fast-changing layer.
Self-check
- There is a live RAG or agent feature with a standing eval set (≥ 50 gold examples checked by hand), and CI turns red when the score drops.
- For any production call, I can state the prompt version, model, tokens, latency, and cost, and I have done one round of error analysis on traces sampled from production.
- When I change a prompt or swap a model, I can give the number for the change in the metric and the number for the change in cost.
- I have written an MCP server or a toolset, and measured the effect of description changes with an agent's tool-selection evals.
- I have drawn a diagram of my own system: where untrusted content comes in, what the model can do, and where data goes out. And I have removed one element.
Translated from the author's Chinese notes by a model; the Chinese page is the original.