ReferenceA–Z

Glossary

52 terms you meet across these 105 plates — what each one actually means, in one sentence and then a short paragraph, with the guides that show it in use.


A 2 terms

Agent

A model put in a loop with tools: it plans, calls a tool, reads the result, and repeats until the task is done or it gives up.

The loop is what makes it an agent — the same model answering one question is just a chat. Each turn it sees the conversation so far plus the output of the last tool call, so a wrong assumption early on compounds. Claude Code, Cursor and Codex CLI are agents with a shell, a file editor and a search tool wired in. What keeps one useful is the bounds around it: permissions, a plan you approved, a place to stop.

See also Tool use Subagent Permissions Plan mode MCP

API key

A secret string that identifies your account to a model provider and gets billed for every request made with it.

It is a bearer credential: anyone holding it can spend your quota, so treat it like a password. Keep it in an environment variable or a secrets manager — never in source control, a screenshot, or client-side JavaScript. Rotate it the moment it leaks. Local runtimes need no key at all, which is one reason people prototype against them.

See also Rate limit OpenAI-compatible API Model ID Ollama

↑ Back to A–Z

B 1 term

Batch API

An asynchronous endpoint that takes many requests at once and returns them within a day, at about half the normal price.

You upload a file of requests, poll for completion, and download the results — nothing streams back. It fits work with nobody waiting: classifying a backlog, writing ten thousand product descriptions, running an eval set. The trade is latency measured in hours and a quota separate from the live API.

See also Rate limit Inference Evals Token

↑ Back to A–Z

C 5 terms

Chain of thought

Asking the model to work through its reasoning before answering, which raises accuracy on multi-step problems.

For arithmetic, logic and anything with several dependent steps, the intermediate text acts as scratch space the model can attend to. Reasoning models now do this internally, so you rarely have to ask. Where you do, keep the working separate from the answer — reasoning first, a clearly marked final line last — so you can strip it before showing anyone.

See also Few-shot Prompt chaining System prompt Hallucination

Chunking

Splitting a long document into passages small enough to embed and retrieve one at a time, before a RAG system can use it.

Chunk too small and a passage loses the context that made it meaningful; too large and the retrieved text wastes context window and buries the answer. A few hundred words with a little overlap is the usual starting point, and splitting on real boundaries — headings, sections — beats splitting every N characters. Store the source and position with each chunk so an answer can point back at where it came from.

See also RAG Embeddings Vector database Context window

CLAUDE.md

A Markdown file in a project that Claude Code reads at the start of every session as standing instructions.

It is project memory: build and test commands, conventions, paths not to touch, where things live. It is prepended to the context of every session in that directory, so keep it short — a page of specifics beats three pages of generalities. A user-level file at ~/.claude/CLAUDE.md holds the rules that apply to every project.

See also System prompt Context window Slash command Agent

Context window

The maximum number of tokens one request can hold — system prompt, conversation, files, tool output and the reply, all together.

It is a hard ceiling, not a suggestion: go over it and the request fails or the oldest turns are dropped. A big window is not free — you pay for every token in it on each turn, and recall of detail buried in the middle of a very long context is worse than at either end. In an agent session the window fills with tool output far faster than with your own words, which is what /clear and /compact are for.

See also Token Prompt caching Chunking RAG CLAUDE.md

Custom GPT

A saved ChatGPT configuration — instructions, files and optional actions — that you can reuse and share; Projects are the lighter version.

The instructions are a system prompt with a form around them, and attached files are retrieved when relevant rather than pasted whole into every chat. Actions let it call an external API you describe with an OpenAPI schema. Projects, in both ChatGPT and Claude, do the instructions-and-files half without the sharing or the actions — usually enough for your own work.

See also System prompt RAG Tool use CLAUDE.md

↑ Back to A–Z

D 1 term

Diffusion model

An image model that starts from noise and removes it step by step, steered at each step by your prompt.

Because generation is iterative and starts from randomness, the same prompt with a different seed gives a different image — fixing the seed makes a result reproducible. Step count trades time for detail; the guidance scale trades prompt fidelity for natural-looking output. This is the family behind Midjourney, Stable Diffusion and Flux, and a different architecture from the transformers behind chat models.

See also Inpainting Multimodal LLM LoRA

↑ Back to A–Z

E 2 terms

Embeddings

Vectors of numbers that place text in a space where nearby points mean similar things.

Similarity is usually the cosine distance between two vectors, which is why embedding search finds a passage about canceling a plan for the query how do I unsubscribe, where keyword search finds nothing. Stored passages and the query must be embedded by the same model, so changing model means re-embedding everything. Embeddings capture topic, not truth: two flatly contradictory sentences on one subject sit close together.

See also Vector database RAG Chunking

Evals

A fixed set of test cases plus a way to score the answers, re-run every time you change a prompt, a model or a retrieval step.

Without them a prompt change is a vibe: you fix one case and quietly break four. Twenty realistic inputs with known-good outputs, scored by exact match, a rule, or a second model, already catch most regressions. Keep the set in version control next to the prompt it tests, and add every bug you find as a new case.

See also Hallucination Guardrails Prompt chaining Batch API

↑ Back to A–Z

F 2 terms

Few-shot

Putting two or three worked examples in the prompt so the model copies their format and level of detail.

It is the fastest way to pin down an output shape that is awkward to describe in words. Examples teach format far more reliably than adjectives do, so make them look exactly like what you want back, edge cases included. Past a handful the returns fall off and you are only paying for tokens; that is the point where fine-tuning starts to make sense.

See also System prompt Structured output Fine-tuning Chain of thought

Fine-tuning

Further training of an existing model on your own examples so it learns behavior or style you cannot get from prompting.

It changes the weights, so it teaches skill and tone — not facts. If you want the model to know your documents, retrieve them at query time instead: that stays current and costs a fraction as much. Budget for hundreds to thousands of clean examples plus a held-out set to measure against, and only start once a good prompt has stopped improving.

See also LoRA RAG Few-shot Evals

↑ Back to A–Z

G 3 terms

GGUF

The single-file model format used by llama.cpp and everything built on it, including Ollama and LM Studio.

One file holds the weights, the tokenizer and the metadata, which is why running a local model is a download rather than an install. The filename usually carries the quantization level — Q4_K_M, Q5_K_M, Q8_0 — and that suffix, not the format, decides size and quality. It replaced the older GGML format; safetensors plays the same role in the GPU and PyTorch world.

See also Quantization llama.cpp Ollama LM Studio Modelfile

Grounding

Giving the model the source text it must answer from, and instructing it to stay inside that text.

It is the practical antidote to invented facts: the model stops recalling and starts reading. Both halves are needed — the passages in the prompt, and an instruction to answer only from them and to say so when the answer is not there. Ask for quotes or citations and you get a cheap way to check whether it actually did.

See also RAG Hallucination System prompt Chunking

Guardrails

Checks around a model — on its input, its output, or the tools it can reach — that catch what prompting alone cannot guarantee.

A system prompt is a request; a guardrail is code that runs whether the model cooperates or not. The useful ones are dull: validate output against a schema and retry on failure, refuse tool calls outside an allowlist, cap spend per session, keep a human in front of anything destructive. Put them outside the model, because anything inside the prompt can be argued with.

See also Permissions Structured output Hooks Evals

↑ Back to A–Z

H 3 terms

Hallucination

A fluent, confident answer that is simply wrong — an invented citation, API, file path or number.

It is not a bug waiting to be patched out: the model predicts plausible text, and plausible is not the same as true. The rate drops sharply when you supply the source material, allow the model to say it does not know, and ask for quotes you can check. Anything load-bearing — a command, a version, a legal fact — still needs verifying against the real thing.

See also Grounding RAG Evals Temperature

Headless mode

Running a coding agent non-interactively — one prompt in, output out — so it can be used from scripts, hooks and CI.

There is no session to steer, so everything the agent needs has to be in the prompt, the project files and its configuration. It is how agents get wired into pre-commit checks, PR reviews and cron jobs. Be strict with permissions here: nobody is watching to approve a command.

See also Agent Hooks Permissions Slash command

Hooks

Shell commands Claude Code runs automatically at defined points — before a tool call, after an edit, when a session ends.

They are configured in settings.json and run as ordinary code, so they do what a prompt can only ask for: format every file after it is written, block edits to a protected path, log what was run. A hook that exits non-zero can stop the action it was attached to. They execute on your machine without asking, so read one as carefully as any other script you run.

See also Permissions Headless mode Slash command Guardrails

↑ Back to A–Z

I 2 terms

Inference

Running a trained model to get an output — the step you pay for per token, as opposed to training.

It has two phases: prefill, which reads your whole prompt at once, and decode, which produces the answer one token at a time. Prefill is compute-bound and fast; decode is limited by memory bandwidth and sets the speed you actually feel. That is why a long prompt costs money but little waiting, while a long answer takes real time.

See also Token Latency vs throughput Quantization Streaming

Inpainting

Regenerating a masked region of an existing image while everything outside the mask stays untouched.

You paint over the part to replace and give a prompt for what belongs there; the model fills it using the surrounding pixels as context. It is the tool for removing an object, fixing a hand, or swapping a background without redrawing the whole image. Feathering the mask edge and describing the existing lighting is most of what makes the seam disappear.

See also Diffusion model Multimodal

↑ Back to A–Z

L 5 terms

Latency vs throughput

Latency is how long one request takes; throughput is how many tokens per second the system delivers across all of them.

They pull in opposite directions: batching more requests together raises throughput and lengthens every individual wait. For a chat interface the number that matters is time to first token, which streaming then covers for the rest. For an overnight job it is total tokens per hour, and nobody cares about the first one.

See also Inference Streaming Batch API Quantization

llama.cpp

The C/C++ inference engine that made running LLMs on ordinary hardware practical, and the origin of the GGUF format.

It runs quantized models on the CPU and offloads as many layers as fit to a GPU, which is how a laptop serves a 7B model at readable speed. Most consumer local-model tools are wrappers around it — Ollama and LM Studio both are. On a GPU server the same role is played by vLLM, which gives up portability for far higher throughput.

See also GGUF Quantization Ollama LM Studio OpenAI-compatible API

LLM

A large language model: a network trained to predict the next token, which turns out to be enough to write, translate, summarize and code.

It has no database and no memory between calls — everything it knows is either baked into the weights at training time or sitting in the current context window. That one mechanism explains both the fluency and the failure mode: it produces what is likely, not what is checked. Prompts, RAG, tools and agents are all ways of steering that single prediction step.

See also Token Context window Inference Hallucination

LM Studio

A desktop app for running local models: a model browser, a chat window, and a one-click OpenAI-compatible server.

It is the no-terminal route into local LLMs — search a model, see whether it fits your RAM, download, chat. The built-in server exposes the same endpoints as the OpenAI API, so existing code points at localhost and works unchanged. Underneath it runs GGUF files through llama.cpp, the same as Ollama.

See also Ollama GGUF llama.cpp OpenAI-compatible API Quantization

LoRA

Low-Rank Adaptation: fine-tuning that trains a small pair of extra matrices instead of touching the whole model.

The result is an adapter of a few dozen megabytes loaded on top of an unchanged base, so one base model can serve many specializations. It is what makes fine-tuning affordable on a single GPU, and the same mechanism behind the style LoRAs traded around image models. An adapter can be merged back into the weights when you want one self-contained model.

See also Fine-tuning Quantization Diffusion model

↑ Back to A–Z

M 6 terms

MCP

Model Context Protocol: an open standard for connecting a model to tools and data, so one integration works in every client.

Before it, every app invented its own plugin format; MCP makes a GitHub or Postgres integration written once usable by any client that speaks the protocol. A server can offer three things: tools the model can call, resources it can read, and prompts the user can pick. The model still decides when to call something — MCP only standardizes how the call travels.

See also MCP server MCP client stdio vs HTTP transport Tool use Agent

MCP client

The side that connects to MCP servers and offers their tools to the model — Claude Code, Claude Desktop, Cursor and others.

The client owns the server list, starts local servers as child processes, and decides which tools are exposed in a given session. Same protocol, different configuration surface: Claude Code takes a CLI command and a per-project scope, Claude Desktop takes a JSON config file. A server written for one client works in the others without changes.

See also MCP MCP server stdio vs HTTP transport Permissions

MCP server

A small program that exposes a set of tools over MCP — read a file, query a database, open a pull request — for any client to use.

It is an ordinary process, usually started by the client on demand and often run straight from npx or uvx, with credentials passed in as environment variables. Every tool it advertises carries a name, a description and a JSON Schema for its arguments, and those descriptions are what the model reads when deciding to call it. Because a server runs with your permissions and its text enters your prompt, install one the way you install a dependency: from a source you trust.

See also MCP MCP client stdio vs HTTP transport Tool use Permissions

Model ID

The exact string you pass to an API to select a model, usually a family name plus a version or a date.

A bare family name is an alias that floats to whatever is newest — convenient in a scratch script, a liability in production, because your prompt can start behaving differently overnight. Pin the dated id for anything you rely on and move it deliberately. Local runtimes use the same idea with a tag: the part after the colon picks the size and the quantization.

See also API key OpenAI-compatible API Quantization Ollama

Modelfile

Ollama’s recipe file: a base model plus a baked-in system prompt and parameters, built into a new named model.

Four directives cover most uses — FROM for the base, SYSTEM for the standing instruction, PARAMETER for settings like temperature, TEMPLATE for the prompt format. Running ollama create turns the file into a model you can call by name, so the behavior travels with the model instead of being repeated in every script. It layers on top of the base rather than copying it, so a variant costs almost no disk.

See also Ollama System prompt Temperature GGUF

Multimodal

A model that accepts more than text — images, audio, PDFs — and reasons over them in the same conversation.

In practice it means you can paste a screenshot of an error, a photo of a whiteboard or a scanned invoice and ask about it directly. Images cost tokens too, roughly in proportion to their pixel count, so a run of screenshots eats a context window fast. Output is usually still text; generating images or speech is a separate model behind the same interface.

See also Token Context window Diffusion model LLM

↑ Back to A–Z

O 2 terms

Ollama

A command-line runtime for local models: one command pulls the weights, one starts a chat, and a local server stays up behind it.

That server listens on port 11434 and speaks both its own REST API and an OpenAI-compatible one, so local models drop into code you already have. Models are GGUF files from its registry, tagged by size and quantization. Custom variants are defined in a Modelfile and built with ollama create.

See also LM Studio GGUF Modelfile OpenAI-compatible API llama.cpp

OpenAI-compatible API

Any server that copies OpenAI’s request and response shape, so a client written for one can point at another by changing the base URL.

It became the de facto interface: Ollama, LM Studio, vLLM and most hosted providers expose a chat completions endpoint of the same shape. Swapping providers is usually a base URL, a key and a model id, with the surrounding code untouched. Compatibility is rarely total, though — tool calling, structured output and streaming details are where it breaks.

See also Model ID API key Streaming Tool use Ollama

↑ Back to A–Z

P 4 terms

Permissions

The allowlist that decides which commands and paths a coding agent may touch without stopping to ask you.

Neither extreme works: ask every time and you train yourself to click yes; allow everything and a stray delete runs unsupervised. Allow the boring reads — tests, linters, git status — and keep the prompt for anything that writes outside the project, installs packages, or reaches the network. Scope matters too: rules can live per project or per user, and the project ones travel with the repo.

See also Agent Hooks Plan mode Guardrails Headless mode

Plan mode

A read-only mode where the agent investigates and proposes a plan but cannot edit anything until you approve it.

It separates understanding the problem from changing the code, which is exactly where agents go wrong in an unfamiliar repo. The output is a plan you can correct cheaply, before a single diff exists. Worth it for anything touching more than a couple of files; skip it for a one-line fix.

See also Permissions Agent CLAUDE.md

Prompt caching

Reusing the provider’s computed state for a prompt prefix you send repeatedly, which cuts both its cost and its latency sharply.

The cache is keyed on an exact prefix, so it only works when the unchanging part — system prompt, tool definitions, the document — comes first and stays byte-identical. Cached input tokens are billed at a steep discount, with a small surcharge for writing the cache, and entries expire after minutes. Long agent sessions and repeated questions about one document are where it pays for itself.

See also Token Context window System prompt Inference

Prompt chaining

Splitting a big task into several smaller prompts, each taking the previous output as its input.

Every step gets a short, focused prompt and can be checked or retried on its own, which beats one enormous instruction that fails somewhere in the middle. It also lets you send the mechanical steps to a cheap model and keep the expensive one for the step that needs judgment. The cost is plumbing: something has to carry state between steps and decide what happens when one fails.

See also Agent Chain of thought Structured output Evals

↑ Back to A–Z

Q 1 term

Quantization

Storing a model’s weights at lower precision — 4, 5 or 8 bits instead of 16 — to shrink it enough to run on your hardware.

The common GGUF levels read as Q4, Q5 and Q8: Q4_K_M lands near a third the size of the 16-bit weights with a quality loss most people never notice, Q5_K_M is the safe middle, Q8_0 is close to lossless at about half of 16-bit. Fit the largest model you can at Q4 before considering a smaller one at Q8 — parameter count buys more than precision does. Below 4 bits the degradation stops being subtle.

See also GGUF llama.cpp Inference LoRA Model ID

↑ Back to A–Z

R 2 terms

RAG

Retrieval-augmented generation: search your own documents first, paste the best passages into the prompt, then ask the question.

The model learns nothing — it reads what you handed it, which is why the answer changes the moment the document does. Quality is set by retrieval, not by the model: if the right passage is missing from the prompt, no amount of prompting recovers it. When the whole corpus fits in the context window, skip the machinery and paste it.

See also Embeddings Vector database Chunking Grounding Context window

Rate limit

A provider’s cap on how much you may send per minute — requests, input tokens, output tokens — enforced with HTTP 429.

Limits are set per model and per account tier, and it is usually the token limit that bites first, not the request count. Handle a 429 by retrying with exponential backoff and jitter, and read the reset headers instead of guessing. If you hit them constantly the answer is batching or a higher tier, not a tighter retry loop.

See also API key Batch API Streaming Latency vs throughput

↑ Back to A–Z

S 6 terms

Slash command

A reusable prompt saved as a Markdown file and invoked by name, with arguments, inside a Claude Code session.

A file in .claude/commands/ becomes /name: the body is the prompt, and $ARGUMENTS is replaced with whatever you type after it. It turns a workflow you keep retyping — write the commit message, review this diff against our rules — into one word. Project commands live in the repo and are shared with everyone working in it; personal ones live in your home directory.

See also CLAUDE.md Subagent Hooks Headless mode

stdio vs HTTP transport

The two ways a client reaches an MCP server: a local child process over stdin and stdout, or a remote one over HTTP.

stdio is the common case for local tools — the client launches the server, they exchange JSON-RPC over the pipe, and it exits with the session; no ports, no auth. HTTP is for servers someone else hosts: a URL, usually a token, one server for many users. Which one a server uses decides how you add it — a command line for stdio, a URL for HTTP.

See also MCP MCP server MCP client API key

Streaming

Receiving the answer token by token as it is generated, instead of waiting for the finished response.

It does not make generation faster; it makes the wait visible, which is why every chat interface uses it. Server-sent events carry the chunks and your code reassembles them — including tool calls, which arrive in fragments. Turn it off for machine-to-machine calls where you only want the finished object.

See also Inference Latency vs throughput Structured output OpenAI-compatible API

Structured output

Forcing the answer into a shape your code can parse — in practice, JSON validated against a JSON Schema.

Asking politely for JSON works most of the time; constrained decoding, where the provider restricts each token to what the schema allows, works every time. Define the schema once and use it for both the request and the validation, and keep it flat — deeply nested schemas raise the error rate. Parse defensively anyway, and retry with the parser error attached to the prompt.

See also Tool use Guardrails Few-shot Evals Streaming

Subagent

A separate agent instance with its own context window and instructions, spawned to do one job and report back.

The point is context isolation: a search that reads forty files fills the subagent window, not the main session, and only the summary comes back. It also lets you match the job to a cheaper or a more capable model. The cost is that it starts blank — everything it needs has to be in the task you hand it.

See also Agent Context window CLAUDE.md Slash command

System prompt

The standing instruction sent ahead of the conversation that sets the model’s role, its rules and its output format.

It carries more weight than the same words typed as a user turn, and it is where anything that must hold for every reply belongs. Four parts cover most of it: who the model is, what it must and must not do, the exact shape of the output, and one example. Every word costs tokens on every turn, so cut it back to the rules that actually change behavior.

See also Few-shot CLAUDE.md Structured output Prompt caching Custom GPT

↑ Back to A–Z

T 4 terms

Temperature

The sampling setting that controls randomness: near 0 the model takes the likeliest token every time, higher it takes risks.

Use 0 to 0.3 for extraction, classification, code and anything factual; 0.7 and above for brainstorming and copy. It does not make the model more or less correct — it widens or narrows the set of tokens it will draw from, and the far tail of that set is where nonsense lives. Tune temperature or top-p, not both at once.

See also Top-p Hallucination Inference Modelfile

Token

The unit a model reads and bills in: a chunk of text averaging about three-quarters of an English word.

The tokenizer splits text before the model sees it, which is why models miscount letters and why code, JSON and non-English text cost more tokens per character than plain prose. Input and output are priced separately, output usually several times higher. Every limit you meet — context window, rate limit, monthly bill — is counted in tokens, not characters.

See also Context window Inference Rate limit Prompt caching

Tool use

Function calling: you describe functions in the request, the model replies asking for one, your code runs it and returns the result.

The model never executes anything — it emits a name and JSON arguments, and your program decides whether to honor them. The description and the parameter schema are the whole interface, which is why a vague description is the usual reason a tool fires at the wrong moment. Feed the result back in and you have an agent; standardize the plumbing and you have MCP.

See also Agent MCP Structured output Permissions

Top-p

Nucleus sampling: consider only the most likely tokens whose probabilities add up to p, and ignore the long tail below that.

At 0.9 the model draws from whatever small set covers 90% of the probability mass, so the cutoff adapts to how certain it is at that moment. It is the other randomness knob beside temperature, and turning both down at once gives flat, repetitive text. Pick one to tune and leave the other at its default.

See also Temperature Inference Modelfile

↑ Back to A–Z

V 1 term

Vector database

A store that indexes embeddings for fast nearest-neighbor search, so you can find the closest passages in milliseconds.

It keeps the vector, the original text and its metadata together, and answers the closest k to this query without comparing against every row. Below roughly ten thousand chunks a plain array and a cosine loop is genuinely enough, and pgvector or an SQLite extension covers most of what is left. You reach for a dedicated one for scale, filtered search, and keeping the index fresh as documents change.

See also Embeddings RAG Chunking

↑ Back to A–Z

52 terms · All 105 plates · Browse by topic

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list