ToolLocal LLMs

Ollama Modelfile builder

Fill in the base model, a system prompt and the parameters you actually want to change. You get a Modelfile you can save, and the two commands that turn it into a model you can run by name. Nothing leaves your browser.


Model

Lowercase letters, digits, - _ . — an optional :tag after it. This is what you type after ollama run.

A name from ollama.com/library (add :tag for a specific size — the model's own page lists them), a Safetensors model directory, or a path to a GGUF file such as ./model.gguf.

The instruction baked into every chat. More than one line is wrapped in the documented """…""" form.

Parameters

Leave a box empty to leave that parameter alone — the model then uses the default shown under it.

float · default 0.8 — higher answers more creatively

int · default 2048 — context window used to generate the next token

float · default 0.9 — higher is more diverse, lower more focused

int · default 40 — 100 is more diverse, 10 more conservative

float · default 1.0 — 1.5 penalizes repetition harder, 0.9 is lenient

int · default 0 — a fixed seed repeats the same text for the same prompt

int · default -1 (infinite) — max tokens to predict

Each line becomes its own PARAMETER stop line. Generation stops when one is produced.

Advanced — template, adapter, license, message history

Leave this empty unless you know the base model's chat format. The template is the exact string sent to the model; the syntax is model specific, so a wrong one breaks chat formatting — the model rambles, never stops, or ignores your system prompt. Empty means your model keeps the base model's template. Copy the real one first with ollama show --modelfile <model>. Whitespace is passed through exactly as you type it, trailing newline included.

Absolute, or relative to the Modelfile. It must match the base model in FROM, or the results are erratic.

The legal license the model is shared under. Only worth filling in if you plan to publish it.

MESSAGE — example history

A short conversation that shows the model how you want it to answer. Rows with no text are skipped.

What each line does

Straight from the Ollama Modelfile reference — only the instructions and parameters this builder can write.

docs.ollama.com/modelfile →
LineType · defaultWhat it does
FROM instruction Defines the base model to use. Required. A library name, a Safetensors model directory, or a path to a .gguf file (absolute, or relative to the Modelfile).
SYSTEM instruction Specifies the system message that will be set in the template — the instruction baked into every chat.
PARAMETER temperature float · 0.8 The temperature of the model. Increasing the temperature will make the model answer more creatively.
PARAMETER num_ctx int · 2048 Sets the size of the context window used to generate the next token.
PARAMETER top_p float · 0.9 Works together with top-k. A higher value (e.g. 0.95) leads to more diverse text; a lower value (e.g. 0.5) is more focused and conservative.
PARAMETER top_k int · 40 Reduces the probability of generating nonsense. A higher value (e.g. 100) gives more diverse answers, a lower value (e.g. 10) is more conservative.
PARAMETER repeat_penalty float · 1.0 Sets how strongly to penalize repetitions. A higher value (e.g. 1.5) penalizes repetitions more strongly, a lower value (e.g. 0.9) is more lenient.
PARAMETER seed int · 0 Sets the random number seed to use for generation. Setting this to a specific number will make the model generate the same text for the same prompt.
PARAMETER num_predict int · -1 Maximum number of tokens to predict when generating text. -1 means infinite generation.
PARAMETER stop string · repeats Sets the stop sequences to use. When this pattern is encountered the LLM will stop generating text and return. Multiple stop patterns are set as multiple separate stop parameters.
TEMPLATE instruction The full prompt template to be sent to the model, in Go template syntax. Variables: .System, .Prompt, .Response. Syntax may be model specific.
ADAPTER instruction Defines the (Q)LoRA adapters to apply to the base model. Absolute path, or relative to the Modelfile. If the base model is not the one the adapter was tuned from, behaviour will be erratic.
LICENSE instruction Specifies the legal license under which the model is shared or distributed.
MESSAGE instruction Specify message history. Repeat it to build a short conversation that guides the model to answer in a similar way. Valid roles: system, user, assistant.

The Modelfile is not case sensitive and instructions can be in any order — uppercase and FROM first are conventions that keep it readable.

Guides that go with it

The plates that cover installing Ollama, calling it from code, and the Modelfile itself.

All Local LLM plates →
  1. 09.02 Ollama Modelfile: a Custom Model With a System Prompt An Ollama Modelfile bakes FROM, SYSTEM and PARAMETER into a named local model with ollama create, so behaviour travels with the model, not your scripts. Local LLMs3 min
  2. 09.01 How to Run Ollama on Windows (Install, GPU, Models) Install Ollama on Windows 10 22H2+, run your first model with ollama run, check it landed on the GPU, and move the model store off your C: drive. Local LLMs3 min
  3. 09.11 How to Run a Local LLM with Ollama on macOS/Linux Install Ollama with one command, pull gemma3:4b, chat offline in the terminal, and keep the daemon running as a background service on port 11434. Local LLMs3 min
  4. 09.03 Call Your Local Ollama Model from curl & Python Ollama serves a local API on port 11434. Call /api/chat with curl or Python, or point any OpenAI client at http://localhost:11434/v1/ instead. Local LLMs3 min

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list