ToolLocal LLMs
Ollama Modelfile builder
Fill in the base model, a system prompt and the parameters you actually want to change. You get a Modelfile you can save, and the two commands that turn it into a model you can run by name. Nothing leaves your browser.
What each line does
Straight from the Ollama Modelfile reference — only the instructions and parameters this builder can write.
| Line | Type · default | What it does |
|---|---|---|
FROM | instruction | Defines the base model to use. Required. A library name, a Safetensors model directory, or a path to a .gguf file (absolute, or relative to the Modelfile). |
SYSTEM | instruction | Specifies the system message that will be set in the template — the instruction baked into every chat. |
PARAMETER temperature | float · 0.8 | The temperature of the model. Increasing the temperature will make the model answer more creatively. |
PARAMETER num_ctx | int · 2048 | Sets the size of the context window used to generate the next token. |
PARAMETER top_p | float · 0.9 | Works together with top-k. A higher value (e.g. 0.95) leads to more diverse text; a lower value (e.g. 0.5) is more focused and conservative. |
PARAMETER top_k | int · 40 | Reduces the probability of generating nonsense. A higher value (e.g. 100) gives more diverse answers, a lower value (e.g. 10) is more conservative. |
PARAMETER repeat_penalty | float · 1.0 | Sets how strongly to penalize repetitions. A higher value (e.g. 1.5) penalizes repetitions more strongly, a lower value (e.g. 0.9) is more lenient. |
PARAMETER seed | int · 0 | Sets the random number seed to use for generation. Setting this to a specific number will make the model generate the same text for the same prompt. |
PARAMETER num_predict | int · -1 | Maximum number of tokens to predict when generating text. -1 means infinite generation. |
PARAMETER stop | string · repeats | Sets the stop sequences to use. When this pattern is encountered the LLM will stop generating text and return. Multiple stop patterns are set as multiple separate stop parameters. |
TEMPLATE | instruction | The full prompt template to be sent to the model, in Go template syntax. Variables: .System, .Prompt, .Response. Syntax may be model specific. |
ADAPTER | instruction | Defines the (Q)LoRA adapters to apply to the base model. Absolute path, or relative to the Modelfile. If the base model is not the one the adapter was tuned from, behaviour will be erratic. |
LICENSE | instruction | Specifies the legal license under which the model is shared or distributed. |
MESSAGE | instruction | Specify message history. Repeat it to build a short conversation that guides the model to answer in a similar way. Valid roles: system, user, assistant. |
The Modelfile is not case sensitive and instructions can be in any order — uppercase and
FROM first are conventions that keep it readable.
Guides that go with it
The plates that cover installing Ollama, calling it from code, and the Modelfile itself.
- 09.02 Ollama Modelfile: a Custom Model With a System Prompt An Ollama Modelfile bakes FROM, SYSTEM and PARAMETER into a named local model with ollama create, so behaviour travels with the model, not your scripts. Local LLMs3 min
- 09.01 How to Run Ollama on Windows (Install, GPU, Models) Install Ollama on Windows 10 22H2+, run your first model with ollama run, check it landed on the GPU, and move the model store off your C: drive. Local LLMs3 min
- 09.11 How to Run a Local LLM with Ollama on macOS/Linux Install Ollama with one command, pull gemma3:4b, chat offline in the terminal, and keep the daemon running as a background service on port 11434. Local LLMs3 min
- 09.03 Call Your Local Ollama Model from curl & Python Ollama serves a local API on port 11434. Call /api/chat with curl or Python, or point any OpenAI client at http://localhost:11434/v1/ instead. Local LLMs3 min