09Topic 9 of 15
Local LLMs
Ollama and LM Studio: run a model locally, pick a quantization, call the API.
- 09.01 How to Run Ollama on Windows (Install, GPU, Models) Install Ollama on Windows 10 22H2+, run your first model with ollama run, check it landed on the GPU, and move the model store off your C: drive. 3 min
- 09.02 Ollama Modelfile: a Custom Model With a System Prompt An Ollama Modelfile bakes FROM, SYSTEM and PARAMETER into a named local model with ollama create, so behaviour travels with the model, not your scripts. 3 min
- 09.03 Call Your Local Ollama Model from curl & Python Ollama serves a local API on port 11434. Call /api/chat with curl or Python, or point any OpenAI client at http://localhost:11434/v1/ instead. 3 min
- 09.04 Choose a GGUF Quantization: Q4 vs Q5 vs Q8 Q4_K_M is the default pick for local LLMs. Real bits-per-weight for each GGUF level, why Q8_0 is a legacy format, and how to match one to your RAM. 3 min
- 09.05 Run a Local LLM with LM Studio (No Terminal) LM Studio downloads and runs GGUF models from a desktop app, then serves them at http://localhost:1234/v1 for any OpenAI-compatible client. No terminal. 3 min
- 09.06 Download Models from Hugging Face with the hf CLI Install the hf CLI, log in, download one GGUF instead of a whole repo with --include, put files where you want with --local-dir, and clean the cache. 3 min
- 09.07 llama.cpp Quickstart: Run a GGUF Model Install llama.cpp, pull a GGUF straight from Hugging Face with llama cli -hf, chat in the terminal, then serve an OpenAI-compatible API on localhost:8080. 4 min
- 09.08 How Much RAM or VRAM a Local LLM Needs (1B to 70B) Real GGUF file sizes at Q4/Q8 for 1B–70B models, the rule of thumb behind them, and what fits in 8, 16, 24 or 64 GB. Pick a model that actually runs. 3 min
- 09.09 How to Run Ollama in Docker with an NVIDIA GPU NVIDIA Container Toolkit setup, the official docker run command, pulling a model, keeping models in a volume, and AMD or CPU fallbacks. 3 min
- 09.10 How to Run Open WebUI with Ollama (Docker) Give your local Ollama models a ChatGPT-style web interface with one Docker command — first login, model picker, and fixes. 3 min
- 09.11 How to Run a Local LLM with Ollama on macOS/Linux Install Ollama with one command, pull gemma3:4b, chat offline in the terminal, and keep the daemon running as a background service on port 11434. 3 min
- 09.12 How to Serve a Model with vLLM Install vLLM, start a server with vllm serve on localhost:8000, and call it with the OpenAI Python SDK. Needs a Linux box with an NVIDIA GPU. 3 min
Keep going