§09.05

Run a Local LLM with LM Studio (No Terminal)

LM Studio downloads and runs GGUF models from a desktop app, then serves them at http://localhost:1234/v1 for any OpenAI-compatible client. No terminal.

published 02 Sept 2026 updated 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Local LLMs Markdown

Step 2 of 5 · Run models locally

On this page6 sections
  1. 1. Install
  2. 2. Download a model
  3. 3. Load it and chat
  4. 4. Start the server
  5. 5. Call it from code
  6. Verify it worked

LM Studio is the shortest path to a local model when you don’t want a terminal: a desktop app that searches, downloads, loads and chats — and then, with one toggle, exposes that model as an OpenAI-compatible API on your machine.

1. Install

Download the installer for macOS, Windows or Linux from lmstudio.ai and run it. Nothing else is required; the app bundles its own inference runtimes.

2. Download a model

Open the search tab and pick a model. Downloads are GGUF files, and the quantization you choose decides whether it fits in memory: an 8B model at Q4 lands around 4.6 GB, so budget roughly 8 GB of free RAM. Work out your ceiling first with RAM and VRAM requirements, then choose a quantization.

3. Load it and chat

In the chat tab, select the downloaded model to load it into memory and start typing. Once the file is on disk nothing leaves your machine — unplug the network and it still answers.

4. Start the server

Go to the Developer tab and toggle Start server. (This moved: older walkthroughs still say “Local Server” tab.) The server listens on port 1234 and speaks the OpenAI wire format, exposing /v1/models, /v1/chat/completions, /v1/completions, /v1/embeddings and /v1/responses.

5. Call it from code

from openai import OpenAI

client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")
r = client.chat.completions.create(
    model="use the model identifier from LM Studio here",
    messages=[{"role": "user", "content": "Say this is a test!"}],
)
print(r.choices[0].message.content)

Sends one chat request to the local server — same SDK as the cloud, only base_url changes, and the key is ignored.

Verify it worked

curl http://localhost:1234/v1/models

Lists the loaded models as JSON. If it fails: an empty list means the server is running but no model is loaded — load one in the chat tab first. Connection refused means the Developer-tab toggle is off. A wrong model string returns a 404, so copy the identifier the app shows. Prefer the command line? Use Ollama instead, or read the LM Studio API docs.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list