Local LLMs

Run a Local LLM with LM Studio (No Terminal)

3 min read

LM Studio is the shortest path to a local model when you don’t want a terminal: a desktop app that searches, downloads, loads and chats — and then, with one toggle, exposes that model as an OpenAI-compatible API on your machine.

1. Install

Download the installer for macOS, Windows or Linux from lmstudio.ai and run it. Nothing else is required; the app bundles its own inference runtimes.

2. Download a model

Open the search tab and pick a model. Downloads are GGUF files, and the quantization you choose decides whether it fits in memory: an 8B model at Q4 lands around 4.6 GB, so budget roughly 8 GB of free RAM. Work out your ceiling first with RAM and VRAM requirements, then choose a quantization.

3. Load it and chat

In the chat tab, select the downloaded model to load it into memory and start typing. Once the file is on disk nothing leaves your machine — unplug the network and it still answers.

4. Start the server

Go to the Developer tab and toggle Start server. (This moved: older walkthroughs still say “Local Server” tab.) The server listens on port 1234 and speaks the OpenAI wire format, exposing /v1/models, /v1/chat/completions, /v1/completions, /v1/embeddings and /v1/responses.

5. Call it from code

from openai import OpenAI

client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")
r = client.chat.completions.create(
    model="use the model identifier from LM Studio here",
    messages=[{"role": "user", "content": "Say this is a test!"}],
)
print(r.choices[0].message.content)

Sends one chat request to the local server — same SDK as the cloud, only base_url changes, and the key is ignored.

Verify it worked

curl http://localhost:1234/v1/models

Lists the loaded models as JSON. If it fails: an empty list means the server is running but no model is loaded — load one in the chat tab first. Connection refused means the Developer-tab toggle is off. A wrong model string returns a 404, so copy the identifier the app shows. Prefer the command line? Use Ollama instead, or read the LM Studio API docs.

Open the full interactive version (with copy buttons) ↗

← All guides