Run a Local LLM with LM Studio (No Terminal)
LM Studio downloads and runs GGUF models from a desktop app, then serves them at http://localhost:1234/v1 for any OpenAI-compatible client. No terminal.
Step 2 of 5 · Run models locally
On this page6 sections
LM Studio is the shortest path to a local model when you don’t want a terminal: a desktop app that searches, downloads, loads and chats — and then, with one toggle, exposes that model as an OpenAI-compatible API on your machine.
1. Install
Download the installer for macOS, Windows or Linux from lmstudio.ai and run it. Nothing else is required; the app bundles its own inference runtimes.
2. Download a model
Open the search tab and pick a model. Downloads are GGUF files, and the quantization you choose decides whether it fits in memory: an 8B model at Q4 lands around 4.6 GB, so budget roughly 8 GB of free RAM. Work out your ceiling first with RAM and VRAM requirements, then choose a quantization.
3. Load it and chat
In the chat tab, select the downloaded model to load it into memory and start typing. Once the file is on disk nothing leaves your machine — unplug the network and it still answers.
4. Start the server
Go to the Developer tab and toggle Start server. (This moved: older walkthroughs still say “Local Server” tab.) The server listens on port 1234 and speaks the OpenAI wire format, exposing /v1/models, /v1/chat/completions, /v1/completions, /v1/embeddings and /v1/responses.
5. Call it from code
from openai import OpenAI
client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")
r = client.chat.completions.create(
model="use the model identifier from LM Studio here",
messages=[{"role": "user", "content": "Say this is a test!"}],
)
print(r.choices[0].message.content)
Sends one chat request to the local server — same SDK as the cloud, only base_url changes, and the key is ignored.
Verify it worked
curl http://localhost:1234/v1/models
Lists the loaded models as JSON. If it fails: an empty list means the server is running but no model is loaded — load one in the chat tab first. Connection refused means the Developer-tab toggle is off. A wrong model string returns a 404, so copy the identifier the app shows. Prefer the command line? Use Ollama instead, or read the LM Studio API docs.