# Run a Local LLM with LM Studio (No Terminal)

> LM Studio downloads and runs GGUF models from a desktop app, then serves them at http://localhost:1234/v1 for any OpenAI-compatible client. No terminal.

- Canonical: https://guides-ai.pages.dev/guides/run-llm-lm-studio/
- Plate 09.05 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 02 Sept 2026 · Updated: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

LM Studio is the shortest path to a local model when you don't want a terminal: a desktop app that searches, downloads, loads and chats — and then, with one toggle, exposes that model as an [OpenAI-compatible API](/glossary/#openai-compatible-api) on your machine.

## 1. Install

Download the installer for macOS, Windows or Linux from **lmstudio.ai** and run it. Nothing else is required; the app bundles its own inference runtimes.

## 2. Download a model

Open the search tab and pick a model. Downloads are [GGUF](/glossary/#gguf) files, and the quantization you choose decides whether it fits in memory: an 8B model at Q4 lands around 4.6 GB, so budget roughly 8 GB of free RAM. Work out your ceiling first with [RAM and VRAM requirements](/guides/local-llm-ram-vram-requirements/), then [choose a quantization](/guides/choose-llm-quantization/).

## 3. Load it and chat

In the chat tab, select the downloaded model to load it into memory and start typing. Once the file is on disk nothing leaves your machine — unplug the network and it still answers.

## 4. Start the server

Go to the **Developer** tab and toggle **Start server**. (This moved: older walkthroughs still say "Local Server" tab.) The server listens on port `1234` and speaks the OpenAI wire format, exposing `/v1/models`, `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings` and `/v1/responses`.

## 5. Call it from code

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")
r = client.chat.completions.create(
    model="use the model identifier from LM Studio here",
    messages=[{"role": "user", "content": "Say this is a test!"}],
)
print(r.choices[0].message.content)
```

Sends one chat request to the local server — same SDK as the cloud, only `base_url` changes, and the key is ignored.

## Verify it worked

```bash
curl http://localhost:1234/v1/models
```

Lists the loaded models as JSON. **If it fails:** an empty list means the server is running but no model is loaded — load one in the chat tab first. Connection refused means the Developer-tab toggle is off. A wrong `model` string returns a 404, so copy the identifier the app shows. Prefer the command line? Use [Ollama instead](/guides/run-local-llm-ollama-macos-linux/), or read the [LM Studio API docs](https://lmstudio.ai/docs/app/api/endpoints/openai).
