# How to Serve a Model with vLLM

> Install vLLM, start a server with vllm serve on localhost:8000, and call it with the OpenAI Python SDK. Needs a Linux box with an NVIDIA GPU.

- Canonical: https://guides-ai.pages.dev/guides/vllm-serve-openai-compatible/
- Plate 09.12 · Topic: Local LLMs (https://guides-ai.pages.dev/topics/local-llm/)
- Published: 06 Sept 2026 · 3 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

vLLM is a throughput-oriented inference server. It is what you reach for when several people or requests hit one model at once — not for chatting on a laptop.

## 1. Check the hardware first

The CUDA build, which is the well-trodden path, needs:

```text
OS:     Linux
Python: 3.10 - 3.13
GPU:    NVIDIA, compute capability 7.5+ (T4, RTX 20xx, A100, L4, H100, B200...)
```

No NVIDIA GPU means no CUDA wheel. There are separate builds for AMD ROCm, Intel XPU, Google TPU, Ascend NPU, and Apple Silicon (vLLM-Metal), each with its own install command — check the vLLM installation docs for those. On a CPU-only machine, use [llama.cpp](/guides/llama-cpp-quickstart/) instead.

## 2. Install

```bash
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
```

Creates a clean virtualenv and installs vLLM; `--torch-backend=auto` inspects your CUDA driver and picks the matching PyTorch build. Use a **fresh** environment — reusing an existing PyTorch install is the most common cause of broken kernels.

No `uv`? `pip install --upgrade uv` inside a conda env, then run the same `uv pip install` line.

## 3. Start the server

```bash
vllm serve Qwen/Qwen2.5-1.5B-Instruct
```

Downloads the model from Hugging Face and serves it at `http://localhost:8000`. First start is slow — weights download plus kernel warm-up. Change the address with `--host` and `--port`. One model per server process.

Confirm it is up:

```bash
curl http://localhost:8000/v1/models
```

Lists the served model id — that exact string is what you pass as `model` in requests.

## 4. Call it with the OpenAI SDK

```python
from openai import OpenAI

client = OpenAI(
    api_key="EMPTY",
    base_url="http://localhost:8000/v1",
)

chat = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Tell me a joke."},
    ],
)
print(chat.choices[0].message.content)
```

Standard OpenAI code with two changes: `base_url` points at your server and the key is a placeholder. Auth is off by default — start vLLM with `--api-key <secret>` (or set `VLLM_API_KEY`) before exposing the port to anything but localhost.

---

Next: [run a lighter local server with llama.cpp](/guides/llama-cpp-quickstart/) or [put a chat UI in front of it](/guides/open-webui-with-ollama/).
