Local LLMs

How to Serve a Model with vLLM

3 min read

vLLM is a throughput-oriented inference server. It is what you reach for when several people or requests hit one model at once — not for chatting on a laptop.

1. Check the hardware first

The CUDA build, which is the well-trodden path, needs:

OS:     Linux
Python: 3.10 - 3.13
GPU:    NVIDIA, compute capability 7.5+ (T4, RTX 20xx, A100, L4, H100, B200...)

No NVIDIA GPU means no CUDA wheel. There are separate builds for AMD ROCm, Intel XPU, Google TPU, Ascend NPU, and Apple Silicon (vLLM-Metal), each with its own install command — check the vLLM installation docs for those. On a CPU-only machine, use llama.cpp instead.

2. Install

uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

Creates a clean virtualenv and installs vLLM; --torch-backend=auto inspects your CUDA driver and picks the matching PyTorch build. Use a fresh environment — reusing an existing PyTorch install is the most common cause of broken kernels.

No uv? pip install --upgrade uv inside a conda env, then run the same uv pip install line.

3. Start the server

vllm serve Qwen/Qwen2.5-1.5B-Instruct

Downloads the model from Hugging Face and serves it at http://localhost:8000. First start is slow — weights download plus kernel warm-up. Change the address with --host and --port. One model per server process.

Confirm it is up:

curl http://localhost:8000/v1/models

Lists the served model id — that exact string is what you pass as model in requests.

4. Call it with the OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key="EMPTY",
    base_url="http://localhost:8000/v1",
)

chat = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Tell me a joke."},
    ],
)
print(chat.choices[0].message.content)

Standard OpenAI code with two changes: base_url points at your server and the key is a placeholder. Auth is off by default — start vLLM with --api-key <secret> (or set VLLM_API_KEY) before exposing the port to anything but localhost.


Next: run a lighter local server with llama.cpp or put a chat UI in front of it.

Open the full interactive version (with copy buttons) ↗

← All guides