§09.12

How to Serve a Model with vLLM

Install vLLM, start a server with vllm serve on localhost:8000, and call it with the OpenAI Python SDK. Needs a Linux box with an NVIDIA GPU.

published 06 Sept 2026 checked against docs 06 Sept 2026 3 min in Local LLMs Markdown

Step 3 of 4 · Serve models yourself

On this page4 sections
  1. 1. Check the hardware first
  2. 2. Install
  3. 3. Start the server
  4. 4. Call it with the OpenAI SDK

vLLM is a throughput-oriented inference server. It is what you reach for when several people or requests hit one model at once — not for chatting on a laptop.

1. Check the hardware first

The CUDA build, which is the well-trodden path, needs:

OS:     Linux
Python: 3.10 - 3.13
GPU:    NVIDIA, compute capability 7.5+ (T4, RTX 20xx, A100, L4, H100, B200...)

No NVIDIA GPU means no CUDA wheel. There are separate builds for AMD ROCm, Intel XPU, Google TPU, Ascend NPU, and Apple Silicon (vLLM-Metal), each with its own install command — check the vLLM installation docs for those. On a CPU-only machine, use llama.cpp instead.

2. Install

uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

Creates a clean virtualenv and installs vLLM; --torch-backend=auto inspects your CUDA driver and picks the matching PyTorch build. Use a fresh environment — reusing an existing PyTorch install is the most common cause of broken kernels.

No uv? pip install --upgrade uv inside a conda env, then run the same uv pip install line.

3. Start the server

vllm serve Qwen/Qwen2.5-1.5B-Instruct

Downloads the model from Hugging Face and serves it at http://localhost:8000. First start is slow — weights download plus kernel warm-up. Change the address with --host and --port. One model per server process.

Confirm it is up:

curl http://localhost:8000/v1/models

Lists the served model id — that exact string is what you pass as model in requests.

4. Call it with the OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key="EMPTY",
    base_url="http://localhost:8000/v1",
)

chat = client.chat.completions.create(
    model="Qwen/Qwen2.5-1.5B-Instruct",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Tell me a joke."},
    ],
)
print(chat.choices[0].message.content)

Standard OpenAI code with two changes: base_url points at your server and the key is a placeholder. Auth is off by default — start vLLM with --api-key <secret> (or set VLLM_API_KEY) before exposing the port to anything but localhost.


Next: run a lighter local server with llama.cpp or put a chat UI in front of it.

← All Local LLMs plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list