How to Serve a Model with vLLM
Install vLLM, start a server with vllm serve on localhost:8000, and call it with the OpenAI Python SDK. Needs a Linux box with an NVIDIA GPU.
Step 3 of 4 · Serve models yourself
On this page4 sections
vLLM is a throughput-oriented inference server. It is what you reach for when several people or requests hit one model at once — not for chatting on a laptop.
1. Check the hardware first
The CUDA build, which is the well-trodden path, needs:
OS: Linux
Python: 3.10 - 3.13
GPU: NVIDIA, compute capability 7.5+ (T4, RTX 20xx, A100, L4, H100, B200...)
No NVIDIA GPU means no CUDA wheel. There are separate builds for AMD ROCm, Intel XPU, Google TPU, Ascend NPU, and Apple Silicon (vLLM-Metal), each with its own install command — check the vLLM installation docs for those. On a CPU-only machine, use llama.cpp instead.
2. Install
uv venv --python 3.12 --seed
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
Creates a clean virtualenv and installs vLLM; --torch-backend=auto inspects your CUDA driver and picks the matching PyTorch build. Use a fresh environment — reusing an existing PyTorch install is the most common cause of broken kernels.
No uv? pip install --upgrade uv inside a conda env, then run the same uv pip install line.
3. Start the server
vllm serve Qwen/Qwen2.5-1.5B-Instruct
Downloads the model from Hugging Face and serves it at http://localhost:8000. First start is slow — weights download plus kernel warm-up. Change the address with --host and --port. One model per server process.
Confirm it is up:
curl http://localhost:8000/v1/models
Lists the served model id — that exact string is what you pass as model in requests.
4. Call it with the OpenAI SDK
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:8000/v1",
)
chat = client.chat.completions.create(
model="Qwen/Qwen2.5-1.5B-Instruct",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Tell me a joke."},
],
)
print(chat.choices[0].message.content)
Standard OpenAI code with two changes: base_url points at your server and the key is a placeholder. Auth is off by default — start vLLM with --api-key <secret> (or set VLLM_API_KEY) before exposing the port to anything but localhost.
Next: run a lighter local server with llama.cpp or put a chat UI in front of it.