To run vLLM you need Linux (or WSL on Windows), an NVIDIA card with compute capability 7.5 or newer and Python 3.10 to 3.13. Then 2 commands: uv pip install vllm --torch-backend=auto and vllm serve. The server answers on port 8000 on every network interface, not only on your own machine: add --host 127.0.0.1.
That last point is what sets vLLM apart. Of the 5 runtimes we compare, it is the only one that listens on all interfaces out of the box; Ollama, LM Studio, llama.cpp and MLX start on 127.0.0.1. It also starts with no key, CORS open to any website, and usage statistics switched on. None of that is hard to change, and this page goes through it in order, with the current stable version, 0.30.0, published on 22 September 2026. Every command and default here was checked against the 0.30.0 code as tagged in its public repository, on 28 September 2026.
What vLLM is, and who it is for
vLLM is an inference engine built to serve language models to many people at once. Its home is a Linux server with NVIDIA GPUs, and that is where its defaults make sense: a service that other machines reach over the network, that reserves most of the card for itself, and that reports to its developers how it is being used. It is open source under the Apache 2.0 licence. The vLLM sheet: what it sends home and whether you can audit it places it at the professional end of the 5 runtimes, and this guide is the practical side of that sheet.
Everything comes in 1 executable, vllm. The subcommands you will use are vllm serve, which starts the OpenAI-compatible server, and vllm chat and vllm complete, which are console clients that talk to that server. There are also vllm bench, vllm collect-env, vllm run-batch and vllm launch, which this page does not need.
What it is not, because each of these gets assumed:
- Not a desktop app. There is no window.
vllm chatis a text client that needs the server already running. - Not a GGUF runner out of the box. GGUF support has moved to a separate plugin, and the model section explains what that means for sizes.
- Not a Mac GPU program. On Apple silicon, vLLM itself runs on the CPU. If you have a Mac, How to set up MLX on a Mac with mlx-lm, and what it sends out uses the GPU directly.
- Not a model switcher. The server hosts 1 model at a time. To change it, you restart the server with another name.
If what you want is a model to chat with on your own computer, How to set up Ollama locally and see what it sends out and How to set up LM Studio and its local server, and what it sends out get you there with fewer decisions. vLLM is for when you want to serve.
What you need: Linux, an NVIDIA card with compute capability 7.5, Python 3.10 to 3.13
| Requirement | What it means in practice |
|---|---|
| Linux | The prebuilt packages on PyPI are Linux wheels only, for x86_64 (314.9 MB) and aarch64 (310.0 MB). Windows goes through WSL |
| NVIDIA GPU, compute capability 7.5 or newer | T4 and the RTX 20 series upwards. The RTX 30 series is 8.6 and the RTX 40 series 8.9. The GTX 10 series is below the line. The package on PyPI is built for CUDA 13.0; a CUDA 12.9 build is attached to the release on GitHub |
| Python 3.10 to 3.13 | The documentation asks for 3.10 to 3.13; the package metadata accepts 3.10 up to, but not including, 3.15. Use 3.12: it sits inside both ranges, and the AMD wheels and the Mac wheel are built for it |
| A fresh environment | vLLM 0.30.0 pins its own PyTorch, torch==2.13.0, plus matching torchaudio and torchvision, which is why its documentation recommends a new environment |
Requirements from vLLM’s installation guide at tag v0.30.0, the PyPI metadata of vllm 0.30.0 and NVIDIA’s compute capability table, read on 28 September 2026.
Windows. vLLM does not run on Windows natively. Its documentation points to the Windows Subsystem for Linux, and inside WSL you follow the Linux steps on this page.
Mac. vLLM supports Apple silicon on the CPU, as experimental support: FP32 and FP16, macOS Sonoma or newer, Xcode 15.4 and its Command Line Tools. PyPI has no macOS wheel, and the documentation describes building from source. The 0.30.0 release on GitHub also attaches a ready-made macOS arm64 CPU wheel for Python 3.12, 29.0 MB; we have not tested it. The Mac’s GPU is a different project: vLLM points to vLLM-Metal, a plugin maintained by the community that uses MLX underneath and loads models from mlx-community. This page does not cover it.
AMD. vLLM supports AMD GPUs with ROCm 6.3 or newer, and for 0.30.0 the prebuilt wheel at https://wheels.vllm.ai/rocm/ is built for ROCm 7.2.3 and needs glibc 2.39 or newer. Those wheels exist for Python 3.12 only, and this is the step to get right: with any other Python, the installer silently falls back to the CUDA wheel from PyPI, which then fails on the AMD card with libcudart.so: cannot open shared object file. The supported list covers MI200, MI300 and MI350, the Radeon RX 7900 (gfx1100/1101) and RX 9000 (gfx1200/1201), and Ryzen AI MAX and AI 300 (gfx1151/1150, ROCm 7.0.2 or newer).
CPU only. On an x86 Linux machine without a suitable GPU, vLLM runs on the processor. AVX512 is recommended; with AVX2 you get limited features. The 0.30.0 release attaches a CPU wheel for x86_64 (147.4 MB) and one for aarch64 (68.3 MB); both need glibc 2.39 or newer.
Which machines from our catalogue can run it
Our catalogue holds 16 machines. Read against vLLM’s own requirements, they fall into 4 groups:
| Machine | Compute capability | How vLLM runs on it |
|---|---|---|
| PC with an RTX 4060 8 GB | 8.9 | CUDA, on Linux or WSL |
| PC with an RTX 3060 12 GB | 8.6 | CUDA, on Linux or WSL |
| PC with an RTX 4070 12 GB | 8.9 | CUDA, on Linux or WSL |
| Workstation with an RTX 4090 24 GB | 8.9 | CUDA, on Linux or WSL |
| Server with 2× RTX 3090 (48 GB) | 8.6 | CUDA across both cards with --tensor-parallel-size 2 |
| The 8 Macs, from the M4 with 16 GB to the Mac Studio with M5 Ultra and 512 GB | — | CPU only, experimental. The GPU only through the vLLM-Metal plugin |
| Office laptop with 16 GB, and the CPU side of the mini-PC with an NPU and 32 GB | — | CPU on Linux. If its processor is a Ryzen AI 300 or Ryzen AI MAX, its integrated GPU is on vLLM’s ROCm list above. vLLM’s main repository has no backend for the NPU |
| Xiaomi AI Cube, 80 GB | — | Not verifiable: an engineering prototype whose processor is not documented against vLLM’s requirements |
Compute capability from NVIDIA’s table, read on 28 September 2026. Routes from vLLM’s installation guide at tag v0.30.0.
The 5 NVIDIA machines are the case vLLM is built for. On the rest, where it runs at all, it runs on the processor.
How to install vLLM, and check which version you got
vLLM’s documentation recommends uv and a new environment. With Python 3.12:
uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
vllm --version
--torch-backend=auto looks at the NVIDIA driver you have installed and picks the matching PyTorch index. --torch-backend=cu129 or cu130 asks for a variant explicitly. With plain pip, the line is simply:
pip install vllm
The vllm wheel itself is 314.9 MB, and it brings PyTorch 2.13.0, torchaudio 2.11.0, torchvision 0.28.0 and FlashInfer with it: the full download is larger than the wheel, and we have not measured it. vllm --version should print 0.30.0.
Pin the version. vLLM moves fast: 32 stable versions in the 12 months to 28 September 2026, and the last 3 arrived 15, 14 and 13 days after the one before (0.28.0 on 26 August, 0.29.0 on 9 September, 0.30.0 on 22 September). To install exactly the version this page describes:
uv pip install vllm==0.30.0 --torch-backend=auto
The Docker route. The official image is vllm/vllm-openai. Tag v0.30.0 was published on 22 September 2026 and weighs 8.1 GiB compressed for amd64 and 9.0 GiB for arm64:
docker run --rm --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 127.0.0.1:8000:8000 \
--ipc=host \
vllm/vllm-openai:v0.30.0 \
--model Qwen/Qwen3-0.6B
That is vLLM’s own command with a fixed tag instead of latest, 127.0.0.1 in front of the port, so Docker publishes it only on your machine, and --rm; the HF_TOKEN line is left out because Qwen/Qwen3-0.6B is public. The volume shares your Hugging Face cache with the container, so models are not downloaded twice. The image runs as root by default; it also includes a built-in vllm user, UID 2000, and the Docker page of vLLM’s documentation shows how to switch to it.
Your first model: vllm serve on port 8000
vllm serve
Without a model name, vLLM loads Qwen/Qwen3-0.6B: 751,632,384 parameters in BF16, 1.41 GiB to download from Hugging Face, a context of up to 40,960 tokens and the Apache 2.0 licence. To pick the model, give its Hugging Face name:
vllm serve Qwen/Qwen3-0.6B --max-model-len 8192
The memory section explains why that second option is there. Once the log shows the server is up, 2 checks from another terminal:
curl http://127.0.0.1:8000/v1/models
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "Qwen/Qwen3-0.6B", "messages": [{"role": "user", "content": "How tall is Mt Everest?"}]}'
The first lists the loaded model; the second returns an answer in OpenAI’s format. The main routes are POST /v1/chat/completions, POST /v1/completions and GET /v1/models, plus GET /health and GET /version. For an OpenAI client, the base URL is http://127.0.0.1:8000/v1. For a quick chat in the terminal:
vllm chat
It connects to http://localhost:8000/v1 unless you give it another --url. The server keeps no record of your requests.
Memory: why vLLM takes 92% of the card, and –max-model-len
vLLM does not load a model and stop there. At startup it reserves a fixed share of the card, 0.92 of its total memory by default, for the model, its working memory and the cache that holds your conversations (the KV cache). 2 consequences follow from how the 0.30.0 code does it.
It counts free memory, not total. If the free memory at startup is less than 92% of the card, vLLM refuses to start with Free memory on device … on startup is less than desired GPU memory utilization (0.92, X GiB). Decrease GPU memory utilization or reduce GPU memory used by other processes. On an 8 GB card that means 7.36 GiB free. If the same card also drives your desktop, the desktop and the CUDA context vLLM opens for itself have to fit in about 0.64 GiB together, or the check fails before any model loads. The fix is the one the message gives:
vllm serve --gpu-memory-utilization 0.85
The value is a limit per process: 2 vLLM servers on 1 card need something like 0.5 each. On the CPU backend the same option reserves system RAM, despite its name.
The context defaults to the model’s maximum. --max-model-len is taken from the model’s own configuration unless you set it, and vLLM refuses to start if the KV cache for that full length does not fit. The error names the numbers: To serve at least one request with the model’s max seq len (N), (X GiB KV cache is needed, which is larger than the available KV cache memory (Y GiB), followed by the estimated maximum length that would fit and the advice to raise gpu_memory_utilization or lower max_model_len. If nothing is left at all, the message is No available memory for the cache blocks.
How large that cache gets, calculated from each model’s architecture with a 16-bit KV cache:
| Model | Maximum context | KV cache at the maximum | KV cache with –max-model-len 8192 |
|---|---|---|---|
| qwen3-1.7b | 40,960 | 4.38 GiB | 0.875 GiB |
| qwen3-8b | 40,960 | 5.63 GiB | 1.125 GiB |
| qwen3-4b-2507 | 262,144 | 36.0 GiB | 1.125 GiB |
Our calculation: 2 × layers × KV heads × head size × 2 bytes per token, from the model data behind our model finder. Gemma 4 and gpt-oss use sliding-window layers, where this formula does not apply, so they are left out.
The 4B model is the one that catches you. At 8,192 tokens, qwen3-4b-2507 needs 1.125 GiB of cache; at its default context it asks for 36.0 GiB of cache alone, more than the whole memory of any single card in our catalogue. Our model finder and hardware sheets calculate everything at 8,192 tokens of context. --max-model-len 8192 (or 8K, with a capital K: a lowercase 8k means 8,000) makes vLLM behave like the figure on our site. With --max-model-len auto, or -1, vLLM takes the longest context that fits.
Which models fit, and why our GGUF sizes are only a guide
Our hardware sheets set 0.8 GiB aside for the system on NVIDIA cards and compare the rest against the 19 models we measure. vLLM draws its own line at 92%, so the 2 budgets differ a little:
| Machine | Our sheet leaves for the model | vLLM’s budget (0.92 × memory), for the model, its working memory and the KV cache | Models that fit (sheet, 8K) | Largest that fits (sheet) |
|---|---|---|---|---|
| RTX 4060 8 GB | 7.20 GiB | 7.36 GiB | 6 of 19 | granite-4.2-8b |
| RTX 3060 12 GB and RTX 4070 12 GB | 11.20 GiB | 11.04 GiB | 7 of 19 | gemma4-12b |
| RTX 4090 24 GB | 23.20 GiB | 22.08 GiB | 15 of 19 | qwen3.6-35b-a3b |
| 2× RTX 3090 48 GB | 47.20 GiB | 2 × 22.08 = 44.16 GiB, with --tensor-parallel-size 2 |
15 of 19 | qwen3.6-35b-a3b |
Sheet figures as served on 28 September 2026, from the machine data of 23 September. vLLM’s budget is our calculation from its 0.92 default.
On the 24 GB card, vLLM leaves itself 1.12 GiB less than our sheet counts; on the 8 GB card, 0.16 GiB more. A model that only just fits on the sheet may need --gpu-memory-utilization raised on a 24 GB card.
The bigger difference is the file. Our sizes are GGUF files at Q4_K_M. vLLM does not load GGUF on its own: that support now lives in a separate plugin, vllm-gguf-plugin, which the documentation calls highly experimental and under-optimised. If you want to try it:
uv pip install vllm-gguf-plugin
Models are then named as repo_id:quant_type, for example with Q4_K_M. The usual route in vLLM is a different compression: AWQ, GPTQ or FP8 files from Hugging Face, or MXFP4 in gpt-oss. We have not measured those sizes for our 19 models, so take the finder’s verdict as a guide for vLLM, not as the same number.
How to lock it down: 0.0.0.0, no key and open CORS
Out of the box, vllm serve listens on 0.0.0.0:8000, and its startup log prints exactly that. It answers on localhost and on every other network interface of the machine: anyone on your network, or on the internet if the port is open, can use it. It checks no key. And CORS is open to any origin, any method and any header, so a web page you open in your browser can send requests to it and read the answers.
What closes it:
export VLLM_API_KEY="a-long-random-string"
vllm serve Qwen/Qwen3-0.6B --host 127.0.0.1 --max-model-len 8192
--host 127.0.0.1keeps other machines out. If one needs the server, reach it through an SSH tunnel (ssh -L 8000:127.0.0.1:8000 you@your-server).--api-key, or theVLLM_API_KEYvariable, makes clients send the key as a Bearer token. The variable keeps the key off the command line, where any user of the machine can read it in the process list.--allowed-origins '["http://localhost:3000"]', a JSON list, names the pages that may call it instead of the default*.
What the key does not cover. The key protects only routes under 4 prefixes: /v1, /v2, /inference and /cohere. Everything else on the same port answers without it: /health, /version, /tokenize, /metrics (Prometheus, always mounted) and /invocations, which reaches the same inference functions as /v1. So do the interactive API pages /docs and /redoc and the schema at /openapi.json, unless you start with --disable-fastapi-docs. vLLM’s security page also lists /pause, which stops generation, and /update_weights; in the 0.30.0 code they are only mounted when VLLM_SERVER_DEV_MODE=1 is set, so leave that variable unset. vLLM’s security page puts it plainly: Do not rely exclusively on --api-key. Its recommendation for anything reachable by others is a reverse proxy in front of vLLM that passes only the routes you intend to expose.
A second listener. vLLM uses PyTorch’s torch.distributed, and when it sets it up over TCP, as it does across several machines or on AMD cards with AITER’s all-reduce, PyTorch’s TCPStore listens on all network interfaces by default. On 1 machine with NVIDIA cards, the 0.30.0 code uses a file instead. vLLM’s own guidance is a firewall that blocks every incoming connection except the port of the API server. --host 127.0.0.1 closes the API; the firewall rule covers what vLLM’s own options do not reach.
What vLLM sends out on its own, and how to switch it off
Usage statistics are on by default and go to https://stats.vllm.ai. In the 0.30.0 code they leave in 2 shapes:
- 1 full report when the engine starts, from a separate thread: a random run ID; the GPU count, model and memory; the CUDA version; the cloud provider, if it detects one; the processor architecture and operating system string; total memory; the processor’s core count, brand, family, model and stepping; the vLLM version; the model’s architecture, not its name; the context length; 5 vLLM environment variables; and startup parameters such as the data type,
gpu_memory_utilization, quantisation, parallelism,max_model_lenandmax_num_seqs. - A heartbeat every 10 minutes for as long as the server runs, carrying only the run ID and a timestamp. That is 144 heartbeats in a day.
Network errors are ignored silently. For comparison, the Ollama guide counts 30 calls a day with its desktop app open.
You can read what was sent. Every report is also appended to a file on your disk:
tail ~/.config/vllm/usage_stats.json
4 ways switch it off, and any one is enough:
export VLLM_NO_USAGE_STATS=1
export DO_NOT_TRACK=1
export VLLM_DO_NOT_TRACK=1
mkdir -p ~/.config/vllm && touch ~/.config/vllm/do_not_track
The value has to be exactly 1. The code compares the text against "1", so DO_NOT_TRACK=true leaves the statistics on. In Docker, pass it with -e VLLM_NO_USAGE_STATS=1.
Hugging Face. The other traffic is the model download. vLLM identifies itself to Hugging Face as vllm, and fetches any model you name that is not already cached. Once your models are downloaded:
export HF_HUB_OFFLINE=1
With that set, no call goes to Hugging Face and a missing model fails instead of downloading. vLLM has no update check and no account.
What stays on disk. Models go to the Hugging Face cache in ~/.cache/huggingface, which HF_HOME moves. vLLM’s own files live in ~/.config/vllm (the statistics file and the do_not_track switch) and in its cache folder, ~/.cache/vllm. There are no chats: the server keeps no copy of your requests.
The order that leaves you with a working and closed vLLM:
- Check the card: NVIDIA with compute capability 7.5 or newer, on Linux or WSL.
- Create a Python 3.12 environment with
uv, installvllm==0.30.0with--torch-backend=auto, and confirm it withvllm --version. - Set
VLLM_NO_USAGE_STATS=1before the firstvllm serve, so not even the startup report leaves. - Start with
--host 127.0.0.1,--max-model-len 8192and a key inVLLM_API_KEY. - If the card also drives your screen, lower
--gpu-memory-utilizationuntil it starts. - Once the model is downloaded, set
HF_HUB_OFFLINE=1.
The 5 runtimes are compared on exactly this in 5 local AI runtimes: telemetry, network calls and audit.