Inference with vLLM on Aurora¶
vLLM is an open-source library designed to optimize the inference and serving. Originally developed at UC Berkeley's Sky Computing Lab, it has evolved into a community-driven project. The library is built around the innovative PagedAttention algorithm, which significantly improves memory management by reducing waste in Key-Value (KV) cache memory.
Provided Installation¶
vLLM is installed and available as part of the frameworks module. Please use the following commands to load the installation and query the version info:
Then, you can import vllm as follows
Known Issue on Aurora¶
CCL_PROCESS_LAUNCHER is set to pmix through the frameworks module, which leads to a warning |CCL_WARN| PMIx_Init failed: PMIX_ERR_UNREACH, but it appears that vllm recovers, and performance is not affected. Cleanest is to set this variable either to none or torchrun. Based on our tests, we have found setting this to be optional.
Tip
Do not forget to set the proxies from the compute node, if performing direct download on the job.
Access Model Weights¶
Model weights for commonly used open-weight models are downloaded and available in the following directory on Aurora.
To ensure your workflows utilize the preloaded model weights and datasets, update the following environment variables in your session. Some models hosted on Hugging Face (HF) may be gated, requiring additional authentication. To access these gated models, you will need a Hugging Face authentication token.
Set the following environment variables to avoid hitting the Hugging Face Hub API, if a model is already cached.
Serving Small Models on a Single Tile¶
For small models that fit within the memory of a single PVC tile (64 GB), no additional configuration is required to serve the model. Simply use the default tensor parallelism size (TP) of 1 when serving the model. This ensures the model is run on a single tile without the need for distributed setup. Models with fewer than 20 billion parameters typically fit within a single tile when using half precision (e.g., bfloat16).
For example, the following command serves meta-llama/Llama-3.1-8B-Instruct on a single tile of a single node:
After the server output says Application startup complete., it is ready to accept prompt requests, for example with the following simple Python code (can background the server process)
Serving Medium Models on Multiple Tiles (Single Node)¶
To serve larger models which require multiple tiles (TP>1) but still only a single node, a more advanced setup is necessary. This involves setting the VLLM_HOST_IP and the TP size. By default, vLLM uses the mp backend, which is sufficient for single node model serving. Models with a few hundred billion parameters can usually fit within a single node utilizing half precision.
The following commands demonstrate how to serve the meta-llama/Llama-3.3-70B-Instruct on 8 tiles on a single node.
Serving Large Models on Multiple Tiles and Nodes¶
To serve models across multiple nodes, vLLM uses Ray to launch processes across nodes. We suggest users take advantage of the provided setup_ray_cluster.sh script to setup a Ray cluster across nodes before running vllm serve.
Setup script
| setup_ray_cluster.sh | |
|---|---|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 | |
The following example serves meta-llama/Llama-3.1-405B-Instruct model using 2 nodes. It sets tensor parallelism TP=8 for intra-node communications and pipeline parallelism PP=2 for inter-node communication, for a total of 16 shards.
Guidelines for vLLM Model Serving
- By default,
setup_ray_cluster.shlaunches a ray cluster with 8 raylets per node by specifingONEAPI_DEVICE_SELECTOR="opencl:gpu;level_zero:0,1,2,3,4,5,6,7"and--num-gpus=8. This matches thevllm servecommand withTP=8. This is the recommended setup, however if users want to change the TP size, remember to also change the number of GPUs in the Ray setup script. - Setting
--max-model-lencan be important in order to fit the model on the GPUs. - Tensor parallelism size must evenly divide the number of attention heads of the model. For example, the
Llama-3.1-70B-Instructmodel has 64 attention heads, so validTPvalues are 1, 2, 4, 8. On Aurora, settingTPsize equal to the number of GPUs on the node (12 PVC tiles per node) is usually not the preferred approach;TP=4,8are preferred instead. - Pipeline parallelism size must evenly divide the number of hidden layers in the model. For example, the
Llama-3.1-70B-Instructmodel has 80 layers, soPPvalues of 1, 2, 4, 5, etc. are valid. Usually,PPis set to the number of nodes used, which was 2 in the case above. -
The product \(\text{TP} \times \text{PP}\) indicates the total number of GPUs used to serve the model, which is usually defined by the number of parameters in the model and the memory of the individual GPUs. As a back of the envelope calculation, when using half precision such as
bfloat16(2 bytes per parameter), the minimum number of GPUs needed to hold the model weights is\[ N_\text{GPU} \geq \frac{2 \times \text{parameters (billions)}}{\text{GPU memory (GB)}} \]For example,
Llama-3.1-405B-Instructon 64 GB PVC tiles needs \(2 \times 405 / 64 \approx 12.7\), so at least 13 GPUs. The command above uses \(\text{TP} \times \text{PP} = 8 \times 2 = 16\), which leaves room for the KV cache.
Scaling vLLM Workflows¶
To scale vLLM workflows on ALCF system there are a few recommended approaches depending on the user's needs and setup. These approaches are described in detail in the GettingStarted repository along with example scripts for each.