What is vLLM and why is it better than naive serving?
vLLM is an LLM inference server that uses PagedAttention — a technique that manages KV cache memory like an OS page table, allowing multiple requests to share GPU memory efficiently. Combined with continuous batching (dynamically grouping in-flight requests), vLLM achieves 10-30x throughput compared to serving one request at a time. It also exposes an OpenAI-compatible API, making migration from the OpenAI API straightforward.