vLLM v0.31.0 adds a vllm preload command that launches its weight-cache daemon

TL;DR

  • On 5 October the vLLM project published v0.31.0, with 717 commits from 307 contributors.
  • A new vllm preload command launches the weight-cache daemon that keeps post-quantized weights in GPU memory across engine restarts, and DeepSeek-V4.1-Flash gets new defaults on SM100 GPUs.
  • Per-request mm_processor_kwargs and media_io_kwargs now return an error unless the server runs with --trust-request-mm-kwargs.

Read the full story

Sign in with your email to read AI News. It’s free.

Share this story

Explain like I’m 15