Back to AI intel
重点

vLLM v0.31.0 Boosts DeepSeek-V4.1-Flash Performance with FlashMLA and NVFP4 KV Cache

AI intel briefing

Core summary

One sentence to understand this update

vLLM v0.31.0 features 717 commits and significantly improves DeepSeek-V4.1-Flash performance using FlashMLA mega attention with NVFP4 compressed KV cache as the SM100 default.

Impact & opportunity

What this could mean

Users of vLLM can now achieve faster and more efficient inference for DeepSeek-V4.1-Flash, particularly on SM100 hardware.