←Hall of Fame

vllm-project/vllm#1348

Implement PagedAttention V2

PagedAttention V2 added sequence-level parallelism to speed up small batches.

6 files·+764/-139·3 lanes

Architecture

Data flow

View the pull request on GitHub

◈ Rendered by PR Lens · crafted with ❤️ by the Coldtea team