Aivora

#parallel-decoding

Parallel Decoding

1 article

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs
arXivLLM

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast Diffusion LLMs

Flash-dLLM:為擴散大語言模型打造的 I/O 感知 KV 快取與平行解碼加速框架

Flash-dLLM is a training-free inference acceleration framework that dramatically speeds up Diffusion LLMs (dLLMs) using an IO-aware fused KV cache and a self-contained draft-and-verify decoding strategy.

2 min read