Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Flash-dLLM optimizes KV caching and parallel decoding for diffusion LLMs via IO-aware techniques, addressing inference bottlenecks in non-autoregressive text generation.