DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
By Tianyi Li · Paper · cs.LG
Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal