DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees

By Tianyi Li · Paper · cs.LG

Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal

Cs.lg

View original

HomeResourceLoading…