When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding

By Aaryam Sharma · Paper · cs.LG

Speculative decoding accelerates language model inference by using a fast drafter to propose candidate tokens that are then verified by a larger target model. Existing theory largely studies the stochastic, distribution-preserving setting, where the goal is to exactly sample from

Cs.lg

View original

HomeResourceLoading…