When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding
By Aaryam Sharma · Paper · cs.LG
Speculative decoding accelerates language model inference by using a fast drafter to propose candidate tokens that are then verified by a larger target model. Existing theory largely studies the stochastic, distribution-preserving setting, where the goal is to exactly sample from