The State-Prediction Separation Hypothesis

By Giovanni Monea · Paper · cs.CL

Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emph{state-prediction separation hypothesis}: disentangling the two roles yields better language modeling performance. We des

Cs.cl

View original

HomeResourceLoading…