Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

By Tate Berenbaum · Paper · cs.DC

Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not enough memory to fit a large model such as a 70B-parameter LLM. We show that a handful of AIPCs, working together over an ordinary network,

Cs.dc

View original

HomeResourceLoading…