Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
By Tate Berenbaum · Paper · cs.DC
Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not enough memory to fit a large model such as a 70B-parameter LLM. We show that a handful of AIPCs, working together over an ordinary network,