Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

By arnav__1 · Show HN

Hey HN, we’re the developers of OpenLake, an open source storage engine for offloading LLM KV caches from GPU memory into a shared tier of RAM and NVMe. We built OpenLake because KV caches are outgrowing GPU memory.A single 256K token conversation on Gemma 4 31B produces approxim

AI Infra · Story 49057767

View original

HomeResourceLoading…