Orchestrating LLM KV-Cache across GPU–DRAM–SSD with Attention-Aware Hotness Scheduling
Published in Targeting SC 2026, 2026
Manuscript and artifact in progress. Project artifact: OrchKvCache.
Recommended citation: Ziqing Li, "Orchestrating LLM KV-Cache across GPU–DRAM–SSD with Attention-Aware Hotness Scheduling." Targeting SC 2026.
