Orchestrating LLM KV-Cache across GPU–DRAM–SSD with Attention-Aware Hotness Scheduling

Published in Targeting SC 2026, 2026

Manuscript and artifact in progress. Project artifact: OrchKvCache.

Recommended citation: Ziqing Li, "Orchestrating LLM KV-Cache across GPU–DRAM–SSD with Attention-Aware Hotness Scheduling." Targeting SC 2026.