Publications and Manuscripts
This page lists my current manuscripts and public research artifacts. Venue labels below indicate current targets or intended submission tracks, not accepted publications.
Manuscripts in progress
OrchKvCache: Orchestrating LLM KV-Cache across GPU-DRAM-SSD with Attention-Aware Hotness Scheduling.
Targeting SC 2026.
Project artifact: OrchKvCache.HALO: Algebraically Identity-Preserving KV Tiering for Memory-Bounded Long-Context Inference.
Targeting EMNLP 2026 (ARR).
Project artifact: HALO.SEER: Schedulable KV-Cache Eviction with a Probabilistic Sizing Bound for Real-Time LLM Decoding.
Targeting IEEE RTSS 2026.
Project artifact: SEER.Mining Attention Dynamics: Auditing KV-Saliency Prediction in LLMs.
Targeting IEEE ICDM 2026.
Project artifact: XQP / KVSalienceBench.Direction Is Not a Knob: Cross-GPU KV-Cache Handoff Contention on NVLink.
Targeting IEEE ICCD 2026.
Project artifact: PeerKV.
Research theme
These manuscripts are part of a single research portfolio on long-context LLM memory systems: tiered KV-cache management, semantic-safe offloading, SLO-aware scheduling, attention-dynamics-based saliency prediction, and interconnect-aware KV movement.
For project-level context, see the Research Systems page.
