Publications and Manuscripts

This page lists my current manuscripts and public research artifacts. Venue labels below indicate current targets or intended submission tracks, not accepted publications.

Manuscripts in progress

  1. OrchKvCache: Orchestrating LLM KV-Cache across GPU-DRAM-SSD with Attention-Aware Hotness Scheduling.
    Targeting SC 2026.
    Project artifact: OrchKvCache.

  2. HALO: Algebraically Identity-Preserving KV Tiering for Memory-Bounded Long-Context Inference.
    Targeting EMNLP 2026 (ARR).
    Project artifact: HALO.

  3. SEER: Schedulable KV-Cache Eviction with a Probabilistic Sizing Bound for Real-Time LLM Decoding.
    Targeting IEEE RTSS 2026.
    Project artifact: SEER.

  4. Mining Attention Dynamics: Auditing KV-Saliency Prediction in LLMs.
    Targeting IEEE ICDM 2026.
    Project artifact: XQP / KVSalienceBench.

  5. Direction Is Not a Knob: Cross-GPU KV-Cache Handoff Contention on NVLink.
    Targeting IEEE ICCD 2026.
    Project artifact: PeerKV.

Research theme

These manuscripts are part of a single research portfolio on long-context LLM memory systems: tiered KV-cache management, semantic-safe offloading, SLO-aware scheduling, attention-dynamics-based saliency prediction, and interconnect-aware KV movement.

For project-level context, see the Research Systems page.