CV
Ziqing Li
Ph.D. Student in Computer Architecture
Summary
Ph.D. student at Huazhong University of Science and Technology working on memory systems for long-context LLM inference, with a focus on KV-cache management, storage systems, GPU memory systems, and MLSys runtime support.
Education
- Ph.D. student in Computer ArchitecturepresentHuazhong University of Science and Technology
- B.S. in Computer Science and Technology2023Beijing Jiaotong University
Work Experience
- Software Engineering Intern, Data Platform / OLAP Engine Team2023-11 - 2024-02Xiaomi GroupMentor: Yaodong Zhang, Senior Software Engineer and Apache Spark committer.
- Implemented column-level encryption for Parquet files in Apache Spark and Trino.
- Helped protect sensitive columns independently in internal OLAP pipelines.
Skills
Programming
- C/C++
- Python
- Java
- Scala
- MATLAB
Systems
- Linux systems programming
- Git
- GDB
- SPDK
- GPUDirect Storage
- HDFS
Research Areas
- Long-context LLM inference
- KV-cache management
- Storage systems
- GPU memory systems
- CXL and memory disaggregation
- MLSys
Publications
- OrchKvCache: Orchestrating LLM KV-Cache across GPU-DRAM-SSD with Attention-Aware Hotness Scheduling2026
- HALO: Algebraically Identity-Preserving KV Tiering for Memory-Bounded Long-Context Inference2026
- SEER: Schedulable KV-Cache Eviction with a Probabilistic Sizing Bound for Real-Time LLM Decoding2026
- Mining Attention Dynamics: Auditing KV-Saliency Prediction in LLMs2026
- Direction Is Not a Knob: Cross-GPU KV-Cache Handoff Contention on NVLink2026
Portfolio
- OrchKvCache2026Tiered kv-cache substrateTiered KV-cache management across GPU HBM, host DRAM, NVM, and SSD.
- HALO2026Semantic-safe offloadingIdentity-preserving KV offloading through chunked attention and log-sum-exp merge.
- SEER2026Slo-aware schedulingSLO-aware KV-cache eviction and prefetch under P99 / P99.9 latency targets.
- XQP2026Saliency measurementMeasurement-driven KV-block saliency prediction from LLM attention dynamics.
- PeerKV2026
Interests
- Memory systems for long-context LLM inferenceKV-cache management, offloading, compression, SLO-aware scheduling, cross-GPU KV handoff
- Storage systems and heterogeneous IOGPU-SSD data paths, GPU interconnects, NVM/SSD tiering, CXL, memory disaggregation